AI coding hallucinations
OngoingIn our ongoing testing, ReGrade reduces AI coding hallucinations by over 79%.
An AI coding hallucination is an unintended change, addition or deletion to working code that the agent was never asked to modify.
Curtail research · 74 models, 23 labs, updated 2026-09-30
Every AI lab we tested produced AI coding hallucinations. All 23 labs. The strongest models included.
With ReGrade AI coding hallucinations dropped to zero for these models.
A hallucination is counted by replaying each trial's saved patch offline and counting the changes that fall outside the bugs the model was asked to fix. Each change is traced back to the edit that caused it, by removing one edit at a time and replaying, so one mistake that surfaces at nine call sites counts once. On the typical model, ReGrade's patented behavioral comparison cut those changes by over 79%. In the main comparisons, no unintended changes were observed in the ReGrade trials on the 20 models below, across 10 labs. Zero observed changes is not a guarantee on future code.
- OpenAIgpt-6-astra, gpt-6-sol, gpt-5.5, gpt-5.6-sol
- Xiaomimimo-v2.6-pro
- xAIgrok-4.5, grok-3-fast
- Metamuse-spark-1.3
- Anthropicfable-5, opus-5.5, fable51, opus-4.8, opus-5, opus-4.6
- Moonshotkimi-k2.7-code
- Googlegemini-3.1-pro-preview
- Sakanafugu-ultra
- Alibabaqwen3.8-max-prime, qwen3.7-max
- Nvidianemotron-3.5-lightning
Every model tested, on AI coding hallucinations
| Hallucinations | Bug Fixes | |||||||
|---|---|---|---|---|---|---|---|---|
| Verdict | ||||||||
| gpt-6-astra | 6.33 | 0 | −100% | 14.7 | 18.0 | OpenAI | 3 | ✓ |
| gpt-6-sol | 3.00 | 0 | −100% | 14.3 | 18.0 | OpenAI | 3 | ✓ |
| muse-spark-1.1 | 2.75 | 0.33 | −88% | 13.2 | 17.7 | Meta | 3 | ✓ |
| gpt-5.4 | 3.16 | 0.24 | −92% | 13.1 | 17.9 | OpenAI | 31 | ✓ |
| gpt-5.5 | 2.93 | 0 | −100% | 13.0 | 18.0 | OpenAI | 28 | ✓ |
| gpt-5.6-sol | 2.92 | 0 | −100% | 12.9 | 18.0 | OpenAI | 24 | ✓ |
| fable51 | 2.00 | 0 | −100% | 12.7 | 18.0 | Anthropic | 3 | ✓ |
| kimi-k3 | 5.33 | 0.33 | −94% | 12.3 | 17.0 | Moonshot | 3 | ✓ |
| grok-4.7 | 11.3 | 1.00 | −91% | 12.3 | 17.3 | xAI | 3 | ✓ |
| opus-5 | 1.83 | 0 | −100% | 11.9 | 18.0 | Anthropic | 30 | ✓ |
| gpt-5.6-terra | 2.32 | 0.07 | −97% | 11.8 | 17.9 | OpenAI | 28 | ✓ |
| fable-5 | 3.31 | 0 | −100% | 11.5 | 18.0 | Anthropic | 16 | ✓ |
| muse-spark-1.3 | 3.40 | 0 | −100% | 11.4 | 17.8 | Meta | 5 | ✓ |
| sonnet-5.5 | 2.00 | 0.67 | −67% | 11.3 | 16.3 | Anthropic | 3 | ✓ |
| grok-4.5 | 3.75 | 0 | −100% | 11.2 | 18.0 | xAI | 4 | ✓ |
| qwen3.8-max-prime | 2.00 | 0 | −100% | 10.7 | 18.0 | Alibaba | 3 | ✓ |
| gemini-3.1-pro-preview | 2.40 | 0 | −100% | 10.5 | 16.0 | 10 | ✓ | |
| opus-5.5 | 3.00 | 0 | −100% | 10.3 | 18.0 | Anthropic | 3 | ✓ |
| fugu-ultra | 2.40 | 0 | −100% | 10.2 | 18.0 | Sakana | 5 | ✓ |
| mimo-v2.6-pro | 4.00 | 0 | −100% | 10.0 | 18.0 | Xiaomi | 3 | ✓ |
| glm-5.3-prime | 2.33 | 1.33 | −43% | 10.0 | 16.7 | Z.ai | 3 | ✓ |
| inkling | 3.67 | 1.67 | −54% | 9.3 | 15.3 | ThinkingMachines | 3 | ✓ |
| opus-4.8 | 1.97 | 0 | −100% | 9.1 | 17.9 | Anthropic | 21 | ✓ |
| qwen3.8-max-0902 | 2.33 | 0.40 | −83% | 9.0 | 17.6 | Alibaba | 3 | ✓ |
| deepseek-v4-pro (Sep 2026) | 4.71 | 0.14 | −97% | 9.0 | 17.7 | DeepSeek | 7 | ✓ |
| gpt-6-luna | 2.67 | 2.00 | −25% | 8.7 | 14.3 | OpenAI | 3 | ✓ |
| glm-5.3 | 2.00 | 1.00 | −50% | 8.7 | 16.3 | Z.ai | 3 | ✓ |
| opus-4.7 | 1.91 | 0.03 | −98% | 8.6 | 17.9 | Anthropic | 33 | ✓ |
| deepseek-v4.1-flash | 2.00 | 0.67 | −67% | 8.3 | 17.3 | DeepSeek | 3 | ✓ |
| gpt-5.6-luna | 1.47 | 0.70 | −52% | 7.5 | 16.8 | OpenAI | 30 | ✓ |
| qwen3.7-max | 1.67 | 0 | −100% | 7.3 | 16.7 | Alibaba | 3 | ✓ |
| minimax-m3 | 1.33 | 0.33 | −75% | 7.3 | 17.7 | MiniMax | 3 | ✓ |
| opus-4.6 | 1.33 | 0 | −100% | 7.0 | 17.7 | Anthropic | 3 | ✓ |
| longcat-2.0 | 3.00 | 0.33 | −89% | 6.7 | 11.0 | Meituan | 3 | ✓ |
| glm-5.2 | 1.85 | 0.74 | −60% | 6.7 | 16.6 | Z.ai | 19 | ✓ |
| mistral-medium-3.5 | 4.00 | 2.10 | −47% | 6.4 | 14.9 | Mistral | 10 | ✓ |
| gemini-3-flash-preview | 1.00 | 0.67 | −33% | 6.3 | 17.5 | 3 | ✓ | |
| hy3 | 2.75 | 1.20 | −56% | 6.0 | 14.0 | Tencent | 4 | ✓ |
| qwen3.6-max-preview | 2.00 | 0.25 | −87% | 5.7 | 12.8 | Alibaba | 3 | ✓ |
| sonnet-5 | 2.00 | 0.27 | −87% | 5.5 | 17.4 | Anthropic | 30 | ✓ |
| devstral-2 | 4.40 | 0.90 | −79% | 5.4 | 13.0 | Mistral | 30 | ✓ |
| deepseek-v4-pro-0423 | 1.00 | 0.50 | −50% | 5.3 | 16.2 | DeepSeek | 3 | ✓ |
| sonnet-4.6 | 1.12 | 0.03 | −97% | 5.2 | 17.7 | Anthropic | 33 | ✓ |
| kimi-k2.7-code | 3.00 | 0 | −100% | 5.0 | 15.7 | Moonshot | 3 | ✓ |
| sakana-namazu | 1.75 | 0.20 | −89% | 5.0 | 12.0 | Sakana | 4 | ✓ |
| qwen3-coder | 3.50 | 0.50 | −86% | 4.7 | 12.5 | Alibaba | 10 | ✓ |
| glm-5.1 | 1.40 | 0.53 | −62% | 4.5 | 17.6 | Z.ai | 30 | ✓ |
| kimi-k26 | 1.57 | 0.50 | −68% | 4.4 | 16.0 | Moonshot | 7 | ✓ |
| deepseek-v4-flash-0731 | 1.00 | 0.60 | −40% | 4.1 | 16.7 | DeepSeek | 9 | ✓ |
| qwen3.6-plus | 0.67 | 0.33 | −50% | 4.0 | 16.0 | Alibaba | 3 | ✓ |
| deepseek-v4-flash-0423 | 2.00 | 0.33 | −83% | 3.8 | 7.0 | DeepSeek | 3 | ✓ |
| solar-pro4 | 1.20 | 1.00 | −17% | 3.8 | 10.6 | Upstage | 4 | ✓ |
| gpt-5.4-mini | 1.60 | 0.60 | −62% | 3.6 | 13.2 | OpenAI | 5 | ✓ |
| step-3.7-flash | 2.86 | 1.04 | −64% | 3.4 | 11.4 | StepFun | 25 | ✓ |
| deepseek-v4-pro-0813 | 1.00 | 0.62 | −38% | 3.2 | 13.5 | DeepSeek | 9 | ✓ |
| grok-4.20-0309-reasoning | 7.00 | 5.60 | −20% | 3.2 | 11.7 | xAI | 10 | ✓ |
| nemotron-3-ultra | 1.60 | 1.40 | −12% | 3.0 | 13.8 | Nvidia | 5 | ✓ |
| seed-2.0-lite | 2.50 | 0.43 | −83% | 2.9 | 14.7 | ByteDance | 16 | ✓ |
| haiku-4.5 (Aug 2026) | 1.70 | 1.60 | −6% | 2.7 | 14.5 | Anthropic | 10 | ✓ |
| laguna-s-2.1 | 1.33 | 0.33 | −75% | 2.7 | 15.3 | Poolside | 3 | ✓ |
| haiku-4.5 | 2.30 | 1.23 | −46% | 2.5 | 12.8 | Anthropic | 30 | ✓ |
| grok-4-1-fast-non-reasoning | 3.30 | 1.11 | −66% | 2.5 | 9.7 | xAI | 9 | ✓ |
| minimax-m2.5 | 1.20 | 1.13 | −6% | 2.1 | 10.9 | MiniMax | 30 | ✓ |
| qwen3-coder-plus | 1.00 | 0.33 | −67% | 2.0 | 12.0 | Alibaba | 3 | ✓ |
| gemini-2.5-flash | 3.00 | 1.67 | −44% | 2.0 | 14.0 | 3 | ✓ | |
| grok-3-fast | 2.67 | 0 | −100% | 2.0 | 8.0 | xAI | 3 | ✓ |
| grok-4.3 | 4.00 | 1.33 | −67% | 2.0 | 6.3 | xAI | 3 | ✓ |
| ling-3.0-flash | 1.08 | 1.00 | −7% | 1.9 | 11.3 | InclusionAI | 13 | ✓ |
| laguna-xs-2.1 | 1.80 | 1.60 | −11% | 1.8 | 10.3 | Poolside | 15 | ✓ |
| ling-2.6-1t | 1.80 | 1.75 | −3% | 1.6 | 7.0 | InclusionAI | 4 | ✓ |
| nemotron-3.5-lightning | 1.25 | 0 | −100% | 1.4 | 5.3 | Nvidia | 3 | ✓ |
| grok-3-mini | 0.67 | 0.33 | −50% | 1.0 | 2.8 | xAI | 3 | ✓ |
| gpt-oss-20b | did not edit | 0.40 | N/A | 0.0 | 3.8 | OpenAI | 5 | no baseline |
| kat-coder-pro-v2.5 | 0.75 | 0.25 | −67% | — | — | Kuaishou | 4 | ✓ |
✓ improved · * got worse · N/A: the model did not edit the file in the control run, so there is no baseline to compare against. Click a column to sort. Measured over 74 models, 23 labs, 1,558 agent-mode trials. The provider-pinned comparisons add 36 separately labelled trials, for 1,594 total. Min. replays/arm is the smaller eligible replay count of the control and treatment arms; individual arms can contain more runs. Scroll sideways to see the lab and sample-count columns. Bug fixes are unaided (mean bugs fixed across the model's control-arm trials) and aided (the same average over the treatment trials), so the table can be sorted by how capable a model is on its own and by how far it gets with help. Both are our own measurements. Repair and damage eligible subsets can differ, so bug-fix averages need not use every replay counted here. A dash means the model sits below the trial floor, so no figure is published for it.
The honest caveats. One model, gpt-oss-20b, broke nothing unaided because it did not edit the file at all, so it is reported as N/A rather than counted as a success. The main comparison results are in the table above, sortable and searchable. The provider-pinned results follow separately, with three trials per arm: July DeepSeek on Baidu reached zero observed changes, Nemotron Ultra's count rose, and GPT-OSS has incomplete damage measurements. Those samples do not establish stable effects.
Provider-pinned comparisons
Six separate cells, 36 trials, three trials per arm. Both arms of each cell used the same provider. These batches remain separate from the historical model comparisons above and their headline statistics. Every arrow reads without ReGrade → with ReGrade.
| Model | Provider and precision | Bugs fixed / 18 | Causal damage per trial |
|---|---|---|---|
| deepseek-v4-flash-0423 | DeepInfra · fp8 | 4.33 → 16.00 | 1.33 → 0.33 |
| deepseek-v4-flash-0731 | Baidu · fp8 | 5.67 → 17.00 | 2.67 to 3.33 → 0.00 |
| nemotron-3-ultra | BaseTen · fp4 | 4.33 → 15.00 | 1.33 → 1.67 |
| gpt-oss-20b | DeepInfra · bf16 | 0.33 → 2.33 | 0.00 → Unavailable (1/3 measured) |
| inkling | Together · precision unknown | 8.67 → 14.33 | 3.67 → 1.67 |
| glm-5.2 | Baidu · fp4 | 7.00 → 16.67 | 2.00 to 3.00 → 1.00 |
Higher repair scores and lower damage are better. Ranges include unresolved hunk-removal variants. Two GPT-OSS treatment patches contained invalid Python: their repair scores remain zero and their damage is unavailable. Nemotron introduced slightly more damage with ReGrade. Three trials per arm do not establish stable model effects.
🛠️ Works with every major coding agent and model API.
5 coding-agent CLIs · 74 models · 23 labs.
Coding-agent CLIs5 CLIs · 25 models
Claude Code
- fable-5
- fable51
- glm-5.1
- haiku-4.5
- haiku-4.5 (Aug 2026)
- opus-4.6
- opus-4.7
- opus-4.8
- opus-5
- opus-5.5
- sonnet-4.6
- sonnet-5
- sonnet-5.5
Codex CLI
- gpt-5.4
- gpt-5.5
- gpt-5.6-luna
- gpt-5.6-sol
- gpt-5.6-terra
- gpt-6-astra
- gpt-6-luna
- gpt-6-sol
qwen-code
- qwen3-coder
- qwen3-coder-plus
Kimi Code CLI
- kimi-k26
Mistral Vibe CLI
- devstral-2
Models tested via API, by lab49 models · 22 labs
xAI
- grok-3-fast
- grok-3-mini
- grok-4-1-fast-non-reasoning
- grok-4.20-0309-reasoning
- grok-4.3
- grok-4.5
- grok-4.7
DeepSeek
- deepseek-v4-flash-0423
- deepseek-v4-flash-0731
- deepseek-v4-pro (Sep 2026)
- deepseek-v4-pro-0423
- deepseek-v4-pro-0813
- deepseek-v4.1-flash
Alibaba
- qwen3.6-max-preview
- qwen3.6-plus
- qwen3.7-max
- qwen3.8-max-0902
- qwen3.8-max-prime
- gemini-2.5-flash
- gemini-3-flash-preview
- gemini-3.1-pro-preview
Z.ai
- glm-5.2
- glm-5.3
- glm-5.3-prime
InclusionAI
- ling-2.6-1t
- ling-3.0-flash
Meta
- muse-spark-1.1
- muse-spark-1.3
MiniMax
- minimax-m2.5
- minimax-m3
Moonshot
- kimi-k2.7-code
- kimi-k3
Nvidia
- nemotron-3-ultra
- nemotron-3.5-lightning
OpenAI
- gpt-5.4-mini
- gpt-oss-20b
Poolside
- laguna-s-2.1
- laguna-xs-2.1
Sakana
- fugu-ultra
- sakana-namazu
ByteDance
- seed-2.0-lite
Kuaishou
- kat-coder-pro-v2.5
Meituan
- longcat-2.0
Mistral
- mistral-medium-3.5
StepFun
- step-3.7-flash
Tencent
- hy3
ThinkingMachines
- inkling
Upstage
- solar-pro4
Xiaomi
- mimo-v2.6-pro
One ReGrade context block in the agent's prompt. No model retraining. No changes to your build or CI pipeline. Tested working across 74 models from 23 labs, US and non-US.
Reduce unintended changes reaching review.
Every lab we tested produced them, all 23 labs. ReGrade cut them by over 79% on the typical model. On 20 models, across 10 labs, no unintended changes were observed in the ReGrade trials.
Reduce unintended code changes.