AI coding hallucinations
OngoingIn our ongoing testing, ReGrade reduces AI coding hallucinations by over 73%.
An AI coding hallucination is an unintended change, addition or deletion to working code that the agent was never asked to modify.
Curtail research · 68 models, 22 labs, updated 2026-09-18
Where AI coding hallucinations stopped.
A hallucination is counted by replaying each trial's saved patch offline and counting the changes that fall outside the bugs the model was asked to fix. Each change is traced back to the edit that caused it, by removing one edit at a time and replaying, so one mistake that surfaces at nine call sites counts once. On the typical model, ReGrade's patented behavioral comparison cut those changes by over 73%. On the 17 models below, across 10 labs, it stopped them completely.
- OpenAIgpt-6-astra, gpt-5.5, gpt-5.6-sol
- xAIgrok-4.5, grok-3-fast
- Metamuse-spark-1.3
- Anthropicfable-5, fable51, opus-4.8, opus-5, opus-4.6
- Moonshotkimi-k2.7-code
- Sakanafugu-ultra
- Googlegemini-3.1-pro-preview
- Alibabaqwen3.7-max
- Tencenthy4-preview
- Nvidianemotron-3.5-lightning
Every AI lab we tested produced AI coding hallucinations. All 22 labs. The strongest models included.
The honest caveats. On 1 model hallucinations rose with ReGrade: deepseek-v4-pro-0813. A further 4 produced none unaided because they did not edit the file at all, so they are reported as N/A rather than counted as successes. Every model we tested is in the table below, sortable and searchable.
Every model tested, on AI coding hallucinations
| Hallucinations | Bug Fixes | |||||||
|---|---|---|---|---|---|---|---|---|
| Verdict | ||||||||
| gpt-6-astra | 10 | 0 | −100% | 15.0 | 18.0 | OpenAI | 3 | ✓ |
| gpt-5.4 | 2.89 | 0.22 | −92% | 13.9 | 17.9 | OpenAI | 35 | ✓ |
| gpt-5.6-sol | 2.33 | 0 | −100% | 13.9 | 18.0 | OpenAI | 30 | ✓ |
| gpt-5.5 | 2.88 | 0 | −100% | 13.3 | 18.0 | OpenAI | 33 | ✓ |
| muse-spark-1.1 | 2.75 | 0.33 | −88% | 13.2 | 17.7 | Meta | 3 | ✓ |
| fable51 | 2.00 | 0 | −100% | 12.7 | 18.0 | Anthropic | 3 | ✓ |
| qwen3.8-max-0902 | 1.40 | 0.40 | −71% | 12.6 | 17.6 | Alibaba | 5 | ✓ |
| kimi-k3 | 5.33 | 0.33 | −94% | 12.3 | 17.0 | Moonshot | 3 | ✓ |
| muse-spark-1.2 | 4.00 | 0.67 | −83% | 12.0 | 17.0 | Meta | 3 | ✓ |
| gpt-5.6-terra | 2.24 | 0.07 | −97% | 12.0 | 17.9 | OpenAI | 29 | ✓ |
| opus-5 | 1.83 | 0 | −100% | 11.9 | 18.0 | Anthropic | 30 | ✓ |
| fable-5 | 3.31 | 0 | −100% | 11.5 | 18.0 | Anthropic | 16 | ✓ |
| muse-spark-1.3 | 3.40 | 0 | −100% | 11.4 | 17.8 | Meta | 5 | ✓ |
| hy4-preview | 1.67 | 0 | −100% | 11.3 | 17.7 | Tencent | 3 | ✓ |
| grok-4.5 | 3.75 | 0 | −100% | 11.2 | 18.0 | xAI | 4 | ✓ |
| opus-4.7 | 1.66 | 0.03 | −98% | 10.6 | 17.8 | Anthropic | 38 | ✓ |
| fugu-ultra | 2.40 | 0 | −100% | 10.2 | 18.0 | Sakana | 5 | ✓ |
| gemini-3-flash-preview | 0.50 | 0.29 | −43% | 9.6 | 17.6 | 6 | ✓ | |
| inkling | 3.67 | 1.67 | −54% | 9.3 | 15.3 | ThinkingMachines | 3 | ✓ |
| opus-4.8 | 1.97 | 0 | −100% | 9.1 | 17.9 | Anthropic | 21 | ✓ |
| glm-5.3 | 2.00 | 1.00 | −50% | 8.7 | 16.3 | Z.ai | 3 | ✓ |
| gemini-3.1-pro-preview | 1.67 | 0 | −100% | 8.2 | 14.5 | 3 | ✓ | |
| gpt-5.6-luna | 1.47 | 0.70 | −52% | 7.5 | 16.8 | OpenAI | 30 | ✓ |
| qwen3.7-max | 1.67 | 0 | −100% | 7.3 | 16.7 | Alibaba | 3 | ✓ |
| sonnet-4.6 | 0.97 | 0.03 | −97% | 7.3 | 17.6 | Anthropic | 38 | ✓ |
| minimax-m3 | 1.33 | 0.33 | −75% | 7.3 | 17.7 | MiniMax | 3 | ✓ |
| opus-4.6 | 1.33 | 0 | −100% | 7.0 | 17.7 | Anthropic | 3 | ✓ |
| deepseek-v4-pro-0813 | 0.85 | 1.00 | +18% | 6.8 | 13.6 | DeepSeek | 13 | * |
| longcat-2.0 | 3.00 | 0.33 | −89% | 6.7 | 11.0 | Meituan | 3 | ✓ |
| glm-5.2 | 1.85 | 0.74 | −60% | 6.7 | 16.6 | Z.ai | 19 | ✓ |
| mistral-medium-3.5 | 4.00 | 2.10 | −47% | 6.4 | 14.9 | Mistral | 10 | ✓ |
| grok-3-fast | 1.33 | 0 | −100% | 6.3 | 5.5 | xAI | 6 | ✓ |
| hy3 | 2.75 | 1.20 | −56% | 6.0 | 14.0 | Tencent | 4 | ✓ |
| qwen3.6-max-preview | 2.00 | 0.25 | −87% | 5.7 | 13.8 | Alibaba | 3 | ✓ |
| sonnet-5 | 2.00 | 0.27 | −87% | 5.5 | 17.4 | Anthropic | 30 | ✓ |
| devstral-2 | 4.40 | 0.90 | −79% | 5.4 | 13.0 | Mistral | 30 | ✓ |
| deepseek-v4-pro | 1.00 | 0.50 | −50% | 5.3 | 16.6 | DeepSeek | 3 | ✓ |
| gemini-2.5-flash | 3.00 | 1.67 | −44% | 5.3 | 5.6 | 3 | ✓ | |
| kimi-k2.7-code | 3.00 | 0 | −100% | 5.0 | 15.7 | Moonshot | 3 | ✓ |
| sakana-namazu | 1.75 | 0.20 | −89% | 5.0 | 12.0 | Sakana | 4 | ✓ |
| deepseek-v4-flash-0731 | 0.91 | 0.60 | −34% | 4.8 | 16.7 | DeepSeek | 10 | ✓ |
| glm-5.1 | 1.40 | 0.53 | −62% | 4.5 | 17.6 | Z.ai | 30 | ✓ |
| kimi-k26 | 1.57 | 0.50 | −68% | 4.4 | 16.0 | Moonshot | 7 | ✓ |
| qwen3.6-plus | 0.67 | 0.33 | −50% | 4.0 | 16.0 | Alibaba | 3 | ✓ |
| deepseek-v4-flash | 2.00 | 0.33 | −83% | 3.8 | 7.0 | DeepSeek | 3 | ✓ |
| solar-pro4 | 1.20 | 1.00 | −17% | 3.8 | 10.6 | Upstage | 4 | ✓ |
| gpt-5.4-mini | 1.60 | 0.60 | −62% | 3.6 | 13.2 | OpenAI | 5 | ✓ |
| step-3.7-flash | 2.86 | 1.00 | −65% | 3.4 | 11.0 | StepFun | 26 | ✓ |
| grok-4.20-0309-reasoning | 7.00 | 5.60 | −20% | 3.2 | 12.3 | xAI | 10 | ✓ |
| nemotron-3-ultra | 1.60 | 1.40 | −12% | 3.0 | 13.8 | Nvidia | 5 | ✓ |
| seed-2.0-lite | 2.50 | 0.43 | −83% | 2.9 | 14.7 | ByteDance | 16 | ✓ |
| ling-3.0-flash | 1.00 | 1.00 | 0% | 2.9 | 11.3 | InclusionAI | 14 | no change |
| haiku-4.5 (Aug 2026) | 1.70 | 1.60 | −6% | 2.7 | 14.5 | Anthropic | 10 | ✓ |
| haiku-4.5 | 2.30 | 1.23 | −46% | 2.5 | 12.8 | Anthropic | 30 | ✓ |
| grok-4-1-fast-non-reasoning | 2.36 | 0.71 | −70% | 2.5 | 9.7 | xAI | 14 | ✓ |
| minimax-m2.5 | 1.20 | 1.13 | −6% | 2.1 | 10.9 | MiniMax | 30 | ✓ |
| qwen3-coder-plus | 1.00 | 0.33 | −67% | 2.0 | 11.8 | Alibaba | 3 | ✓ |
| grok-4.3 | 4.00 | 1.33 | −67% | 2.0 | 6.3 | xAI | 3 | ✓ |
| grok-3-mini | 0.67 | 0.33 | −50% | 1.9 | 1.9 | xAI | 3 | ✓ |
| laguna-xs-2.1 | 1.80 | 1.60 | −11% | 1.8 | 10.3 | Poolside | 15 | ✓ |
| qwen3-coder | 1.33 | 0.67 | −50% | 1.7 | 5.0 | Alibaba | 3 | ✓ |
| ling-2.6-1t | 1.80 | 1.75 | −3% | 1.6 | 7.0 | InclusionAI | 4 | ✓ |
| nemotron-3.5-lightning | 1.25 | 0 | −100% | 1.4 | 5.3 | Nvidia | 4 | ✓ |
| gemini-3.5-flash | did not edit | 0.33 | N/A | 0.0 | 16.7 | 3 | no baseline | |
| gemini-3.6-flash | did not edit | 0.17 | N/A | 0.0 | 8.2 | 6 | no baseline | |
| gpt-oss-20b | did not edit | – | N/A | 0.0 | 3.8 | OpenAI | 5 | no baseline |
| laguna-s-2.1 | did not edit | 0.33 | N/A | 0.0 | 10.0 | Poolside | 3 | no baseline |
| kat-coder-pro-v2.5 | 0.75 | 0.25 | −67% | — | — | Kuaishou | 4 | ✓ |
✓ improved · * got worse · N/A: the model did not edit the file in the control run, so there is no baseline to compare against. Click a column to sort. Measured over 68 models, 22 labs, 1,571 agent-mode trials. Bug fixes are unaided (mean bugs fixed across the model's control-arm trials) and aided (the same average over the treatment trials), so the table can be sorted by how capable a model is on its own and by how far it gets with help. Both are our own measurement from the same trials as the rest of this table, not a third party's score. — means the model sits below the trial floor, so no figure is published for it.
🛠️ Works with every major coding agent and model API.
5 coding-agent CLIs · 68 models · 22 labs.
Coding-agent CLIs5 CLIs · 21 models
Claude Code
- fable-5
- fable51
- glm-5.1
- haiku-4.5
- haiku-4.5 (Aug 2026)
- opus-4.6
- opus-4.7
- opus-4.8
- opus-5
- sonnet-4.6
- sonnet-5
Codex CLI
- gpt-5.4
- gpt-5.5
- gpt-5.6-luna
- gpt-5.6-sol
- gpt-5.6-terra
- gpt-6-astra
qwen-code
- qwen3-coder
- qwen3-coder-plus
Kimi Code CLI
- kimi-k26
Mistral Vibe CLI
- devstral-2
Models tested via API, by lab47 models · 21 labs
xAI
- grok-3-fast
- grok-3-mini
- grok-4-1-fast-non-reasoning
- grok-4.20-0309-reasoning
- grok-4.3
- grok-4.5
- gemini-2.5-flash
- gemini-3-flash-preview
- gemini-3.1-pro-preview
- gemini-3.5-flash
- gemini-3.6-flash
Alibaba
- qwen3.6-max-preview
- qwen3.6-plus
- qwen3.7-max
- qwen3.8-max-0902
DeepSeek
- deepseek-v4-flash
- deepseek-v4-flash-0731
- deepseek-v4-pro
- deepseek-v4-pro-0813
Meta
- muse-spark-1.1
- muse-spark-1.2
- muse-spark-1.3
InclusionAI
- ling-2.6-1t
- ling-3.0-flash
MiniMax
- minimax-m2.5
- minimax-m3
Moonshot
- kimi-k2.7-code
- kimi-k3
Nvidia
- nemotron-3-ultra
- nemotron-3.5-lightning
OpenAI
- gpt-5.4-mini
- gpt-oss-20b
Poolside
- laguna-s-2.1
- laguna-xs-2.1
Sakana
- fugu-ultra
- sakana-namazu
Tencent
- hy3
- hy4-preview
Z.ai
- glm-5.2
- glm-5.3
ByteDance
- seed-2.0-lite
Kuaishou
- kat-coder-pro-v2.5
Meituan
- longcat-2.0
Mistral
- mistral-medium-3.5
StepFun
- step-3.7-flash
ThinkingMachines
- inkling
Upstage
- solar-pro4
One ReGrade context block in the agent's prompt. No model retraining. No changes to your build or CI pipeline. Tested working across 68 models from 22 labs, US and non-US.
🌍 The newest labs, too.
The behavioral-diff effect is not a quirk of the big US providers. We ran the same test on the newest releases from five more labs, at 30 trials per model, and the rescue held every time.
| Lab | Model | Without | With | Δ |
|---|---|---|---|---|
| ByteDance | Seed-2.0-Lite | 2.9 / 18 | 14.7 / 18 | +408% |
| MiniMax | M2.5 | 2.1 / 18 | 10.9 / 18 | +427% |
| Zhipu | GLM-5.1 | 4.5 / 18 | 17.6 / 18 | +292% |
| StepFun | Step 3.7-flash | 3.4 / 18 | 11.0 / 18 | +221% |
| Mistral | Devstral 2 | 5.4 / 18 | 13.0 / 18 | +142% |
68 models across 22 labs, US and non-US, from frontier flagships down to free tiers. 65 of the 67 we could measure fixed more silent regressions with ReGrade.
Stop hallucinations reaching your codebase.
Every lab we tested produced them, all 22 labs. ReGrade cut them by over 73% on the typical model. On 17 models, across 10 labs, it stopped them completely.
Stop hallucinations in your codebase.