The Research
ReGrade reduces AI coding hallucinations by over 83%.
An AI coding hallucination is an unintended change, addition or deletion to working code that the agent was never asked to modify.
Every AI lab we tested produced them. All 22 labs. With ReGrade's patented behavioral comparison, they fell on 50 of the 58 models we could score, and stopped completely on 13 models across 8 labs. Measured over 1,431 agent-mode trials on 63 AI models from 22 labs. Every model is marked below, including the 8 that got worse.
-83%
on the typical model
The model stops changing code it was not asked to touch.
Every model we tested, with its hallucinations per trial before and after ReGrade. Sort any column, or search for a model or lab. The 8 models that got worse are marked, not hidden.
How we tested →| Verdict | |||||||
|---|---|---|---|---|---|---|---|
| gpt-5.6-sol | 4.90 | 0 | −100% | 78 | OpenAI | 30 | ✓ |
| opus-5 | 1.80 | 0 | −100% | 78 | Anthropic | 30 | ✓ |
| fable-5 | 3.33 | 0 | −100% | 77 | Anthropic | 12 | ✓ |
| kimi-k3 | 7.00 | 0 | −100% | 76 | Moonshot | 2 | ✓ |
| gpt-5.5 | 3.76 | 0 | −100% | 75 | OpenAI | 33 | ✓ |
| opus-4.8 | 2.23 | 0 | −100% | 74 | Anthropic | 21 | ✓ |
| grok-4.5 | 9.50 | 0 | −100% | 72 | xAI | 4 | ✓ |
| qwen3.8-max | 1.50 | 0 | −100% | 72 | Alibaba | 2 | ✓ |
| gemini-3.1-pro-preview | 3.00 | 0 | −100% | 69 | 3 | ✓ | |
| qwen3.7-max | 1.67 | 0 | −100% | 66 | Alibaba | 3 | ✓ |
| kimi-k2.7-code | 7.33 | 0 | −100% | 61 | Moonshot | 3 | ✓ |
| minimax-m3 | 2.00 | 0 | −100% | 59 | MiniMax | 1 | ✓ |
| nemotron-3.5-lightning | 2.75 | 0 | −100% | 27 | Nvidia | 4 | ✓ |
| opus-4.6 | 1.33 | 0 | −100% | — | Anthropic | 3 | ✓ |
| fugu-ultra | 2.40 | 0 | −100% | — | Sakana | 5 | ✓ |
| grok-3-fast | 5.33 | 0 | −100% | — | xAI | 6 | ✓ |
| grok-3-mini | 2.50 | 0 | −100% | — | xAI | 2 | ✓ |
| grok-4-1-fast-non-reasoning | 1.50 | 0 | −100% | — | xAI | 2 | ✓ |
| opus-4.7 | 2.13 | 0.03 | −99% | 74 | Anthropic | 38 | ✓ |
| sonnet-4.6 | 2.13 | 0.03 | −99% | 63 | Anthropic | 38 | ✓ |
| gpt-5.6-terra | 4.55 | 0.13 | −97% | 77 | OpenAI | 29 | ✓ |
| sakana-namazu | 4.25 | 0.20 | −95% | — | Sakana | 4 | ✓ |
| muse-spark-1.2 | 11.3 | 0.67 | −94% | 72 | Meta | 3 | ✓ |
| qwen3.6-max-preview | 3.00 | 0.25 | −92% | — | Alibaba | 2 | ✓ |
| grok-4.3 | 15 | 1.33 | −91% | 42 | xAI | 3 | ✓ |
| muse-spark-1.1 | 3.33 | 0.33 | −90% | 71 | Meta | 3 | ✓ |
| gpt-5.4 | 5.66 | 0.58 | −90% | 71 | OpenAI | 35 | ✓ |
| qwen3-coder-plus | 3.00 | 0.33 | −89% | — | Alibaba | 3 | ✓ |
| deepseek-v4-flash | 2.00 | 0.33 | −83% | 56 | DeepSeek | 3 | ✓ |
| kat-coder-pro-v2.5 | 3.00 | 0.50 | −83% | — | Kuaishou | 2 | ✓ |
| devstral-2 | 9.33 | 1.57 | −83% | 31 | Mistral | 30 | ✓ |
| glm-5.1 | 3.07 | 0.53 | −83% | 56 | Z.ai | 30 | ✓ |
| kimi-k26 | 2.71 | 0.50 | −82% | 62 | Moonshot | 7 | ✓ |
| gemini-3-flash-preview | 1.17 | 0.29 | −75% | — | 6 | ✓ | |
| gpt-5.4-mini | 3.60 | 1.00 | −72% | 56 | OpenAI | 5 | ✓ |
| seed-2.0-lite | 2.62 | 0.77 | −71% | — | ByteDance | 16 | ✓ |
| sonnet-5 | 4.07 | 1.80 | −56% | 72 | Anthropic | 30 | ✓ |
| qwen3-coder | 1.33 | 0.67 | −50% | — | Alibaba | 3 | ✓ |
| gpt-5.6-luna | 3.77 | 1.90 | −50% | 71 | OpenAI | 30 | ✓ |
| glm-5.2 | 2.60 | 1.58 | −39% | 69 | Z.ai | 19 | ✓ |
| gemini-2.5-flash | 11.3 | 7.00 | −38% | — | 1 | ✓ | |
| longcat-2.0 | 3.50 | 2.33 | −33% | 45 | Meituan | 2 | ✓ |
| step-3.7-flash | 5.17 | 3.54 | −32% | 40 | StepFun | 26 | ✓ |
| nemotron-3-ultra | 3.20 | 2.20 | −31% | 49 | Nvidia | 5 | ✓ |
| ling-2.6-1t | 4.20 | 3.25 | −23% | — | InclusionAI | 4 | ✓ |
| laguna-xs-2.1 | 4.67 | 3.67 | −21% | — | Poolside | 3 | ✓ |
| deepseek-v4-pro | 2.33 | 2.00 | −14% | 59 | DeepSeek | 3 | ✓ |
| minimax-m2.5 | 4.07 | 3.53 | −13% | — | MiniMax | 29 | ✓ |
| solar-pro4 | 2.20 | 2.00 | −9% | 53 | Upstage | 4 | ✓ |
| inkling | 7.67 | 7.33 | −4% | 52 | ThinkingMachines | 3 | ✓ |
| grok-4.20-0309-reasoning | 7.50 | 8.10 | +8% | — | xAI | 10 | * |
| hy3 | 3.75 | 4.20 | +12% | 59 | Tencent | 4 | * |
| deepseek-v4-flash-0731 | 2.20 | 3.10 | +41% | 69 | DeepSeek | 10 | * |
| haiku-4.5 | 4.20 | 5.93 | +41% | 44 | Anthropic | 30 | * |
| haiku-4.5 (Aug 2026) | 4.30 | 7.20 | +67% | 44 | Anthropic | 10 | * |
| ling-3.0-flash | 2.20 | 3.79 | +72% | 51 | InclusionAI | 14 | * |
| qwen3.6-plus | 1.33 | 2.33 | +75% | 55 | Alibaba | 3 | * |
| mistral-medium-3.5 | 2.20 | 6.00 | +173% | 47 | Mistral | 5 | * |
| gemini-3.5-flash | did not edit | 0.33 | N/A | 70 | 3 | no baseline | |
| gemini-3.6-flash | did not edit | 0.17 | N/A | 69 | 6 | no baseline | |
| gpt-oss-20b | did not edit | – | N/A | 21 | OpenAI | 5 | no baseline |
| grok-4.1-fast | did not edit | – | N/A | — | xAI | 3 | no baseline |
| laguna-s-2.1 | did not edit | 0.33 | N/A | — | Poolside | 3 | no baseline |
✓ improved · * got worse · N/A: the model did not edit the file in the control run, so there is no baseline to compare against. Click a column to sort. Measured over 63 models, 22 labs, 1,431 agent-mode trials. Capability is the Artificial Analysis Coding Index, retrieved 2026-08-18, best listed configuration per model. It is not measured by us; it is here so the table can be sorted by how capable a model is. — means the model is not listed there.
Where AI coding hallucinations stopped.
A hallucination is counted by replaying each trial's saved patch offline and counting the changes that fall outside the bugs the model was asked to fix. Counts are deduplicated to root causes, so one mistake that surfaces at nine call sites counts once. On the typical model, ReGrade's patented behavioral comparison cut those changes by over 83%. On the 13 models below, across 8 labs, it stopped them completely.
- xAIgrok-4.5, grok-3-fast
- Moonshotkimi-k2.7-code
- OpenAIgpt-5.6-sol, gpt-5.5
- Anthropicfable-5, opus-4.8, opus-5, opus-4.6
- Googlegemini-3.1-pro-preview
- Nvidianemotron-3.5-lightning
- Sakanafugu-ultra
- Alibabaqwen3.7-max
Every AI lab we tested produced AI coding hallucinations. All 22 labs. The strongest models included.
The honest caveats. On 8 models hallucinations rose with ReGrade: deepseek-v4-flash-0731, grok-4.20-0309-reasoning, haiku-4.5, haiku-4.5 (Aug 2026), hy3, ling-3.0-flash, mistral-medium-3.5, qwen3.6-plus. A further 5 produced none unaided because they did not edit the file at all, so they are reported as N/A rather than counted as successes. Every model we tested is in the table at the top of the page, sortable and searchable.
How we tested.
ReGrade is a behavioral-diff context block your AI coding agent reads alongside its existing prompt. The block shows the agent how the new code's runtime behavior differs from the old, the way a human checks for regressions at code review. No model retraining. No changes to your build or CI pipeline.
How we measured it. At Curtail® we tested 63 AI models on one task: fix 18 real bugs in a working 300,000 line Python service. Every model ran that task twice. On the first run the model worked blind, the way a coding agent normally works today. On the second run we gave the model ReGrade's patented behavioral comparison, which runs the same live traffic against both versions of the service and shows the model exactly what its last changes did to the running system. 1,431 agent-mode trials across 63 AI models from 22 labs, US and non-US, from frontier flagships down to free tiers.
What counts as an AI coding hallucination. An unintended change, addition or deletion to working code the agent was never asked to modify. The code is not what hallucinates; the AI writing it is. We measure it by replaying the agent's patch against the original behavior and counting what changed that nobody requested, deduplicated to root causes so one mistake surfacing at nine call sites counts once.
63
AI models
22
labs
1,431
agent-mode trials
18
known bugs
300,000-line Python
codebase
-83%
AI coding hallucinations
🛠️ Works with every major coding agent and model API.
5 coding-agent CLIs · 63 models · 22 labs.
Coding-agent CLIs5 CLIs · 19 models
Claude Code
- fable-5
- glm-5.1
- haiku-4.5
- haiku-4.5 (Aug 2026)
- opus-4.6
- opus-4.7
- opus-4.8
- opus-5
- sonnet-4.6
- sonnet-5
Codex CLI
- gpt-5.4
- gpt-5.5
- gpt-5.6-luna
- gpt-5.6-sol
- gpt-5.6-terra
qwen-code
- qwen3-coder
- qwen3-coder-plus
Kimi Code CLI
- kimi-k26
Mistral Vibe CLI
- devstral-2
Models tested via API, by lab44 models · 21 labs
xAI
- grok-3-fast
- grok-3-mini
- grok-4-1-fast-non-reasoning
- grok-4.1-fast
- grok-4.20-0309-reasoning
- grok-4.3
- grok-4.5
- gemini-2.5-flash
- gemini-3-flash-preview
- gemini-3.1-pro-preview
- gemini-3.5-flash
- gemini-3.6-flash
Alibaba
- qwen3.6-max-preview
- qwen3.6-plus
- qwen3.7-max
- qwen3.8-max
DeepSeek
- deepseek-v4-flash
- deepseek-v4-flash-0731
- deepseek-v4-pro
InclusionAI
- ling-2.6-1t
- ling-3.0-flash
Meta
- muse-spark-1.1
- muse-spark-1.2
MiniMax
- minimax-m2.5
- minimax-m3
Moonshot
- kimi-k2.7-code
- kimi-k3
Nvidia
- nemotron-3-ultra
- nemotron-3.5-lightning
OpenAI
- gpt-5.4-mini
- gpt-oss-20b
Poolside
- laguna-s-2.1
- laguna-xs-2.1
Sakana
- fugu-ultra
- sakana-namazu
ByteDance
- seed-2.0-lite
Kuaishou
- kat-coder-pro-v2.5
Meituan
- longcat-2.0
Mistral
- mistral-medium-3.5
StepFun
- step-3.7-flash
Tencent
- hy3
ThinkingMachines
- inkling
Upstage
- solar-pro4
Z.ai
- glm-5.2
One ReGrade context block in the agent's prompt. No model retraining. No changes to your build or CI pipeline. Tested working across 63 models from 22 labs, US and non-US.
Want ReGrade for your AI coding agent?
Drop the context block into your agent's prompt. No retraining. No CI changes.