Skip to main content
Curtail

AI coding hallucinations

Ongoing

In our ongoing testing, ReGrade reduces AI coding hallucinations by over 79%.

An AI coding hallucination is an unintended change, addition or deletion to working code that the agent was never asked to modify.

Curtail research · 74 models, 23 labs, updated 2026-09-30

Every AI lab we tested produced AI coding hallucinations. All 23 labs. The strongest models included.

With ReGrade AI coding hallucinations dropped to zero for these models.

A hallucination is counted by replaying each trial's saved patch offline and counting the changes that fall outside the bugs the model was asked to fix. Each change is traced back to the edit that caused it, by removing one edit at a time and replaying, so one mistake that surfaces at nine call sites counts once. On the typical model, ReGrade's patented behavioral comparison cut those changes by over 79%. In the main comparisons, no unintended changes were observed in the ReGrade trials on the 20 models below, across 10 labs. Zero observed changes is not a guarantee on future code.

  • OpenAIgpt-6-astra, gpt-6-sol, gpt-5.5, gpt-5.6-sol
  • Xiaomimimo-v2.6-pro
  • xAIgrok-4.5, grok-3-fast
  • Metamuse-spark-1.3
  • Anthropicfable-5, opus-5.5, fable51, opus-4.8, opus-5, opus-4.6
  • Moonshotkimi-k2.7-code
  • Googlegemini-3.1-pro-preview
  • Sakanafugu-ultra
  • Alibabaqwen3.8-max-prime, qwen3.7-max
  • Nvidianemotron-3.5-lightning

Every model tested, on AI coding hallucinations

74 models · 23 labs
HallucinationsBug Fixes
Verdict
gpt-6-astra6.330−100%14.718.0OpenAI3✓
gpt-6-sol3.000−100%14.318.0OpenAI3✓
muse-spark-1.12.750.33−88%13.217.7Meta3✓
gpt-5.43.160.24−92%13.117.9OpenAI31✓
gpt-5.52.930−100%13.018.0OpenAI28✓
gpt-5.6-sol2.920−100%12.918.0OpenAI24✓
fable512.000−100%12.718.0Anthropic3✓
kimi-k35.330.33−94%12.317.0Moonshot3✓
grok-4.711.31.00−91%12.317.3xAI3✓
opus-51.830−100%11.918.0Anthropic30✓
gpt-5.6-terra2.320.07−97%11.817.9OpenAI28✓
fable-53.310−100%11.518.0Anthropic16✓
muse-spark-1.33.400−100%11.417.8Meta5✓
sonnet-5.52.000.67−67%11.316.3Anthropic3✓
grok-4.53.750−100%11.218.0xAI4✓
qwen3.8-max-prime2.000−100%10.718.0Alibaba3✓
gemini-3.1-pro-preview2.400−100%10.516.0Google10✓
opus-5.53.000−100%10.318.0Anthropic3✓
fugu-ultra2.400−100%10.218.0Sakana5✓
mimo-v2.6-pro4.000−100%10.018.0Xiaomi3✓
glm-5.3-prime2.331.33−43%10.016.7Z.ai3✓
inkling3.671.67−54%9.315.3ThinkingMachines3✓
opus-4.81.970−100%9.117.9Anthropic21✓
qwen3.8-max-09022.330.40−83%9.017.6Alibaba3✓
deepseek-v4-pro (Sep 2026)4.710.14−97%9.017.7DeepSeek7✓
gpt-6-luna2.672.00−25%8.714.3OpenAI3✓
glm-5.32.001.00−50%8.716.3Z.ai3✓
opus-4.71.910.03−98%8.617.9Anthropic33✓
deepseek-v4.1-flash2.000.67−67%8.317.3DeepSeek3✓
gpt-5.6-luna1.470.70−52%7.516.8OpenAI30✓
qwen3.7-max1.670−100%7.316.7Alibaba3✓
minimax-m31.330.33−75%7.317.7MiniMax3✓
opus-4.61.330−100%7.017.7Anthropic3✓
longcat-2.03.000.33−89%6.711.0Meituan3✓
glm-5.21.850.74−60%6.716.6Z.ai19✓
mistral-medium-3.54.002.10−47%6.414.9Mistral10✓
gemini-3-flash-preview1.000.67−33%6.317.5Google3✓
hy32.751.20−56%6.014.0Tencent4✓
qwen3.6-max-preview2.000.25−87%5.712.8Alibaba3✓
sonnet-52.000.27−87%5.517.4Anthropic30✓
devstral-24.400.90−79%5.413.0Mistral30✓
deepseek-v4-pro-04231.000.50−50%5.316.2DeepSeek3✓
sonnet-4.61.120.03−97%5.217.7Anthropic33✓
kimi-k2.7-code3.000−100%5.015.7Moonshot3✓
sakana-namazu1.750.20−89%5.012.0Sakana4✓
qwen3-coder3.500.50−86%4.712.5Alibaba10✓
glm-5.11.400.53−62%4.517.6Z.ai30✓
kimi-k261.570.50−68%4.416.0Moonshot7✓
deepseek-v4-flash-07311.000.60−40%4.116.7DeepSeek9✓
qwen3.6-plus0.670.33−50%4.016.0Alibaba3✓
deepseek-v4-flash-04232.000.33−83%3.87.0DeepSeek3✓
solar-pro41.201.00−17%3.810.6Upstage4✓
gpt-5.4-mini1.600.60−62%3.613.2OpenAI5✓
step-3.7-flash2.861.04−64%3.411.4StepFun25✓
deepseek-v4-pro-08131.000.62−38%3.213.5DeepSeek9✓
grok-4.20-0309-reasoning7.005.60−20%3.211.7xAI10✓
nemotron-3-ultra1.601.40−12%3.013.8Nvidia5✓
seed-2.0-lite2.500.43−83%2.914.7ByteDance16✓
haiku-4.5 (Aug 2026)1.701.60−6%2.714.5Anthropic10✓
laguna-s-2.11.330.33−75%2.715.3Poolside3✓
haiku-4.52.301.23−46%2.512.8Anthropic30✓
grok-4-1-fast-non-reasoning3.301.11−66%2.59.7xAI9✓
minimax-m2.51.201.13−6%2.110.9MiniMax30✓
qwen3-coder-plus1.000.33−67%2.012.0Alibaba3✓
gemini-2.5-flash3.001.67−44%2.014.0Google3✓
grok-3-fast2.670−100%2.08.0xAI3✓
grok-4.34.001.33−67%2.06.3xAI3✓
ling-3.0-flash1.081.00−7%1.911.3InclusionAI13✓
laguna-xs-2.11.801.60−11%1.810.3Poolside15✓
ling-2.6-1t1.801.75−3%1.67.0InclusionAI4✓
nemotron-3.5-lightning1.250−100%1.45.3Nvidia3✓
grok-3-mini0.670.33−50%1.02.8xAI3✓
gpt-oss-20bdid not edit0.40N/A0.03.8OpenAI5no baseline
kat-coder-pro-v2.50.750.25−67%——Kuaishou4✓

✓ improved · * got worse · N/A: the model did not edit the file in the control run, so there is no baseline to compare against. Click a column to sort. Measured over 74 models, 23 labs, 1,558 agent-mode trials. The provider-pinned comparisons add 36 separately labelled trials, for 1,594 total. Min. replays/arm is the smaller eligible replay count of the control and treatment arms; individual arms can contain more runs. Scroll sideways to see the lab and sample-count columns. Bug fixes are unaided (mean bugs fixed across the model's control-arm trials) and aided (the same average over the treatment trials), so the table can be sorted by how capable a model is on its own and by how far it gets with help. Both are our own measurements. Repair and damage eligible subsets can differ, so bug-fix averages need not use every replay counted here. A dash means the model sits below the trial floor, so no figure is published for it.

The honest caveats. One model, gpt-oss-20b, broke nothing unaided because it did not edit the file at all, so it is reported as N/A rather than counted as a success. The main comparison results are in the table above, sortable and searchable. The provider-pinned results follow separately, with three trials per arm: July DeepSeek on Baidu reached zero observed changes, Nemotron Ultra's count rose, and GPT-OSS has incomplete damage measurements. Those samples do not establish stable effects.

Provider-pinned comparisons

Six separate cells, 36 trials, three trials per arm. Both arms of each cell used the same provider. These batches remain separate from the historical model comparisons above and their headline statistics. Every arrow reads without ReGrade → with ReGrade.

Provider-pinned repair and causal damage measurements
ModelProvider and precisionBugs fixed / 18Causal damage per trial
deepseek-v4-flash-0423DeepInfra · fp84.33 → 16.001.33 → 0.33
deepseek-v4-flash-0731Baidu · fp85.67 → 17.002.67 to 3.33 → 0.00
nemotron-3-ultraBaseTen · fp44.33 → 15.001.33 → 1.67
gpt-oss-20bDeepInfra · bf160.33 → 2.330.00 → Unavailable (1/3 measured)
inklingTogether · precision unknown8.67 → 14.333.67 → 1.67
glm-5.2Baidu · fp47.00 → 16.672.00 to 3.00 → 1.00

Higher repair scores and lower damage are better. Ranges include unresolved hunk-removal variants. Two GPT-OSS treatment patches contained invalid Python: their repair scores remain zero and their damage is unavailable. Nemotron introduced slightly more damage with ReGrade. Three trials per arm do not establish stable model effects.

🛠️ Works with every major coding agent and model API.

5 coding-agent CLIs · 74 models · 23 labs.

Coding-agent CLIs5 CLIs · 25 models

Claude Code

  • fable-5
  • fable51
  • glm-5.1
  • haiku-4.5
  • haiku-4.5 (Aug 2026)
  • opus-4.6
  • opus-4.7
  • opus-4.8
  • opus-5
  • opus-5.5
  • sonnet-4.6
  • sonnet-5
  • sonnet-5.5

Codex CLI

  • gpt-5.4
  • gpt-5.5
  • gpt-5.6-luna
  • gpt-5.6-sol
  • gpt-5.6-terra
  • gpt-6-astra
  • gpt-6-luna
  • gpt-6-sol

qwen-code

  • qwen3-coder
  • qwen3-coder-plus

Kimi Code CLI

  • kimi-k26

Mistral Vibe CLI

  • devstral-2
Models tested via API, by lab49 models · 22 labs

xAI

  • grok-3-fast
  • grok-3-mini
  • grok-4-1-fast-non-reasoning
  • grok-4.20-0309-reasoning
  • grok-4.3
  • grok-4.5
  • grok-4.7

DeepSeek

  • deepseek-v4-flash-0423
  • deepseek-v4-flash-0731
  • deepseek-v4-pro (Sep 2026)
  • deepseek-v4-pro-0423
  • deepseek-v4-pro-0813
  • deepseek-v4.1-flash

Alibaba

  • qwen3.6-max-preview
  • qwen3.6-plus
  • qwen3.7-max
  • qwen3.8-max-0902
  • qwen3.8-max-prime

Google

  • gemini-2.5-flash
  • gemini-3-flash-preview
  • gemini-3.1-pro-preview

Z.ai

  • glm-5.2
  • glm-5.3
  • glm-5.3-prime

InclusionAI

  • ling-2.6-1t
  • ling-3.0-flash

Meta

  • muse-spark-1.1
  • muse-spark-1.3

MiniMax

  • minimax-m2.5
  • minimax-m3

Moonshot

  • kimi-k2.7-code
  • kimi-k3

Nvidia

  • nemotron-3-ultra
  • nemotron-3.5-lightning

OpenAI

  • gpt-5.4-mini
  • gpt-oss-20b

Poolside

  • laguna-s-2.1
  • laguna-xs-2.1

Sakana

  • fugu-ultra
  • sakana-namazu

ByteDance

  • seed-2.0-lite

Kuaishou

  • kat-coder-pro-v2.5

Meituan

  • longcat-2.0

Mistral

  • mistral-medium-3.5

StepFun

  • step-3.7-flash

Tencent

  • hy3

ThinkingMachines

  • inkling

Upstage

  • solar-pro4

Xiaomi

  • mimo-v2.6-pro

One ReGrade context block in the agent's prompt. No model retraining. No changes to your build or CI pipeline. Tested working across 74 models from 23 labs, US and non-US.

Reduce unintended changes reaching review.

Every lab we tested produced them, all 23 labs. ReGrade cut them by over 79% on the typical model. On 20 models, across 10 labs, no unintended changes were observed in the ReGrade trials.