Skip to main content
Curtail

AI coding hallucinations

Curtail research · 63 models, 22 labs, updated 2026-08-29

Every model tested, on AI coding hallucinations

63 models · 22 labs
Verdict
gpt-5.6-sol4.900−100%78OpenAI30
opus-51.800−100%78Anthropic30
fable-53.330−100%77Anthropic12
kimi-k37.000−100%76Moonshot2
gpt-5.53.760−100%75OpenAI33
opus-4.82.230−100%74Anthropic21
grok-4.59.500−100%72xAI4
qwen3.8-max1.500−100%72Alibaba2
gemini-3.1-pro-preview3.000−100%69Google3
qwen3.7-max1.670−100%66Alibaba3
kimi-k2.7-code7.330−100%61Moonshot3
minimax-m32.000−100%59MiniMax1
nemotron-3.5-lightning2.750−100%27Nvidia4
opus-4.61.330−100%Anthropic3
fugu-ultra2.400−100%Sakana5
grok-3-fast5.330−100%xAI6
grok-3-mini2.500−100%xAI2
grok-4-1-fast-non-reasoning1.500−100%xAI2
opus-4.72.130.03−99%74Anthropic38
sonnet-4.62.130.03−99%63Anthropic38
gpt-5.6-terra4.550.13−97%77OpenAI29
sakana-namazu4.250.20−95%Sakana4
muse-spark-1.211.30.67−94%72Meta3
qwen3.6-max-preview3.000.25−92%Alibaba2
grok-4.3151.33−91%42xAI3
muse-spark-1.13.330.33−90%71Meta3
gpt-5.45.660.58−90%71OpenAI35
qwen3-coder-plus3.000.33−89%Alibaba3
deepseek-v4-flash2.000.33−83%56DeepSeek3
kat-coder-pro-v2.53.000.50−83%Kuaishou2
devstral-29.331.57−83%31Mistral30
glm-5.13.070.53−83%56Z.ai30
kimi-k262.710.50−82%62Moonshot7
gemini-3-flash-preview1.170.29−75%Google6
gpt-5.4-mini3.601.00−72%56OpenAI5
seed-2.0-lite2.620.77−71%ByteDance16
sonnet-54.071.80−56%72Anthropic30
qwen3-coder1.330.67−50%Alibaba3
gpt-5.6-luna3.771.90−50%71OpenAI30
glm-5.22.601.58−39%69Z.ai19
gemini-2.5-flash11.37.00−38%Google1
longcat-2.03.502.33−33%45Meituan2
step-3.7-flash5.173.54−32%40StepFun26
nemotron-3-ultra3.202.20−31%49Nvidia5
ling-2.6-1t4.203.25−23%InclusionAI4
laguna-xs-2.14.673.67−21%Poolside3
deepseek-v4-pro2.332.00−14%59DeepSeek3
minimax-m2.54.073.53−13%MiniMax29
solar-pro42.202.00−9%53Upstage4
inkling7.677.33−4%52ThinkingMachines3
grok-4.20-0309-reasoning7.508.10+8%xAI10*
hy33.754.20+12%59Tencent4*
deepseek-v4-flash-07312.203.10+41%69DeepSeek10*
haiku-4.54.205.93+41%44Anthropic30*
haiku-4.5 (Aug 2026)4.307.20+67%44Anthropic10*
ling-3.0-flash2.203.79+72%51InclusionAI14*
qwen3.6-plus1.332.33+75%55Alibaba3*
mistral-medium-3.52.206.00+173%47Mistral5*
gemini-3.5-flashdid not edit0.33N/A70Google3no baseline
gemini-3.6-flashdid not edit0.17N/A69Google6no baseline
gpt-oss-20bdid not editN/A21OpenAI5no baseline
grok-4.1-fastdid not editN/AxAI3no baseline
laguna-s-2.1did not edit0.33N/APoolside3no baseline

improved · * got worse · N/A: the model did not edit the file in the control run, so there is no baseline to compare against. Click a column to sort. Measured over 63 models, 22 labs, 1,431 agent-mode trials. Capability is the Artificial Analysis Coding Index, retrieved 2026-08-18, best listed configuration per model. It is not measured by us; it is here so the table can be sorted by how capable a model is. — means the model is not listed there.

Where AI coding hallucinations stopped.

A hallucination is counted by replaying each trial's saved patch offline and counting the changes that fall outside the bugs the model was asked to fix. Counts are deduplicated to root causes, so one mistake that surfaces at nine call sites counts once. On the typical model, ReGrade's patented behavioral comparison cut those changes by over 83%. On the 13 models below, across 8 labs, it stopped them completely.

  • xAIgrok-4.5, grok-3-fast
  • Moonshotkimi-k2.7-code
  • OpenAIgpt-5.6-sol, gpt-5.5
  • Anthropicfable-5, opus-4.8, opus-5, opus-4.6
  • Googlegemini-3.1-pro-preview
  • Nvidianemotron-3.5-lightning
  • Sakanafugu-ultra
  • Alibabaqwen3.7-max

Every AI lab we tested produced AI coding hallucinations. All 22 labs. The strongest models included.

The honest caveats. On 8 models hallucinations rose with ReGrade: deepseek-v4-flash-0731, grok-4.20-0309-reasoning, haiku-4.5, haiku-4.5 (Aug 2026), hy3, ling-3.0-flash, mistral-medium-3.5, qwen3.6-plus. A further 5 produced none unaided because they did not edit the file at all, so they are reported as N/A rather than counted as successes. Every model we tested is in the table at the top of the page, sortable and searchable.

🛠️ Works with every major coding agent and model API.

5 coding-agent CLIs · 63 models · 22 labs.

Coding-agent CLIs5 CLIs · 19 models

Claude Code

  • fable-5
  • glm-5.1
  • haiku-4.5
  • haiku-4.5 (Aug 2026)
  • opus-4.6
  • opus-4.7
  • opus-4.8
  • opus-5
  • sonnet-4.6
  • sonnet-5

Codex CLI

  • gpt-5.4
  • gpt-5.5
  • gpt-5.6-luna
  • gpt-5.6-sol
  • gpt-5.6-terra

qwen-code

  • qwen3-coder
  • qwen3-coder-plus

Kimi Code CLI

  • kimi-k26

Mistral Vibe CLI

  • devstral-2
Models tested via API, by lab44 models · 21 labs

xAI

  • grok-3-fast
  • grok-3-mini
  • grok-4-1-fast-non-reasoning
  • grok-4.1-fast
  • grok-4.20-0309-reasoning
  • grok-4.3
  • grok-4.5

Google

  • gemini-2.5-flash
  • gemini-3-flash-preview
  • gemini-3.1-pro-preview
  • gemini-3.5-flash
  • gemini-3.6-flash

Alibaba

  • qwen3.6-max-preview
  • qwen3.6-plus
  • qwen3.7-max
  • qwen3.8-max

DeepSeek

  • deepseek-v4-flash
  • deepseek-v4-flash-0731
  • deepseek-v4-pro

InclusionAI

  • ling-2.6-1t
  • ling-3.0-flash

Meta

  • muse-spark-1.1
  • muse-spark-1.2

MiniMax

  • minimax-m2.5
  • minimax-m3

Moonshot

  • kimi-k2.7-code
  • kimi-k3

Nvidia

  • nemotron-3-ultra
  • nemotron-3.5-lightning

OpenAI

  • gpt-5.4-mini
  • gpt-oss-20b

Poolside

  • laguna-s-2.1
  • laguna-xs-2.1

Sakana

  • fugu-ultra
  • sakana-namazu

ByteDance

  • seed-2.0-lite

Kuaishou

  • kat-coder-pro-v2.5

Meituan

  • longcat-2.0

Mistral

  • mistral-medium-3.5

StepFun

  • step-3.7-flash

Tencent

  • hy3

ThinkingMachines

  • inkling

Upstage

  • solar-pro4

Z.ai

  • glm-5.2

One ReGrade context block in the agent's prompt. No model retraining. No changes to your build or CI pipeline. Tested working across 63 models from 22 labs, US and non-US.

🌍 The newest labs, too.

The behavioral-diff effect is not a quirk of the big US providers. We ran the same test on the newest releases from five more labs, at 30 trials per model, and the rescue held every time.

LabModelWithoutWithΔ
ByteDanceSeed-2.0-Lite2.9 / 1814.7 / 18+408%
MiniMaxM2.52.1 / 1810.9 / 18+427%
ZhipuGLM-5.14.5 / 1817.6 / 18+292%
StepFunStep 3.7-flash3.4 / 1811.0 / 18+221%
MistralDevstral 25.4 / 1813.0 / 18+142%

63 models across 22 labs, US and non-US, from frontier flagships down to free tiers. 55 of the 58 we could measure fixed more silent regressions with ReGrade.

Want ReGrade for your AI coding agent?

Drop the context block into your agent's prompt. No retraining. No CI changes.