More bugs fixed
OngoingIn our ongoing testing, ReGrade helps AI coding agents fix 136% more bugs.
Measured on the typical model on DriftBench, across 74 models from 23 labs. 73 of the 73 we scored fixed more silent regressions with ReGrade.
Curtail research · 74 models, 23 labs, updated 2026-09-30
The AFRL study measured this across 17 models in April to May 2026. These figures cover 74 models. Read the study as published.
Every model tested, on bugs fixed
| Bug Fixes | ||||||
|---|---|---|---|---|---|---|
| Verdict | ||||||
| gemini-2.5-flash | 2.00 | 14 | +600% | 6 | ✓ | |
| qwen3-coder-plus | 2.00 | 12 | +500% | Alibaba | 6 | ✓ |
| ling-3.0-flash | 1.90 | 11.3 | +488% | InclusionAI | 27 | ✓ |
| laguna-s-2.1 | 2.70 | 15.3 | +475% | Poolside | 6 | ✓ |
| laguna-xs-2.1 | 1.80 | 10.3 | +474% | Poolside | 30 | ✓ |
| haiku-4.5 (Aug 2026) | 2.70 | 14.5 | +437% | Anthropic | 20 | ✓ |
| minimax-m2.5 | 2.10 | 10.9 | +427% | MiniMax | 60 | ✓ |
| haiku-4.5 | 2.50 | 12.8 | +413% | Anthropic | 60 | ✓ |
| seed-2.0-lite | 2.90 | 14.7 | +408% | ByteDance | 60 | ✓ |
| nemotron-3-ultra | 3.00 | 13.8 | +360% | Nvidia | 10 | ✓ |
| ling-2.6-1t | 1.60 | 7.00 | +338% | InclusionAI | 10 | ✓ |
| deepseek-v4-pro-0813 | 3.20 | 13.5 | +323% | DeepSeek | 25 | ✓ |
| deepseek-v4-flash-0731 | 4.10 | 16.7 | +307% | DeepSeek | 20 | ✓ |
| qwen3.6-plus | 4.00 | 16 | +300% | Alibaba | 6 | ✓ |
| grok-3-fast | 2.00 | 8.00 | +300% | xAI | 6 | ✓ |
| grok-4-1-fast-non-reasoning | 2.50 | 9.70 | +295% | xAI | 21 | ✓ |
| glm-5.1 | 4.50 | 17.6 | +292% | Z.ai | 60 | ✓ |
| nemotron-3.5-lightning | 1.40 | 5.30 | +281% | Nvidia | 8 | ✓ |
| gpt-5.4-mini | 3.60 | 13.2 | +267% | OpenAI | 10 | ✓ |
| grok-4.20-0309-reasoning | 3.20 | 11.7 | +266% | xAI | 20 | ✓ |
| kimi-k26 | 4.40 | 16 | +261% | Moonshot | 13 | ✓ |
| sonnet-4.6 | 5.20 | 17.7 | +238% | Anthropic | 72 | ✓ |
| step-3.7-flash | 3.40 | 11.4 | +232% | StepFun | 59 | ✓ |
| grok-4.3 | 2.00 | 6.30 | +217% | xAI | 6 | ✓ |
| sonnet-5 | 5.50 | 17.4 | +216% | Anthropic | 60 | ✓ |
| kimi-k2.7-code | 5.00 | 15.7 | +213% | Moonshot | 6 | ✓ |
| deepseek-v4-pro-0423 | 5.30 | 16.2 | +205% | DeepSeek | 7 | ✓ |
| solar-pro4 | 3.80 | 10.6 | +179% | Upstage | 10 | ✓ |
| gemini-3-flash-preview | 6.30 | 17.5 | +177% | 32 | ✓ | |
| grok-3-mini | 1.00 | 2.80 | +175% | xAI | 8 | ✓ |
| qwen3-coder | 4.70 | 12.5 | +166% | Alibaba | 20 | ✓ |
| opus-4.6 | 7.00 | 17.7 | +152% | Anthropic | 6 | ✓ |
| glm-5.2 | 6.70 | 16.6 | +150% | Z.ai | 39 | ✓ |
| devstral-2 | 5.40 | 13 | +142% | Mistral | 60 | ✓ |
| minimax-m3 | 7.30 | 17.7 | +141% | MiniMax | 6 | ✓ |
| sakana-namazu | 5.00 | 12 | +140% | Sakana | 7 | ✓ |
| hy3 | 6.00 | 14 | +133% | Tencent | 9 | ✓ |
| mistral-medium-3.5 | 6.40 | 14.9 | +133% | Mistral | 20 | ✓ |
| qwen3.7-max | 7.30 | 16.7 | +127% | Alibaba | 6 | ✓ |
| qwen3.6-max-preview | 5.70 | 12.8 | +125% | Alibaba | 7 | ✓ |
| gpt-5.6-luna | 7.50 | 16.8 | +123% | OpenAI | 60 | ✓ |
| opus-4.7 | 8.60 | 17.9 | +109% | Anthropic | 66 | ✓ |
| deepseek-v4.1-flash | 8.30 | 17.3 | +108% | DeepSeek | 6 | ✓ |
| deepseek-v4-pro (Sep 2026) | 9.00 | 17.7 | +97% | DeepSeek | 14 | ✓ |
| opus-4.8 | 9.10 | 17.9 | +96% | Anthropic | 60 | ✓ |
| qwen3.8-max-0902 | 9.00 | 17.6 | +96% | Alibaba | 8 | ✓ |
| glm-5.3 | 8.70 | 16.3 | +89% | Z.ai | 6 | ✓ |
| deepseek-v4-flash-0423 | 3.80 | 7.00 | +84% | DeepSeek | 8 | ✓ |
| mimo-v2.6-pro | 10 | 18 | +80% | Xiaomi | 6 | ✓ |
| fugu-ultra | 10.2 | 18 | +77% | Sakana | 10 | ✓ |
| opus-5.5 | 10.3 | 18 | +74% | Anthropic | 6 | ✓ |
| qwen3.8-max-prime | 10.7 | 18 | +69% | Alibaba | 6 | ✓ |
| glm-5.3-prime | 10 | 16.7 | +67% | Z.ai | 6 | ✓ |
| gpt-6-luna | 8.70 | 14.3 | +65% | OpenAI | 6 | ✓ |
| longcat-2.0 | 6.70 | 11 | +65% | Meituan | 6 | ✓ |
| inkling | 9.30 | 15.3 | +64% | ThinkingMachines | 6 | ✓ |
| grok-4.5 | 11.2 | 18 | +60% | xAI | 8 | ✓ |
| fable-5 | 11.5 | 18 | +57% | Anthropic | 32 | ✓ |
| muse-spark-1.3 | 11.4 | 17.8 | +56% | Meta | 10 | ✓ |
| gemini-3.1-pro-preview | 10.5 | 16 | +52% | 20 | ✓ | |
| gpt-5.6-terra | 11.8 | 17.9 | +52% | OpenAI | 57 | ✓ |
| opus-5 | 11.9 | 18 | +52% | Anthropic | 60 | ✓ |
| sonnet-5.5 | 11.3 | 16.3 | +44% | Anthropic | 6 | ✓ |
| fable51 | 12.7 | 18 | +42% | Anthropic | 6 | ✓ |
| grok-4.7 | 12.3 | 17.3 | +41% | xAI | 6 | ✓ |
| gpt-5.6-sol | 12.9 | 18 | +40% | OpenAI | 51 | ✓ |
| gpt-5.5 | 13 | 18 | +38% | OpenAI | 61 | ✓ |
| kimi-k3 | 12.3 | 17 | +38% | Moonshot | 6 | ✓ |
| gpt-5.4 | 13.1 | 17.9 | +36% | OpenAI | 64 | ✓ |
| muse-spark-1.1 | 13.2 | 17.7 | +33% | Meta | 7 | ✓ |
| gpt-6-sol | 14.3 | 18 | +26% | OpenAI | 6 | ✓ |
| gpt-6-astra | 14.7 | 18 | +23% | OpenAI | 6 | ✓ |
| gpt-oss-20b | fixed none | 3.80 | N/A | OpenAI | 10 | no baseline |
✓ improved · * got worse · N/A: the model fixed no bugs without ReGrade, so there is no baseline to compare against. Click a column to sort. Measured over 74 models, 23 labs, 1,558 agent-mode trials. The provider-pinned comparisons add 36 separately labelled trials, for 1,594 total. Runs, both arms counts control and treatment runs together. Scroll sideways to see the lab and sample-count columns. Bug fixes are unaided (mean bugs fixed across the model's control-arm trials) and aided (the same average over the treatment trials), so the table can be sorted by how capable a model is on its own and by how far it gets with help. Both are our own measurements. Repair and damage eligible subsets can differ, so bug-fix averages need not use every replay counted here. A dash means the model sits below the trial floor, so no figure is published for it.
Provider-pinned comparisons
Six separate cells, 36 trials, three trials per arm. Both arms of each cell used the same provider. These batches remain separate from the historical model comparisons above and their headline statistics. Every arrow reads without ReGrade → with ReGrade.
| Model | Provider and precision | Bugs fixed / 18 | Causal damage per trial |
|---|---|---|---|
| deepseek-v4-flash-0423 | DeepInfra · fp8 | 4.33 → 16.00 | 1.33 → 0.33 |
| deepseek-v4-flash-0731 | Baidu · fp8 | 5.67 → 17.00 | 2.67 to 3.33 → 0.00 |
| nemotron-3-ultra | BaseTen · fp4 | 4.33 → 15.00 | 1.33 → 1.67 |
| gpt-oss-20b | DeepInfra · bf16 | 0.33 → 2.33 | 0.00 → Unavailable (1/3 measured) |
| inkling | Together · precision unknown | 8.67 → 14.33 | 3.67 → 1.67 |
| glm-5.2 | Baidu · fp4 | 7.00 → 16.67 | 2.00 to 3.00 → 1.00 |
Higher repair scores and lower damage are better. Ranges include unresolved hunk-removal variants. Two GPT-OSS treatment patches contained invalid Python: their repair scores remain zero and their damage is unavailable. Nemotron introduced slightly more damage with ReGrade. Three trials per arm do not establish stable model effects.
How we tested.
What ReGrade does. ReGrade compares recorded HTTP traffic across software versions and reports field-level changes. A team records baseline requests, replays them against a candidate build, and reviews the differences. Coding agents can use the report as context without model retraining.
How we measured it. At Curtail® we tested 74 model/version entries on DriftBench, a Python benchmark with 18 scored, seeded regressions. About 1,800 functional lines are surrounded by roughly 300,000 lines of synthetic padding, which does not affect the scored endpoints. Agents were told which eight files to inspect. Each model was evaluated in two conditions, repeated across trials: without a behavioral report, and with the report supplied in the initial prompt. Agents could inspect and edit the candidate code, but could not read the reference version's source.
Scope of the results. The main comparisons use 1,558 trials. Another 36 provider-pinned trials are reported separately, for 1,594 total. The typical-model result is the median change across scored model/version entries on this benchmark. The entries share a task and scoring system; the lab and model counts do not represent independent production deployments. Sample sizes vary, and small samples do not establish reliable prevention. Evaluate ReGrade on your own service before setting a release policy.
What counts as an AI coding hallucination. An unintended change, addition or deletion to working code the agent was never asked to modify. We replay each saved patch against reference behavior and trace changes to the edits that caused them by removing one edit at a time and replaying. One mistake surfacing at nine call sites counts once. The measurement covers the replayed requests, not every possible behavior.
74
AI models
23
labs
1,594
agent-mode trials
18
seeded regressions
~1,800 lines
functional Python
-79%
AI coding hallucinations
🕵️ It catches what your tests can't.
Functional tests check status codes and required fields. They routinely miss the behavioral drift that actually breaks production: a CORS header that quietly changed, a list that's now an object, a timestamp format that broke a downstream consumer. ReGrade compares runtime behavior directly, so the agent sees the drift the test suite waved through.
35%
of AI patches that passed every functional test still silently broke behavior
191 of 548 test-passing patches, across 14 of 26 tasks
0 → 9
multi-file bug sites fixed, weaker models fixed 0 of 9 on their own
with ReGrade, agents clear all 9 sites of a cross-file regression
18
classes of silent drift planted, headers, encoding, ordering, response shape
the behavioral defects that pass type-checks, lint, and existing tests
Your test suite says “green.” ReGrade says “but the API’s behavior drifted.” That gap is exactly where production incidents come from.
Measured on BaxBench and the multi-site experiment in July 2026. These three figures come from that experiment rather than from the running corpus, so they do not move as more models are tested.
🔎 Better answers, and better detection.
Finding a bug is only half the job, a reviewer still has to understand and trust the fix. Even when a model already spots a regression, ReGrade’s behavioral evidence helps it explain exactly what changed and where.
2×
more precise bug explanations on a production codebase (Ghost CMS)
explanation precision rose from 0.40 to 0.91
9 of 10
models explained the bug better on a 291,000-line codebase (NetBox)
+3.2 points on a 20-point rubric, even where the model already found it
The difference between “something looks off” and a root-cause a human can act on with confidence.
In collaboration with AFRL · figures as published
Distribution (A) Approved for public release; distribution is unlimited; AFRL-2026-3793, AUG 2026
Is it the evidence, or just the extra words?
The obvious objection to the whole study is that any five thousand tokens would help, and the behavioral report just happens to be five thousand tokens. So we pre-registered a test of exactly that: the same agents, the same bugs, but the report swapped for pseudo-random tokens of the same length, and again for unrelated source code of the same length.
| Model | No extra contextNothing | Random tokensRandom | Unrelated sourceSource | ReGradeReGrade |
|---|---|---|---|---|
| Haiku 4.5 | 2.4 | 0.8 | 1.9 | 12.5 |
| Sonnet 4.6 | 5.5 | 5.9 | 5.3 | 17.7 |
| Opus 4.7 | 9.1 | 9.2 | 8.7 | 18.0 |
| GPT-5.4 | 13.3 | 14.1 | 12.4 | 18.0 |
| GPT-5.5 | 12.3 | 15.2 | 14.1 | 18.0 |
Bugs fixed out of 18, mean of ten trials per model per arm.
The rescue does not survive the substitution. Padding of either kind lands within about a bug of no extra context at all, while the behavioral report lands eight bugs above it. All three pre-registered hypotheses reject at p below 1e-4.
- +8.32 bugs against no extra context (95% CI 7.32 to 9.32)
- +7.80 bugs against random tokens of the same length (95% CI 6.66 to 8.90)
- +8.36 bugs against unrelated source of the same length (95% CI 7.40 to 9.30)
Two models are worth reading closely. Haiku gets worse under random tokens than with nothing at all, 0.8 against 2.4. The test measures repair performance, not the model's attention mechanism. The OpenAI models move the other way: GPT-5.5 gains from padding of either kind, 15.2 and 14.1 against a control of 12.3, so some of the lift there is volume rather than signal. The behavioral report still beats both, at 18.0.
Get more fixed from the agents you already run.
73 of the 73 models we scored fixed more silent regressions with ReGrade, on the traffic a team already has.
Fix more bugs with the agents you run.