Cost reduction
OngoingIn our ongoing testing, ReGrade cuts the cost per bug fixed by over 54%.
Across 57 models in Curtail’s current corpus, the typical model cut its cost per bug fixed by 54.4% with ReGrade. 50 of the 57 got cheaper per bug fixed. Cost per bug fixed is what a run costs, divided by the bugs the agent actually eliminated, at each provider’s list price.
Curtail research · 68 models, 22 labs, updated 2026-09-18
The AFRL study measured this across 17 models in April to May 2026, where the median model cut cost per bug fixed by 42%. These figures cover 68 models. Read the study as published.
Adding the behavioral report raises what a run costs, and the agent finds enough more that the cost per bug fixed falls.
Every model tested, on cost per bug fixed
| Cost | Bug Fixes | |||||||
|---|---|---|---|---|---|---|---|---|
| Verdict | ||||||||
| ling-2.6-1t | $0.09 | $0.006 | −93% | 1.6 | 7.0 | InclusionAI | 10 | ✓ |
| solar-pro4 | $0.02 | $0.002 | −91% | 3.8 | 10.6 | Upstage | 10 | ✓ |
| qwen3.6-max-preview | $0.34 | $0.05 | −87% | 5.7 | 13.8 | Alibaba | 8 | ✓ |
| deepseek-v4-flash-0731 | $0.03 | $0.004 | −86% | 4.8 | 16.7 | DeepSeek | 22 | ✓ |
| laguna-xs-2.1 | $0.05 | $0.007 | −84% | 1.8 | 10.3 | Poolside | 30 | ✓ |
| deepseek-v4-pro-0813 | $0.11 | $0.02 | −83% | 6.8 | 13.6 | DeepSeek | 30 | ✓ |
| minimax-m2.5 | $0.02 | $0.003 | −83% | 2.1 | 10.9 | MiniMax | 60 | ✓ |
| kimi-k2.7-code | $0.05 | $0.010 | −81% | 5.0 | 15.7 | Moonshot | 6 | ✓ |
| minimax-m3 | $0.06 | $0.01 | −80% | 7.3 | 17.7 | MiniMax | 6 | ✓ |
| sakana-namazu | $0.16 | $0.04 | −77% | 5.0 | 12.0 | Sakana | 7 | ✓ |
| seed-2.0-lite | $0.01 | $0.003 | −75% | 2.9 | 14.7 | ByteDance | 60 | ✓ |
| deepseek-v4-flash | $0.02 | $0.005 | −71% | 3.8 | 7.0 | DeepSeek | 8 | ✓ |
| haiku-4.5 | $0.18 | $0.05 | −71% | 2.5 | 12.8 | Anthropic | 60 | ✓ |
| hy3 | $0.02 | $0.004 | −71% | 6.0 | 14.0 | Tencent | 9 | ✓ |
| grok-4-1-fast-non-reasoning | $0.02 | $0.005 | −69% | 2.5 | 9.7 | xAI | 21 | ✓ |
| glm-5.2 | $0.03 | $0.008 | −69% | 6.7 | 16.6 | Z.ai | 39 | ✓ |
| haiku-4.5 (Aug 2026) | $0.15 | $0.05 | −68% | 2.7 | 14.5 | Anthropic | 20 | ✓ |
| opus-4.6 | $0.29 | $0.10 | −67% | 7.0 | 17.7 | Anthropic | 6 | ✓ |
| step-3.7-flash | $0.04 | $0.01 | −64% | 3.4 | 11.0 | StepFun | 60 | ✓ |
| qwen3.6-plus | $0.06 | $0.02 | −64% | 4.0 | 16.0 | Alibaba | 7 | ✓ |
| gpt-5.4-mini | $0.02 | $0.009 | −63% | 3.6 | 13.2 | OpenAI | 10 | ✓ |
| grok-4.20-0309-reasoning | $0.27 | $0.10 | −62% | 3.2 | 12.3 | xAI | 21 | ✓ |
| qwen3.8-max-0902 | $0.21 | $0.09 | −58% | 12.6 | 17.6 | Alibaba | 10 | ✓ |
| fable-5 | $0.59 | $0.25 | −57% | 11.5 | 18.0 | Anthropic | 32 | ✓ |
| gpt-5.6-luna | $0.01 | $0.005 | −57% | 7.5 | 16.8 | OpenAI | 60 | ✓ |
| glm-5.1 | $0.05 | $0.02 | −56% | 4.5 | 17.6 | Z.ai | 60 | ✓ |
| fugu-ultra | $0.49 | $0.21 | −56% | 10.2 | 18.0 | Sakana | 10 | ✓ |
| qwen3-coder-plus | $0.07 | $0.03 | −55% | 2.0 | 11.8 | Alibaba | 7 | ✓ |
| muse-spark-1.2 | $0.07 | $0.03 | −54% | 12.0 | 17.0 | Meta | 6 | ✓ |
| grok-4.5 | $0.03 | $0.01 | −53% | 11.2 | 18.0 | xAI | 8 | ✓ |
| sonnet-5 | $0.26 | $0.13 | −52% | 5.5 | 17.4 | Anthropic | 60 | ✓ |
| gpt-5.5 | $0.17 | $0.09 | −50% | 13.3 | 18.0 | OpenAI | 67 | ✓ |
| muse-spark-1.1 | $0.04 | $0.02 | −49% | 13.2 | 17.7 | Meta | 7 | ✓ |
| opus-4.8 | $0.40 | $0.21 | −49% | 9.1 | 17.9 | Anthropic | 60 | ✓ |
| qwen3.7-max | $0.05 | $0.03 | −48% | 7.3 | 16.7 | Alibaba | 6 | ✓ |
| kat-coder-pro-v2.5 | $0.05 | $0.03 | −46% | — | — | Kuaishou | 5 | ✓ |
| ling-3.0-flash | $0.004 | $0.002 | −45% | 2.9 | 11.3 | InclusionAI | 29 | ✓ |
| gemini-3-flash-preview | $0.02 | $0.01 | −45% | 9.6 | 17.6 | 47 | ✓ | |
| gpt-5.6-terra | $0.02 | $0.01 | −43% | 12.0 | 17.9 | OpenAI | 60 | ✓ |
| fable51 | $0.25 | $0.15 | −38% | 12.7 | 18.0 | Anthropic | 6 | ✓ |
| kimi-k3 | $0.06 | $0.04 | −37% | 12.3 | 17.0 | Moonshot | 6 | ✓ |
| grok-4.3 | $0.02 | $0.01 | −35% | 2.0 | 6.3 | xAI | 6 | ✓ |
| opus-5 | $0.25 | $0.17 | −31% | 11.9 | 18.0 | Anthropic | 60 | ✓ |
| longcat-2.0 | $0.01 | $0.008 | −30% | 6.7 | 11.0 | Meituan | 6 | ✓ |
| gpt-5.6-sol | $0.07 | $0.05 | −29% | 13.9 | 18.0 | OpenAI | 60 | ✓ |
| gpt-5.4 | $0.03 | $0.02 | −28% | 13.9 | 17.9 | OpenAI | 79 | ✓ |
| inkling | $0.03 | $0.02 | −21% | 9.3 | 15.3 | ThinkingMachines | 6 | ✓ |
| sonnet-4.6 | $0.11 | $0.09 | −20% | 7.3 | 17.6 | Anthropic | 87 | ✓ |
| nemotron-3-ultra | $0.07 | $0.06 | −20% | 3.0 | 13.8 | Nvidia | 10 | ✓ |
| opus-4.7 | $0.20 | $0.17 | −18% | 10.6 | 17.8 | Anthropic | 84 | ✓ |
| gemini-3.1-pro-preview | $0.12 | $0.12 | +2% | 8.2 | 14.5 | 8 | * | |
| mistral-medium-3.5 | $0.05 | $0.06 | +20% | 6.4 | 14.9 | Mistral | 20 | * |
| qwen3-coder | $0.10 | $0.12 | +21% | 1.7 | 5.0 | Alibaba | 15 | * |
| gemini-2.5-flash | $0.008 | $0.010 | +30% | 5.3 | 5.6 | 15 | * | |
| grok-3-fast | $0.05 | $0.07 | +40% | 6.3 | 5.5 | xAI | 21 | * |
| deepseek-v4-pro | $0.01 | $0.02 | +44% | 5.3 | 16.6 | DeepSeek | 8 | * |
| grok-3-mini | $0.005 | $0.009 | +88% | 1.9 | 1.9 | xAI | 16 | * |
✓ improved · * got worse · N/A: the model did not edit the file in the control run, so there is no baseline to compare against. Click a column to sort. Measured over 68 models, 22 labs, 1,571 agent-mode trials. Bug fixes are unaided (mean bugs fixed across the model's control-arm trials) and aided (the same average over the treatment trials), so the table can be sorted by how capable a model is on its own and by how far it gets with help. Both are our own measurement from the same trials as the rest of this table, not a third party's score. — means the model sits below the trial floor, so no figure is published for it.
🏆 Three recommended models, ranked by cost with ReGrade.
For teams choosing their AI coding agent today: best results-per-dollar with ReGrade.
Ranked from Curtail's current corpus of 68 models, by cost per bug fixed with ReGrade among those fixing at least 16 of 18. This list moves as later testing continues.
🥇
deepseek-v4-flash-0731
16.7 / 18
bugs fixed
$0.004
per bug
🥈
gpt-5.6-luna
16.8 / 18
bugs fixed
$0.005
per bug
🥉
glm-5.2
16.6 / 18
bugs fixed
$0.008
per bug
24 more model and ReGrade combinations clear the same threshold, at least 16 of 18 bugs fixed. Works equally well across providers.
ReGrade is a single addition to the agent's prompt, you pay only the marginal LLM tokens for that context, with no separate ReGrade fee per bug.
Cut your cost per bug fixed.
Across 57 models, 50 got cheaper per bug fixed, and the typical model cut it by 54.4%. Ask us what that looks like on your stack.
Cut your cost per bug fixed.