Cost reduction
Curtail research · 63 models, 22 labs, updated 2026-08-29
Across 55 model cells in Curtail’s current corpus, the typical cell cut its cost per bug fixed by 51.6% with ReGrade. 47 of the 55 got cheaper per bug fixed. Cost per bug fixed is what a run costs, divided by the bugs the agent actually eliminated, at each provider’s list price.
Adding the behavioral report raises what a run costs, and the agent finds enough more that the cost per bug fixed falls.
Every model tested, on cost per bug fixed
| Verdict | |||||||
|---|---|---|---|---|---|---|---|
| laguna-xs-2.1 | $0.15 | $0.009 | −94% | — | Poolside | 6 | ✓ |
| ling-2.6-1t | $0.09 | $0.006 | −93% | — | InclusionAI | 10 | ✓ |
| solar-pro4 | $0.02 | $0.002 | −91% | 53 | Upstage | 10 | ✓ |
| qwen3.6-max-preview | $0.34 | $0.05 | −87% | — | Alibaba | 8 | ✓ |
| minimax-m2.5 | $0.02 | $0.003 | −83% | — | MiniMax | 60 | ✓ |
| deepseek-v4-flash-0731 | $0.02 | $0.004 | −83% | 69 | DeepSeek | 22 | ✓ |
| kimi-k2.7-code | $0.05 | $0.010 | −81% | 61 | Moonshot | 6 | ✓ |
| minimax-m3 | $0.06 | $0.01 | −80% | 59 | MiniMax | 6 | ✓ |
| sakana-namazu | $0.16 | $0.04 | −77% | — | Sakana | 7 | ✓ |
| seed-2.0-lite | $0.01 | $0.003 | −75% | — | ByteDance | 60 | ✓ |
| deepseek-v4-flash | $0.02 | $0.005 | −71% | 56 | DeepSeek | 8 | ✓ |
| haiku-4.5 | $0.18 | $0.05 | −71% | 44 | Anthropic | 60 | ✓ |
| hy3 | $0.02 | $0.004 | −71% | 59 | Tencent | 9 | ✓ |
| glm-5.2 | $0.03 | $0.008 | −69% | 69 | Z.ai | 39 | ✓ |
| haiku-4.5 (Aug 2026) | $0.15 | $0.05 | −68% | 44 | Anthropic | 20 | ✓ |
| opus-4.6 | $0.29 | $0.10 | −67% | — | Anthropic | 6 | ✓ |
| step-3.7-flash | $0.04 | $0.01 | −64% | 40 | StepFun | 60 | ✓ |
| qwen3.6-plus | $0.06 | $0.02 | −64% | 55 | Alibaba | 7 | ✓ |
| gpt-5.4-mini | $0.02 | $0.009 | −63% | 56 | OpenAI | 10 | ✓ |
| grok-4.20-0309-reasoning | $0.27 | $0.10 | −62% | — | xAI | 21 | ✓ |
| fable-5 | $0.59 | $0.25 | −57% | 77 | Anthropic | 32 | ✓ |
| gpt-5.6-luna | $0.01 | $0.005 | −57% | 71 | OpenAI | 60 | ✓ |
| glm-5.1 | $0.05 | $0.02 | −56% | 56 | Z.ai | 60 | ✓ |
| fugu-ultra | $0.49 | $0.21 | −56% | — | Sakana | 10 | ✓ |
| qwen3-coder-plus | $0.07 | $0.03 | −55% | — | Alibaba | 7 | ✓ |
| muse-spark-1.2 | $0.07 | $0.03 | −54% | 72 | Meta | 6 | ✓ |
| grok-4.5 | $0.03 | $0.01 | −53% | 72 | xAI | 8 | ✓ |
| sonnet-5 | $0.26 | $0.13 | −52% | 72 | Anthropic | 60 | ✓ |
| gpt-5.5 | $0.17 | $0.09 | −50% | 75 | OpenAI | 67 | ✓ |
| muse-spark-1.1 | $0.04 | $0.02 | −49% | 71 | Meta | 7 | ✓ |
| opus-4.8 | $0.40 | $0.21 | −49% | 74 | Anthropic | 60 | ✓ |
| qwen3.7-max | $0.05 | $0.03 | −48% | 66 | Alibaba | 6 | ✓ |
| kat-coder-pro-v2.5 | $0.05 | $0.03 | −46% | — | Kuaishou | 5 | ✓ |
| ling-3.0-flash | $0.004 | $0.002 | −45% | 51 | InclusionAI | 29 | ✓ |
| qwen3.8-max | $0.09 | $0.05 | −45% | 72 | Alibaba | 6 | ✓ |
| gemini-3-flash-preview | $0.02 | $0.01 | −45% | — | 47 | ✓ | |
| gpt-5.6-terra | $0.02 | $0.01 | −43% | 77 | OpenAI | 60 | ✓ |
| kimi-k3 | $0.06 | $0.04 | −37% | 76 | Moonshot | 6 | ✓ |
| grok-4.3 | $0.02 | $0.01 | −35% | 42 | xAI | 6 | ✓ |
| opus-5 | $0.25 | $0.17 | −31% | 78 | Anthropic | 60 | ✓ |
| longcat-2.0 | $0.01 | $0.008 | −30% | 45 | Meituan | 6 | ✓ |
| gpt-5.6-sol | $0.07 | $0.05 | −29% | 78 | OpenAI | 60 | ✓ |
| gpt-5.4 | $0.03 | $0.02 | −28% | 71 | OpenAI | 79 | ✓ |
| inkling | $0.03 | $0.02 | −21% | 52 | ThinkingMachines | 6 | ✓ |
| sonnet-4.6 | $0.11 | $0.09 | −20% | 63 | Anthropic | 87 | ✓ |
| nemotron-3-ultra | $0.07 | $0.06 | −20% | 49 | Nvidia | 10 | ✓ |
| opus-4.7 | $0.20 | $0.17 | −18% | 74 | Anthropic | 84 | ✓ |
| gemini-3.1-pro-preview | $0.12 | $0.12 | +2% | 69 | 8 | * | |
| qwen3-coder | $0.10 | $0.12 | +21% | — | Alibaba | 15 | * |
| gemini-2.5-flash | $0.008 | $0.010 | +30% | — | 15 | * | |
| grok-3-fast | $0.05 | $0.07 | +40% | — | xAI | 21 | * |
| deepseek-v4-pro | $0.01 | $0.02 | +44% | 59 | DeepSeek | 8 | * |
| grok-3-mini | $0.005 | $0.010 | +106% | — | xAI | 14 | * |
| mistral-medium-3.5 | $0.02 | $0.05 | +107% | 47 | Mistral | 10 | * |
| grok-4-1-fast-non-reasoning | $0.02 | $0.06 | +200% | — | xAI | 6 | * |
✓ improved · * got worse · N/A: the model did not edit the file in the control run, so there is no baseline to compare against. Click a column to sort. Measured over 63 models, 22 labs, 1,431 agent-mode trials. Capability is the Artificial Analysis Coding Index, retrieved 2026-08-18, best listed configuration per model. It is not measured by us; it is here so the table can be sorted by how capable a model is. — means the model is not listed there.
In collaboration with AFRL · figures as published
Distribution (A) Approved for public release; distribution is unlimited; AFRL-2026-3793, AUG 2026
🏆 Three recommended models, ranked by cost with ReGrade.
For teams choosing their AI coding agent today: best results-per-dollar with ReGrade.
Ranked from the cleared study's 17 models, measured April to May 2026. These three rows are the study as published and do not move as later testing continues.
🥇
Gemini 3 Flash Preview
17.7 / 18
bugs fixed
$0.008
per bug
🥈
GPT-5.4
17.7 / 18
bugs fixed
$0.015
per bug
🥉
DeepSeek V4 Pro
16.2 / 18
bugs fixed
$0.017
per bug
Five more model + ReGrade combinations clear the cheap-AND-effective threshold. Works equally well across providers.
ReGrade is a single addition to the agent's prompt, you pay only the marginal LLM tokens for that context, with no separate ReGrade fee per bug.
Want ReGrade for your AI coding agent?
Drop the context block into your agent's prompt. No retraining. No CI changes.