Skip to main content
Curtail

Cost reduction

Ongoing

In our ongoing testing, ReGrade cuts the cost per bug fixed by over 54%.

Across 57 models in Curtail’s current corpus, the typical model cut its cost per bug fixed by 54.4% with ReGrade. 50 of the 57 got cheaper per bug fixed. Cost per bug fixed is what a run costs, divided by the bugs the agent actually eliminated, at each provider’s list price.

Curtail research · 68 models, 22 labs, updated 2026-09-18

The AFRL study measured this across 17 models in April to May 2026, where the median model cut cost per bug fixed by 42%. These figures cover 68 models. Read the study as published.

Adding the behavioral report raises what a run costs, and the agent finds enough more that the cost per bug fixed falls.

Every model tested, on cost per bug fixed

57 models · 22 labs
CostBug Fixes
Verdict
ling-2.6-1t$0.09$0.006−93%1.67.0InclusionAI10
solar-pro4$0.02$0.002−91%3.810.6Upstage10
qwen3.6-max-preview$0.34$0.05−87%5.713.8Alibaba8
deepseek-v4-flash-0731$0.03$0.004−86%4.816.7DeepSeek22
laguna-xs-2.1$0.05$0.007−84%1.810.3Poolside30
deepseek-v4-pro-0813$0.11$0.02−83%6.813.6DeepSeek30
minimax-m2.5$0.02$0.003−83%2.110.9MiniMax60
kimi-k2.7-code$0.05$0.010−81%5.015.7Moonshot6
minimax-m3$0.06$0.01−80%7.317.7MiniMax6
sakana-namazu$0.16$0.04−77%5.012.0Sakana7
seed-2.0-lite$0.01$0.003−75%2.914.7ByteDance60
deepseek-v4-flash$0.02$0.005−71%3.87.0DeepSeek8
haiku-4.5$0.18$0.05−71%2.512.8Anthropic60
hy3$0.02$0.004−71%6.014.0Tencent9
grok-4-1-fast-non-reasoning$0.02$0.005−69%2.59.7xAI21
glm-5.2$0.03$0.008−69%6.716.6Z.ai39
haiku-4.5 (Aug 2026)$0.15$0.05−68%2.714.5Anthropic20
opus-4.6$0.29$0.10−67%7.017.7Anthropic6
step-3.7-flash$0.04$0.01−64%3.411.0StepFun60
qwen3.6-plus$0.06$0.02−64%4.016.0Alibaba7
gpt-5.4-mini$0.02$0.009−63%3.613.2OpenAI10
grok-4.20-0309-reasoning$0.27$0.10−62%3.212.3xAI21
qwen3.8-max-0902$0.21$0.09−58%12.617.6Alibaba10
fable-5$0.59$0.25−57%11.518.0Anthropic32
gpt-5.6-luna$0.01$0.005−57%7.516.8OpenAI60
glm-5.1$0.05$0.02−56%4.517.6Z.ai60
fugu-ultra$0.49$0.21−56%10.218.0Sakana10
qwen3-coder-plus$0.07$0.03−55%2.011.8Alibaba7
muse-spark-1.2$0.07$0.03−54%12.017.0Meta6
grok-4.5$0.03$0.01−53%11.218.0xAI8
sonnet-5$0.26$0.13−52%5.517.4Anthropic60
gpt-5.5$0.17$0.09−50%13.318.0OpenAI67
muse-spark-1.1$0.04$0.02−49%13.217.7Meta7
opus-4.8$0.40$0.21−49%9.117.9Anthropic60
qwen3.7-max$0.05$0.03−48%7.316.7Alibaba6
kat-coder-pro-v2.5$0.05$0.03−46%Kuaishou5
ling-3.0-flash$0.004$0.002−45%2.911.3InclusionAI29
gemini-3-flash-preview$0.02$0.01−45%9.617.6Google47
gpt-5.6-terra$0.02$0.01−43%12.017.9OpenAI60
fable51$0.25$0.15−38%12.718.0Anthropic6
kimi-k3$0.06$0.04−37%12.317.0Moonshot6
grok-4.3$0.02$0.01−35%2.06.3xAI6
opus-5$0.25$0.17−31%11.918.0Anthropic60
longcat-2.0$0.01$0.008−30%6.711.0Meituan6
gpt-5.6-sol$0.07$0.05−29%13.918.0OpenAI60
gpt-5.4$0.03$0.02−28%13.917.9OpenAI79
inkling$0.03$0.02−21%9.315.3ThinkingMachines6
sonnet-4.6$0.11$0.09−20%7.317.6Anthropic87
nemotron-3-ultra$0.07$0.06−20%3.013.8Nvidia10
opus-4.7$0.20$0.17−18%10.617.8Anthropic84
gemini-3.1-pro-preview$0.12$0.12+2%8.214.5Google8*
mistral-medium-3.5$0.05$0.06+20%6.414.9Mistral20*
qwen3-coder$0.10$0.12+21%1.75.0Alibaba15*
gemini-2.5-flash$0.008$0.010+30%5.35.6Google15*
grok-3-fast$0.05$0.07+40%6.35.5xAI21*
deepseek-v4-pro$0.01$0.02+44%5.316.6DeepSeek8*
grok-3-mini$0.005$0.009+88%1.91.9xAI16*

improved · * got worse · N/A: the model did not edit the file in the control run, so there is no baseline to compare against. Click a column to sort. Measured over 68 models, 22 labs, 1,571 agent-mode trials. Bug fixes are unaided (mean bugs fixed across the model's control-arm trials) and aided (the same average over the treatment trials), so the table can be sorted by how capable a model is on its own and by how far it gets with help. Both are our own measurement from the same trials as the rest of this table, not a third party's score. — means the model sits below the trial floor, so no figure is published for it.

Cut your cost per bug fixed.

Across 57 models, 50 got cheaper per bug fixed, and the typical model cut it by 54.4%. Ask us what that looks like on your stack.