Skip to main content
Curtail

Cost reduction

Curtail research · 63 models, 22 labs, updated 2026-08-29

Across 55 model cells in Curtail’s current corpus, the typical cell cut its cost per bug fixed by 51.6% with ReGrade. 47 of the 55 got cheaper per bug fixed. Cost per bug fixed is what a run costs, divided by the bugs the agent actually eliminated, at each provider’s list price.

Adding the behavioral report raises what a run costs, and the agent finds enough more that the cost per bug fixed falls.

Every model tested, on cost per bug fixed

55 models · 22 labs
Verdict
laguna-xs-2.1$0.15$0.009−94%Poolside6
ling-2.6-1t$0.09$0.006−93%InclusionAI10
solar-pro4$0.02$0.002−91%53Upstage10
qwen3.6-max-preview$0.34$0.05−87%Alibaba8
minimax-m2.5$0.02$0.003−83%MiniMax60
deepseek-v4-flash-0731$0.02$0.004−83%69DeepSeek22
kimi-k2.7-code$0.05$0.010−81%61Moonshot6
minimax-m3$0.06$0.01−80%59MiniMax6
sakana-namazu$0.16$0.04−77%Sakana7
seed-2.0-lite$0.01$0.003−75%ByteDance60
deepseek-v4-flash$0.02$0.005−71%56DeepSeek8
haiku-4.5$0.18$0.05−71%44Anthropic60
hy3$0.02$0.004−71%59Tencent9
glm-5.2$0.03$0.008−69%69Z.ai39
haiku-4.5 (Aug 2026)$0.15$0.05−68%44Anthropic20
opus-4.6$0.29$0.10−67%Anthropic6
step-3.7-flash$0.04$0.01−64%40StepFun60
qwen3.6-plus$0.06$0.02−64%55Alibaba7
gpt-5.4-mini$0.02$0.009−63%56OpenAI10
grok-4.20-0309-reasoning$0.27$0.10−62%xAI21
fable-5$0.59$0.25−57%77Anthropic32
gpt-5.6-luna$0.01$0.005−57%71OpenAI60
glm-5.1$0.05$0.02−56%56Z.ai60
fugu-ultra$0.49$0.21−56%Sakana10
qwen3-coder-plus$0.07$0.03−55%Alibaba7
muse-spark-1.2$0.07$0.03−54%72Meta6
grok-4.5$0.03$0.01−53%72xAI8
sonnet-5$0.26$0.13−52%72Anthropic60
gpt-5.5$0.17$0.09−50%75OpenAI67
muse-spark-1.1$0.04$0.02−49%71Meta7
opus-4.8$0.40$0.21−49%74Anthropic60
qwen3.7-max$0.05$0.03−48%66Alibaba6
kat-coder-pro-v2.5$0.05$0.03−46%Kuaishou5
ling-3.0-flash$0.004$0.002−45%51InclusionAI29
qwen3.8-max$0.09$0.05−45%72Alibaba6
gemini-3-flash-preview$0.02$0.01−45%Google47
gpt-5.6-terra$0.02$0.01−43%77OpenAI60
kimi-k3$0.06$0.04−37%76Moonshot6
grok-4.3$0.02$0.01−35%42xAI6
opus-5$0.25$0.17−31%78Anthropic60
longcat-2.0$0.01$0.008−30%45Meituan6
gpt-5.6-sol$0.07$0.05−29%78OpenAI60
gpt-5.4$0.03$0.02−28%71OpenAI79
inkling$0.03$0.02−21%52ThinkingMachines6
sonnet-4.6$0.11$0.09−20%63Anthropic87
nemotron-3-ultra$0.07$0.06−20%49Nvidia10
opus-4.7$0.20$0.17−18%74Anthropic84
gemini-3.1-pro-preview$0.12$0.12+2%69Google8*
qwen3-coder$0.10$0.12+21%Alibaba15*
gemini-2.5-flash$0.008$0.010+30%Google15*
grok-3-fast$0.05$0.07+40%xAI21*
deepseek-v4-pro$0.01$0.02+44%59DeepSeek8*
grok-3-mini$0.005$0.010+106%xAI14*
mistral-medium-3.5$0.02$0.05+107%47Mistral10*
grok-4-1-fast-non-reasoning$0.02$0.06+200%xAI6*

improved · * got worse · N/A: the model did not edit the file in the control run, so there is no baseline to compare against. Click a column to sort. Measured over 63 models, 22 labs, 1,431 agent-mode trials. Capability is the Artificial Analysis Coding Index, retrieved 2026-08-18, best listed configuration per model. It is not measured by us; it is here so the table can be sorted by how capable a model is. — means the model is not listed there.

In collaboration with AFRL · figures as published

Distribution (A) Approved for public release; distribution is unlimited; AFRL-2026-3793, AUG 2026

Read the study as published.

Want ReGrade for your AI coding agent?

Drop the context block into your agent's prompt. No retraining. No CI changes.