Skip to main content
Curtail

Cost reduction

Ongoing

In our ongoing testing, ReGrade cuts LLM API cost per bug fixed by over 58%.

Across 64 models in Curtail’s current corpus, the typical model cut its LLM API cost per bug fixed by 58.5% with ReGrade. 61 of the 64 got cheaper per bug fixed. LLM API cost per bug fixed is the model API spend divided by the bugs the agent actually eliminated, at the rate each provider billed, or its published price where no bill exists. Subscription fees, deployment work and recording/replay infrastructure are excluded.

Curtail research · 74 models, 23 labs, updated 2026-09-30

The AFRL study measured this across 17 models in April to May 2026, where the median model cut cost per bug fixed by 42%. These figures cover 74 models. Read the study as published.

Adding the behavioral report raises what a run costs, and the agent finds enough more that the cost per bug fixed falls.

Models with LLM API cost measurements

64 models have comparable costs. 1 more has no baseline cost per fix. The table excludes models without cost measurements; the full corpus contains 74 models.

65 models · 21 labs
CostBug Fixes
Verdict
deepseek-v4.1-flash$0.10$0.0037−96%8.317.3DeepSeek6✓
ling-2.6-1t$0.0873$0.0064−93%1.67.0InclusionAI10✓
deepseek-v4-pro-0813$0.32$0.0273−92%3.213.5DeepSeek25✓
solar-pro4$0.0217$0.0019−91%3.810.6Upstage10✓
qwen3.6-max-preview$0.34$0.0361−89%5.712.8Alibaba7✓
deepseek-v4-flash-0731$0.0216$0.0029−86%4.116.7DeepSeek20✓
laguna-xs-2.1$0.0475$0.0074−84%1.810.3Poolside30✓
minimax-m2.5$0.0339$0.0056−83%2.110.9MiniMax60✓
kimi-k2.7-code$0.0513$0.0096−81%5.015.7Moonshot6✓
deepseek-v4-pro (Sep 2026)$0.18$0.0362−80%9.017.7DeepSeek14✓
minimax-m3$0.0596$0.0121−80%7.317.7MiniMax6✓
muse-spark-1.3$0.52$0.12−77%11.417.8Meta10✓
sakana-namazu$0.16$0.0357−77%5.012.0Sakana7✓
seed-2.0-lite$0.0215$0.0056−74%2.914.7ByteDance60✓
deepseek-v4-flash-0423$0.0156$0.0042−73%3.87.0DeepSeek8✓
glm-5.3$0.24$0.0675−72%8.716.3Z.ai6✓
nemotron-3.5-lightning$0.0693$0.0197−72%1.45.3Nvidia8✓
haiku-4.5$0.18$0.0511−71%2.512.8Anthropic60✓
hy3$0.0153$0.0045−71%6.014.0Tencent9✓
mimo-v2.6-pro$0.0174$0.0053−69%10.018.0Xiaomi6✓
grok-4-1-fast-non-reasoning$0.0158$0.0048−69%2.59.7xAI21✓
glm-5.2$0.0218$0.007−68%6.716.6Z.ai39✓
haiku-4.5 (Aug 2026)$0.15$0.0471−68%2.714.5Anthropic20✓
gemini-2.5-flash$0.0255$0.0084−67%2.014.0Google6✓
opus-4.6$0.29$0.0972−67%7.017.7Anthropic6✓
qwen3.8-max-0902$0.27$0.0883−67%9.017.6Alibaba8✓
step-3.7-flash$0.0398$0.0142−64%3.411.4StepFun59✓
gpt-5.4-mini$0.0229$0.0085−63%3.613.2OpenAI10✓
gemini-3-flash-preview$0.0333$0.0125−63%6.317.5Google32✓
qwen3.6-plus$0.0609$0.0238−61%4.016.0Alibaba6✓
grok-4.20-0309-reasoning$0.27$0.11−60%3.211.7xAI20✓
glm-5.3-prime$0.17$0.0722−58%10.016.7Z.ai6✓
ling-3.0-flash$0.0094$0.0039−58%1.911.3InclusionAI27✓
fable-5$0.59$0.25−57%11.518.0Anthropic32✓
gpt-5.6-luna$0.012$0.0052−57%7.516.8OpenAI60✓
glm-5.1$0.0455$0.0198−56%4.517.6Z.ai60✓
fugu-ultra$0.49$0.21−56%10.218.0Sakana10✓
qwen3.8-max-prime$0.29$0.13−56%10.718.0Alibaba6✓
qwen3-coder-plus$0.0692$0.032−54%2.012.0Alibaba6✓
grok-4.5$0.032$0.0149−53%11.218.0xAI8✓
sonnet-5$0.26$0.13−52%5.517.4Anthropic60✓
muse-spark-1.1$0.0376$0.019−49%13.217.7Meta7✓
gpt-5.5$0.17$0.0862−49%13.018.0OpenAI61✓
opus-4.8$0.40$0.21−49%9.117.9Anthropic60✓
qwen3.7-max$0.0538$0.0278−48%7.316.7Alibaba6✓
kat-coder-pro-v2.5$0.0529$0.0283−46%——Kuaishou5✓
gpt-5.6-terra$0.0211$0.0119−44%11.817.9OpenAI57✓
grok-3-fast$0.10$0.0597−43%2.08.0xAI6✓
sonnet-4.6$0.14$0.0838−38%5.217.7Anthropic72✓
fable51$0.25$0.15−38%12.718.0Anthropic6✓
gpt-6-sol$0.0153$0.0096−37%14.318.0OpenAI6✓
kimi-k3$0.0589$0.0371−37%12.317.0Moonshot6✓
grok-4.3$0.0157$0.0102−35%2.06.3xAI6✓
opus-4.7$0.26$0.17−33%8.617.9Anthropic66✓
gpt-5.4$0.0379$0.0259−32%13.117.9OpenAI64✓
gpt-5.6-sol$0.0699$0.048−31%12.918.0OpenAI51✓
opus-5$0.25$0.17−31%11.918.0Anthropic60✓
gpt-6-luna$0.0012$0.0009−28%8.714.3OpenAI6✓
opus-5.5$0.0391$0.0345−12%10.318.0Anthropic6✓
sonnet-5.5$0.015$0.0138−8%11.316.3Anthropic6✓
gpt-6-astra$0.0465$0.0441−5%14.718.0OpenAI6✓
qwen3-coder$0.0347$0.0377+9%4.712.5Alibaba20*
mistral-medium-3.5$0.0503$0.0605+20%6.414.9Mistral20*
gemini-3.1-pro-preview$0.0902$0.11+26%10.516.0Google20*
gpt-oss-20bfixed none$0.0035N/A0.03.8OpenAI10no baseline

✓ improved · * got worse · N/A: the model fixed no bugs without ReGrade, so there is no cost per bug to compare against. Click a column to sort. Measured over 74 models, 23 labs, 1,558 agent-mode trials. The provider-pinned comparisons add 36 separately labelled trials, for 1,594 total. Runs, both arms counts control and treatment runs together. Scroll sideways to see the lab and sample-count columns. Bug fixes are unaided (mean bugs fixed across the model's control-arm trials) and aided (the same average over the treatment trials), so the table can be sorted by how capable a model is on its own and by how far it gets with help. Both are our own measurements. Repair and damage eligible subsets can differ, so bug-fix averages need not use every replay counted here. A dash means the model sits below the trial floor, so no figure is published for it.

Evaluate LLM API savings on your service.

Across 64 models, 61 got cheaper per bug fixed, and the typical model cut it by 58.5%. Ask us what that looks like on your stack.