Cost reduction
OngoingIn our ongoing testing, ReGrade cuts LLM API cost per bug fixed by over 58%.
Across 64 models in Curtail’s current corpus, the typical model cut its LLM API cost per bug fixed by 58.5% with ReGrade. 61 of the 64 got cheaper per bug fixed. LLM API cost per bug fixed is the model API spend divided by the bugs the agent actually eliminated, at the rate each provider billed, or its published price where no bill exists. Subscription fees, deployment work and recording/replay infrastructure are excluded.
Curtail research · 74 models, 23 labs, updated 2026-09-30
The AFRL study measured this across 17 models in April to May 2026, where the median model cut cost per bug fixed by 42%. These figures cover 74 models. Read the study as published.
Adding the behavioral report raises what a run costs, and the agent finds enough more that the cost per bug fixed falls.
Models with LLM API cost measurements
64 models have comparable costs. 1 more has no baseline cost per fix. The table excludes models without cost measurements; the full corpus contains 74 models.
| Cost | Bug Fixes | |||||||
|---|---|---|---|---|---|---|---|---|
| Verdict | ||||||||
| deepseek-v4.1-flash | $0.10 | $0.0037 | −96% | 8.3 | 17.3 | DeepSeek | 6 | ✓ |
| ling-2.6-1t | $0.0873 | $0.0064 | −93% | 1.6 | 7.0 | InclusionAI | 10 | ✓ |
| deepseek-v4-pro-0813 | $0.32 | $0.0273 | −92% | 3.2 | 13.5 | DeepSeek | 25 | ✓ |
| solar-pro4 | $0.0217 | $0.0019 | −91% | 3.8 | 10.6 | Upstage | 10 | ✓ |
| qwen3.6-max-preview | $0.34 | $0.0361 | −89% | 5.7 | 12.8 | Alibaba | 7 | ✓ |
| deepseek-v4-flash-0731 | $0.0216 | $0.0029 | −86% | 4.1 | 16.7 | DeepSeek | 20 | ✓ |
| laguna-xs-2.1 | $0.0475 | $0.0074 | −84% | 1.8 | 10.3 | Poolside | 30 | ✓ |
| minimax-m2.5 | $0.0339 | $0.0056 | −83% | 2.1 | 10.9 | MiniMax | 60 | ✓ |
| kimi-k2.7-code | $0.0513 | $0.0096 | −81% | 5.0 | 15.7 | Moonshot | 6 | ✓ |
| deepseek-v4-pro (Sep 2026) | $0.18 | $0.0362 | −80% | 9.0 | 17.7 | DeepSeek | 14 | ✓ |
| minimax-m3 | $0.0596 | $0.0121 | −80% | 7.3 | 17.7 | MiniMax | 6 | ✓ |
| muse-spark-1.3 | $0.52 | $0.12 | −77% | 11.4 | 17.8 | Meta | 10 | ✓ |
| sakana-namazu | $0.16 | $0.0357 | −77% | 5.0 | 12.0 | Sakana | 7 | ✓ |
| seed-2.0-lite | $0.0215 | $0.0056 | −74% | 2.9 | 14.7 | ByteDance | 60 | ✓ |
| deepseek-v4-flash-0423 | $0.0156 | $0.0042 | −73% | 3.8 | 7.0 | DeepSeek | 8 | ✓ |
| glm-5.3 | $0.24 | $0.0675 | −72% | 8.7 | 16.3 | Z.ai | 6 | ✓ |
| nemotron-3.5-lightning | $0.0693 | $0.0197 | −72% | 1.4 | 5.3 | Nvidia | 8 | ✓ |
| haiku-4.5 | $0.18 | $0.0511 | −71% | 2.5 | 12.8 | Anthropic | 60 | ✓ |
| hy3 | $0.0153 | $0.0045 | −71% | 6.0 | 14.0 | Tencent | 9 | ✓ |
| mimo-v2.6-pro | $0.0174 | $0.0053 | −69% | 10.0 | 18.0 | Xiaomi | 6 | ✓ |
| grok-4-1-fast-non-reasoning | $0.0158 | $0.0048 | −69% | 2.5 | 9.7 | xAI | 21 | ✓ |
| glm-5.2 | $0.0218 | $0.007 | −68% | 6.7 | 16.6 | Z.ai | 39 | ✓ |
| haiku-4.5 (Aug 2026) | $0.15 | $0.0471 | −68% | 2.7 | 14.5 | Anthropic | 20 | ✓ |
| gemini-2.5-flash | $0.0255 | $0.0084 | −67% | 2.0 | 14.0 | 6 | ✓ | |
| opus-4.6 | $0.29 | $0.0972 | −67% | 7.0 | 17.7 | Anthropic | 6 | ✓ |
| qwen3.8-max-0902 | $0.27 | $0.0883 | −67% | 9.0 | 17.6 | Alibaba | 8 | ✓ |
| step-3.7-flash | $0.0398 | $0.0142 | −64% | 3.4 | 11.4 | StepFun | 59 | ✓ |
| gpt-5.4-mini | $0.0229 | $0.0085 | −63% | 3.6 | 13.2 | OpenAI | 10 | ✓ |
| gemini-3-flash-preview | $0.0333 | $0.0125 | −63% | 6.3 | 17.5 | 32 | ✓ | |
| qwen3.6-plus | $0.0609 | $0.0238 | −61% | 4.0 | 16.0 | Alibaba | 6 | ✓ |
| grok-4.20-0309-reasoning | $0.27 | $0.11 | −60% | 3.2 | 11.7 | xAI | 20 | ✓ |
| glm-5.3-prime | $0.17 | $0.0722 | −58% | 10.0 | 16.7 | Z.ai | 6 | ✓ |
| ling-3.0-flash | $0.0094 | $0.0039 | −58% | 1.9 | 11.3 | InclusionAI | 27 | ✓ |
| fable-5 | $0.59 | $0.25 | −57% | 11.5 | 18.0 | Anthropic | 32 | ✓ |
| gpt-5.6-luna | $0.012 | $0.0052 | −57% | 7.5 | 16.8 | OpenAI | 60 | ✓ |
| glm-5.1 | $0.0455 | $0.0198 | −56% | 4.5 | 17.6 | Z.ai | 60 | ✓ |
| fugu-ultra | $0.49 | $0.21 | −56% | 10.2 | 18.0 | Sakana | 10 | ✓ |
| qwen3.8-max-prime | $0.29 | $0.13 | −56% | 10.7 | 18.0 | Alibaba | 6 | ✓ |
| qwen3-coder-plus | $0.0692 | $0.032 | −54% | 2.0 | 12.0 | Alibaba | 6 | ✓ |
| grok-4.5 | $0.032 | $0.0149 | −53% | 11.2 | 18.0 | xAI | 8 | ✓ |
| sonnet-5 | $0.26 | $0.13 | −52% | 5.5 | 17.4 | Anthropic | 60 | ✓ |
| muse-spark-1.1 | $0.0376 | $0.019 | −49% | 13.2 | 17.7 | Meta | 7 | ✓ |
| gpt-5.5 | $0.17 | $0.0862 | −49% | 13.0 | 18.0 | OpenAI | 61 | ✓ |
| opus-4.8 | $0.40 | $0.21 | −49% | 9.1 | 17.9 | Anthropic | 60 | ✓ |
| qwen3.7-max | $0.0538 | $0.0278 | −48% | 7.3 | 16.7 | Alibaba | 6 | ✓ |
| kat-coder-pro-v2.5 | $0.0529 | $0.0283 | −46% | — | — | Kuaishou | 5 | ✓ |
| gpt-5.6-terra | $0.0211 | $0.0119 | −44% | 11.8 | 17.9 | OpenAI | 57 | ✓ |
| grok-3-fast | $0.10 | $0.0597 | −43% | 2.0 | 8.0 | xAI | 6 | ✓ |
| sonnet-4.6 | $0.14 | $0.0838 | −38% | 5.2 | 17.7 | Anthropic | 72 | ✓ |
| fable51 | $0.25 | $0.15 | −38% | 12.7 | 18.0 | Anthropic | 6 | ✓ |
| gpt-6-sol | $0.0153 | $0.0096 | −37% | 14.3 | 18.0 | OpenAI | 6 | ✓ |
| kimi-k3 | $0.0589 | $0.0371 | −37% | 12.3 | 17.0 | Moonshot | 6 | ✓ |
| grok-4.3 | $0.0157 | $0.0102 | −35% | 2.0 | 6.3 | xAI | 6 | ✓ |
| opus-4.7 | $0.26 | $0.17 | −33% | 8.6 | 17.9 | Anthropic | 66 | ✓ |
| gpt-5.4 | $0.0379 | $0.0259 | −32% | 13.1 | 17.9 | OpenAI | 64 | ✓ |
| gpt-5.6-sol | $0.0699 | $0.048 | −31% | 12.9 | 18.0 | OpenAI | 51 | ✓ |
| opus-5 | $0.25 | $0.17 | −31% | 11.9 | 18.0 | Anthropic | 60 | ✓ |
| gpt-6-luna | $0.0012 | $0.0009 | −28% | 8.7 | 14.3 | OpenAI | 6 | ✓ |
| opus-5.5 | $0.0391 | $0.0345 | −12% | 10.3 | 18.0 | Anthropic | 6 | ✓ |
| sonnet-5.5 | $0.015 | $0.0138 | −8% | 11.3 | 16.3 | Anthropic | 6 | ✓ |
| gpt-6-astra | $0.0465 | $0.0441 | −5% | 14.7 | 18.0 | OpenAI | 6 | ✓ |
| qwen3-coder | $0.0347 | $0.0377 | +9% | 4.7 | 12.5 | Alibaba | 20 | * |
| mistral-medium-3.5 | $0.0503 | $0.0605 | +20% | 6.4 | 14.9 | Mistral | 20 | * |
| gemini-3.1-pro-preview | $0.0902 | $0.11 | +26% | 10.5 | 16.0 | 20 | * | |
| gpt-oss-20b | fixed none | $0.0035 | N/A | 0.0 | 3.8 | OpenAI | 10 | no baseline |
✓ improved · * got worse · N/A: the model fixed no bugs without ReGrade, so there is no cost per bug to compare against. Click a column to sort. Measured over 74 models, 23 labs, 1,558 agent-mode trials. The provider-pinned comparisons add 36 separately labelled trials, for 1,594 total. Runs, both arms counts control and treatment runs together. Scroll sideways to see the lab and sample-count columns. Bug fixes are unaided (mean bugs fixed across the model's control-arm trials) and aided (the same average over the treatment trials), so the table can be sorted by how capable a model is on its own and by how far it gets with help. Both are our own measurements. Repair and damage eligible subsets can differ, so bug-fix averages need not use every replay counted here. A dash means the model sits below the trial floor, so no figure is published for it.
🏆 Three recommended models, ranked by LLM API cost per bug fixed.
Measured on DriftBench among models fixing at least 16 of 18 bugs with ReGrade.
Ranked from Curtail's current corpus of 74 models, by LLM API cost per bug fixed with ReGrade among those fixing at least 16 of 18. This list moves as later testing continues.
🥇
deepseek-v4-flash-0731
16.7 / 18
bugs fixed
$0.003
LLM API cost per bug
🥈
deepseek-v4.1-flash
17.3 / 18
bugs fixed
$0.004
LLM API cost per bug
🥉
gpt-5.6-luna
16.8 / 18
bugs fixed
$0.005
LLM API cost per bug
34 more model and ReGrade combinations clear the same threshold, at least 16 of 18 bugs fixed. Models from multiple providers met this threshold.
These figures measure LLM API cost per bug fixed. ReGrade subscription fees, recording and replay infrastructure, deployment work and human review are excluded. Use your own service to measure total cost.
See ReGrade subscription pricingEvaluate LLM API savings on your service.
Across 64 models, 61 got cheaper per bug fixed, and the typical model cut it by 58.5%. Ask us what that looks like on your stack.
Evaluate LLM API savings on your service.