Joint study with AFRL
FrozenReGrade helps AI coding agents fix 152% more bugs.
In a recently published joint study, Curtail and the U.S. Air Force Research Laboratory tested 17 language models from six providers.
Every one of the seventeen models improved, from 33% on GPT-5.4 to 825% on Haiku 4.5. Cost per bug fixed dropped on 14 of the 17, by a median of 42% and by as much as 89%.
The research, page by page
In collaboration with AFRL
FrozenJoint study with AFRL
The cleared paper and its executive summary, as published.
Curtail, current testing
OngoingMore bugs fixed
How much more AI agents catch with ReGrade, on current data.
Curtail, current testing
OngoingCost reduction
What each bug costs to find, with and without ReGrade.
Curtail, current testing
OngoingAI coding hallucinations
How much unintended change ReGrade removes.
Ongoing pages are recomputed as we test more models. The study’s figures are frozen as published.
See what changed in your own traffic.
Choose one service and define a baseline, candidate build and replay traffic. Evaluate repair results and total operating cost before expanding to your team.
See what ReGrade finds in your traffic.