Skip to main content
Curtail

Joint study with AFRL

Frozen

ReGrade helps AI coding agents fix 152% more bugs.

In a recently published joint study, Curtail and the U.S. Air Force Research Laboratory tested 17 language models from six providers.

Every one of the seventeen models improved, from 33% on GPT-5.4 to 825% on Haiku 4.5. Cost per bug fixed dropped on 14 of the 17, by a median of 42% and by as much as 89%.

See what changed in your own traffic.

Choose one service and define a baseline, candidate build and replay traffic. Evaluate repair results and total operating cost before expanding to your team.