Skip to main content
Curtail

Review AI-written code with evidence

Joint study with AFRL

Frozen

ReGrade helps AI coding agents fix 152% more bugs.

Every model we tested fixed more bugs.

A joint study with the Air Force Research Laboratory, cleared for public release. 17 language models from 6 providers, and the median model fixed 152% more bugs at a median 42% lower LLM API cost per bug fixed.

The study supplied behavioral evidence in the agent's initial prompt.

See the evidence for this →

Why teams use ReGrade

Three minutes on the problem it solves: catching the changes no test asserts, knowing a release holds up under load, and proving a refactor behaves like the code it replaced.

Want to watch one run start to finish? See the runnable demos. Each one has a repo you can clone.

The code is changing. The risks are growing.

84%

of survey respondents are using or planning to use AI tools

Stack Overflow 2025

1.7×

more issues per AI-co-authored PR in a study of 470 open-source PRs

CodeRabbit 2025

96%

of developers don't fully trust AI-generated code

ShiftMag / Sonar 2025

41%

of tracked AI-introduced security issues remained in the studied repositories

arXiv 2026

Proven in the Wild

A vulnerability hid for 7 years. Zero tests caught it. ReGrade found it on the first replay.

A widely-used open-source platform shipped a password hash disclosure bug for 7 years. Standard tests passed every day. ReGrade caught it on the first replay, without knowing the vulnerability existed.

example ReGrade merge request comments

example ReGrade merge request comments

7 Years

Undetected

First Replay

Found Immediately

Zero Prior Knowledge

Used existing tests

Powered by NCAST Technology

Compare behavior before you approve a release.

ReGrade sends identical real-world requests to your current and candidate versions, then compares the responses field by field. It runs on the traffic and the tests you already have.

1

Record

Capture traffic from any source: production, CI tests, or security scanners. ReGrade works with whatever you already have.

2

Replay

Replay the same traffic against your candidate version. The original responses were already captured in step one.

3

Compare

Review field-level and performance differences to identify unintended changes before release.

What testing and code review miss

Tests verify what you expect. ReGrade catches what you don't.

Slow before anyone complains

ReGrade compares P95 and P99 latency across versions, so a release that quietly got slower is a number you see on the merge request rather than a support ticket next week.

Standout differentiator

Zero-Day Discovery

By comparing actual responses against a known-good baseline, ReGrade surfaces vulnerabilities that no test was written to find, because no one knew they existed.

Proven: Mattermost case study

Give your agents evidence for repairs

ReGrade supplies behavioral findings to the coding agent so it can attempt repairs. Your engineers review the changes and decide which differences are acceptable.

Agent feedback

Works with your AI coding tools via MCP (Model Context Protocol)

It meets your team where they already work

MCP Server

Connect your coding agents through MCP so they can investigate behavioral differences and attempt repairs.

Merge Request Analysis

Automatically analyze candidate builds before merge. ReGrade comments appear alongside your code review with specific field-level findings.

Preview Before Deploy

Replay production traffic against your candidate version in a safe environment. Inspect differences on the requests you replay before approving a release.

Plan a ReGrade evaluation

Choose one service and compare a baseline with a candidate build. Discuss what changed, how your team will review findings, and what recording and replay will cost.