Lotu Radar About

Copilot tops GitHub’s own AI code review benchmark. An independent one tells a different story.

The New Stack Cloud & Infrastructure Score 7/10
Copilot tops GitHub’s own AI code review benchmark. An independent one tells a different story.

Summary

If there’s one thing the AI software engineering world doesn’t lack, it’s benchmarks. Want to know whether an agent can The post Copilot tops GitHub’s own AI code review benchmark. An independent one tells a different story. appeared first on The New Stack .

Original Text

If there’s one thing the AI software engineering world doesn’t lack, it’s benchmarks. Want to know whether an agent can resolve real-world GitHub issues? There’s SWE-bench. Or how about seeing how well it can operate inside a terminal? Terminal-Bench has you covered.

Now, GitHub has decided AI code review needs another yardstick. On Monday, the company unveiled ReviewBench, an open benchmark developed with Microsoft for measuring how well AI code review agents find useful problems in pull requests. And sitting at the top of its inaugural leaderboard is — perhaps unsurprisingly — GitHub Copilot code review.

Putting code reviewers to the test

GitHub debuted Copilot code review in October 2024, making it generally available to all paid Copilot subscribers the following April. In the intervening months, GitHub has moved the reviewer onto an agentic architecture that pulls in wider context from the repository, started billing it in GitHub Actions minutes on private repositories, let it approve pull requests, and made a more thorough “Balanced” review mode the default on Sept. 28.

Code review has always been a crucial part of software development, and GitHub has plenty of company trying to automate more of it. Qodo, Greptile, Cubic, Devin and Cursor are among the products now offering AI-assisted reviews, while CodeRabbit also features prominently on other code-review leaderboards.

And now GitHub wants to provide another common test for comparing how well those systems actually perform.

ReviewBench takes 219 public pull requests from 187 public repositories spanning 19 programming languages, and asks competing reviewers to inspect the same changes. GitHub says it selected the corpus after analyzing 103.9 million pull requests: its language and repository-size distributions closely mirror GitHub overall, while pull request size is deliberately weighted away from tiny, single-file changes.

To establish what the reviewers ought to find, the benchmark builds a reference set from several sources, including human review comments, subsequent changes made by authors, static-analysis tools and LLM reviewers. Claude Sonnet 5 classifies findings, while a separate LLM matcher determines whether candidate findings correspond to the same underlying issues in that reference set. The methodology says 47 findings initially classified as true positives were manually corrected, while human and classifier judgments agreed 96.6% of the time on whether findings were true or false positives.

Copilot code review, tested in its Balanced configuration, leads the leaderboard with a 40.1% grounded F1 score, which combines precision and recall against the benchmark’s known findings.

A snapshot of the ReviewBench leaderboard

There are some sizeable caveats attached, however. GitHub says its team generated the initial entries itself by running the publicly available versions of each product — the vendors neither conducted nor verified those tests. And because the products were tested on different dates, some results are substantially older than others: Copilot was tested on Oct. 1, while Cubic and Greptile were tested back in June. ReviewBench cautions that the products may have changed since then, and that performance on its corpus may not translate to a company’s own code.

A vendor-published benchmark topped by that vendor’s own product makes those qualifications particularly notable. GitHub does, however, publish the dataset, methodology and judging setup, and allows vendors to submit their own runs.

“We’ve been using ReviewBench to improve GitHub Copilot Code Review, and its offline results have consistently anticipated the direction of later production experiments.”

GitHub is also using its own product development as evidence that ReviewBench has some value beyond leaderboard metrics. Taking to LinkedIn on Monday, Alejandro Carderera de Diego, a staff applied engineer at GitHub, says the benchmark has proved useful as an early signal for how changes to Copilot Code Review will fare when tested in production.

“We’ve been using ReviewBench to improve GitHub Copilot Code Review, and its offline results have consistently anticipated the direction of later production experiments,” he writes.

Ask another benchmark, get another winner

In a blog post published on Monday, Carderera de Diego and Michelle Zhou, an applied scientist at Microsoft, argue that existing approaches force compromises over what gets measured and how closely the results resemble real-world reviewing.

“Existing benchmarks often make tradeoffs between label quality, coverage, and how well they represent real-world code review, leaving a gap for a rigorous and reproducible evaluation methodology that brings these pieces together,” they write.

“Existing benchmarks often make tradeoffs between label quality, coverage, and how well they represent real-world code review.”

One such existing attempt comes from Martian, an AI research company, which launched its Code Review Bench in February. It combines an offline test with an online tracker based on how developers respond to review comments across real open source repositories. At launch, Martian said its offline results diverged from what it observed in real-world use, and made the online benchmark its headline metric.

As of Oct. 6, Martian’s online leaderboard puts Cubic first with a 64.9% F1 score, followed by Greptile and CodeRabbit, while GitHub Copilot ranks fourth at 60.9%. F1 is a standard metric that combines precision and recall into a single score, giving both equal weight.

Martian’s leaderboard (online)

The picture changes on Martian’s offline leaderboard. Qodo Deep ranks first, followed by Cubic and Augment, while GitHub Copilot sits fifth with a 58% F2 score. F2 is a variation of the same metric that gives recall more weight than precision — in other words, it rewards systems more heavily for catching a greater share of relevant issues.

Martian’s leaderboard (offline)

These scores still aren’t directly comparable with ReviewBench. Although ReviewBench also uses an F1-style measure for its default ranking, it draws on a different dataset and scoring methodology, while Martian’s online and offline tests measure different kinds of evidence.

Still, the contrasting results help illustrate how much benchmark design can influence the picture of which tools are performing best.

Martian has also made independence a core part of its pitch for Code Review Bench. In its February launch post, the authors argued that maintaining a credible benchmark takes considerable money and effort, particularly when both the tools being tested and the evidence used to judge them keep changing. It says that has historically pushed benchmark development toward either academia or the vendors themselves.

“We’re trying a third option: a well-funded research lab that doesn’t train models or sell coding tools, and has no stake in which tool wins,” the authors wrote at the time.

“We’re trying a third option: a well-funded research lab that doesn’t train models or sell coding tools, and has no stake in which tool wins.”

The benchmark and methodology are open source under an MIT license, and Martian has invited tool builders, model makers and researchers to review and contribute to the project.

Separately, it’s worth noting that back in July, LangChain published its own AI code-review benchmark called ReviewBench. That version is a smaller, more model-focused evaluation built from 59 tasks drawn from LangChain’s LangSmith codebase, and unlike GitHub or Martian’s versions, it doesn’t come with a public vendor leaderboard. LangChain ran different models through the same Deep Agents setup, so it’s really testing model performance under a common agent configuration rather than ranking commercial code-review products.

That all said, GitHub’s ReviewBench has some scope to become more representative over time. Vendors can submit their own runs through a self-service system, with new results added to the leaderboard as they are scored — though they will, of course, also choose which configurations to put forward. Whether enough of them do so to turn the current GitHub-produced snapshot into a broader vendor-tested leaderboard will be the more interesting test.

The post Copilot tops GitHub’s own AI code review benchmark. An independent one tells a different story. appeared first on The New Stack.

CloudInfrastructure

Lotu Radar provides attributed news summaries and links to the original publisher. Full reporting and copyright remain with the source.