Particle.news

GitHub Releases ReviewBench to Standardize AI Code-Review Evaluation

The research-preview benchmark uses a validated reference set and an AI judge to let teams compare reviewers and improve production models.

Overview

  • GitHub made ReviewBench publicly available in research preview on Tuesday to offer a reproducible testbed for AI code-review agents.
  • The benchmark runs all agents on a fixed corpus of 219 pull requests from 187 public repositories spanning 19 programming languages and reports six metrics focused on grounded recall.
  • GitHub built a validated 'golden set' of known issues from human reviews, follow-up changes, static tools and model outputs and uses an AI judge to credit valid findings outside that set.
  • GitHub applied ReviewBench to iterate Copilot code review and reported an 8% rise in comments that prompted code changes, a 13.6% lift in recall, and an 8% cut in cost per review in a cited production test.
  • Users can submit agents via a container image and model key, run a 25-PR trial then three full runs with scores private until maintainer approval, and GitHub is inviting researchers to test the methodology while noting the 219-PR corpus may not cover every project type.