JudgeBench: A Benchmark for Evaluating LLM-based Judges

JudgeBench evaluates LLM-based judges on 350 pairwise response comparisons (GPT-4o-generated split) across knowledge (MMLU-Pro), reasoning (LiveBench), math (LiveBench), and coding (LiveCodeBench). Each pair contains one objectively correct and one incorrect response generated by a single strong model. Judges select the better response; accuracy against ground-truth labels is the metric. Positional bias is measured by running each pair twice with swapped order and aggregating verdicts.

Overview

⚠️ External evaluation. Code lives in an upstream repository. inspect_evals lists it for discoverability; review the upstream repo and pinned commit before running.

Source: theruviparambil/judgebench-inspect@9bb5f35

JudgeBench evaluates LLM-based judges on 350 pairwise response comparisons (GPT-4o-generated split) across knowledge (MMLU-Pro), reasoning (LiveBench), math (LiveBench), and coding (LiveCodeBench). Each pair contains one objectively correct and one incorrect response generated by a single strong model. Judges select the better response; accuracy against ground-truth labels is the metric. Positional bias is measured by running each pair twice with swapped order and aggregating verdicts.

Usage

Installation

This is an externally-maintained evaluation. Clone the upstream repository at the pinned commit and install its dependencies:

git clone https://github.com/theruviparambil/judgebench-inspect
cd judgebench-inspect
git checkout 9bb5f357e4f4c1784a52d250f2313dcba5e33183
uv sync

Running evaluations

CLI

uv run inspect eval src/judgebench_inspect/task.py@judgebench_gpt_positional --model openai/gpt-5-nano

Python

from inspect_ai import eval
from judgebench_inspect.task import judgebench_gpt_positional

eval(judgebench_gpt_positional(), model="openai/gpt-5-nano")

View logs

Log viewing (inspect view) and default-model setup are documented in the Inspect Evals README.

More information

For the dataset, scorer, task parameters, and validation, see the upstream repo: theruviparambil/judgebench-inspect.

Options

You can control a variety of options from the command line. For example:

uv run inspect eval src/judgebench_inspect/task.py@judgebench_gpt_positional --limit 10 --sample-shuffle
uv run inspect eval src/judgebench_inspect/task.py@judgebench_gpt_positional --max-connections 10
uv run inspect eval src/judgebench_inspect/task.py@judgebench_gpt_positional --temperature 0.5

See uv run inspect eval --help for all available options.

More command-line options: Inspect docs ↗

Static checks

Results of inspect-evals-lint at the registered commit. Rule names link to their documentation; findings link to the file at that commit. These checks describe structure and conventions, not whether the evaluation measures what it claims.

Loading lint results…