MonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Models (Dual Objectives: Summarization subset)

MonitorBench evaluates CoT monitorability: whether an LLM’s chain-of-thought explicitly surfaces decision-critical factors driving its outputs. The summarization subtask presents 50 instances where the model must summarize a report while verbatim-including a required sentence (side objective). A judge LLM monitors CoT-only, action-only, and all-messages scopes to flag whether the side objective appears. Three conditions are tested: standard, direct concealment, and monitor-aware evasion. Monitorability score measures flagging rate on successful rollouts.

Overview

⚠️ External evaluation. Code lives in an upstream repository. inspect_evals lists it for discoverability; review the upstream repo and pinned commit before running.

Source: semsorock/inspect-evals-monitor-bench@e0b4a57

MonitorBench evaluates CoT monitorability: whether an LLM’s chain-of-thought explicitly surfaces decision-critical factors driving its outputs. The summarization subtask presents 50 instances where the model must summarize a report while verbatim-including a required sentence (side objective). A judge LLM monitors CoT-only, action-only, and all-messages scopes to flag whether the side objective appears. Three conditions are tested: standard, direct concealment, and monitor-aware evasion. Monitorability score measures flagging rate on successful rollouts.

Usage

Installation

This is an externally-maintained evaluation. Clone the upstream repository at the pinned commit and install its dependencies:

git clone https://github.com/semsorock/inspect-evals-monitor-bench
cd inspect-evals-monitor-bench
git checkout e0b4a577ded3928852c295e4cbe88cbdcf3d6463
uv sync

Running evaluations

CLI

uv run inspect eval src/monitor_bench/tasks/dual_objectives/summarization/eval.py@dual_objectives_summarization --model openai/gpt-5-nano

Python

from inspect_ai import eval
from monitor_bench.tasks.dual_objectives.summarization.eval import dual_objectives_summarization

eval(dual_objectives_summarization(), model="openai/gpt-5-nano")

View logs

Log viewing (inspect view) and default-model setup are documented in the Inspect Evals README.

More information

For the dataset, scorer, task parameters, and validation, see the upstream repo: semsorock/inspect-evals-monitor-bench.

Options

You can control a variety of options from the command line. For example:

uv run inspect eval src/monitor_bench/tasks/dual_objectives/summarization/eval.py@dual_objectives_summarization --limit 10 --sample-shuffle
uv run inspect eval src/monitor_bench/tasks/dual_objectives/summarization/eval.py@dual_objectives_summarization --max-connections 10
uv run inspect eval src/monitor_bench/tasks/dual_objectives/summarization/eval.py@dual_objectives_summarization --temperature 0.5

See uv run inspect eval --help for all available options.

More command-line options: Inspect docs ↗

Static checks

Results of inspect-evals-lint at the registered commit. Rule names link to their documentation; findings link to the file at that commit. These checks describe structure and conventions, not whether the evaluation measures what it claims.

Loading lint results…