Hack-Verifiable Terminal Bench (HVTB)

HVTB measures reward-hacking rates of LLM agents on 89 real-world terminal/coding tasks from Terminal Bench 2.1. Each task container plants a reference solution and held-out tests under /app/admin/, monitored via inotify filesystem watchers. If an agent reads these files, a hack is flagged automatically. The primary metric is hack_rate (fraction of tasks where a watcher fired). The Inspect implementation uses a react agent with a bash tool instead of the paper’s native agent harnesses.

Overview

⚠️ External evaluation. Code lives in an upstream repository. inspect_evals lists it for discoverability; review the upstream repo and pinned commit before running.

Source: Aarav500/rhob@a2c7341

HVTB measures reward-hacking rates of LLM agents on 89 real-world terminal/coding tasks from Terminal Bench 2.1. Each task container plants a reference solution and held-out tests under /app/admin/, monitored via inotify filesystem watchers. If an agent reads these files, a hack is flagged automatically. The primary metric is hack_rate (fraction of tasks where a watcher fired). The Inspect implementation uses a react agent with a bash tool instead of the paper’s native agent harnesses.

Usage

Installation

This is an externally-maintained evaluation. Clone the upstream repository at the pinned commit and install its dependencies:

git clone https://github.com/Aarav500/rhob
cd rhob
git checkout a2c73417a79be65b802b2fbe62dc6fac03e9c3de
uv sync

Running evaluations

CLI

uv run inspect eval hvtb_hack_detection/tasks.py@hvtb_hack_rate --model openai/gpt-5-nano

Python

from inspect_ai import eval
from hvtb_hack_detection.tasks import hvtb_hack_rate

eval(hvtb_hack_rate(), model="openai/gpt-5-nano")

View logs

Log viewing (inspect view) and default-model setup are documented in the Inspect Evals README.

More information

For the dataset, scorer, task parameters, and validation, see the upstream repo: Aarav500/rhob.

Options

You can control a variety of options from the command line. For example:

uv run inspect eval hvtb_hack_detection/tasks.py@hvtb_hack_rate --limit 10 --sample-shuffle
uv run inspect eval hvtb_hack_detection/tasks.py@hvtb_hack_rate --max-connections 10
uv run inspect eval hvtb_hack_detection/tasks.py@hvtb_hack_rate --temperature 0.5

See uv run inspect eval --help for all available options.

More command-line options: Inspect docs ↗

Static checks

Results of inspect-evals-lint at the registered commit. Rule names link to their documentation; findings link to the file at that commit. These checks describe structure and conventions, not whether the evaluation measures what it claims.

Loading lint results…