HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals

HarvestBench measures revealed animal-welfare preferences of LLM agents in a gridworld farm simulation built on Inspect. Two LLM-controlled tractors harvest corn while encountering animals, hay bales, and rocks. At each obstacle, the model chooses: continue (free), swerve (small fuel cost), or reroute (large fuel cost). Kill rate and price elasticity of avoidance are the primary metrics. Scoring is purely event-based (no LLM grader). Conditions vary system prompt (morality vs. neutral briefing) and swerve price.

Overview

⚠️ External evaluation. Code lives in an upstream repository. inspect_evals lists it for discoverability; review the upstream repo and pinned commit before running.

Source: CompassionML/harvestbench@e6be10a

HarvestBench measures revealed animal-welfare preferences of LLM agents in a gridworld farm simulation built on Inspect. Two LLM-controlled tractors harvest corn while encountering animals, hay bales, and rocks. At each obstacle, the model chooses: continue (free), swerve (small fuel cost), or reroute (large fuel cost). Kill rate and price elasticity of avoidance are the primary metrics. Scoring is purely event-based (no LLM grader). Conditions vary system prompt (morality vs. neutral briefing) and swerve price.

Usage

Installation

This is an externally-maintained evaluation. Clone the upstream repository at the pinned commit and install its dependencies:

git clone https://github.com/CompassionML/harvestbench
cd harvestbench
git checkout e6be10a23f3d4a2448fb18ae1d3a60af63222ad6
uv sync

Running evaluations

CLI

uv run inspect eval harvest/contact_task.py@harvest_contact --model openai/gpt-5-nano

Python

from inspect_ai import eval
from harvest.contact_task import harvest_contact

eval(harvest_contact(), model="openai/gpt-5-nano")

View logs

Log viewing (inspect view) and default-model setup are documented in the Inspect Evals README.

More information

For the dataset, scorer, task parameters, and validation, see the upstream repo: CompassionML/harvestbench.

Options

You can control a variety of options from the command line. For example:

uv run inspect eval harvest/contact_task.py@harvest_contact --limit 10 --sample-shuffle
uv run inspect eval harvest/contact_task.py@harvest_contact --max-connections 10
uv run inspect eval harvest/contact_task.py@harvest_contact --temperature 0.5

See uv run inspect eval --help for all available options.

More command-line options: Inspect docs ↗

Static checks

Results of inspect-evals-lint at the registered commit. Rule names link to their documentation; findings link to the file at that commit. These checks describe structure and conventions, not whether the evaluation measures what it claims.

Loading lint results…