PRISM Instruction-Order Sensitivity (MATH-500 variant)

Measures how permuting a fixed set of reasoning instruction modules affects a model’s mathematical accuracy on MATH-500 hard problems (levels 3-5). Eight modules are permuted across up to 40,320 orderings (8!); the default samples 50 orderings over 100 questions. Scoring uses canonical answer extraction with deterministic decoding (temperature=0). The construct measured is ordering-induced accuracy variance, isolating permutation effects from wording or content changes.

Overview

⚠️ External evaluation. Code lives in an upstream repository. inspect_evals lists it for discoverability; review the upstream repo and pinned commit before running.

Source: bleymambwe/PRISM@2d578d7

Measures how permuting a fixed set of reasoning instruction modules affects a model’s mathematical accuracy on MATH-500 hard problems (levels 3-5). Eight modules are permuted across up to 40,320 orderings (8!); the default samples 50 orderings over 100 questions. Scoring uses canonical answer extraction with deterministic decoding (temperature=0). The construct measured is ordering-induced accuracy variance, isolating permutation effects from wording or content changes.

Usage

Installation

This is an externally-maintained evaluation. Clone the upstream repository at the pinned commit and install its dependencies:

git clone https://github.com/bleymambwe/PRISM
cd PRISM
git checkout 2d578d761f70e2fef3a381a6243265247b297759
uv sync

Running evaluations

CLI

uv run inspect eval evals/instruction_order/task.py@instruction_order_math500 --model openai/gpt-5-nano

Python

from inspect_ai import eval
from evals.instruction_order.task import instruction_order_math500

eval(instruction_order_math500(), model="openai/gpt-5-nano")

View logs

Log viewing (inspect view) and default-model setup are documented in the Inspect Evals README.

More information

For the dataset, scorer, task parameters, and validation, see the upstream repo: bleymambwe/PRISM.

Options

You can control a variety of options from the command line. For example:

uv run inspect eval evals/instruction_order/task.py@instruction_order_math500 --limit 10 --sample-shuffle
uv run inspect eval evals/instruction_order/task.py@instruction_order_math500 --max-connections 10
uv run inspect eval evals/instruction_order/task.py@instruction_order_math500 --temperature 0.5

See uv run inspect eval --help for all available options.

More command-line options: Inspect docs ↗

Static checks

Results of inspect-evals-lint at the registered commit. Rule names link to their documentation; findings link to the file at that commit. These checks describe structure and conventions, not whether the evaluation measures what it claims.

Loading lint results…