FailSafeQA: Robustness and Context Grounding Benchmark for Long-Context Financial QA

FailSafeQA evaluates LLM robustness and context grounding on 220 long-context (4k-27k token) 10-K SEC filings. Each item is tested under seven conditions: baseline, three query perturbations (misspelled, incomplete, out-of-domain), OCR-degraded context, missing context, and irrelevant context. Scoring uses LLM-as-a-Judge (Whilst the paper used Qwen2.5-72B-Instruct, this implementation sets openai/gpt-5.6-luna as the default, with an optional secondary_judge) with a 1-6 relevance scale, binarized at 4 for compliance. Robustness is the per-item minimum compliance across perturbation conditions; Context Grounding measures refusal accuracy under missing/irrelevant context.

Overview

⚠️ External evaluation. Code lives in an upstream repository. inspect_evals lists it for discoverability; review the upstream repo and pinned commit before running.

Source: theruviparambil/failsafeqa-inspect@515c2b9

FailSafeQA evaluates LLM robustness and context grounding on 220 long-context (4k-27k token) 10-K SEC filings. Each item is tested under seven conditions: baseline, three query perturbations (misspelled, incomplete, out-of-domain), OCR-degraded context, missing context, and irrelevant context. Scoring uses LLM-as-a-Judge (Whilst the paper used Qwen2.5-72B-Instruct, this implementation sets openai/gpt-5.6-luna as the default, with an optional secondary_judge) with a 1-6 relevance scale, binarized at 4 for compliance. Robustness is the per-item minimum compliance across perturbation conditions; Context Grounding measures refusal accuracy under missing/irrelevant context.

Usage

Installation

This is an externally-maintained evaluation. Clone the upstream repository at the pinned commit and install its dependencies:

git clone https://github.com/theruviparambil/failsafeqa-inspect
cd failsafeqa-inspect
git checkout 515c2b9b0d195904b5f3f7949e40eb45a073a858
uv sync

Running evaluations

CLI

uv run inspect eval src/failsafeqa_inspect/task.py@failsafeqa --model openai/gpt-5-nano

Python

from inspect_ai import eval
from failsafeqa_inspect.task import failsafeqa

eval(failsafeqa(), model="openai/gpt-5-nano")

View logs

Log viewing (inspect view) and default-model setup are documented in the Inspect Evals README.

More information

For the dataset, scorer, task parameters, and validation, see the upstream repo: theruviparambil/failsafeqa-inspect.

Options

You can control a variety of options from the command line. For example:

uv run inspect eval src/failsafeqa_inspect/task.py@failsafeqa --limit 10 --sample-shuffle
uv run inspect eval src/failsafeqa_inspect/task.py@failsafeqa --max-connections 10
uv run inspect eval src/failsafeqa_inspect/task.py@failsafeqa --temperature 0.5

See uv run inspect eval --help for all available options.

More command-line options: Inspect docs ↗

Static checks

Results of inspect-evals-lint at the registered commit. Rule names link to their documentation; findings link to the file at that commit. These checks describe structure and conventions, not whether the evaluation measures what it claims.

Loading lint results…