CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs

CharXiv evaluates multimodal LLMs on chart understanding using 2,323 real-world charts from arXiv papers. It contains two question types: descriptive (19 template-based questions per chart testing basic element extraction, with 25% unanswerable) and reasoning (one hand-crafted free-form question per chart requiring numerical/visual synthesis). Answers are graded by GPT-4o using binary correctness scoring. The implementation supports filtering by subset (descriptive/reasoning) and field of study.

Overview

⚠️ External evaluation. Code lives in an upstream repository. inspect_evals lists it for discoverability; review the upstream repo and pinned commit before running.

Source: meridianlabs-ai/inspect_charxiv@317e4ff

CharXiv evaluates multimodal LLMs on chart understanding using 2,323 real-world charts from arXiv papers. It contains two question types: descriptive (19 template-based questions per chart testing basic element extraction, with 25% unanswerable) and reasoning (one hand-crafted free-form question per chart requiring numerical/visual synthesis). Answers are graded by GPT-4o using binary correctness scoring. The implementation supports filtering by subset (descriptive/reasoning) and field of study.

Usage

Installation

This is an externally-maintained evaluation. Clone the upstream repository at the pinned commit and install its dependencies:

git clone https://github.com/meridianlabs-ai/inspect_charxiv
cd inspect_charxiv
git checkout 317e4ff4335d7d4242d8e3c803898b7345973bd2
uv sync

Running evaluations

CLI

uv run inspect eval src/inspect_charxiv/tasks.py@charxiv --model openai/gpt-5-nano

Python

from inspect_ai import eval
from inspect_charxiv.tasks import charxiv

eval(charxiv(), model="openai/gpt-5-nano")

View logs

Log viewing (inspect view) and default-model setup are documented in the Inspect Evals README.

More information

For the dataset, scorer, task parameters, and validation, see the upstream repo: meridianlabs-ai/inspect_charxiv.

Options

You can control a variety of options from the command line. For example:

uv run inspect eval src/inspect_charxiv/tasks.py@charxiv --limit 10 --sample-shuffle
uv run inspect eval src/inspect_charxiv/tasks.py@charxiv --max-connections 10
uv run inspect eval src/inspect_charxiv/tasks.py@charxiv --temperature 0.5

See uv run inspect eval --help for all available options.

More command-line options: Inspect docs ↗

Static checks

Results of inspect-evals-lint at the registered commit. Rule names link to their documentation; findings link to the file at that commit. These checks describe structure and conventions, not whether the evaluation measures what it claims.

Loading lint results…