CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs
CharXiv evaluates multimodal LLMs on chart understanding using 2,323 real-world charts from arXiv papers. It contains two question types: descriptive (19 template-based questions per chart testing basic element extraction, with 25% unanswerable) and reasoning (one hand-crafted free-form question per chart requiring numerical/visual synthesis). Answers are graded by GPT-4o using binary correctness scoring. The implementation supports filtering by subset (descriptive/reasoning) and field of study.
Overview
⚠️ External evaluation. Code lives in an upstream repository. inspect_evals lists it for discoverability; review the upstream repo and pinned commit before running.
Source: meridianlabs-ai/inspect_charxiv@317e4ff
CharXiv evaluates multimodal LLMs on chart understanding using 2,323 real-world charts from arXiv papers. It contains two question types: descriptive (19 template-based questions per chart testing basic element extraction, with 25% unanswerable) and reasoning (one hand-crafted free-form question per chart requiring numerical/visual synthesis). Answers are graded by GPT-4o using binary correctness scoring. The implementation supports filtering by subset (descriptive/reasoning) and field of study.
Usage
Installation
This is an externally-maintained evaluation. Clone the upstream repository at the pinned commit and install its dependencies:
git clone https://github.com/meridianlabs-ai/inspect_charxiv
cd inspect_charxiv
git checkout 317e4ff4335d7d4242d8e3c803898b7345973bd2
uv syncRunning evaluations
CLI
uv run inspect eval src/inspect_charxiv/tasks.py@charxiv --model openai/gpt-5-nanoPython
from inspect_ai import eval
from inspect_charxiv.tasks import charxiv
eval(charxiv(), model="openai/gpt-5-nano")View logs
Log viewing (inspect view) and default-model setup are documented in the Inspect Evals README.
More information
For the dataset, scorer, task parameters, and validation, see the upstream repo: meridianlabs-ai/inspect_charxiv.
Options
You can control a variety of options from the command line. For example:
uv run inspect eval src/inspect_charxiv/tasks.py@charxiv --limit 10 --sample-shuffle
uv run inspect eval src/inspect_charxiv/tasks.py@charxiv --max-connections 10
uv run inspect eval src/inspect_charxiv/tasks.py@charxiv --temperature 0.5See uv run inspect eval --help for all available options.
More command-line options: Inspect docs ↗
Static checks
Results of inspect-evals-lint at the registered commit. Rule names link to their documentation; findings link to the file at that commit. These checks describe structure and conventions, not whether the evaluation measures what it claims.
Loading lint results…