# Inspect Evals > A collection of evaluation tasks for the Inspect AI framework ## Pages - [task-configurability](https://ukgovernmentbeis.github.io/inspect_evals/task-configurability.llms.md) - [methodology](https://ukgovernmentbeis.github.io/inspect_evals/methodology.llms.md) - [ORIGINAL_IMPLEMENTATION_ISSUES](https://ukgovernmentbeis.github.io/inspect_evals/evals/theagentcompany/ORIGINAL_IMPLEMENTATION_ISSUES.llms.md) - [EVALUATION](https://ukgovernmentbeis.github.io/inspect_evals/evals/tac/EVALUATION.llms.md) - [report](https://ukgovernmentbeis.github.io/inspect_evals/evals/persistbench/report.llms.md) - [appendix](https://ukgovernmentbeis.github.io/inspect_evals/evals/mask/appendix.llms.md) - [documentation](https://ukgovernmentbeis.github.io/inspect_evals/documentation.llms.md) - [Inspect Evals](https://ukgovernmentbeis.github.io/inspect_evals/index.llms.md) - [Comparing Evaluation Results Over Time](https://ukgovernmentbeis.github.io/inspect_evals/guides/comparisons-across-time.llms.md) - [XSTest: A benchmark for identifying exaggerated safety behaviours in LLM's](https://ukgovernmentbeis.github.io/inspect_evals/evals/xstest/index.llms.md) - [WorldSense: Grounded Reasoning Benchmark](https://ukgovernmentbeis.github.io/inspect_evals/evals/worldsense/index.llms.md) - [WINOGRANDE: An Adversarial Winograd Schema Challenge at Scale](https://ukgovernmentbeis.github.io/inspect_evals/evals/winogrande/index.llms.md) - [V*Bench: A Visual QA Benchmark with Detailed High-resolution Images](https://ukgovernmentbeis.github.io/inspect_evals/evals/vstar_bench/index.llms.md) - [VimGolf: Evaluating LLMs in Vim Editing Proficiency](https://ukgovernmentbeis.github.io/inspect_evals/evals/vimgolf_challenges/index.llms.md) - [Uganda Cultural and Cognitive Benchmark (UCCB)](https://ukgovernmentbeis.github.io/inspect_evals/evals/uccb/index.llms.md) - [Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities](https://ukgovernmentbeis.github.io/inspect_evals/evals/threecb/index.llms.md) - [Tau2](https://ukgovernmentbeis.github.io/inspect_evals/evals/tau2/index.llms.md) - [TAC: Animal Welfare Awareness in AI Ticket Agents](https://ukgovernmentbeis.github.io/inspect_evals/evals/tac/index.llms.md) - [SycoBench-600: Measuring Sycophancy and Correction Selectivity in LLM Assistants](https://ukgovernmentbeis.github.io/inspect_evals/evals/sycobench-600/index.llms.md) - [SWE-bench Verified: Resolving Real-World GitHub Issues](https://ukgovernmentbeis.github.io/inspect_evals/evals/swe_bench/index.llms.md) - [StereoSet: Measuring stereotypical bias in pretrained language models](https://ukgovernmentbeis.github.io/inspect_evals/evals/stereoset/index.llms.md) - [SOS BENCH: Benchmarking Safety Alignment on Scientific Knowledge](https://ukgovernmentbeis.github.io/inspect_evals/evals/sosbench/index.llms.md) - [SEvenLLM: A benchmark to elicit, and improve cybersecurity incident analysis and response abilities in LLMs for Security Events.](https://ukgovernmentbeis.github.io/inspect_evals/evals/sevenllm/index.llms.md) - [SciKnowEval: Evaluating Multi-level Scientific Knowledge of Large Language Models](https://ukgovernmentbeis.github.io/inspect_evals/evals/sciknoweval/index.llms.md) - [scBench: A Benchmark for Single-Cell RNA-seq Analysis](https://ukgovernmentbeis.github.io/inspect_evals/evals/scbench/index.llms.md) - [RACE-H: A benchmark for testing reading comprehension and reasoning abilities of neural models](https://ukgovernmentbeis.github.io/inspect_evals/evals/race_h/index.llms.md) - [Pre-Flight: Aviation Operations Knowledge Evaluation](https://ukgovernmentbeis.github.io/inspect_evals/evals/pre_flight/index.llms.md) - [PinchBench: Skill Composition for Coding Agents](https://ukgovernmentbeis.github.io/inspect_evals/evals/pinchbench/index.llms.md) - [Personality](https://ukgovernmentbeis.github.io/inspect_evals/evals/personality/index.llms.md) - [PAWS: Paraphrase Adversaries from Word Scrambling](https://ukgovernmentbeis.github.io/inspect_evals/evals/paws/index.llms.md) - [OSWorld: Multimodal Computer Interaction Tasks](https://ukgovernmentbeis.github.io/inspect_evals/evals/osworld/index.llms.md) - [OpenBookQA](https://ukgovernmentbeis.github.io/inspect_evals/evals/openbookqa/index.llms.md) - [NoveltyBench: Evaluating Language Models for Humanlike Diversity](https://ukgovernmentbeis.github.io/inspect_evals/evals/novelty_bench/index.llms.md) - [Neural Activation Reading for Collusion Benchmark (NARCBench) -- Black-Box Monitor Variant](https://ukgovernmentbeis.github.io/inspect_evals/evals/narcbench/index.llms.md) - [MORU: Moral Reasoning under Uncertainty](https://ukgovernmentbeis.github.io/inspect_evals/evals/moru/index.llms.md) - [MMLU-Pro: Advanced Multitask Knowledge and Reasoning Evaluation](https://ukgovernmentbeis.github.io/inspect_evals/evals/mmlu_pro/index.llms.md) - [MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models](https://ukgovernmentbeis.github.io/inspect_evals/evals/mmiu/index.llms.md) - [MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering](https://ukgovernmentbeis.github.io/inspect_evals/evals/mle_bench/index.llms.md) - [Mind2Web: Towards a Generalist Agent for the Web](https://ukgovernmentbeis.github.io/inspect_evals/evals/mind2web/index.llms.md) - [MedQA: Medical exam Q&A benchmark](https://ukgovernmentbeis.github.io/inspect_evals/evals/medqa/index.llms.md) - [MCPTox: Tool Poisoning Attacks on Real-World MCP Servers](https://ukgovernmentbeis.github.io/inspect_evals/evals/mcptox/index.llms.md) - [MathVista: Visual Math Problem-Solving](https://ukgovernmentbeis.github.io/inspect_evals/evals/mathvista/index.llms.md) - [MASK: Disentangling Honesty from Accuracy in AI Systems](https://ukgovernmentbeis.github.io/inspect_evals/evals/mask/index.llms.md) - [MakeMeSay](https://ukgovernmentbeis.github.io/inspect_evals/evals/makemesay/index.llms.md) - [MACHIAVELLI](https://ukgovernmentbeis.github.io/inspect_evals/evals/machiavelli/index.llms.md) - [LiveCodeBench-Pro: Competitive Programming Benchmark](https://ukgovernmentbeis.github.io/inspect_evals/evals/livecodebench_pro/index.llms.md) - [LingOly](https://ukgovernmentbeis.github.io/inspect_evals/evals/lingoly/index.llms.md) - [LAB-Bench: Measuring Capabilities of Language Models for Biology Research](https://ukgovernmentbeis.github.io/inspect_evals/evals/lab_bench/index.llms.md) - [CodeIPI: Indirect Prompt Injection for Coding Agents](https://ukgovernmentbeis.github.io/inspect_evals/evals/ipi_coding_agent/index.llms.md) - [∞Bench: Extending Long Context Evaluation Beyond 100K Tokens](https://ukgovernmentbeis.github.io/inspect_evals/evals/infinite_bench/index.llms.md) - [IFEval: Instruction-Following Evaluation](https://ukgovernmentbeis.github.io/inspect_evals/evals/ifeval/index.llms.md) - [Humanity's Last Exam](https://ukgovernmentbeis.github.io/inspect_evals/evals/hle/index.llms.md) - [HealthBench: Evaluating Large Language Models Towards Improved Human Health](https://ukgovernmentbeis.github.io/inspect_evals/evals/healthbench/index.llms.md) - [GSM8K: Grade School Math Word Problems](https://ukgovernmentbeis.github.io/inspect_evals/evals/gsm8k/index.llms.md) - [GDPval](https://ukgovernmentbeis.github.io/inspect_evals/evals/gdpval/index.llms.md) - [GDM Dangerous Capabilities: Self-reasoning](https://ukgovernmentbeis.github.io/inspect_evals/evals/gdm_self_reasoning/index.llms.md) - [InterCode: Security and Coding Capture-the-Flag Challenges](https://ukgovernmentbeis.github.io/inspect_evals/evals/gdm_intercode_ctf/index.llms.md) - [GAIA: A Benchmark for General AI Assistants](https://ukgovernmentbeis.github.io/inspect_evals/evals/gaia/index.llms.md) - [Frontier-CS: Benchmarking LLMs on Computer Science Problems](https://ukgovernmentbeis.github.io/inspect_evals/evals/frontier_cs/index.llms.md) - [FORTRESS](https://ukgovernmentbeis.github.io/inspect_evals/evals/fortress/index.llms.md) - [DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation](https://ukgovernmentbeis.github.io/inspect_evals/evals/ds1000/index.llms.md) - [DocVQA: A Dataset for VQA on Document Images](https://ukgovernmentbeis.github.io/inspect_evals/evals/docvqa/index.llms.md) - [Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs](https://ukgovernmentbeis.github.io/inspect_evals/evals/do_not_answer/index.llms.md) - [CyberSecEval 4: Advanced Cybersecurity Evaluation Benchmarks](https://ukgovernmentbeis.github.io/inspect_evals/evals/cyberseceval_4/index.llms.md) - [CyberSecEval_2: Cybersecurity Risk and Vulnerability Evaluation](https://ukgovernmentbeis.github.io/inspect_evals/evals/cyberseceval_2/index.llms.md) - [CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale](https://ukgovernmentbeis.github.io/inspect_evals/evals/cybergym/index.llms.md) - [CVEBench: Benchmark for AI Agents Ability to Exploit Real-World Web Application Vulnerabilities](https://ukgovernmentbeis.github.io/inspect_evals/evals/cve_bench/index.llms.md) - [CORE-Bench](https://ukgovernmentbeis.github.io/inspect_evals/evals/core_bench/index.llms.md) - [ComputeEval: CUDA Code Generation Benchmark](https://ukgovernmentbeis.github.io/inspect_evals/evals/compute_eval/index.llms.md) - [The Art of Saying No: Contextual Noncompliance in Language Models](https://ukgovernmentbeis.github.io/inspect_evals/evals/coconot/index.llms.md) - [ChipBench (Verilog Generation subset)](https://ukgovernmentbeis.github.io/inspect_evals/evals/chipbench_verilog_gen/index.llms.md) - [ChipBench (Verilog Debugging subset)](https://ukgovernmentbeis.github.io/inspect_evals/evals/chipbench_debug/index.llms.md) - [BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents](https://ukgovernmentbeis.github.io/inspect_evals/evals/browse_comp/index.llms.md) - [BOLD: Bias in Open-ended Language Generation Dataset](https://ukgovernmentbeis.github.io/inspect_evals/evals/bold/index.llms.md) - [BFCL: Berkeley Function-Calling Leaderboard](https://ukgovernmentbeis.github.io/inspect_evals/evals/bfcl/index.llms.md) - [BBH: Challenging BIG-Bench Tasks](https://ukgovernmentbeis.github.io/inspect_evals/evals/bbh/index.llms.md) - [b3: Backbone Breaker Benchmark](https://ukgovernmentbeis.github.io/inspect_evals/evals/b3/index.llms.md) - [ArxivRollBench](https://ukgovernmentbeis.github.io/inspect_evals/evals/arxivrollbench/index.llms.md) - [AraTrust: An Evaluation of Trustworthiness for LLMs in Arabic](https://ukgovernmentbeis.github.io/inspect_evals/evals/aratrust/index.llms.md) - [APPS: Automated Programming Progress Standard](https://ukgovernmentbeis.github.io/inspect_evals/evals/apps/index.llms.md) - [ANIMA: Animal Norms In Moral Assessment](https://ukgovernmentbeis.github.io/inspect_evals/evals/anima/index.llms.md) - [AIR Bench: AI Risk Benchmark](https://ukgovernmentbeis.github.io/inspect_evals/evals/air_bench/index.llms.md) - [AIME 2025: Problems from the American Invitational Mathematics Examination](https://ukgovernmentbeis.github.io/inspect_evals/evals/aime2025/index.llms.md) - [AHB: Adversarial Humanities Benchmark](https://ukgovernmentbeis.github.io/inspect_evals/evals/ahb/index.llms.md) - [Agentic Misalignment: How LLMs could be insider threats](https://ukgovernmentbeis.github.io/inspect_evals/evals/agentic_misalignment/index.llms.md) - [AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents](https://ukgovernmentbeis.github.io/inspect_evals/evals/agentdojo/index.llms.md) - [AgentBench: Evaluate LLMs as Agents](https://ukgovernmentbeis.github.io/inspect_evals/evals/agent_bench/index.llms.md) - [About the Register](https://ukgovernmentbeis.github.io/inspect_evals/eval-register/index.llms.md) - [index](https://ukgovernmentbeis.github.io/inspect_evals/contributing/index.llms.md) - [AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions](https://ukgovernmentbeis.github.io/inspect_evals/evals/abstention_bench/index.llms.md) - [AgentThreatBench: Evaluating LLM Agent Resilience to OWASP Agentic Threats](https://ukgovernmentbeis.github.io/inspect_evals/evals/agent_threat_bench/index.llms.md) - [AgentHarm: Harmfulness Potential in AI Agents](https://ukgovernmentbeis.github.io/inspect_evals/evals/agentharm/index.llms.md) - [AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models](https://ukgovernmentbeis.github.io/inspect_evals/evals/agieval/index.llms.md) - [AIME 2024: Problems from the American Invitational Mathematics Examination](https://ukgovernmentbeis.github.io/inspect_evals/evals/aime2024/index.llms.md) - [AIME 2026: Problems from the American Invitational Mathematics Examination](https://ukgovernmentbeis.github.io/inspect_evals/evals/aime2026/index.llms.md) - [Alignment Faking in Large Language Models](https://ukgovernmentbeis.github.io/inspect_evals/evals/alignment_faking/index.llms.md) - [APE: Attempt to Persuade Eval](https://ukgovernmentbeis.github.io/inspect_evals/evals/ape/index.llms.md) - [AppWorld Benchmark (Interactive Coding Agents over App APIs)](https://ukgovernmentbeis.github.io/inspect_evals/evals/appworld/index.llms.md) - [ARC: AI2 Reasoning Challenge](https://ukgovernmentbeis.github.io/inspect_evals/evals/arc/index.llms.md) - [AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?](https://ukgovernmentbeis.github.io/inspect_evals/evals/assistant_bench/index.llms.md) - [BIG-Bench Extra Hard](https://ukgovernmentbeis.github.io/inspect_evals/evals/bbeh/index.llms.md) - [BBQ: Bias Benchmark for Question Answering](https://ukgovernmentbeis.github.io/inspect_evals/evals/bbq/index.llms.md) - [BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions](https://ukgovernmentbeis.github.io/inspect_evals/evals/bigcodebench/index.llms.md) - [BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions](https://ukgovernmentbeis.github.io/inspect_evals/evals/boolq/index.llms.md) - [ChemBench: Are large language models superhuman chemists?](https://ukgovernmentbeis.github.io/inspect_evals/evals/chembench/index.llms.md) - [ChipBench (Reference Model Generation variant)](https://ukgovernmentbeis.github.io/inspect_evals/evals/chipbench_refmodel/index.llms.md) - [ClassEval: A Manually-Crafted Benchmark for Evaluating LLMs on Class-level Code Generation](https://ukgovernmentbeis.github.io/inspect_evals/evals/class_eval/index.llms.md) - [CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge](https://ukgovernmentbeis.github.io/inspect_evals/evals/commonsense_qa/index.llms.md) - [ContractBench: LLM Agent Observation Contract Compliance Benchmark](https://ukgovernmentbeis.github.io/inspect_evals/evals/contractbench/index.llms.md) - [CTI-REALM: Cyber Threat Intelligence Detection Rule Development Benchmark](https://ukgovernmentbeis.github.io/inspect_evals/evals/cti_realm/index.llms.md) - [Cybench: Capture-The-Flag Cybersecurity Challenges](https://ukgovernmentbeis.github.io/inspect_evals/evals/cybench/index.llms.md) - [CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge](https://ukgovernmentbeis.github.io/inspect_evals/evals/cybermetric/index.llms.md) - [CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models](https://ukgovernmentbeis.github.io/inspect_evals/evals/cyberseceval_3/index.llms.md) - [DeceptionBench: Belief-vs-Behaviour Deception under Scenario Pressure](https://ukgovernmentbeis.github.io/inspect_evals/evals/deceptionbench/index.llms.md) - [Do-Not-Answer Adversarial (jailbreak variant of Do-Not-Answer)](https://ukgovernmentbeis.github.io/inspect_evals/evals/do_not_answer_adversarial/index.llms.md) - [DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs](https://ukgovernmentbeis.github.io/inspect_evals/evals/drop/index.llms.md) - [ExploitBench: Capability Ladder Benchmark for LLM Cybersecurity Agents (V8 exploit development)](https://ukgovernmentbeis.github.io/inspect_evals/evals/exploitbench/index.llms.md) - [FRAMES: Factual Retrieval and Multi-hop Evaluation Set](https://ukgovernmentbeis.github.io/inspect_evals/evals/frames/index.llms.md) - [FrontierScience: Expert-Level Scientific Reasoning](https://ukgovernmentbeis.github.io/inspect_evals/evals/frontierscience/index.llms.md) - [GDM Dangerous Capabilities: Capture the Flag](https://ukgovernmentbeis.github.io/inspect_evals/evals/gdm_in_house_ctf/index.llms.md) - [GDM Dangerous Capabilities: Self-proliferation](https://ukgovernmentbeis.github.io/inspect_evals/evals/gdm_self_proliferation/index.llms.md) - [GDM Dangerous Capabilities: Stealth](https://ukgovernmentbeis.github.io/inspect_evals/evals/gdm_stealth/index.llms.md) - [GPQA: Graduate-Level STEM Knowledge Challenge](https://ukgovernmentbeis.github.io/inspect_evals/evals/gpqa/index.llms.md) - [Hangman Bench](https://ukgovernmentbeis.github.io/inspect_evals/evals/hangman-bench/index.llms.md) - [HellaSwag: Commonsense Event Continuation](https://ukgovernmentbeis.github.io/inspect_evals/evals/hellaswag/index.llms.md) - [HumanEval: Python Function Generation from Instructions](https://ukgovernmentbeis.github.io/inspect_evals/evals/humaneval/index.llms.md) - [IFEvalCode: Controlled Code Generation](https://ukgovernmentbeis.github.io/inspect_evals/evals/ifevalcode/index.llms.md) - [InstrumentalEval - Evaluating the Paperclip Maximizer: Are RL-Based Language Models More Likely to Pursue Instrumental Goals?](https://ukgovernmentbeis.github.io/inspect_evals/evals/instrumentaleval/index.llms.md) - [KernelBench: Can LLMs Write Efficient GPU Kernels?](https://ukgovernmentbeis.github.io/inspect_evals/evals/kernelbench/index.llms.md) - [LABBench2: An Improved Benchmark for AI Systems Performing Biology Research](https://ukgovernmentbeis.github.io/inspect_evals/evals/lab_bench_2/index.llms.md) - [LiveBench: A Challenging, Contamination-Free LLM Benchmark](https://ukgovernmentbeis.github.io/inspect_evals/evals/livebench/index.llms.md) - [MaCBench: Probing the limitations of multimodal language models for chemistry and materials research](https://ukgovernmentbeis.github.io/inspect_evals/evals/macbench/index.llms.md) - [Make Me Pay](https://ukgovernmentbeis.github.io/inspect_evals/evals/make_me_pay/index.llms.md) - [MANTA: Multi-Turn Adversarial Benchmark for Animal Welfare Reasoning](https://ukgovernmentbeis.github.io/inspect_evals/evals/manta/index.llms.md) - [MATH: Measuring Mathematical Problem Solving](https://ukgovernmentbeis.github.io/inspect_evals/evals/math/index.llms.md) - [MBPP: Basic Python Coding Challenges](https://ukgovernmentbeis.github.io/inspect_evals/evals/mbpp/index.llms.md) - [MedCalc-Bench](https://ukgovernmentbeis.github.io/inspect_evals/evals/medcalc-bench/index.llms.md) - [MGSM: Multilingual Grade School Math](https://ukgovernmentbeis.github.io/inspect_evals/evals/mgsm/index.llms.md) - [Mind2Web-SC](https://ukgovernmentbeis.github.io/inspect_evals/evals/mind2web_sc/index.llms.md) - [MLRC-Bench: Can Language Agents Solve Machine Learning Research Challenges?](https://ukgovernmentbeis.github.io/inspect_evals/evals/mlrc_bench/index.llms.md) - [MMLU: Measuring Massive Multitask Language Understanding](https://ukgovernmentbeis.github.io/inspect_evals/evals/mmlu/index.llms.md) - [MMMU: Multimodal College-Level Understanding and Reasoning](https://ukgovernmentbeis.github.io/inspect_evals/evals/mmmu/index.llms.md) - [MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning](https://ukgovernmentbeis.github.io/inspect_evals/evals/musr/index.llms.md) - [Needle in a Haystack (NIAH): In-Context Retrieval Benchmark for Long Context LLMs](https://ukgovernmentbeis.github.io/inspect_evals/evals/niah/index.llms.md) - [O-NET: A high-school student knowledge test](https://ukgovernmentbeis.github.io/inspect_evals/evals/onet/index.llms.md) - [OR-Bench Hard-1K (subset of OR-Bench: Over-Refusal Benchmark)](https://ukgovernmentbeis.github.io/inspect_evals/evals/or_bench/index.llms.md) - [PaperBench: Evaluating AI's Ability to Replicate AI Research (Work In Progress)](https://ukgovernmentbeis.github.io/inspect_evals/evals/paperbench/index.llms.md) - [PersistBench: When Should Long-Term Memories Be Forgotten by LLMs?](https://ukgovernmentbeis.github.io/inspect_evals/evals/persistbench/index.llms.md) - [PerspectiveGap Role-Fragment Assignment](https://ukgovernmentbeis.github.io/inspect_evals/evals/perspective_gap/index.llms.md) - [PIQA: Physical Commonsense Reasoning Test](https://ukgovernmentbeis.github.io/inspect_evals/evals/piqa/index.llms.md) - [PubMedQA: A Dataset for Biomedical Research Question Answering](https://ukgovernmentbeis.github.io/inspect_evals/evals/pubmedqa/index.llms.md) - [SAD: Situational Awareness Dataset](https://ukgovernmentbeis.github.io/inspect_evals/evals/sad/index.llms.md) - [SciCode: A Research Coding Benchmark Curated by Scientists](https://ukgovernmentbeis.github.io/inspect_evals/evals/scicode/index.llms.md) - [SecQA: A Concise Question-Answering Dataset for Evaluating Large Language Models in Computer Security](https://ukgovernmentbeis.github.io/inspect_evals/evals/sec_qa/index.llms.md) - [SimpleQA/SimpleQA Verified: Measuring short-form factuality in large language models](https://ukgovernmentbeis.github.io/inspect_evals/evals/simpleqa/index.llms.md) - [SQuAD: A Reading Comprehension Benchmark requiring reasoning over Wikipedia articles](https://ukgovernmentbeis.github.io/inspect_evals/evals/squad/index.llms.md) - [StrongREJECT: Measuring LLM susceptibility to jailbreak attacks](https://ukgovernmentbeis.github.io/inspect_evals/evals/strong_reject/index.llms.md) - [SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?](https://ukgovernmentbeis.github.io/inspect_evals/evals/swe_lancer/index.llms.md) - [Sycophancy Eval](https://ukgovernmentbeis.github.io/inspect_evals/evals/sycophancy/index.llms.md) - [TarantuBench: A Web Security Benchmark generated by the TarantuLabs engine](https://ukgovernmentbeis.github.io/inspect_evals/evals/tarantubench/index.llms.md) - [The Agent Company: Evaluating multi-tool autonomous agents in a synthetic company](https://ukgovernmentbeis.github.io/inspect_evals/evals/theagentcompany/index.llms.md) - [TruthfulQA: Measuring How Models Mimic Human Falsehoods](https://ukgovernmentbeis.github.io/inspect_evals/evals/truthfulqa/index.llms.md) - [USACO: USA Computing Olympiad](https://ukgovernmentbeis.github.io/inspect_evals/evals/usaco/index.llms.md) - [VQA-RAD: Visual Question Answering for Radiology](https://ukgovernmentbeis.github.io/inspect_evals/evals/vqa_rad/index.llms.md) - [WildClawBench: Productivity & Safety Evaluation](https://ukgovernmentbeis.github.io/inspect_evals/evals/wildclawbench/index.llms.md) - [WMDP: Measuring and Reducing Malicious Use With Unlearning](https://ukgovernmentbeis.github.io/inspect_evals/evals/wmdp/index.llms.md) - [WritingBench: A Comprehensive Benchmark for Generative Writing](https://ukgovernmentbeis.github.io/inspect_evals/evals/writingbench/index.llms.md) - [ZeroBench](https://ukgovernmentbeis.github.io/inspect_evals/evals/zerobench/index.llms.md) - [Running a Specific Version of an Evaluation Task](https://ukgovernmentbeis.github.io/inspect_evals/guides/run-specific-eval-version.llms.md) - [index](https://ukgovernmentbeis.github.io/inspect_evals/register/index.llms.md) - [CTI_REALM_TRANSPARENCY](https://ukgovernmentbeis.github.io/inspect_evals/evals/cti_realm/CTI_REALM_TRANSPARENCY.llms.md) - [SCORING_DESIGN](https://ukgovernmentbeis.github.io/inspect_evals/evals/paperbench/SCORING_DESIGN.llms.md) - [decision_log](https://ukgovernmentbeis.github.io/inspect_evals/evals/sad/decision_log.llms.md) - [report](https://ukgovernmentbeis.github.io/inspect_evals/evals/tau2/report.llms.md) - [TRAJECTORY_ANALYSIS](https://ukgovernmentbeis.github.io/inspect_evals/evals/theagentcompany/TRAJECTORY_ANALYSIS.llms.md) - [Privacy Notice](https://ukgovernmentbeis.github.io/inspect_evals/privacy.llms.md)