GPQA
ReasoningGraduate-level multiple-choice questions written by domain experts in biology, physics, and chemistry. Questions are Google-proof and extremely difficult.
View model performance across different capability dimensions to help you choose the best AI model
Graduate-level multiple-choice questions written by domain experts in biology, physics, and chemistry. Questions are Google-proof and extremely difficult.
Humanity's Last Exam — extremely challenging questions designed to test the upper limits of AI capability across diverse domains.
Science coding benchmark measuring AI ability to solve scientific computing tasks.
Long Context Retrieval benchmark testing ability to find and use information in very long documents.
Instruction Following Benchmark measuring LLM ability to adhere to nuanced writing constraints and formatting requirements.
Tau2 benchmark testing multi-turn agent capabilities in airline and retail domains.
Terminal-based benchmark testing AI ability to interact with command-line interfaces and solve system tasks.
Massive Multitask Language Understanding benchmark testing knowledge across 57 diverse subjects including STEM, humanities, social sciences, and professional domains.
Real-world coding benchmark with problems from competitive programming contests, testing code generation and problem-solving abilities.
American Invitational Mathematics Examination 2025 problems testing olympiad-level mathematical reasoning.
AGIEval English — human-level reasoning tasks from standardized exams like SAT, LSAT, and civil service exams.
Big-Bench Hard — challenging subset of BIG-Bench focusing on tasks where language models previously underperformed.
Competition mathematics problems requiring multi-step reasoning, covering algebra, geometry, number theory, and calculus.
American Invitational Mathematics Examination 2024 problems testing olympiad-level mathematical reasoning.
Massive Multitask Language Understanding — tests knowledge across 57 subjects.
Multimodal Understanding benchmark testing vision-language models on expert-level tasks.
OpenAI HumanEval benchmark measuring Python code generation from function docstrings.
Software Engineering benchmark testing ability to resolve real GitHub issues.
Mostly Basic Python Problems Plus — tests Python code generation with enhanced test cases.
Accounting and audit benchmark testing financial reasoning capabilities.
Grade School Math 8K — 8,500 high quality grade school math word problems.
Simple question answering benchmark testing factual accuracy and knowledge retrieval.
AI2 Reasoning Challenge (Easy set) — grade-school science questions.
AI2 Reasoning Challenge (Challenge set) — grade-school science questions requiring complex reasoning.
Big Bench Extra Hard — even more challenging reasoning tasks pushing the limits of language model capabilities.
BIRD-CRITIC — multi-turn benchmark testing SQL generation and database interaction.
AGIEval Chinese — reasoning tasks from Chinese standardized exams (Gaokao, civil service).
Mathematics benchmark covering algebra, geometry, number theory, and calculus problems.
Medical question answering benchmark from USMLE-style questions.
Berkeley Function Calling Leaderboard v3 — testing function/tool calling accuracy.
Instruction Following Evaluation benchmark testing how well LLMs follow detailed formatting and content constraints.
Logic puzzle benchmark based on knights (truth-tellers) and knaves (liars) puzzles.
Stock market benchmark testing financial analysis capabilities.
Extended formal logic benchmark testing deductive and propositional reasoning.
GAIA — General AI Assistants benchmark testing multi-step real-world tasks.
Weapons of Mass Destruction Proxy — benchmark testing knowledge safety boundaries.
DarkBench — benchmark testing model safety and resistance to adversarial attacks.
Workshop on Machine Translation 2014 — multilingual translation quality benchmark.
MultiChallenge — general knowledge benchmark with diverse challenge types.
This page aggregates results from public benchmark platforms. Task coverage, scoring methods, and dataset versions differ between benchmarks, so scores are best compared within the same test and should be validated against your real use case.