TokenAcornTokenAcorn

Subscribe to model updates

Enter your email and we'll notify you when model prices, rankings, or new releases change.

Public evaluation data

Model Performance Benchmarks

View model performance across different capability dimensions to help you choose the best AI model

REA

GPQA

Reasoning

Graduate-level multiple-choice questions written by domain experts in biology, physics, and chemistry. Questions are Google-proof and extremely difficult.

Leading models
4 Results
1007550250
Gemini 3.1 Pro Preview
94
GPT-5.6 Sol
94
GPT-5.5
94
Grok 4.5
93
REA

Humanity's Last Exam

Reasoning

Humanity's Last Exam — extremely challenging questions designed to test the upper limits of AI capability across diverse domains.

Leading models
4 Results
1007550250
Claude Fable 5
53
GPT-5.6 Sol
47
Claude Opus 4.8
46
Gemini 3.1 Pro Preview
45
DEV

SciCode

Coding

Science coding benchmark measuring AI ability to solve scientific computing tasks.

Leading models
4 Results
1007550250
Claude Fable 5
60
Gemini 3.1 Pro Preview
59
GPT-5.4
57
Gemini 3 Pro Preview
56
LONG

LCR

Long Context

Long Context Retrieval benchmark testing ability to find and use information in very long documents.

Leading models
4 Results
1007550250
GPT-5.2-Codex
76
GPT-5
76
GPT-5.1
75
GPT-5.5
74
REA

IFBench

Reasoning

Instruction Following Benchmark measuring LLM ability to adhere to nuanced writing constraints and formatting requirements.

Leading models
4 Results
1007550250
MiniMax M3
83
Nemotron 3 Ultra 550B A55B
81
Qwen3.7 Max
81
MiMo v2.5 Pro
80
ALL

Tau2

Comprehensive

Tau2 benchmark testing multi-turn agent capabilities in airline and retail domains.

Leading models
4 Results
1007550250
GLM-5.2
99
GLM-4.7-Flash
99
GLM-5 Turbo
99
GLM-5V Turbo
99
DEV

TerminalBench

Coding

Terminal-based benchmark testing AI ability to interact with command-line interfaces and solve system tasks.

Leading models
4 Results
1007550250
GPT-5.6 Sol
66
Claude Fable 5
63
GPT-5.5
61
Claude Opus 4.8
58
ALL

MMLU-Pro

Comprehensive

Massive Multitask Language Understanding benchmark testing knowledge across 57 diverse subjects including STEM, humanities, social sciences, and professional domains.

Leading models
3 Results
1007550250
Gemini 3 Pro Preview
90
Claude Opus 4.5
90
Gemini 3 Flash Preview
89
DEV

LiveCodeBench

Coding

Real-world coding benchmark with problems from competitive programming contests, testing code generation and problem-solving abilities.

Leading models
4 Results
1007550250
Gemini 3 Pro Preview
92
Gemini 3 Flash Preview
91
DeepSeek V3.2 Speciale
90
GPT-5.2
89
MATH

AIME 2025

Math

American Invitational Mathematics Examination 2025 problems testing olympiad-level mathematical reasoning.

Leading models
4 Results
1007550250
GPT-5.2 Pro
99
GPT-5 Codex
99
Gemini 3 Flash Preview
97
DeepSeek V3.2 Speciale
97
REA

AGIEval English

Reasoning

AGIEval English — human-level reasoning tasks from standardized exams like SAT, LSAT, and civil service exams.

Leading models
4 Results
1007550250
Gemini 3.1 Pro Preview
94
MiniMax M2.1
94
Gemini 3.5 Flash
93
Grok 4.5
93
REA

BBH

Reasoning

Big-Bench Hard — challenging subset of BIG-Bench focusing on tasks where language models previously underperformed.

Leading models
4 Results
1007550250
Gemini 3.1 Pro Preview
96
Qwen3.7 Max
96
Qwen3.6 Max Preview
95
Claude Sonnet 4.5
94
MATH

MATH-500

Math

Competition mathematics problems requiring multi-step reasoning, covering algebra, geometry, number theory, and calculus.

Leading models
4 Results
1007550250
GPT-5
99
Grok 3 Mini
99
o3
99
Claude Sonnet 4
99
MATH

AIME 2024

Math

American Invitational Mathematics Examination 2024 problems testing olympiad-level mathematical reasoning.

Leading models
4 Results
1007550250
GPT-5
96
Grok 4
94
o4 Mini
94
Qwen3 235B A22B Thinking 2507
94
ALL

MMLU

Comprehensive

Massive Multitask Language Understanding — tests knowledge across 57 subjects.

Leading models
4 Results
1007550250
Qwen3.7 Max
94
GPT-5
93
o3
93
Qwen3.5 397B A17B
93
ALL

MMMU

Comprehensive

Multimodal Understanding benchmark testing vision-language models on expert-level tasks.

Leading models
4 Results
1007550250
Gemini 3.1 Pro Preview
84
Gemini 3.5 Flash
83
GPT-5.5
83
Gemini 3 Flash Preview
83
DEV

HumanEval

Coding

OpenAI HumanEval benchmark measuring Python code generation from function docstrings.

Leading models
3 Results
1007550250
Claude Sonnet 4.5
98
R1
97
Claude Opus 4.5
97
DEV

SWE-bench Lite

Coding

Software Engineering benchmark testing ability to resolve real GitHub issues.

Leading models
3 Results
1007550250
Claude Opus 4.6
63
MiniMax M2.5
56
GPT-5
54
DEV

MBPP Plus

Coding

Mostly Basic Python Problems Plus — tests Python code generation with enhanced test cases.

Leading models
3 Results
1007550250
Kimi K2.6
67
Qwen3 235B A22B
66
R1
65
REA

Accounting Audit

Reasoning

Accounting and audit benchmark testing financial reasoning capabilities.

Leading models
3 Results
1007550250
Gemini 2.5 Pro Preview 05-06
87
Gemini 2.5 Flash
83
GPT-4.1
83
MATH

GSM8K

Math

Grade School Math 8K — 8,500 high quality grade school math word problems.

Leading models
3 Results
1007550250
Claude Opus 4
96
R1
96
o4 Mini High
96
DEV

SimpleQA

Coding

Simple question answering benchmark testing factual accuracy and knowledge retrieval.

Leading models
4 Results
1007550250
Gemini 2.5 Pro
53
Qwen3 235B A22B Instruct 2507
51
Qwen3 VL 235B A22B Instruct
47
GPT-4.1
40
REA

ARC Easy

Reasoning

AI2 Reasoning Challenge (Easy set) — grade-school science questions.

Leading models
2 Results
1007550250
Claude Opus 4
100
Qwen3 32B
99
REA

ARC Challenge

Reasoning

AI2 Reasoning Challenge (Challenge set) — grade-school science questions requiring complex reasoning.

Leading models
3 Results
1007550250
GPT-5
96
GLM 5
96
DeepSeek V4 Flash (Non-Reasoning)
95
REA

BBEH

Reasoning

Big Bench Extra Hard — even more challenging reasoning tasks pushing the limits of language model capabilities.

Leading models
4 Results
1007550250
DeepSeek V4 Flash (Non-Reasoning)
67
DeepSeek V4 Pro
65
GPT-5
64
GPT-5 Mini
55
DEV

BIRD-CRITIC

Coding

BIRD-CRITIC — multi-turn benchmark testing SQL generation and database interaction.

Leading models
3 Results
1007550250
Claude Opus 4.7
36
Claude Opus 4.6
34
Kimi K2.6
33
LANG

AGIEval Chinese

Multilingual

AGIEval Chinese — reasoning tasks from Chinese standardized exams (Gaokao, civil service).

Leading models
3 Results
1007550250
DeepSeek V3.2 Exp
90
Qwen3 235B A22B
89
GLM 4.6
88
MATH

Mathematics

Math

Mathematics benchmark covering algebra, geometry, number theory, and calculus problems.

Leading models
3 Results
1007550250
Claude Opus 4.6
96
o4 Mini High
95
GLM 5
94
DEV

MedQA

Coding

Medical question answering benchmark from USMLE-style questions.

Leading models
4 Results
1007550250
o4 Mini High
95
Gemini 2.5 Pro
95
R1
92
o3 Mini
91
ALL

BFCL v3

Comprehensive

Berkeley Function Calling Leaderboard v3 — testing function/tool calling accuracy.

Leading models
4 Results
1007550250
GLM 4.5
77
Claude Opus 4.7
77
Gemini 3.1 Flash Lite Preview
76
Qwen3 32B
76
REA

IFEval

Reasoning

Instruction Following Evaluation benchmark testing how well LLMs follow detailed formatting and content constraints.

Leading models
3 Results
1007550250
Kimi K2.5
93
Gemini 2.5 Pro
91
GLM 4.7
91
REA

Knights and Knaves

Reasoning

Logic puzzle benchmark based on knights (truth-tellers) and knaves (liars) puzzles.

Leading models
4 Results
1007550250
DeepSeek V4 Pro
100
o3 Mini
100
o4 Mini High
100
R1 0528
98
REA

Stock BCS

Reasoning

Stock market benchmark testing financial analysis capabilities.

Leading models
4 Results
1007550250
Qwen2.5 72B Instruct
100
GPT-4.1
92
R1
92
o3 Mini
83
REA

Formal Logic Extended

Reasoning

Extended formal logic benchmark testing deductive and propositional reasoning.

Leading models
4 Results
1007550250
o3 Mini
100
R1 0528
99
R1
98
GPT-4.1
91
ALL

GAIA

Comprehensive

GAIA — General AI Assistants benchmark testing multi-step real-world tasks.

Leading models
4 Results
1007550250
GPT-5 Mini
45
Gemini 2.5 Pro
33
R1 0528
28
Mistral Medium 3.1
23
REA

WMDP

Reasoning

Weapons of Mass Destruction Proxy — benchmark testing knowledge safety boundaries.

Leading models
3 Results
1007550250
Gemini 3 Flash Preview
87
o3 Mini
80
DeepSeek V3
72
REA

DarkBench

Reasoning

DarkBench — benchmark testing model safety and resistance to adversarial attacks.

Leading models
2 Results
1007550250
Qwen3 32B
49
Gemini 3 Flash Preview
49
LANG

WMT 2014

Multilingual

Workshop on Machine Translation 2014 — multilingual translation quality benchmark.

Leading models
4 Results
1007550250
Llama 4 Maverick
38
GPT-4.1
38
Llama 4 Scout
37
Nova Pro 1.0
37
ALL

MultiChallenge

Comprehensive

MultiChallenge — general knowledge benchmark with diverse challenge types.

Leading models
2 Results
1007550250
Mistral Medium 3.1
37
Nova Pro 1.0
19
How to read benchmark results

Data source: Public benchmark platforms

This page aggregates results from public benchmark platforms. Task coverage, scoring methods, and dataset versions differ between benchmarks, so scores are best compared within the same test and should be validated against your real use case.