HELM Capabilities
Language models scored on five scenarios by the Stanford HELM Capabilities leaderboard.
Aggregated scores from the HELM Capabilities leaderboard, published by Stanford's Center for Research on Foundation Models (crfm.stanford.edu/helm). Each model row carries the release's mean score and per-scenario scores for MMLU-Pro, GPQA, IFEval, WildBench, and Omni-MATH on a 0-1 scale (higher is better), exactly as published. HELM entered maintenance mode in June 2026, so these results are final.
Results
mean_score per subject over 68 measurements, ranked by average.
| Subject | Avg | Median | Min | Max | P95 | P99 | Count |
|---|---|---|---|---|---|---|---|
| GPT-5 mini (2025-08-07) | 0.819 | 0.819 | 0.819 | 0.819 | 0.819 | 0.819 | 1 |
| openai-o4-mini-2025-04-16 | 0.8117 | 0.8117 | 0.8117 | 0.8117 | 0.8117 | 0.8117 | 1 |
| openai-o3-2025-04-16 | 0.8112 | 0.8112 | 0.8112 | 0.8112 | 0.8112 | 0.8112 | 1 |
| GPT-5 (2025-08-07) | 0.8068 | 0.8068 | 0.8068 | 0.8068 | 0.8068 | 0.8068 | 1 |
| Gemini 3 Pro (Preview) | 0.7993 | 0.7993 | 0.7993 | 0.7993 | 0.7993 | 0.7993 | 1 |
| qwen-qwen3-235b-a22b-instruct-2507-fp8 | 0.798 | 0.798 | 0.798 | 0.798 | 0.798 | 0.798 | 1 |
| xai-grok-4-0709 | 0.7853 | 0.7853 | 0.7853 | 0.7853 | 0.7853 | 0.7853 | 1 |
| Claude 4 Opus (20250514, extended thinking) | 0.78 | 0.78 | 0.78 | 0.78 | 0.78 | 0.78 | 1 |
| openai-gpt-oss-120b | 0.7696 | 0.7696 | 0.7696 | 0.7696 | 0.7696 | 0.7696 | 1 |
| Kimi K2 Instruct | 0.7675 | 0.7675 | 0.7675 | 0.7675 | 0.7675 | 0.7675 | 1 |
| Claude 4 Sonnet (20250514, extended thinking) | 0.7659 | 0.7659 | 0.7659 | 0.7659 | 0.7659 | 0.7659 | 1 |
| Claude 4.5 Sonnet (20250929) | 0.7624 | 0.7624 | 0.7624 | 0.7624 | 0.7624 | 0.7624 | 1 |
| Claude 4 Opus (20250514) | 0.7574 | 0.7574 | 0.7574 | 0.7574 | 0.7574 | 0.7574 | 1 |
| openai-gpt-5-nano-2025-08-07 | 0.7483 | 0.7483 | 0.7483 | 0.7483 | 0.7483 | 0.7483 | 1 |
| Gemini 2.5 Pro (03-25 preview) | 0.7449 | 0.7449 | 0.7449 | 0.7449 | 0.7449 | 0.7449 | 1 |
| Claude 4 Sonnet (20250514) | 0.7325 | 0.7325 | 0.7325 | 0.7325 | 0.7325 | 0.7325 | 1 |
| xai-grok-3-beta | 0.7272 | 0.7272 | 0.7272 | 0.7272 | 0.7272 | 0.7272 | 1 |
| GPT-4.1 (2025-04-14) | 0.7267 | 0.7267 | 0.7267 | 0.7267 | 0.7267 | 0.7267 | 1 |
| qwen-qwen3-235b-a22b-fp8-tput | 0.7264 | 0.7264 | 0.7264 | 0.7264 | 0.7264 | 0.7264 | 1 |
| GPT-4.1 mini (2025-04-14) | 0.7261 | 0.7261 | 0.7261 | 0.7261 | 0.7261 | 0.7261 | 1 |
| Llama 4 Maverick (17Bx128E) Instruct FP8 | 0.718 | 0.718 | 0.718 | 0.718 | 0.718 | 0.718 | 1 |
| Claude 4.5 Haiku (20251001) | 0.7168 | 0.7168 | 0.7168 | 0.7168 | 0.7168 | 0.7168 | 1 |
| qwen-qwen3-next-80b-a3b-thinking | 0.7 | 0.7 | 0.7 | 0.7 | 0.7 | 0.7 | 1 |
| DeepSeek-R1-0528 | 0.6989 | 0.6989 | 0.6989 | 0.6989 | 0.6989 | 0.6989 | 1 |
| writer-palmyra-x5 | 0.6965 | 0.6965 | 0.6965 | 0.6965 | 0.6965 | 0.6965 | 1 |
…and 43 more subjects.
Metrics
- mean_score — recorded on each measurement. Mean of the model's five scenario scores, each normalized to a 0-1 scale. The leaderboard's headline number; higher is better
- mmlu_pro — recorded on each measurement. MMLU-Pro: fraction of correct answers, with chain-of-thought reasoning, on graduate-level multiple-choice questions spanning 14 subject areas (0-1, higher is better)
- gpqa — recorded on each measurement. GPQA: fraction of correct answers, with chain-of-thought reasoning, on graduate-level science questions written to be search-resistant (0-1, higher is better)
- ifeval — recorded on each measurement. IFEval: fraction of responses satisfying every verifiable instruction in the prompt, under strict checking (0-1, higher is better)
- wildbench — recorded on each measurement. WildBench: judged response quality on challenging real-world user queries, rescaled to 0-1; higher is better
- omni_math — recorded on each measurement. Omni-MATH: fraction of correct answers, with chain-of-thought reasoning, on Olympiad-level mathematics problems (0-1, higher is better)
Subjects (68)
- OLMo 2 32B Instruct March 2025
- OLMo 2 13B Instruct November 2024
- OLMo 2 7B Instruct November 2024
- OLMoE 1B-7B Instruct January 2025
- Amazon Nova Lite
- Amazon Nova Micro
- Amazon Nova Premier
- Amazon Nova Pro
- Claude 3.5 Haiku (20241022)
- Claude 3.5 Sonnet (20241022)
- Claude 3.7 Sonnet (20250219)
- Claude 4.5 Haiku (20251001)
- Claude 4 Opus (20250514)
- Claude 4 Opus (20250514, extended thinking)
- Claude 4 Sonnet (20250514)
- Claude 4 Sonnet (20250514, extended thinking)
- Claude 4.5 Sonnet (20250929)
- DeepSeek-R1-0528
- DeepSeek v3
- Gemini 1.5 Flash (002)
- Gemini 1.5 Pro (002)
- Gemini 2.0 Flash
- Gemini 2.0 Flash Lite (02-05 preview)
- Gemini 2.5 Flash-Lite
- Gemini 2.5 Flash (04-17 preview)
- Gemini 2.5 Pro (03-25 preview)
- Gemini 3 Pro (Preview)
- IBM Granite 3.3 8B Instruct
- IBM Granite 4.0 Small
- IBM Granite 4.0 Micro
- Marin 8B Instruct
- Llama 3.1 Instruct Turbo (405B)
- Llama 3.1 Instruct Turbo (70B)
- Llama 3.1 Instruct Turbo (8B)
- Llama 4 Maverick (17Bx128E) Instruct FP8
- Llama 4 Scout (17Bx16E) Instruct
- Mistral Instruct v0.3 (7B)
- Mistral Large (2411)
- Mistral Small 3.1 (2503)
- Mixtral Instruct (8x22B)
- Mixtral Instruct (8x7B)
- Kimi K2 Instruct
- GPT-4.1 (2025-04-14)
- GPT-4.1 mini (2025-04-14)
- GPT-4.1 nano (2025-04-14)
- GPT-4o (2024-11-20)
- GPT-4o mini (2024-07-18)
- GPT-5.1 (2025-11-13)
- GPT-5 (2025-08-07)
- GPT-5 mini (2025-08-07)
…and 18 more.
Published by Stanford HELM.
No observations in this window.