HELM Capabilities

Language models scored on five scenarios by the Stanford HELM Capabilities leaderboard.

Aggregated scores from the HELM Capabilities leaderboard, published by Stanford's Center for Research on Foundation Models (crfm.stanford.edu/helm). Each model row carries the release's mean score and per-scenario scores for MMLU-Pro, GPQA, IFEval, WildBench, and Omni-MATH on a 0-1 scale (higher is better), exactly as published. HELM entered maintenance mode in June 2026, so these results are final.

Results

mean_score per subject over 68 measurements, ranked by average.

SubjectAvgMedianMinMaxP95P99Count
GPT-5 mini (2025-08-07)0.8190.8190.8190.8190.8190.8191
openai-o4-mini-2025-04-160.81170.81170.81170.81170.81170.81171
openai-o3-2025-04-160.81120.81120.81120.81120.81120.81121
GPT-5 (2025-08-07)0.80680.80680.80680.80680.80680.80681
Gemini 3 Pro (Preview)0.79930.79930.79930.79930.79930.79931
qwen-qwen3-235b-a22b-instruct-2507-fp80.7980.7980.7980.7980.7980.7981
xai-grok-4-07090.78530.78530.78530.78530.78530.78531
Claude 4 Opus (20250514, extended thinking)0.780.780.780.780.780.781
openai-gpt-oss-120b0.76960.76960.76960.76960.76960.76961
Kimi K2 Instruct0.76750.76750.76750.76750.76750.76751
Claude 4 Sonnet (20250514, extended thinking)0.76590.76590.76590.76590.76590.76591
Claude 4.5 Sonnet (20250929)0.76240.76240.76240.76240.76240.76241
Claude 4 Opus (20250514)0.75740.75740.75740.75740.75740.75741
openai-gpt-5-nano-2025-08-070.74830.74830.74830.74830.74830.74831
Gemini 2.5 Pro (03-25 preview)0.74490.74490.74490.74490.74490.74491
Claude 4 Sonnet (20250514)0.73250.73250.73250.73250.73250.73251
xai-grok-3-beta0.72720.72720.72720.72720.72720.72721
GPT-4.1 (2025-04-14)0.72670.72670.72670.72670.72670.72671
qwen-qwen3-235b-a22b-fp8-tput0.72640.72640.72640.72640.72640.72641
GPT-4.1 mini (2025-04-14)0.72610.72610.72610.72610.72610.72611
Llama 4 Maverick (17Bx128E) Instruct FP80.7180.7180.7180.7180.7180.7181
Claude 4.5 Haiku (20251001)0.71680.71680.71680.71680.71680.71681
qwen-qwen3-next-80b-a3b-thinking0.70.70.70.70.70.71
DeepSeek-R1-05280.69890.69890.69890.69890.69890.69891
writer-palmyra-x50.69650.69650.69650.69650.69650.69651

…and 43 more subjects.

Metrics

  • mean_score — recorded on each measurement. Mean of the model's five scenario scores, each normalized to a 0-1 scale. The leaderboard's headline number; higher is better
  • mmlu_pro — recorded on each measurement. MMLU-Pro: fraction of correct answers, with chain-of-thought reasoning, on graduate-level multiple-choice questions spanning 14 subject areas (0-1, higher is better)
  • gpqa — recorded on each measurement. GPQA: fraction of correct answers, with chain-of-thought reasoning, on graduate-level science questions written to be search-resistant (0-1, higher is better)
  • ifeval — recorded on each measurement. IFEval: fraction of responses satisfying every verifiable instruction in the prompt, under strict checking (0-1, higher is better)
  • wildbench — recorded on each measurement. WildBench: judged response quality on challenging real-world user queries, rescaled to 0-1; higher is better
  • omni_math — recorded on each measurement. Omni-MATH: fraction of correct answers, with chain-of-thought reasoning, on Olympiad-level mathematics problems (0-1, higher is better)

Subjects (68)

  • OLMo 2 32B Instruct March 2025
  • OLMo 2 13B Instruct November 2024
  • OLMo 2 7B Instruct November 2024
  • OLMoE 1B-7B Instruct January 2025
  • Amazon Nova Lite
  • Amazon Nova Micro
  • Amazon Nova Premier
  • Amazon Nova Pro
  • Claude 3.5 Haiku (20241022)
  • Claude 3.5 Sonnet (20241022)
  • Claude 3.7 Sonnet (20250219)
  • Claude 4.5 Haiku (20251001)
  • Claude 4 Opus (20250514)
  • Claude 4 Opus (20250514, extended thinking)
  • Claude 4 Sonnet (20250514)
  • Claude 4 Sonnet (20250514, extended thinking)
  • Claude 4.5 Sonnet (20250929)
  • DeepSeek-R1-0528
  • DeepSeek v3
  • Gemini 1.5 Flash (002)
  • Gemini 1.5 Pro (002)
  • Gemini 2.0 Flash
  • Gemini 2.0 Flash Lite (02-05 preview)
  • Gemini 2.5 Flash-Lite
  • Gemini 2.5 Flash (04-17 preview)
  • Gemini 2.5 Pro (03-25 preview)
  • Gemini 3 Pro (Preview)
  • IBM Granite 3.3 8B Instruct
  • IBM Granite 4.0 Small
  • IBM Granite 4.0 Micro
  • Marin 8B Instruct
  • Llama 3.1 Instruct Turbo (405B)
  • Llama 3.1 Instruct Turbo (70B)
  • Llama 3.1 Instruct Turbo (8B)
  • Llama 4 Maverick (17Bx128E) Instruct FP8
  • Llama 4 Scout (17Bx16E) Instruct
  • Mistral Instruct v0.3 (7B)
  • Mistral Large (2411)
  • Mistral Small 3.1 (2503)
  • Mixtral Instruct (8x22B)
  • Mixtral Instruct (8x7B)
  • Kimi K2 Instruct
  • GPT-4.1 (2025-04-14)
  • GPT-4.1 mini (2025-04-14)
  • GPT-4.1 nano (2025-04-14)
  • GPT-4o (2024-11-20)
  • GPT-4o mini (2024-07-18)
  • GPT-5.1 (2025-11-13)
  • GPT-5 (2025-08-07)
  • GPT-5 mini (2025-08-07)

…and 18 more.

Published by Stanford HELM.