Open LLM Leaderboard (archived)

Final scores from the retired Hugging Face Open LLM Leaderboard: open-weight language models on six standardized benchmarks.

The final leaderboard table of Hugging Face's Open LLM Leaderboard, which evaluated open-weight language models on six benchmarks (IFEval, BBH, MATH level 5, GPQA, MuSR, MMLU-Pro) with one standardized harness until its retirement in March 2025. Scores are the leaderboard's published normalized values on a 0-100 scale (higher is better); the average is the mean of the six. Community-flagged submissions are excluded, mirroring the leaderboard's own default view.

Results

average per subject over 4,575 measurements, ranked by average.

SubjectAvgMedianMinMaxP95P99Count
maziyarpanahi-calme-3-2-instruct-78b-bfloat1652.0852.0852.0852.0852.0852.081
maziyarpanahi-calme-3-1-instruct-78b-bfloat1651.2951.2951.2951.2951.2951.291
dfurman-calmerys-78b-orpo-v0-1-bfloat1651.2351.2351.2351.2351.2351.231
maziyarpanahi-calme-2-4-rys-78b-bfloat1650.7750.7750.7750.7750.7750.771
huihui-ai-qwen2-5-72b-instruct-abliterated-bfloat1648.1148.1148.1148.1148.1148.111
qwen-qwen2-5-72b-instruct-bfloat1647.9847.9847.9847.9847.9847.981
maziyarpanahi-calme-2-1-qwen2-5-72b-bfloat1647.8647.8647.8647.8647.8647.861
newsbang-homer-v1-0-qwen2-5-72b-bfloat1647.4647.4647.4647.4647.4647.461
ehristoforu-qwen2-5-test-32b-it-bfloat1647.3747.3747.3747.3747.3747.371
saxo-linkbricks-horizon-ai-avengers-v1-32b-bfloat1647.3447.3447.3447.3447.3447.341
maziyarpanahi-calme-2-2-qwen2-5-72b-bfloat1647.2247.2247.2247.2247.2247.221
fluently-lm-fluentlylm-prinum-bfloat1647.2247.2247.2247.2247.2247.221
jungzoona-t3q-qwen2-5-14b-instruct-1m-e3-bfloat1647.0947.0947.0947.0947.0947.091
jungzoona-t3q-qwen2-5-14b-v1-0-e3-bfloat1647.0947.0947.0947.0947.0947.091
zetasepic-qwen2-5-32b-instruct-abliterated-v2-bfloat1646.8946.8946.8946.8946.8946.891
rubenroy-gilgamesh-72b-float1646.7946.7946.7946.7946.7946.791
sakalti-ultiima-72b-float1646.7746.7746.7746.7746.7746.771
combinhorizon-zetasepic-abliteratedv2-qwen2-5-32b-inst-basemerge-ties-bfloat1646.7646.7646.7646.7646.7646.761
maldv-awqward2-5-32b-instruct-bfloat1646.7546.7546.7546.7546.7546.751
raphgg-test-2-5-72b-bfloat1646.7446.7446.7446.7446.7446.741
shuttleai-shuttle-3-float1646.746.746.746.746.746.71
qwen-qwen2-5-32b-instruct-bfloat1646.646.646.646.646.646.61
mistralai-mistral-large-instruct-2411-float1646.5246.5246.5246.5246.5246.521
rombodawg-rombos-llm-v2-5-qwen-72b-bfloat1646.546.546.546.546.546.51
saxo-linkbricks-horizon-ai-avengers-v3-32b-bfloat1646.3746.3746.3746.3746.3746.371

…and 4,550 more subjects.

Metrics

  • average — recorded on each measurement. Mean of the six normalized benchmark scores, on a 0–100 scale. The leaderboard's headline ranking number; higher is better
  • ifeval — recorded on each measurement. IFEval: how reliably the model follows verifiable formatting instructions. Normalized to 0–100; higher is better
  • bbh — recorded on each measurement. BBH (Big-Bench Hard): a suite of challenging reasoning tasks. Normalized to 0–100 above the random baseline; higher is better
  • math_lvl_5 — recorded on each measurement. MATH level 5: the hardest tier of competition mathematics problems. Normalized to 0–100; higher is better
  • gpqa — recorded on each measurement. GPQA: graduate-level science questions written to resist lookup. Normalized to 0–100 above the random baseline; higher is better
  • musr — recorded on each measurement. MuSR: multistep soft reasoning over long narrative problems. Normalized to 0–100 above the random baseline; higher is better
  • mmlu_pro — recorded on each measurement. MMLU-Pro: a harder, ten-choice revision of the MMLU knowledge benchmark. Normalized to 0–100 above the random baseline; higher is better

Subjects (4575)

  • 0-hero/Matter-0.2-7B-DPO
  • 01-ai/Yi-1.5-34B-32K
  • 01-ai/Yi-1.5-34B
  • 01-ai/Yi-1.5-34B-Chat-16K
  • 01-ai/Yi-1.5-34B-Chat
  • 01-ai/Yi-1.5-6B
  • 01-ai/Yi-1.5-6B-Chat
  • 01-ai/Yi-1.5-9B-32K
  • 01-ai/Yi-1.5-9B
  • 01-ai/Yi-1.5-9B-Chat-16K
  • 01-ai/Yi-1.5-9B-Chat
  • 01-ai/Yi-34B-200K
  • 01-ai/Yi-34B
  • 01-ai/Yi-34B-Chat
  • 01-ai/Yi-6B-200K
  • 01-ai/Yi-6B
  • 01-ai/Yi-6B-Chat
  • 01-ai/Yi-9B-200K
  • 01-ai/Yi-9B
  • 01-ai/Yi-Coder-9B-Chat
  • 1-800-LLMs/Qwen-2.5-14B-Hindi-Custom-Instruct
  • 1-800-LLMs/Qwen-2.5-14B-Hindi
  • 1024m/PHI-4-Hindi
  • 1024m/QWEN-14B-B100
  • 152334H/miqu-1-70b-sf
  • 1TuanPham/T-VisStar-7B-v0.1
  • 1TuanPham/T-VisStar-v0.1
  • 3rd-Degree-Burn/L-3.1-Science-Writer-8B
  • 3rd-Degree-Burn/Llama-3.1-8B-Squareroot
  • 3rd-Degree-Burn/Llama-3.1-8B-Squareroot-v1
  • 3rd-Degree-Burn/Llama-Squared-8B
  • 4season/final_model_test_v2
  • aaditya/Llama3-OpenBioLLM-70B
  • AALF/FuseChat-Llama-3.1-8B-Instruct-preview
  • AALF/FuseChat-Llama-3.1-8B-SFT-preview
  • AALF/gemma-2-27b-it-SimPO-37K-100steps
  • AALF/gemma-2-27b-it-SimPO-37K
  • Aashraf995/Creative-7B-nerd
  • Aashraf995/Gemma-Evo-10B
  • Aashraf995/Qwen-Evo-7B
  • Aashraf995/QwenStock-14B
  • abacusai/bigstral-12b-32k
  • abacusai/bigyi-15b
  • abacusai/Dracarys-72B-Instruct
  • abacusai/Liberated-Qwen1.5-14B
  • abacusai/Llama-3-Smaug-8B
  • abacusai/Smaug-34B-v0.1
  • abacusai/Smaug-72B-v0.1
  • abacusai/Smaug-Llama-3-70B-Instruct-32K
  • abacusai/Smaug-Mixtral-v0.1

…and 4525 more.

Published by Hugging Face Open LLM Leaderboard.