Open LLM Leaderboard (archived)
Final scores from the retired Hugging Face Open LLM Leaderboard: open-weight language models on six standardized benchmarks.
The final leaderboard table of Hugging Face's Open LLM Leaderboard, which evaluated open-weight language models on six benchmarks (IFEval, BBH, MATH level 5, GPQA, MuSR, MMLU-Pro) with one standardized harness until its retirement in March 2025. Scores are the leaderboard's published normalized values on a 0-100 scale (higher is better); the average is the mean of the six. Community-flagged submissions are excluded, mirroring the leaderboard's own default view.
Results
average per subject over 4,575 measurements, ranked by average.
| Subject | Avg | Median | Min | Max | P95 | P99 | Count |
|---|---|---|---|---|---|---|---|
| maziyarpanahi-calme-3-2-instruct-78b-bfloat16 | 52.08 | 52.08 | 52.08 | 52.08 | 52.08 | 52.08 | 1 |
| maziyarpanahi-calme-3-1-instruct-78b-bfloat16 | 51.29 | 51.29 | 51.29 | 51.29 | 51.29 | 51.29 | 1 |
| dfurman-calmerys-78b-orpo-v0-1-bfloat16 | 51.23 | 51.23 | 51.23 | 51.23 | 51.23 | 51.23 | 1 |
| maziyarpanahi-calme-2-4-rys-78b-bfloat16 | 50.77 | 50.77 | 50.77 | 50.77 | 50.77 | 50.77 | 1 |
| huihui-ai-qwen2-5-72b-instruct-abliterated-bfloat16 | 48.11 | 48.11 | 48.11 | 48.11 | 48.11 | 48.11 | 1 |
| qwen-qwen2-5-72b-instruct-bfloat16 | 47.98 | 47.98 | 47.98 | 47.98 | 47.98 | 47.98 | 1 |
| maziyarpanahi-calme-2-1-qwen2-5-72b-bfloat16 | 47.86 | 47.86 | 47.86 | 47.86 | 47.86 | 47.86 | 1 |
| newsbang-homer-v1-0-qwen2-5-72b-bfloat16 | 47.46 | 47.46 | 47.46 | 47.46 | 47.46 | 47.46 | 1 |
| ehristoforu-qwen2-5-test-32b-it-bfloat16 | 47.37 | 47.37 | 47.37 | 47.37 | 47.37 | 47.37 | 1 |
| saxo-linkbricks-horizon-ai-avengers-v1-32b-bfloat16 | 47.34 | 47.34 | 47.34 | 47.34 | 47.34 | 47.34 | 1 |
| maziyarpanahi-calme-2-2-qwen2-5-72b-bfloat16 | 47.22 | 47.22 | 47.22 | 47.22 | 47.22 | 47.22 | 1 |
| fluently-lm-fluentlylm-prinum-bfloat16 | 47.22 | 47.22 | 47.22 | 47.22 | 47.22 | 47.22 | 1 |
| jungzoona-t3q-qwen2-5-14b-instruct-1m-e3-bfloat16 | 47.09 | 47.09 | 47.09 | 47.09 | 47.09 | 47.09 | 1 |
| jungzoona-t3q-qwen2-5-14b-v1-0-e3-bfloat16 | 47.09 | 47.09 | 47.09 | 47.09 | 47.09 | 47.09 | 1 |
| zetasepic-qwen2-5-32b-instruct-abliterated-v2-bfloat16 | 46.89 | 46.89 | 46.89 | 46.89 | 46.89 | 46.89 | 1 |
| rubenroy-gilgamesh-72b-float16 | 46.79 | 46.79 | 46.79 | 46.79 | 46.79 | 46.79 | 1 |
| sakalti-ultiima-72b-float16 | 46.77 | 46.77 | 46.77 | 46.77 | 46.77 | 46.77 | 1 |
| combinhorizon-zetasepic-abliteratedv2-qwen2-5-32b-inst-basemerge-ties-bfloat16 | 46.76 | 46.76 | 46.76 | 46.76 | 46.76 | 46.76 | 1 |
| maldv-awqward2-5-32b-instruct-bfloat16 | 46.75 | 46.75 | 46.75 | 46.75 | 46.75 | 46.75 | 1 |
| raphgg-test-2-5-72b-bfloat16 | 46.74 | 46.74 | 46.74 | 46.74 | 46.74 | 46.74 | 1 |
| shuttleai-shuttle-3-float16 | 46.7 | 46.7 | 46.7 | 46.7 | 46.7 | 46.7 | 1 |
| qwen-qwen2-5-32b-instruct-bfloat16 | 46.6 | 46.6 | 46.6 | 46.6 | 46.6 | 46.6 | 1 |
| mistralai-mistral-large-instruct-2411-float16 | 46.52 | 46.52 | 46.52 | 46.52 | 46.52 | 46.52 | 1 |
| rombodawg-rombos-llm-v2-5-qwen-72b-bfloat16 | 46.5 | 46.5 | 46.5 | 46.5 | 46.5 | 46.5 | 1 |
| saxo-linkbricks-horizon-ai-avengers-v3-32b-bfloat16 | 46.37 | 46.37 | 46.37 | 46.37 | 46.37 | 46.37 | 1 |
…and 4,550 more subjects.
Metrics
- average — recorded on each measurement. Mean of the six normalized benchmark scores, on a 0–100 scale. The leaderboard's headline ranking number; higher is better
- ifeval — recorded on each measurement. IFEval: how reliably the model follows verifiable formatting instructions. Normalized to 0–100; higher is better
- bbh — recorded on each measurement. BBH (Big-Bench Hard): a suite of challenging reasoning tasks. Normalized to 0–100 above the random baseline; higher is better
- math_lvl_5 — recorded on each measurement. MATH level 5: the hardest tier of competition mathematics problems. Normalized to 0–100; higher is better
- gpqa — recorded on each measurement. GPQA: graduate-level science questions written to resist lookup. Normalized to 0–100 above the random baseline; higher is better
- musr — recorded on each measurement. MuSR: multistep soft reasoning over long narrative problems. Normalized to 0–100 above the random baseline; higher is better
- mmlu_pro — recorded on each measurement. MMLU-Pro: a harder, ten-choice revision of the MMLU knowledge benchmark. Normalized to 0–100 above the random baseline; higher is better
Subjects (4575)
- 0-hero/Matter-0.2-7B-DPO
- 01-ai/Yi-1.5-34B-32K
- 01-ai/Yi-1.5-34B
- 01-ai/Yi-1.5-34B-Chat-16K
- 01-ai/Yi-1.5-34B-Chat
- 01-ai/Yi-1.5-6B
- 01-ai/Yi-1.5-6B-Chat
- 01-ai/Yi-1.5-9B-32K
- 01-ai/Yi-1.5-9B
- 01-ai/Yi-1.5-9B-Chat-16K
- 01-ai/Yi-1.5-9B-Chat
- 01-ai/Yi-34B-200K
- 01-ai/Yi-34B
- 01-ai/Yi-34B-Chat
- 01-ai/Yi-6B-200K
- 01-ai/Yi-6B
- 01-ai/Yi-6B-Chat
- 01-ai/Yi-9B-200K
- 01-ai/Yi-9B
- 01-ai/Yi-Coder-9B-Chat
- 1-800-LLMs/Qwen-2.5-14B-Hindi-Custom-Instruct
- 1-800-LLMs/Qwen-2.5-14B-Hindi
- 1024m/PHI-4-Hindi
- 1024m/QWEN-14B-B100
- 152334H/miqu-1-70b-sf
- 1TuanPham/T-VisStar-7B-v0.1
- 1TuanPham/T-VisStar-v0.1
- 3rd-Degree-Burn/L-3.1-Science-Writer-8B
- 3rd-Degree-Burn/Llama-3.1-8B-Squareroot
- 3rd-Degree-Burn/Llama-3.1-8B-Squareroot-v1
- 3rd-Degree-Burn/Llama-Squared-8B
- 4season/final_model_test_v2
- aaditya/Llama3-OpenBioLLM-70B
- AALF/FuseChat-Llama-3.1-8B-Instruct-preview
- AALF/FuseChat-Llama-3.1-8B-SFT-preview
- AALF/gemma-2-27b-it-SimPO-37K-100steps
- AALF/gemma-2-27b-it-SimPO-37K
- Aashraf995/Creative-7B-nerd
- Aashraf995/Gemma-Evo-10B
- Aashraf995/Qwen-Evo-7B
- Aashraf995/QwenStock-14B
- abacusai/bigstral-12b-32k
- abacusai/bigyi-15b
- abacusai/Dracarys-72B-Instruct
- abacusai/Liberated-Qwen1.5-14B
- abacusai/Llama-3-Smaug-8B
- abacusai/Smaug-34B-v0.1
- abacusai/Smaug-72B-v0.1
- abacusai/Smaug-Llama-3-70B-Instruct-32K
- abacusai/Smaug-Mixtral-v0.1
…and 4525 more.
Published by Hugging Face Open LLM Leaderboard.