# MMLU Benchmark

**Canonical:** https://apis.io/apis/evals/mmlu-benchmark/  
**Provider:** Evals — https://apis.io/providers/evals/  
**Base URL:** https://github.com/hendrycks/test  
**Documentation:** https://github.com/hendrycks/test

MMLU Benchmark is one of 20 APIs that [Evals](https://apis.io/providers/evals/) publishes on the [APIs.io](https://apis.io/) network. Tagged areas include Benchmarks, Knowledge, Multiple Choice, Multitask, and Reference-Based. The published artifact set on APIs.io includes a GitHub repository.

MMLU (Measuring Massive Multitask Language Understanding) is a multiple-choice benchmark spanning 57 subjects from STEM and international law to nutrition and religion. It contains 15,908 multiple-choice questions (four options each), of which 1,540 are reserved for hyperparameter tuning. Per its overview, "It was one of the most commonly used benchmarks for comparing the capabilities of large language models, with over 100 million downloads as of July 2024."

## Machine-readable artifacts (3)

- **GitHubRepository** — https://github.com/hendrycks/test
- **Paper** — https://arxiv.org/abs/2009.03300
- **Dataset** — https://huggingface.co/datasets/cais/mmlu

## Other Evals APIs (12)

- [OpenAI Evals](https://apis.io/apis/evals/openai-evals/)
- [Inspect AI](https://apis.io/apis/evals/inspect-ai/)
- [Braintrust](https://apis.io/apis/evals/braintrust/)
- [LangSmith Evaluation](https://apis.io/apis/evals/langsmith-evaluation/)
- [Promptfoo](https://apis.io/apis/evals/promptfoo/)
- [Helicone](https://apis.io/apis/evals/helicone/)
- [Patronus AI](https://apis.io/apis/evals/patronus-ai/)
- [DeepEval (Confident AI)](https://apis.io/apis/evals/deepeval-confident-ai/)
- [Arize AI (Phoenix)](https://apis.io/apis/evals/arize-ai-phoenix/)
- [Galileo](https://apis.io/apis/evals/galileo/)
- [Humanloop](https://apis.io/apis/evals/humanloop/)
- [TruLens](https://apis.io/apis/evals/trulens/)

## Tags

Benchmarks, Knowledge, Multiple Choice, Multitask, Reference-Based

---

Profiled by [API Evangelist](https://apievangelist.com) and published on [APIs.io](https://apis.io/apis/evals/mmlu-benchmark/). The API's provider profile, Kin Score and agent-readiness rating are at https://apis.io/providers/evals/.
