# HumanEval Benchmark

**Canonical:** https://apis.io/apis/evals/humaneval-benchmark/  
**Provider:** Evals — https://apis.io/providers/evals/  
**Base URL:** https://github.com/openai/human-eval  
**Documentation:** https://github.com/openai/human-eval

HumanEval Benchmark is one of 20 APIs that [Evals](https://apis.io/providers/evals/) publishes on the [APIs.io](https://apis.io/) network. Tagged areas include Benchmarks, Code Generation, Functional Correctness, Pass@k, and Reference-Based. The published artifact set on APIs.io includes a GitHub repository.

HumanEval is OpenAI's evaluation harness for code-generation models, described in its README as "an evaluation harness for the HumanEval problem solving dataset described in the paper 'Evaluating Large Language Models Trained on Code'." Functional correctness is measured by executing model-generated code against unit tests, reported as pass@1, pass@10, and pass@100 by default.

## Machine-readable artifacts (3)

- **GitHubRepository** — https://github.com/openai/human-eval
- **Paper** — https://arxiv.org/abs/2107.03374
- **Dataset** — https://huggingface.co/datasets/openai/openai_humaneval

## Other Evals APIs (12)

- [OpenAI Evals](https://apis.io/apis/evals/openai-evals/)
- [Inspect AI](https://apis.io/apis/evals/inspect-ai/)
- [Braintrust](https://apis.io/apis/evals/braintrust/)
- [LangSmith Evaluation](https://apis.io/apis/evals/langsmith-evaluation/)
- [Promptfoo](https://apis.io/apis/evals/promptfoo/)
- [Helicone](https://apis.io/apis/evals/helicone/)
- [Patronus AI](https://apis.io/apis/evals/patronus-ai/)
- [DeepEval (Confident AI)](https://apis.io/apis/evals/deepeval-confident-ai/)
- [Arize AI (Phoenix)](https://apis.io/apis/evals/arize-ai-phoenix/)
- [Galileo](https://apis.io/apis/evals/galileo/)
- [Humanloop](https://apis.io/apis/evals/humanloop/)
- [TruLens](https://apis.io/apis/evals/trulens/)

## Tags

Benchmarks, Code Generation, Functional Correctness, Pass@k, Reference-Based

---

Profiled by [API Evangelist](https://apievangelist.com) and published on [APIs.io](https://apis.io/apis/evals/humaneval-benchmark/). The API's provider profile, Kin Score and agent-readiness rating are at https://apis.io/providers/evals/.
