# OpenAI Evals

**Canonical:** https://apis.io/apis/evals/openai-evals/  
**Provider:** Evals — https://apis.io/providers/evals/  
**Base URL:** https://github.com/openai/evals  
**Documentation:** https://github.com/openai/evals

OpenAI Evals is one of 20 APIs that [Evals](https://apis.io/providers/evals/) publishes on the [APIs.io](https://apis.io/) network. Tagged areas include OpenAI, Open-Source, Model Graded, Benchmark Registry, and Python. The published artifact set on APIs.io includes a GitHub repository and API documentation.

OpenAI Evals is the open-source framework released by OpenAI for evaluating large language models and LLM-based systems. The README states "Evals provide a framework for evaluating large language models (LLMs) or systems built using LLMs." The repo bundles a registry of benchmark evals, support for model-graded grading without writing custom code, private eval data via Snowflake logging, and templates for prompt chains and tool-using agents. Written primarily in Python, the project sits at roughly 18.5k stars / 3k forks.

## Machine-readable artifacts (4)

- **GitHubRepository** — https://github.com/openai/evals
- **Documentation** — https://github.com/openai/evals/tree/main/docs
- **License** — https://github.com/openai/evals/blob/main/LICENSE.md
- **APIsJSON** — https://raw.githubusercontent.com/api-evangelist/evals/refs/heads/main/apis.yml

## Other Evals APIs (12)

- [Inspect AI](https://apis.io/apis/evals/inspect-ai/)
- [Braintrust](https://apis.io/apis/evals/braintrust/)
- [LangSmith Evaluation](https://apis.io/apis/evals/langsmith-evaluation/)
- [Promptfoo](https://apis.io/apis/evals/promptfoo/)
- [Helicone](https://apis.io/apis/evals/helicone/)
- [Patronus AI](https://apis.io/apis/evals/patronus-ai/)
- [DeepEval (Confident AI)](https://apis.io/apis/evals/deepeval-confident-ai/)
- [Arize AI (Phoenix)](https://apis.io/apis/evals/arize-ai-phoenix/)
- [Galileo](https://apis.io/apis/evals/galileo/)
- [Humanloop](https://apis.io/apis/evals/humanloop/)
- [TruLens](https://apis.io/apis/evals/trulens/)
- [Weights and Biases Weave](https://apis.io/apis/evals/weights-and-biases-weave/)

## Tags

OpenAI, Open-Source, Model Graded, Benchmark Registry, Python

---

Profiled by [API Evangelist](https://apievangelist.com) and published on [APIs.io](https://apis.io/apis/evals/openai-evals/). The API's provider profile, Kin Score and agent-readiness rating are at https://apis.io/providers/evals/.
