OpenAI Evals

OpenAI Evals is the open-source framework released by OpenAI for evaluating large language models and LLM-based systems. The README states "Evals provide a framework for evaluating large language models (LLMs) or systems built using LLMs." The repo bundles a registry of benchmark evals, support for model-graded grading without writing custom code, private eval data via Snowflake logging, and templates for prompt chains and tool-using agents. Written primarily in Python, the project sits at roughly 18.5k stars / 3k forks.

Work with this as data

Every API here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for apis

7 MCP tools reach this
  • find_apisBrowse and filter every API in the catalog.
  • get_api_artifactsOne API's artifacts, grouped by type.
  • get_openapiThe primary OpenAPI for this API.
  • find_similar_apisAPIs that look like this one.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools

Call it yourself

curl for this page
This API
curl "https://apis.io/api/v1/apis/openai-evals"
All apis
curl "https://apis.io/api/v1/apis?limit=25"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no email required.

A second provider on the same verified email joins the account you already have.

API entry from apis.yml

apis.yml Raw ↑
name: OpenAI Evals
description: OpenAI Evals is the open-source framework released by OpenAI for evaluating large language
  models and LLM-based systems. The README states "Evals provide a framework for evaluating large language
  models (LLMs) or systems built using LLMs." The repo bundles a registry of benchmark evals, support
  for model-graded grading without writing custom code, private eval data via Snowflake logging, and templates
  for prompt chains and tool-using agents. Written primarily in Python, the project sits at roughly 18.5k
  stars / 3k forks.
humanURL: https://github.com/openai/evals
baseURL: https://github.com/openai/evals
tags:
- OpenAI
- Open-Source
- Model Graded
- Benchmark Registry
- Python
tags_raw:
- OpenAI
- Open Source
- Model Graded
- Benchmark Registry
- Python
properties:
- type: GitHubRepository
  url: https://github.com/openai/evals
- type: Documentation
  url: https://github.com/openai/evals/tree/main/docs
- type: License
  url: https://github.com/openai/evals/blob/main/LICENSE.md