Evals
Variants seen in the corpus:
EVALSEvalsevals
Providers using this tag (15)
Ranked by API Evangelist rating — Exemplar and Strong are expanded by default.
Strong 3 Solid coverage with minor gaps
OpenAIAPIs for accessing OpenAI's artificial intelligence models including GPT, DALL-E, Whisper, and Embeddings.SeekrSeekr Technologies builds explainable, auditable, sovereign AI for regulated industries and high-stakes gov...Prime IntellectPrime Intellect is a San Francisco–based startup building an open and decentralized stack for developing an...
Developing 7 Usable, with meaningful gaps to close
AgnoAgno (formerly Phidata) is a high-performance open-source Python framework for building multi-modal, multi-...Altimate AIAltimate AI is an agentic data engineering platform that gives data teams AI "Datamates" — trustworthy agen...tessl.ioTessl is an agent-enablement platform for spec-driven and agentic software development. It provides a regis...HeliconeHelicone is an open-source LLM observability platform and AI gateway that helps developers monitor, evaluat...Surge AISurge AI is a human-data company that provides large-scale, expert-quality labeled data for training and ev...BraintrustBraintrust (braintrust.dev) is an end-to-end platform for building, evaluating, and observing AI applicatio...BraintrustBraintrust is an enterprise-grade AI observability and evaluation platform for teams building LLM applicati...
Thin 2 Limited public surface area
Emerging 3 Early or largely undocumented
MercorMercor is an AI-powered talent and human-intelligence marketplace that organizes expert humans to power fro...TouchmarkTouchmark is solving AI pricing - instead of AI being priced per token regardless of quality, Touchmark pri...EvalsA landscape catalog of the platforms, frameworks, libraries, and benchmark suites used to evaluate large la...
APIs with this tag (16)
Ranked by the provider's API Evangelist rating — the Kin Score is scored per provider, not per API, so every API of a provider shares its band. How the rating works →
Strong 3 Solid coverage with minor gaps
Developing 8 Usable, with meaningful gaps to close
Agno Evals APIThe Evals API from Agno — 2 operation(s) for evals.Altimate AI EVALS APIThe EVALS API from Altimate AI — 18 operation(s) for evals.tessl.io Evals APIEval runs, scenarios, and solution scoring.Helicone Evals APIThe Evals API from Helicone — 4 operation(s) for evals.Surge RL Environments and AgentsSurge's product surface for delivering complex reinforcement-learning environments and agents that challeng...Surge Rubrics and VerifiersScoring rubrics and automated verifiers for grading AI outputs across domains.Braintrust Evals APIThe Evals API from Braintrust — 1 operation(s) for evals.Braintrust Evals APIThe Evals API from Braintrust — 1 operation(s) for evals.
Thin 2 Limited public surface area
Emerging 3 Early or largely undocumented
APEX Benchmarks (AI Productivity Index)Mercor's public AI productivity benchmark and research surface. APEX measures how well AI models perform re...APEX-Agents LeaderboardPublic leaderboard for AI agent performance run by Mercor's research team.APEX-SWE LeaderboardPublic leaderboard for AI software-engineering performance run by Mercor's research team.
Score breakdown
Frequency
44.6
log-scaled weighted occurrences
Breadth
0.4
spread across providers
Quality lift
50.3
mean composite of providers using it
Cohesion
2.7
strength of nearest seed neighbor
Related tags
LLM 8 co-occurrences
Models 5 co-occurrences
Team 5 co-occurrences
Observability 5 co-occurrences
Agents 9 co-occurrences
Project 5 co-occurrences
Open-Source 7 co-occurrences
User 6 co-occurrences
Where this tag comes from
Api tag17
Provider tag2
Openapi tag9
Openapi op tag63
Work with this as data
Every tag here is available over the APIs.io API and to AI agents over MCP.
MCP server
One button, every client — Claude, Cursor, VS Code and the rest.
https://apis.io/mcp
Tools for tags
7 MCP tools reach this
find_tagsBrowse and filter every tag in the catalog.get_cohortThis tag as a scored cohort — every provider carrying it, with scores.cohort_statsPRO — the distribution across this tag: mean, median, band split, adoption rates.cohort_rankingsPRO — the leaderboard, on composite AND agent-readiness axes.apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.resolveTurn a domain, URL or GitHub org into the provider it belongs to.find_cohortsEvery scored population of providers in the catalog.
Call it yourself
curl for this page
This tag
curl "https://apis.io/api/v1/tags/evals"
All tags
curl "https://apis.io/api/v1/tags?limit=25"
As a scored cohort
curl "https://apis.io/api/v1/cohorts/tag/evals"
The distribution (Pro)
curl "https://apis.io/api/v1/cohorts/tag/evals/stats" \
-H "X-API-Key: $APIS_IO_KEY"
Discovery needs no key. Ratings and market analysis are Pro.