emerging · 0.0 marketcapability domainresource

Evaluation

Score 23.2 / 100 48 providers +40 via APIs 59 APIs Search apis.io →
Variants seen in the corpus: EvaluationsEvaluation

Providers using this tag (48)

Ranked by API Evangelist rating — Exemplar and Strong are expanded by default.

Strong 2 Solid coverage with minor gaps
Developing 20 Usable, with meaningful gaps to close
Thin 13 Limited public surface area
Emerging 12 Early or largely undocumented
Minimal 1 Almost no public developer surface

APIs with this tag (59)

Ranked by the provider's API Evangelist rating — the Kin Score is scored per provider, not per API, so every API of a provider shares its band. How the rating works →

Exemplar 6 Complete, well-documented, and agent-ready
Strong 6 Solid coverage with minor gaps
Developing 24 Usable, with meaningful gaps to close
Zoom Quality Management APIThe Zoom Quality Management API is designed to help contact centers track and analyze customer interactions...54.0Observe.AI Evaluations APIEvaluations API can be used to pull all the Evaluations(Manual and Auto QA) done on Observe AI platform. Pl...53.8Convai Evaluation APIThe Evaluation API from Convai — 1 operation(s) for evaluation.53.2Avito Evaluation APIThe Evaluation API from Avito — 5 operation(s) for evaluation.52.3Buk Evaluations APIThe Evaluations API from Buk — 3 operation(s) for evaluations.50.9Microsoft Windows 10 Evaluation APIThe Evaluation API from Microsoft Windows 10 — 1 operation(s) for evaluation.48.6Braintrust APIThe Braintrust REST API provides programmatic access to projects, experiments, datasets, prompts, functions...48.2Trakstar Evaluations APIManage candidate evaluations46.3Rapidata Evaluation APIThe Evaluation API from Rapidata — 1 operation(s) for evaluation.45.7BigML Evaluations APIEvaluate model performance against a test dataset45.4Laminar Evaluations APIThe lower-level LaminarClient.evals surface for wiring evaluations into an existing pipeline: create an eva...45.3AmeriCorps Open Data SODA APIThe AmeriCorps Open Data portal provides programmatic access to AmeriCorps research, evaluation, and progra...43.5Vellum LLM Platform APIThe Vellum REST API exposes prompts, workflows, evaluations, datasets, document indexes, deployments, and e...43.4W&B Weave (LLM Observability)LLM observability and evaluation platform providing tracing, output evaluation, cost estimation, prompt pla...43.4Understudy Control Plane APIREST surface behind the dashboard and CLI, served from the gateway host under /admin/v1 and /customer/v1, o...43.2AlyxAlyx is Arize's AI engineering agent that helps developers debug traces, create evaluators, build dashboard...43.1Arize AXArize AX is the commercial AI engineering platform covering tracing, evaluation, experiments, prompt manage...43.1PhoenixPhoenix is Arize's open-source LLM observability platform offering local tracing, evaluation, experiments, ...43.1Eigenpal Evaluation APIManage datasets, examples, evaluators, experiment batches, evaluator scores, and run promotion workflows.41.6HashiCorp Nomad Evaluations APIEndpoints for querying evaluations. Evaluations are the mechanism by which Nomad makes scheduling decisions.41.4Log10 Evaluation APIAPI for running automated evaluations and benchmarking logged completions across multiple LLM providers, ge...40.7Log10 Feedback APIFeedback40.7Humanloop LLM Platform APIThe Humanloop REST API and SDKs covered prompts, tools, datasets, evaluations, evaluators, and logs for LLM...40.1Athina AI Evaluations APIRun evaluations against datasets and logged inferences.39.4
Thin 18 Limited public surface area
Split Evaluation APIEndpoints for evaluating feature flags and retrieving treatment values for given keys and feature flag names.39.0GentraceGentrace was an AI evaluation and observability product; the company has shut down and its codebase is now ...38.7Agenta Evaluations APIRun evaluations of variants against testsets.38.4Oxen evaluations APIThe evaluations API from Oxen — 2 operation(s) for evaluations.38.4Extend Evaluations APIThe Evaluations API from Extend — 4 operation(s) for evaluations.38.3Docling EvalEnd-to-end evaluation framework for document parsing models and services. Provides standard datasets and me...37.9Workable Evaluations APISubmit and read interviewer evaluations and scorecards aligned to the job's interview kit.37.0Tonic ValidateTonic Validate is an open-source RAG evaluation framework and metrics platform for measuring retrieval-augm...36.5Together AI evaluation APIThe evaluation API from Together AI — 4 operation(s) for evaluation.36.3EvaluationPlatform and SDK capabilities for assessing model performance across data splits, running error and slice a...34.2UpTrain Evaluation APIThe Evaluation API from UpTrain — 3 operation(s) for evaluation.34.2LiteLLM Evals APIProvides /evals endpoints for the Evaluations API, enabling measurement and benchmarking of model performan...33.0Lunary APIREST API for Lunary covering ingestion (logs/traces), prompts (template management with versions and labels...31.5Confident AI PlatformConfident AI is the hosted platform that complements DeepEval with observability, centralized reporting, re...29.3AIMon APIREST API for AIMon LLM monitoring and evaluation — manage users, models, applications, evaluations and eval...28.5ARC Research and Data APIAnonymous read access to the Appalachian Regional Commission's research surface as JSON: 304 research repor...28.3Opik (GenAI Observability)Opik is Comet's open-source GenAI observability product. It provides spans-based tracing, evaluations, prom...28.2USAID Development Experience Clearinghouse APIThe USAID Development Experience Clearinghouse (DEC) API provides programmatic access to the largest online...27.0
Emerging 5 Early or largely undocumented

Companies reaching this through an API (40)

These companies publish an API, specification or operation carrying “Evaluation” but do not classify their business under it. Listed unranked and kept out of the count above, because one tagged operation is not a statement about what a company does.

8x8 Agenta AI Gateway Alloy AmeriCorps Amigo Amplitude Appalachian Regional Commission Authenticx Avito BigML Buk Comet Confident AI Convai Docling Eigenpal Extend Google Dialogflow Gumloop Harness HashiCorp Nomad LiteLLM Lunary Microsoft Windows 10 n8n Observe.AI OpenAI Oxen Patronus AI Qlik Sense Ragas Rapidata Split Together AI Tonic.ai Trakstar USAID Workable Zoom

Score breakdown

Frequency
58.9
log-scaled weighted occurrences
Breadth
1.3
spread across providers
Quality lift
40.7
mean composite of providers using it
Cohesion
0.0
strength of nearest seed neighbor

Related tags

Tracing 15 co-occurrences Prompts 14 co-occurrences Experiments 12 co-occurrences Prompt Management 8 co-occurrences LLM 36 co-occurrences Traces 10 co-occurrences Benchmarks 9 co-occurrences

Where this tag comes from

Api tag59
Provider tag48
Openapi tag27
Openapi op tag104

Cohort brief

Auto-generated

The 48 providers in the APIs.io catalog tagged Evaluation, scored on the Kin Score. Every figure below is computed from the catalog — nothing here is written.

Providers
48
all scored
Mean Kin Score
35.9
+11.9 vs catalog 24
Mean Agent Readiness
21.6
+9.2 vs catalog 12.4
Spread
5–61
median 38.4 · σ 12.6
How the 48 split by band
Exemplar 0Strong 2Developing 20Thin 13Emerging 12Minimal 1
Facet averages, against the whole catalog
FacetThis cohortCatalogDifferenceScored
Contract Quality 39.7 20.1 +19.6 48
Developer Ergonomics 37.9 23.4 +14.5 48
Operational Transparency 24.2 14 +10.2 48
Access Clarity 35.9 26.5 +9.4 48
Discoverability 69 61 +8 48
Contract Governance 4.9 6.5 -1.6 48

A facet is averaged over the members that carry it, not over the whole cohort — the “Scored” column is that count. Averaging an absent facet as zero would score our own coverage gaps as the providers’ posture.

What this cohort publishes
ArtifactThis cohortCatalogDifference
MCP server (any) 21% 10% +11
MCP server (first-party) 29% 13% +16
Agent Skills 0% 0% 0
OAuth scopes 2% 11% -9
Security 100% 96% +4
Arazzo workflows 6% 2% +4
Governance rules 13% 13% 0

mcp_pct counts any mcp/ artifact including ones API Evangelist derived from the provider OpenAPI; mcp_first_party_pct counts only servers the provider publishes. Prefer the latter.

Top by Kin Score
  1. 1 Runloop 61
  2. 2 Prime Intellect 56.5
  3. 3 Scorecard 52.8
  4. 4 Coval 52.1
  5. 5 Comet 51.5
  6. 6 Opik 50.1
  7. 7 Bluejay 48.3
  8. 8 Braintrust 48.2
  9. 9 Vijil 47.8
  10. 10 Traceloop 47.5
Top by Agent Readiness
  1. 1 Coval 59.8
  2. 2 Scorecard 50.9
  3. 3 Braintrust 49.6
  4. 4 Openlayer 44.6
  5. 5 Bluejay 38.6
  6. 6 Langfuse 36.7
  7. 7 Prime Intellect 33.7
  8. 8 LangChain 33.4
  9. 9 Archal 31
  10. 10 LangSmith 29.8
All 48 members of this roster resolve to a scored provider in the catalog.Generated from the catalog build of 20 September 2026, across 27620 providers, using the same computation served by the APIs.io cohort API. Briefs are published for rosters of 5 or more scored providers; below that a distribution is not meaningful.

Work with this as data

Every tag here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for tags

7 MCP tools reach this
  • find_tagsBrowse and filter every tag in the catalog.
  • get_cohortThis tag as a scored cohort — every provider carrying it, with scores.
  • cohort_statsPRO — the distribution across this tag: mean, median, band split, adoption rates.
  • cohort_rankingsPRO — the leaderboard, on composite AND agent-readiness axes.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools →

Call it yourself

curl for this page
This tag
curl "https://apis.io/api/v1/tags/evaluation"
All tags
curl "https://apis.io/api/v1/tags?limit=25"
As a scored cohort
curl "https://apis.io/api/v1/cohorts/tag/evaluation"
The distribution (Pro)
curl "https://apis.io/api/v1/cohorts/tag/evaluation/stats" \
  -H "X-API-Key: $APIS_IO_KEY"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no form to fill in. Signing in shares your email address with us — we store it to create your key and to recognise you if you sign in with another provider. See our Privacy Policy and Terms.

A second provider on the same verified email joins the account you already have.