Benchmarks
Variants seen in the corpus:
BenchmarkBenchmarks
Providers using this tag (15)
Ranked by API Evangelist rating — Exemplar and Strong are expanded by default.
Strong 1 Solid coverage with minor gaps
Thin 2 Limited public surface area
Emerging 9 Early or largely undocumented
Minimal 3 Almost no public developer surface
APIs with this tag (29)
Ranked by the provider's API Evangelist rating — the Kin Score is scored per provider, not per API, so every API of a provider shares its band. How the rating works →
Exemplar 1 Complete, well-documented, and agent-ready
Strong 5 Solid coverage with minor gaps
Scope3 Benchmarks APIThe Benchmarks API from Scope3 — 1 operation(s) for benchmarks.Runloop Benchmark APIThe Benchmark API from Runloop — 25 operation(s) for benchmark.Visier Benchmarks APIGet benchmark values.Mastercard Benchmarks APIThe Benchmarks API from Mastercard — 1 operation(s) for benchmarks.Kaiko IndicesRegulated benchmark crypto indices designed for fund administrators and ETF/ETP issuers.
Developing 8 Usable, with meaningful gaps to close
tpuf-benchmarkOpen-source general-purpose benchmarking tool for turbopuffer deployments. Useful for validating recall, la...Aleph Alpha Benchmarks APIEndpoints for handling benchmarks. A Benchmark provides a way to compare and evaluate the quality of your t...CME Term SOFR APIJSON-over-REST API delivering CME Term SOFR Reference Rates - the IOSCO-compliant forward-looking term Secu...Rapidata Benchmark APIThe Benchmark API from Rapidata — 19 operation(s) for benchmark.Bloomberg Buyside Enterprise Solutions Benchmarks APIBenchmark assignment and comparisonCMS — Centers for Medicare & Medicaid Services Benchmarks APIThe Benchmarks API from CMS — Centers for Medicare & Medicaid Services — 1 operation(s) for benchmarks.Workera Benchmarks APIThe Benchmarks API from Workera — 2 operation(s) for benchmarks.Xignite Global Indices APIStock market index and benchmark values with chart bars and index metadata from the globalindices.xignite.c...
Thin 8 Limited public surface area
Docling EvalEnd-to-end evaluation framework for document parsing models and services. Provides standard datasets and me...Serif Health Payer InventoryLive public inventory of 200+ payers with network-quality scoring, updated monthly, exposing data coverage ...EvaluationPlatform and SDK capabilities for assessing model performance across data splits, running error and slice a...LiteLLM Evals APIProvides /evals endpoints for the Evaluations API, enabling measurement and benchmarking of model performan...NOAA CO-OPS Benchmarks APIThe Benchmarks API from NOAA CO-OPS — 1 operation(s) for benchmarks.PennyLane DatasetsA curated catalogue of pre-computed quantum datasets — molecular Hamiltonians, spin systems, QML benchmarks...APEX Benchmarks (AI Productivity Index)Mercor's public AI productivity benchmark and research surface. APEX measures how well AI models perform re...Terminal-BenchPublic benchmark / task-submission framework published by Mercor (terminal-bench-3 on GitHub) for evaluatin...
Emerging 7 Early or largely undocumented
CoinAPI Indexes APIThe Indexes API aggregates data from many exchanges to compute reference rates and benchmark indexes that s...FinanceBenchFinanceBench is an open benchmark of 10,000 financial question-answer pairs grounded in public filings, use...AgentBenchAgentBench is the first benchmark designed to evaluate LLM-as-Agent across a diverse spectrum of environmen...BIG-BenchThe Beyond the Imitation Game Benchmark (BIG-Bench) is "a collaborative benchmark intended to probe large l...GAIA BenchmarkGAIA is "a benchmark for General AI Assistants," published in 2023 (arXiv 2311.12983). It tests general-pur...HumanEval BenchmarkHumanEval is OpenAI's evaluation harness for code-generation models, described in its README as "an evaluat...MMLU BenchmarkMMLU (Measuring Massive Multitask Language Understanding) is a multiple-choice benchmark spanning 57 subjec...
Companies reaching this through an API (21)
These companies publish an API, specification or operation carrying “Benchmarks” but do not classify their business under it. Listed unranked and kept out of the count above, because one tagged operation is not a statement about what a company does.
Aleph Alpha
Bloomberg Buyside Enterprise Solutions
CME Group
CMS — Centers for Medicare & Medicaid Services
CoinAPI
Docling
Factset
Kaiko
LiteLLM
Mastercard
Mercor
NOAA CO-OPS
Rapidata
Scope3
Serif Health
Snorkel AI
turbopuffer
Visier
Workera
Xanadu
Xignite
Score breakdown
Frequency
47.4
log-scaled weighted occurrences
Breadth
0.4
spread across providers
Quality lift
24.5
mean composite of providers using it
Cohesion
3.6
strength of nearest seed neighbor
Related tags
Evaluation 9 co-occurrences
Datasets 9 co-occurrences
Python 6 co-occurrences
Market Data 6 co-occurrences
Data 7 co-occurrences
Research 6 co-occurrences
Real-Time 5 co-occurrences
Agents 9 co-occurrences
Where this tag comes from
Provider tag15
Api tag29
Openapi tag10
Openapi op tag32
Work with this as data
Every tag here is available over the APIs.io API and to AI agents over MCP.