LLM Evaluation
Providers using this tag (13)
Ranked by API Evangelist rating — Exemplar and Strong are expanded by default.
Strong 1 Solid coverage with minor gaps
Emerging 4 Early or largely undocumented
APIs with this tag (4)
Ranked by the provider's API Evangelist rating — the Kin Score is scored per provider, not per API, so every API of a provider shares its band. How the rating works →
Thin 3 Limited public surface area
Promptfoo CLIThe Promptfoo CLI is the primary entry point for running prompt and model evaluations from the command line...Promptfoo Node.js LibraryThe Promptfoo Node.js package exposes the same evaluation engine programmatically so developers can embed e...DeepEvalDeepEval is an open-source Python framework for evaluating LLM applications as unit tests. It ships with re...
Emerging 1 Early or largely undocumented
Score breakdown
Frequency
33.7
log-scaled weighted occurrences
Breadth
0.4
spread across providers
Quality lift
38.7
mean composite of providers using it
Cohesion
2.4
strength of nearest seed neighbor
Related tags
Guardrails 5 co-occurrences
Evaluation 5 co-occurrences
Python 5 co-occurrences
Datasets 5 co-occurrences
Observability 5 co-occurrences
Open-Source 8 co-occurrences
Artificial Intelligence 5 co-occurrences
Where this tag comes from
Provider tag13
Api tag4
Work with this as data
Every tag here is available over the APIs.io API and to AI agents over MCP.