Confident AI is the company behind DeepEval, the widely adopted open-source LLM evaluation framework, and the Confident AI cloud platform that layers observability, dataset management, regression testing, and red teaming on top of the local framework. DeepEval treats LLM evaluation as unit testing with research-backed metrics such as GEval, AnswerRelevancy, and Faithfulness, while DeepTeam provides an open-source red teaming framework. The hosted platform is SOC 2 Type II, HIPAA, and GDPR compliant with self-hosting available for regulated customers.
Confident AI publishes 3 APIs on the APIs.io network. Tagged areas include LLM Evaluation, Open-Source, Observability, Red Teaming, and Guardrails.
Confident AI’s developer surface includes documentation, engineering blog, pricing, and 15 more developer resources.
Open Source Surface applies to this provider. This product is open source and we
read its repository directly, so Open Source Surface carries
10 points of the composite. It is scored from what the repository actually
publishes — a security policy, a contribution guide, a release history, a code of conduct — read live from the
provider rather than inferred from our own catalog pointers.
This facet adds; nothing was taken away to make room for it. An open-source project is not excused from
the commercial facets, because exemption would strip it of the points it does earn.
If we have the wrong repository, or this product is not open source, say so on your
provider repo and we
will drop the facet rather than have you publish against it.
Create-or-Update Ergonomics could not be measured. We hold no machine-readable contract for
this provider to read, so there is nothing to measure a write surface against. Excluded rather than scored zero:
never-measured and measured-empty are different facts. Publishing an OpenAPI is what makes this facet — and
several others — scorable at all.
The six quality facets above are damped to 90 points between them,
because the conditional facet above carries the other
10. That is why each facet's contribution is shown against a damped
maximum: raising a quality facet moves the composite by 90% of its nominal
weight, not 100%. The full arithmetic is at apis.io/rating/.
Improve this rating by publishing the missing artifacts — every area above can be raised, and the full rubric is at apis.io/rating/. Every facet and dimension name above is a link: it opens that measurement's own page — what it means, the exact checks that feed it, how the whole catalog distributes on it, and the providers at the top of it. This rating is computed from github.com/api-evangelist/confident-ai: open an issue to ask a question, or submit a pull request to add artifacts.
Submit an artifact on GitHub — free →Manage your own listing — the Influence plan, $499/mo →
DeepEval is an open-source Python framework for evaluating LLM applications as unit tests. It ships with research-backed metrics including GEval, AnswerRelevancyMetric, Faithful...
Confident AI is the hosted platform that complements DeepEval with observability, centralized reporting, regression testing, prompt versioning, dataset management, trace ingesti...
DeepTeam is Confident AI's open-source red teaming framework for stress-testing LLM applications against adversarial attacks including prompt injection, jailbreaks, PII leakage,...
Discovery needs no key. Ratings and market analysis are Pro.
Get an API key
Free tier, no form to fill in. Signing in shares your email address with us — we
store it to create your key and to recognise you if you sign in with another
provider. See our Privacy Policy and
Terms.