# DeepEval

**Canonical:** https://apis.io/providers/deepeval/  
**Website:** https://www.confident-ai.com  
**APIs profiled:** 0

DeepEval is an open-source LLM evaluation framework — built and maintained by Confident AI — for testing and benchmarking large language model applications. It is structured like Pytest but specialized for LLM systems, providing 40+ research-backed metrics (G-Eval, DAG, RAG metrics, agent metrics, multi-turn conversation metrics, multimodal metrics, MCP metrics, hallucination, bias, toxicity, summarization, JSON correctness) that run locally against any LLM provider (OpenAI, Anthropic, Gemini, Bedrock, Vertex AI, Ollama, OpenRouter, vLLM, LM Studio, LiteLLM, Azure OpenAI, DeepSeek, Grok, Moonshot, Portkey). DeepEval supports end-to-end and component-level evaluation via the `@observe()` decorator, synthetic dataset generation, multi-turn conversation simulation, CI/CD integration, automatic prompt optimization, and one-line LLM benchmarking (MMLU, HellaSwag, DROP, BIG-Bench Hard, TruthfulQA, HumanEval, GSM8K). DeepEval ships as the `deepeval` Python package on PyPI together with a `deepeval` command-line tool. The framework integrates natively with pytest, LangChain, LangGraph, LlamaIndex, OpenAI Agents, CrewAI, Pydantic AI, AWS AgentCore, Google ADK, and Strands. DeepEval is open source under Apache 2.0 and is the engine that powers Confident AI's commercial LLM evaluation, observability, and red-teaming platform; `deepeval login` connects local test runs to the Confident AI cloud for shared regression reports, dataset management, production tracing, and prompt versioning. A sibling open-source framework, DeepTeam (`deepteam`), targets adversarial / red-team testing of LLM apps.

## Kin Score — 26.2 / 100 (emerging)

Scored 2026-08-17 under rubric 0.11.0. Trend: flat (+0.0 from 26.2).

| Facet | Score |
|---|---|
| Discoverability | 50.0 |
| Contract Quality | 0.0 |
| Governance | 0.0 |
| Operational Transparency | 21.1 |
| Developer Ergonomics | 47.8 |
| Commercial Clarity | 44.7 |

## Agent readiness — 1.6 (human-only)

| Dimension | Value |
|---|---|
| Spec Presence | no |
| Agentic Access | no |
| MCP Server | no |
| Auth Clarity | no |
| Idempotency | no |
| Error Semantics | no |
| OpenAPI Examples | documented |
| Rate Limit Signal | no |
| Event Surface Described | no |
| Agent Skills | no |
| Well Known Catalog | no |
| Consent Identity | no |
| Agent Card | no |
| Dry Run Mode | no |

## Access

Unknown — onboarding: unknown, pricing: unknown, trial: no (confidence: low).

## Security (1)

- **Deepeval Domain Security** — TLSv1.3 · HSTS · DNSSEC · DMARC

## Tags

LLM Evaluation, LLM Testing, Evaluation Framework, Evaluation Metrics, LLM Observability, LLM as a Judge, G-Eval, RAG Evaluation, Agent Evaluation, Hallucination Detection, Bias Detection, Toxicity Detection, Red Teaming, Benchmarks, MMLU, Synthetic Data Generation, Prompt Optimization, CI CD, Pytest, Python, Open Source, Apache 2.0, MCP

---

Profiled by [API Evangelist](https://apievangelist.com) and published on [APIs.io](https://apis.io/providers/deepeval/). Scores are computed from the provider's own public artifacts under a published rubric.
