# Patronus AI

**Canonical:** https://apis.io/providers/patronus-ai/  
**Website:** https://www.patronus.ai/  
**APIs profiled:** 7

Patronus AI is an evaluation and guardrails platform for production LLM applications and AI agents. It combines an API-first evaluation service with Python and TypeScript SDKs, in-house judge models (Lynx for hallucination detection, Glider for reasoning evaluation, Percival for agent debugging), and a portfolio of open benchmarks and datasets including FinanceBench, BLUR, and RL environments. Customers use Patronus for experimentation, production monitoring, RAG and agent evaluation, dataset generation, and human-in-the-loop annotation.

## Kin Score — 20.6 / 100 (emerging)

Scored 2026-08-25 under rubric 0.14.0. Trend: flat (+0.0 from 20.6).

| Facet | Score |
|---|---|
| Discoverability | 64.8 |
| Contract Quality | 0.0 |
| Governance | 0.0 |
| Contract Governance | 0.0 |
| Operational Transparency | 34.2 |
| Developer Ergonomics | 2.4 |
| Commercial Clarity | 46.1 |
| Access Clarity | 46.1 |

## Agent readiness — 2.5 (human-only)

| Dimension | Value |
|---|---|
| Spec Presence | no |
| Agentic Access | no |
| Reversibility Documented | no |
| MCP Server | no |
| Auth Clarity | no |
| Idempotency | no |
| Error Semantics | no |
| OpenAPI Examples | no |
| Rate Limit Signal | documented |
| Event Surface Described | no |
| Agent Skills | no |
| Well Known Catalog | no |
| Consent Identity | no |
| Agent Card | no |
| Dry Run Mode | no |
| Delegated Identity | no |
| Protected Resource Metadata | no |
| Dynamic Client Registration | no |
| Agentic Commerce | no |

## Access

Free — onboarding: unknown, pricing: free, trial: no (confidence: medium).

## APIs (7)

- **Patronus Evaluation API** — The Patronus Evaluation API scores LLM outputs against built-in and custom evaluators covering hallucination, answer relevance, context utilization, safety, and PII. Evaluators ...
- **Patronus Python SDK** — The Patronus Python SDK provides decorators and clients for instrumenting LLM applications, running evaluators inline, recording traces, and pushing experiments to the Patronus ...
- **Patronus TypeScript SDK** — The Patronus TypeScript SDK brings the same evaluation, tracing, and experiment workflows to Node.js and browser environments used by JavaScript-first AI applications.
- **Lynx** — Lynx is Patronus's open-weights hallucination detection model published on Hugging Face. It is positioned as state-of-the-art on hallucination benchmarks and is available both a...
- **Glider** — Glider is Patronus's small judge model for evaluating reasoning chains and rubric-based scoring with low latency and cost relative to large frontier judges.
- **Percival** — Percival is Patronus's agent debugging product that ingests agent traces and surfaces failure modes, tool misuse, and reasoning errors across multi-step runs.
- **FinanceBench** — FinanceBench is an open benchmark of 10,000 financial question-answer pairs grounded in public filings, used to evaluate LLM performance on financial document understanding.

## Security (1)

- **Patronus Ai Domain Security** — TLSv1.3 · HSTS · DMARC

## Plans (1)

- **Patronus Ai Plans Pricing**

## Use cases (5)

- **RAG Evaluation** — Score retrieval and generation quality in RAG applications across faithfulness, relevance, and context.
- **Agent Debugging** — Trace and diagnose failures in multi-step agentic systems using Percival.
- **Model Benchmarking** — Benchmark candidate models against domain-specific datasets such as FinanceBench.
- **Guardrails** — Apply Patronus judges as runtime guardrails on LLM responses.
- **Regression Testing** — Detect quality regressions across prompt, model, and configuration changes.

## Tags

LLM Evaluation, Guardrails, Judges, Hallucination Detection, AI Research, Benchmarks

---

Profiled by [API Evangelist](https://apievangelist.com) and published on [APIs.io](https://apis.io/providers/patronus-ai/). Scores are computed from the provider's own public artifacts under a published rubric.
