Patronus AI
Patronus AI is an evaluation and guardrails platform for production LLM applications and AI agents. It combines an API-first evaluation service with Python and TypeScript SDKs, in-house judge models (Lynx for hallucination detection, Glider for reasoning evaluation, Percival for agent debugging), and a portfolio of open benchmarks and datasets including FinanceBench, BLUR, and RL environments. Customers use Patronus for experimentation, production monitoring, RAG and agent evaluation, dataset generation, and human-in-the-loop annotation.
Patronus AI publishes 7 APIs on the APIs.io network. Tagged areas include LLM Evaluation, Guardrails, Judges, Hallucination Detection, and AI Research.
Patronus AI’s developer surface includes documentation, API reference, engineering blog, pricing, and 7 more developer resources.
Kin Score
APIs 7
Individual APIs this provider publishes, each with its own machine-readable definition.
Patronus Evaluation API
The Patronus Evaluation API scores LLM outputs against built-in and custom evaluators covering hallucination, answer relevance, context utilization, safety, and PII. Evaluators ...
Patronus Python SDK
The Patronus Python SDK provides decorators and clients for instrumenting LLM applications, running evaluators inline, recording traces, and pushing experiments to the Patronus ...
Patronus TypeScript SDK
The Patronus TypeScript SDK brings the same evaluation, tracing, and experiment workflows to Node.js and browser environments used by JavaScript-first AI applications.
Lynx
Lynx is Patronus's open-weights hallucination detection model published on Hugging Face. It is positioned as state-of-the-art on hallucination benchmarks and is available both a...
Glider
Glider is Patronus's small judge model for evaluating reasoning chains and rubric-based scoring with low latency and cost relative to large frontier judges.
Percival
Percival is Patronus's agent debugging product that ingests agent traces and surfaces failure modes, tool misuse, and reasoning errors across multi-step runs.
FinanceBench
FinanceBench is an open benchmark of 10,000 financial question-answer pairs grounded in public filings, used to evaluate LLM performance on financial document understanding.
Scroll for all 7
Pricing Plans 1
Published pricing tiers and plan structures.
Rate Limits 1
Documented rate limits and quota policies.
Patronus Ai Rate Limits
RATE LIMITSFinOps 1
Cost, billing, and metering signals for API financial operations.
Patronus Ai Finops
FINOPSFeatures 8
Notable capabilities this provider offers.
Evaluation API
Hosted API for running built-in and custom evaluators on LLM inputs and outputs.
Lynx Hallucination Detection
State-of-the-art open-weights hallucination judge available as a hosted evaluator.
Glider Judge
Small reasoning-focused judge for rubric-based evaluation at production latency.
Percival Agent Debugger
Agent trace analysis surfacing failure modes, tool misuse, and reasoning errors.
Experimentation
Compare prompts, models, and configurations across datasets with side-by-side outputs.
Production Monitoring
Real-time alerts, tracing, and dashboards for live LLM applications.
Dataset Generation
Synthetic dataset creation including red-teaming sets for RAG and agent systems.
Human Annotation
Workflows for human-in-the-loop labeling and reviewer agreement tracking.
Scroll for all 8
Security Posture 1
Authentication, domain security, vulnerability disclosure, and trust-center signals.
Use Cases 5
What developers build with this provider.
RAG Evaluation
Score retrieval and generation quality in RAG applications across faithfulness, relevance, and context.
Agent Debugging
Trace and diagnose failures in multi-step agentic systems using Percival.
Model Benchmarking
Benchmark candidate models against domain-specific datasets such as FinanceBench.
Guardrails
Apply Patronus judges as runtime guardrails on LLM responses.
Regression Testing
Detect quality regressions across prompt, model, and configuration changes.
Integrations 6
Pre-built integrations with other platforms and tools.
OpenAI
Score outputs from OpenAI models inside Patronus experiments and monitoring.
Anthropic
Evaluate Anthropic Claude outputs using Patronus judges.
LangChain
SDK integrations for LangChain chains and agents.
LlamaIndex
Evaluate LlamaIndex RAG pipelines with Patronus evaluators.
OpenTelemetry
Ingest OTel-compatible LLM traces for evaluation and monitoring.
Hugging Face
Lynx and Glider weights are distributed via Hugging Face for self-hosting.
Resources
Get Started 1
Portal, sign-up, and the first successful call
Documentation 2
Reference material describing how the API behaves
Build 1
SDKs, sample code, and the tooling you integrate with
Access & Security 2
Authentication, authorization, and security posture
Operate 1
Status, limits, changes, and where to get help
Commercial 1
Pricing, plans, and the legal terms of use
Company 2
The organization behind the API
Other 1
Properties that don't map to a standard resource type