# AgentBench

**Canonical:** https://apis.io/apis/evals/agentbench/  
**Provider:** Evals — https://apis.io/providers/evals/  
**Base URL:** https://github.com/THUDM/AgentBench  
**Documentation:** https://github.com/THUDM/AgentBench

AgentBench is one of 20 APIs that [Evals](https://apis.io/providers/evals/) publishes on the [APIs.io](https://apis.io/) network. Tagged areas include Benchmarks, AI Agents, Multi-Environment, LLM-as-Agent, and Tsinghua. The published artifact set on APIs.io includes a GitHub repository.

AgentBench is the first benchmark designed to evaluate LLM-as-Agent across a diverse spectrum of environments. It bundles 8 environments — 5 newly created (Operating System, Database, Knowledge Graph, Digital Card Game, Lateral Thinking Puzzles) and 3 adapted (House-Holding from ALFWorld, Web Shopping from WebShop, Web Browsing from Mind2Web). The benchmark requires roughly 4,000 dev-set and 13,000 test-set interactions per model.

## Machine-readable artifacts (3)

- **GitHubRepository** — https://github.com/THUDM/AgentBench
- **Paper** — https://arxiv.org/abs/2308.03688
- **Leaderboard** — https://llmbench.ai/agent

## Other Evals APIs (12)

- [OpenAI Evals](https://apis.io/apis/evals/openai-evals/)
- [Inspect AI](https://apis.io/apis/evals/inspect-ai/)
- [Braintrust](https://apis.io/apis/evals/braintrust/)
- [LangSmith Evaluation](https://apis.io/apis/evals/langsmith-evaluation/)
- [Promptfoo](https://apis.io/apis/evals/promptfoo/)
- [Helicone](https://apis.io/apis/evals/helicone/)
- [Patronus AI](https://apis.io/apis/evals/patronus-ai/)
- [DeepEval (Confident AI)](https://apis.io/apis/evals/deepeval-confident-ai/)
- [Arize AI (Phoenix)](https://apis.io/apis/evals/arize-ai-phoenix/)
- [Galileo](https://apis.io/apis/evals/galileo/)
- [Humanloop](https://apis.io/apis/evals/humanloop/)
- [TruLens](https://apis.io/apis/evals/trulens/)

## Tags

Benchmarks, AI Agents, Multi-Environment, LLM-as-Agent, Tsinghua

---

Profiled by [API Evangelist](https://apievangelist.com) and published on [APIs.io](https://apis.io/apis/evals/agentbench/). The API's provider profile, Kin Score and agent-readiness rating are at https://apis.io/providers/evals/.
