# GAIA Benchmark

**Canonical:** https://apis.io/apis/evals/gaia-benchmark/  
**Provider:** Evals — https://apis.io/providers/evals/  
**Base URL:** https://huggingface.co/gaia-benchmark  
**Documentation:** https://huggingface.co/gaia-benchmark

GAIA Benchmark is one of 20 APIs that [Evals](https://apis.io/providers/evals/) publishes on the [APIs.io](https://apis.io/) network. Tagged areas include Benchmarks, AI Agents, Reasoning, Tool Use, and Leaderboard.

GAIA is "a benchmark for General AI Assistants," published in 2023 (arXiv 2311.12983). It tests general-purpose AI agent capability across reasoning, tool use, multi-modality, and web browsing, with a public leaderboard hosted on Hugging Face for community submissions. The benchmark has become a reference point for evaluating agentic systems that combine an LLM with tools and a browser.

## Machine-readable artifacts (3)

- **Dataset** — https://huggingface.co/datasets/gaia-benchmark/GAIA
- **Paper** — https://arxiv.org/abs/2311.12983
- **Leaderboard** — https://huggingface.co/spaces/gaia-benchmark/leaderboard

## Other Evals APIs (12)

- [OpenAI Evals](https://apis.io/apis/evals/openai-evals/)
- [Inspect AI](https://apis.io/apis/evals/inspect-ai/)
- [Braintrust](https://apis.io/apis/evals/braintrust/)
- [LangSmith Evaluation](https://apis.io/apis/evals/langsmith-evaluation/)
- [Promptfoo](https://apis.io/apis/evals/promptfoo/)
- [Helicone](https://apis.io/apis/evals/helicone/)
- [Patronus AI](https://apis.io/apis/evals/patronus-ai/)
- [DeepEval (Confident AI)](https://apis.io/apis/evals/deepeval-confident-ai/)
- [Arize AI (Phoenix)](https://apis.io/apis/evals/arize-ai-phoenix/)
- [Galileo](https://apis.io/apis/evals/galileo/)
- [Humanloop](https://apis.io/apis/evals/humanloop/)
- [TruLens](https://apis.io/apis/evals/trulens/)

## Tags

Benchmarks, AI Agents, Reasoning, Tool Use, Leaderboard

---

Profiled by [API Evangelist](https://apievangelist.com) and published on [APIs.io](https://apis.io/apis/evals/gaia-benchmark/). The API's provider profile, Kin Score and agent-readiness rating are at https://apis.io/providers/evals/.
