GAIA Benchmark

GAIA is "a benchmark for General AI Assistants," published in 2023 (arXiv 2311.12983). It tests general-purpose AI agent capability across reasoning, tool use, multi-modality, and web browsing, with a public leaderboard hosted on Hugging Face for community submissions. The benchmark has become a reference point for evaluating agentic systems that combine an LLM with tools and a browser.

Work with this as data

Every API here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for apis

7 MCP tools reach this
  • find_apisBrowse and filter every API in the catalog.
  • get_api_artifactsOne API's artifacts, grouped by type.
  • get_openapiThe primary OpenAPI for this API.
  • find_similar_apisAPIs that look like this one.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools →

Call it yourself

curl for this page
This API
curl "https://apis.io/api/v1/apis/gaia-benchmark"
All apis
curl "https://apis.io/api/v1/apis?limit=25"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no form to fill in. Signing in shares your email address with us — we store it to create your key and to recognise you if you sign in with another provider. See our Privacy Policy and Terms.

A second provider on the same verified email joins the account you already have.

API entry from apis.yml

apis.yml Raw ↑
name: GAIA Benchmark
description: GAIA is "a benchmark for General AI Assistants," published in 2023 (arXiv 2311.12983). It
  tests general-purpose AI agent capability across reasoning, tool use, multi-modality, and web browsing,
  with a public leaderboard hosted on Hugging Face for community submissions. The benchmark has become
  a reference point for evaluating agentic systems that combine an LLM with tools and a browser.
humanURL: https://huggingface.co/gaia-benchmark
baseURL: https://huggingface.co/gaia-benchmark
tags:
- Benchmarks
- AI Agents
- Reasoning
- Tool Use
- Leaderboard
tags_raw:
- Benchmark
- AI Agents
- Reasoning
- Tool Use
- Leaderboard
properties:
- type: Dataset
  url: https://huggingface.co/datasets/gaia-benchmark/GAIA
- type: Paper
  url: https://arxiv.org/abs/2311.12983
- type: Leaderboard
  url: https://huggingface.co/spaces/gaia-benchmark/leaderboard