AI Crawler Index Crawlers API

One record per crawler.

Operations 57

GET /crawler/{slug}.json One crawler record by slug #
GET /crawler/gptbot.json Record for GPTBot #
GET /crawler/oai-searchbot.json Record for OAI-SearchBot #
GET /crawler/chatgpt-user.json Record for ChatGPT-User #
GET /crawler/claudebot.json Record for ClaudeBot #
GET /crawler/claude-searchbot.json Record for Claude-SearchBot #
GET /crawler/claude-user.json Record for Claude-User #
GET /crawler/anthropic-ai.json Record for anthropic-ai #
GET /crawler/claude-web.json Record for Claude-Web #
GET /crawler/google-extended.json Record for Google-Extended #
GET /crawler/googlebot.json Record for Googlebot #
GET /crawler/googleother.json Record for GoogleOther #
GET /crawler/google-cloudvertexbot.json Record for Google-CloudVertexBot #
GET /crawler/google-inspectiontool.json Record for Google-InspectionTool #
GET /crawler/googlebot-image.json Record for Googlebot-Image #
GET /crawler/googlebot-news.json Record for Googlebot-News #
GET /crawler/storebot-google.json Record for Storebot-Google #
GET /crawler/bingbot.json Record for bingbot #
GET /crawler/applebot.json Record for Applebot #
GET /crawler/applebot-extended.json Record for Applebot-Extended #
GET /crawler/perplexitybot.json Record for PerplexityBot #
GET /crawler/perplexity-user.json Record for Perplexity-User #
GET /crawler/ccbot.json Record for CCBot #
GET /crawler/bytespider.json Record for Bytespider #
GET /crawler/tiktokspider.json Record for TikTokSpider #
GET /crawler/meta-externalagent.json Record for meta-externalagent #
GET /crawler/meta-externalfetcher.json Record for meta-externalfetcher #
GET /crawler/facebookexternalhit.json Record for facebookexternalhit #
GET /crawler/facebookbot.json Record for FacebookBot #
GET /crawler/amazonbot.json Record for Amazonbot #
GET /crawler/duckassistbot.json Record for DuckAssistBot #
GET /crawler/duckduckbot.json Record for DuckDuckBot #
GET /crawler/ai2bot.json Record for AI2Bot #
GET /crawler/ai2bot-dolma.json Record for Ai2Bot-Dolma #
GET /crawler/cohere-ai.json Record for cohere-ai #
GET /crawler/cohere-training-data-crawler.json Record for cohere-training-data-crawler #
GET /crawler/mistralai-user.json Record for MistralAI-User #
GET /crawler/youbot.json Record for YouBot #
GET /crawler/diffbot.json Record for Diffbot #
GET /crawler/omgilibot.json Record for omgilibot #
GET /crawler/omgili.json Record for omgili #
GET /crawler/webzio-extended.json Record for Webzio-Extended #
GET /crawler/imagesiftbot.json Record for ImagesiftBot #
GET /crawler/timpibot.json Record for Timpibot #
GET /crawler/semrushbot.json Record for SemrushBot #
GET /crawler/semrushbot-ocob.json Record for SemrushBot-OCOB #
GET /crawler/ahrefsbot.json Record for AhrefsBot #
GET /crawler/archive-org-bot.json Record for archive.org_bot #
GET /crawler/ia-archiver.json Record for ia_archiver #
GET /crawler/yandexbot.json Record for YandexBot #
GET /crawler/baiduspider.json Record for Baiduspider #
GET /crawler/seznambot.json Record for SeznamBot #
GET /crawler/yeti.json Record for Yeti #
GET /crawler/petalbot.json Record for PetalBot #
GET /crawler/firecrawlagent.json Record for FirecrawlAgent #
GET /crawler/scrapy.json Record for Scrapy #
GET /crawler/img2dataset.json Record for img2dataset #

Documentation

Specifications

Schemas & Data

Other Resources

🔗
x-openapi-yaml
https://www.pathwren.workers.dev/openapi.yaml
🔗
Swagger
https://www.pathwren.workers.dev/swagger.json
🔗
Feed
https://www.pathwren.workers.dev/feed.xml
🔗
Feed
https://www.pathwren.workers.dev/feed.json
🔗
x-mcp-server
https://www.pathwren.workers.dev/mcp
🔗
x-mcp-manifest
https://www.pathwren.workers.dev/.well-known/mcp.json
🔗
x-onboarding
https://www.pathwren.workers.dev/.well-known/api-onboarding
🔗
x-plugin-manifest
https://www.pathwren.workers.dev/.well-known/ai-plugin.json
🔗
x-api-catalog
https://www.pathwren.workers.dev/.well-known/api-catalog
🔗
x-llms-txt
https://www.pathwren.workers.dev/llms.txt
🔗
x-bulk-data
https://www.pathwren.workers.dev/data/agents.json
🔗
x-ip-ranges
https://www.pathwren.workers.dev/ip-ranges/all.json
🔗
x-security-txt
https://www.pathwren.workers.dev/.well-known/security.txt
🔗
x-sitemap
https://www.pathwren.workers.dev/sitemap.xml
🔗
StatusPage
https://www.pathwren.workers.dev/status.html
🔗
License
https://creativecommons.org/publicdomain/zero/1.0/
🔗
Conventions
https://raw.githubusercontent.com/api-evangelist/pathwren/refs/heads/main/conventions/pathwren-conventions.yml
🔗
ErrorCatalog
https://raw.githubusercontent.com/api-evangelist/pathwren/refs/heads/main/errors/pathwren-problem-types.yml
🔗
DataModel
https://raw.githubusercontent.com/api-evangelist/pathwren/refs/heads/main/data-model/pathwren-data-model.yml
🔗
ToolCrosswalk
https://raw.githubusercontent.com/api-evangelist/pathwren/refs/heads/main/mcp/pathwren-tool-crosswalk.yml
🔗
ChangeLog
https://www.pathwren.workers.dev/changelog.html

Work with this as data

Every API here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for apis

7 MCP tools reach this
  • find_apisBrowse and filter every API in the catalog.
  • get_api_artifactsOne API's artifacts, grouped by type.
  • get_openapiThe primary OpenAPI for this API.
  • find_similar_apisAPIs that look like this one.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools →

Call it yourself

curl for this page
This API
curl "https://apis.io/api/v1/apis/pathwren-crawlers-api"
All apis
curl "https://apis.io/api/v1/apis?limit=25"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no form to fill in. Signing in shares your email address with us — we store it to create your key and to recognise you if you sign in with another provider. See our Privacy Policy and Terms.

A second provider on the same verified email joins the account you already have.

OpenAPI Specification

pathwren-crawlers-api-openapi.yml Raw ↑
openapi: 3.2.0
info:
  title: AI Crawler Index Crawlers API
  version: '2026-09-01'
  summary: Every AI crawler on the web, what it is for, what blocking it costs you, and the IP ranges its operator publishes — as JSON, CSV, robots.txt and regex.
  description: 'A read-only, static, keyless index of 56 web crawlers operated by 30 companies and projects: what each one is for, the exact robots.txt token and user-agent substring, whether the operator says it obeys robots.txt, how to verify it is genuine, and — the part nobody else publishes — what you lose by blocking it.'
  license:
    name: CC0-1.0
    url: https://creativecommons.org/publicdomain/zero/1.0/
  contact:
    url: https://www.pathwren.workers.dev/about.html
servers:
- url: https://www.pathwren.workers.dev
tags:


# --- truncated at 32 KB (33 KB total) ---
# Full source: https://raw.githubusercontent.com/api-evangelist/pathwren/refs/heads/main/openapi/pathwren-crawlers-api-openapi.yml