AI Crawler Index · Schema

Crawler

One crawler record from the Pathwren AI Crawler Index. Extracted verbatim from components.schemas.Crawler in the provider's own OpenAPI 3.1 document at https://www.pathwren.workers.dev/openapi.json (fetched 2026-09-01). No fields added or renamed.

AI crawlersweb crawlersrobots.txtuser agentsbot detectionGPTBotClaudeBotcrawler IP rangesllms.txtopen data

Properties

Name Type Description
slug string Stable identifier used in URLs.
name string
operator string
operator_slug string
category string
robots_token string Exact User-agent value for robots.txt.
user_agent_substring string Substring that reliably identifies it in a UA header. A match is a claim, not a proof.
user_agent_example string
respects_robots_txt string
verification_method string
published_ip_ranges_url stringnull
ipv4_prefix_count integer
ipv6_prefix_count integer
what_it_is string
cost_of_blocking string What you lose by disallowing it. This index's own assessment, not the operator's.
operator_docs string
last_reviewed string
View JSON Schema on GitHub

JSON Schema

pathwren-crawler.schema.json Raw ↑
{
 "$schema": "https://json-schema.org/draft/2020-12/schema",
 "$id": "https://www.pathwren.workers.dev/schemas/crawler.schema.json",
 "title": "Crawler",
 "description": "One crawler record from the Pathwren AI Crawler Index. Extracted verbatim from components.schemas.Crawler in the provider's own OpenAPI 3.1 document at https://www.pathwren.workers.dev/openapi.json (fetched 2026-09-01). No fields added or renamed.",
 "type": "object",
 "required": [
  "slug",
  "name",
  "operator",
  "category",
  "robots_token"
 ],
 "properties": {
  "slug": {
   "type": "string",
   "description": "Stable identifier used in URLs."
  },
  "name": {
   "type": "string"
  },
  "operator": {
   "type": "string"
  },
  "operator_slug": {
   "type": "string"
  },
  "category": {
   "type": "string",
   "enum": [
    "ai-search",
    "ai-training",
    "archive",
    "dataset",
    "preview",
    "search",
    "seo",
    "tool",
    "user-fetch"
   ]
  },
  "robots_token": {
   "type": "string",
   "description": "Exact User-agent value for robots.txt."
  },
  "user_agent_substring": {
   "type": "string",
   "description": "Substring that reliably identifies it in a UA header. A match is a claim, not a proof."
  },
  "user_agent_example": {
   "type": "string"
  },
  "respects_robots_txt": {
   "type": "string",
   "enum": [
    "documented",
    "by-design-no",
    "disputed",
    "n-a"
   ]
  },
  "verification_method": {
   "type": "string",
   "enum": [
    "published-ranges",
    "reverse-dns",
    "none"
   ]
  },
  "published_ip_ranges_url": {
   "type": [
    "string",
    "null"
   ],
   "format": "uri"
  },
  "ipv4_prefix_count": {
   "type": "integer"
  },
  "ipv6_prefix_count": {
   "type": "integer"
  },
  "what_it_is": {
   "type": "string"
  },
  "cost_of_blocking": {
   "type": "string",
   "description": "What you lose by disallowing it. This index's own assessment, not the operator's."
  },
  "operator_docs": {
   "type": "string",
   "format": "uri"
  },
  "last_reviewed": {
   "type": "string",
   "format": "date"
  }
 }
}

Work with this as data

Every JSON Schema here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for schemas

4 MCP tools reach this
  • find_json_schemasBrowse and filter every JSON Schema in the catalog.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools →

Call it yourself

curl for this page
This JSON Schema
curl "https://apis.io/api/v1/json-schemas/pathwren-crawler.schema"
All schemas
curl "https://apis.io/api/v1/json-schemas?limit=25"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no email required.

A second provider on the same verified email joins the account you already have.