Crawlbase Crawling API

Fetch any URL through the rotating proxy network, optionally rendered.

Operations 2

GET / Crawl a URL #
POST / Crawl a URL (POST) #

Work with this as data

Every API here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for apis

7 MCP tools reach this
  • find_apisBrowse and filter every API in the catalog.
  • get_api_artifactsOne API's artifacts, grouped by type.
  • get_openapiThe primary OpenAPI for this API.
  • find_similar_apisAPIs that look like this one.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools →

Call it yourself

curl for this page
This API
curl "https://apis.io/api/v1/apis/crawlbase-crawling-api-api"
All apis
curl "https://apis.io/api/v1/apis?limit=25"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no form to fill in. Signing in shares your email address with us — we store it to create your key and to recognise you if you sign in with another provider. See our Privacy Policy and Terms.

A second provider on the same verified email joins the account you already have.

OpenAPI Specification

crawlbase-crawling-api-api-openapi.yml Raw ↑
openapi: 3.2.0
info:
  title: Crawlbase Crawling API
  description: Crawlbase (formerly ProxyCrawl) is a web crawling and scraping platform.
  version: '1.0'
  contact:
    name: Crawlbase
    url: https://crawlbase.com
servers:
- url: https://api.crawlbase.com
  description: Crawlbase API host (all products share this host)
security:
- tokenAuth: []
tags:
- name: Crawling API
  description: Fetch any URL through the rotating proxy network, optionally rendered.
paths:
  /:
    get:
      operationId: crawl
      tags:
      - Crawling API
      summary: Crawl a URL
      description: 'Fetches the target `url` through Crawlbase''s proxy network and returns the page. Use the Normal (TCP) token for static HTML/JSON, or the JavaScript token to render the page in headless Chrome before returning it. Crawlbase status is surfaced in the `pc_status` response header; requests that do not return `pc_status: 200` are not charged.'
      parameters:
      - $ref: '#/components/parameters/Url'
      - $ref: '#/components/parameters/Format'
      - $ref: '#/components/parameters/Country'
      - $ref: '#/components/parameters/Device'
      - $ref: '#/components/parameters/UserAgent'
      - $ref: '#/components/parameters/Scraper'
      - name: autoparse
        in: query
        required: false
        description: Auto-detect the page type and return parsed JSON.
        schema:
          type: boolean
      - name: page_wait
        in: query
        required: false
        description: Milliseconds to wait after page load before capture (JavaScript token only).
        schema:
          type: integer
      - name: ajax_wait
        in: query
        required: false
        description: Wait for AJAX / network idle before capture (JavaScript token only).
        schema:
          type: boolean
      - name: css_click_selector
        in: query
        required: false
        description: CSS selector to click before capture (JavaScript token only).
        schema:
          type: string
      - name: scroll
        in: query
        required: false
        description: Scroll toward the bottom of the page before capture (JavaScript token only).
        schema:
          type: boolean
      - name: scroll_interval
        in: query
        required: false
        description: Maximum seconds to scroll (default 10, max 60).
        schema:
          type: integer
      - name: screenshot
        in: query
        required: false
        description: Capture a JPEG screenshot of the rendered page (JavaScript token only).
        schema:
          type: boolean
      - name: get_headers
        in: query
        required: false
        description: Return the target's response headers alongside the body.
        schema:
          type: boolean
      - name: get_cookies
        in: query
        required: false
        description: Return the target's Set-Cookie values.
        schema:
          type: boolean
      - name: cookies_session
        in: query
        required: false
        description: Sticky session identifier to reuse cookies across requests.
        schema:
          type: string
      - name: store
        in: query
        required: false
        description: Persist the response in Crawlbase Cloud Storage.
        schema:
          type: boolean
      - name: async
        in: query
        required: false
        description: Queue the request and return immediately with a request id (rid).
        schema:
          type: boolean
      - name: callback
        in: query
        required: false
        description: Webhook URL that Crawlbase posts async results to.
        schema:
          type: string
          format: uri
      responses:
        '200':
          description: The crawled page. Body is raw upstream content (HTML, Markdown, JSON, or image) depending on `format`, `scraper`, and `screenshot`. Crawlbase-specific metadata (pc_status, original_status, rid, url) is returned in response headers, or inline when `format=json`.
          content:
            text/html:
              schema:
                type: string
            application/json:
              schema:
                type: object
                additionalProperties: true
        '401':
          $ref: '#/components/responses/Unauthorized'
    post:
      operationId: crawlPost
      tags:
      - Crawling API
      summary: Crawl a URL (POST)
      description: Same as the GET crawl operation, but the request payload (for example, POST form data to submit to the target) is sent in the request body. The `token` and `url` remain query parameters.
      parameters:
      - $ref: '#/components/parameters/Url'
      - $ref: '#/components/parameters/Format'
      - $ref: '#/components/parameters/Country'
      - $ref: '#/components/parameters/Device'
      requestBody:
        required: false
        content:
          application/x-www-form-urlencoded:
            schema:
              type: object
              additionalProperties: true
      responses:
        '200':
          description: The crawled page.
          content:
            text/html:
              schema:
                type: string
            application/json:
              schema:
                type: object
                additionalProperties: true
        '401':
          $ref: '#/components/responses/Unauthorized'
components:
  parameters:
    Url:
      name: url
      in: query
      required: true
      description: Fully URL-encoded target URL, including http:// or https://.
      schema:
        type: string
    Scraper:
      name: scraper
      in: query
      required: false
      description: Name of a built-in extractor to apply to the crawled page (returns parsed JSON).
      schema:
        type: string
    UserAgent:
      name: user_agent
      in: query
      required: false
      description: Override the User-Agent header sent to the target.
      schema:
        type: string
    Format:
      name: format
      in: query
      required: false
      description: Response format - html (default), json, or md (Markdown).
      schema:
        type: string
        enum:
        - html
        - json
        - md
        default: html
    Device:
      name: device
      in: query
      required: false
      description: Device profile - desktop (default), tablet, or mobile.
      schema:
        type: string
        enum:
        - desktop
        - tablet
        - mobile
        default: desktop
    Country:
      name: country
      in: query
      required: false
      description: Two-letter ISO 3166 country code to geolocate the request (for example US, DE, GB).
      schema:
        type: string
  responses:
    Unauthorized:
      description: Missing or invalid token.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
  schemas:
    Error:
      type: object
      properties:
        pc_status:
          type: integer
        error:
          type: string
  securitySchemes:
    tokenAuth:
      type: apiKey
      in: query
      name: token
      description: 'Crawlbase authentication token passed as the `token` query parameter. Each account has two tokens: a Normal (TCP) token for static content and a JavaScript token for headless-Chrome rendering.'