Search1API Crawl API

The Crawl API from Search1API — 5 operation(s) for crawl.

Operations 5

POST /crawl Crawl a URL and extract its content #
POST /sitemap Extract sitemap URLs from a website #
POST /extract Extract structured content from a URL #
POST /deepcrawl Deep crawl a website across multiple pages #
GET /deepcrawl/status/{taskId} Check deepcrawl task status #

Documentation

Specifications

Schemas & Data

Other Resources

Work with this as data

Every API here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for apis

7 MCP tools reach this
  • find_apisBrowse and filter every API in the catalog.
  • get_api_artifactsOne API's artifacts, grouped by type.
  • get_openapiThe primary OpenAPI for this API.
  • find_similar_apisAPIs that look like this one.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools →

Call it yourself

curl for this page
This API
curl "https://apis.io/api/v1/apis/s1-dev:s1-dev-crawl-api"
All apis
curl "https://apis.io/api/v1/apis?limit=25"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no form to fill in. Signing in shares your email address with us — we store it to create your key and to recognise you if you sign in with another provider. See our Privacy Policy and Terms.

A second provider on the same verified email joins the account you already have.

OpenAPI Specification

s1-dev-crawl-api-openapi.yml Raw ↑
openapi: 3.2.0
info:
  title: S1 Dev Crawl API
  version: 1.0.0
  x-guidance: Search1API provides web search, news aggregation, URL crawling, webpage screenshots, sitemap extraction, trending topics, content extraction, and deep crawling. All paid endpoints accept POST with a JSON body. Use POST /search with a "query" field for web search. Use POST /news with a "query" field for news. Use POST /ask with a natural-language "query" to have Search1API choose the engines and time window and return only relevant results (API key only). Use POST /crawl with a "url" field to crawl a page. Use POST /screenshot with a "url" field to render a PNG, JPEG, or WebP image. Use POST /sitemap with a "url" field to extract sitemap URLs. Use POST /trending with a "search_service" field for trends. Use POST /extract with a "url" field for structured content extraction. Use POST /deepcrawl with a "url" field for deep multi-page crawling.
  description: 'Operations tagged Crawl across 2 of this provider''s published API definitions: search1api-openapi.json, s1-dev-openapi.yml. Each path carries the servers of the definition it was published in.'
servers:
- url: https://api.search1api.com
tags:
- name: Crawl
paths:
  /crawl:
    post:
      operationId: crawl
      summary: Crawl a URL and extract its content
      description: 'Fetch one public URL and return its readable title and body as clean text, with navigation, boilerplate, and scripts stripped. Use it on a URL the user supplied or one returned by POST /search. To ingest a whole site rather than a single page, use POST /deepcrawl. Send an array of request objects to crawl several URLs in one call: the response is an array containing only the URLs that succeeded, in no guaranteed order, so match items back to your input by `crawlParameters.url`; failed URLs are omitted and not charged. `enable_fallback` is accepted as a snake_case alias of `enableFallback`. Costs 1 credit per request.'
      tags:
      - Crawl
      x-codeSamples:
      - id: js
        lang: ts
        label: TypeScript SDK
        source: 'import { Search1API } from ''@search1api/client'';


          const client = new Search1API();

          const response = await client.crawl(''https://example.com'');


          console.log(response);'
      - id: python
        lang: python
        label: Python SDK
        source: 'from search1api import Search1API


          client = Search1API()

          response = client.crawl("https://example.com")


          print(response)'
      responses:
        '200':
          description: Successful response
          content:
            application/json:
              schema:
                oneOf:
                - $ref: '#/components/schemas/CrawlResponse'
                - type: array
                  items:
                    $ref: '#/components/schemas/CrawlResponse'
        '400':
          description: Bad Request
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
        '401':
          description: Unauthorized
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
        '402':
          description: Payment Required
          content:
            application/problem+json:
              schema:
                $ref: '#/components/schemas/ApiError'
        '404':
          description: Not Found
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
        '410':
          description: Gone
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
        '422':
          description: Validation Error
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
        '429':
          description: Too Many Requests
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
        '500':
          description: Internal Server Error
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
        '502':
          description: Bad Gateway
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
      security:
      - bearerAuth: []
      x-payment-info:
        protocols:
        - mpp
        pricingMode: fixed
        price: '0.003'
      requestBody:
        required: true
        content:
          application/json:
            schema:
              anyOf:
              - type: object
                properties:
                  url:
                    type: string
                    format: uri
                    minLength: 1
                  enableFallback:
                    type: boolean
                    default: true
                required:
                - url
                additionalProperties: false
              - type: array
                items:
                  type: object
                  properties:
                    url:
                      type: string
                      format: uri
                      minLength: 1
                    enableFallback:
                      type: boolean
                      default: true
                  required:
                  - url
                  additionalProperties: false
                minItems: 1
    servers:
    - url: https://api.search1api.com
  /sitemap:
    post:
      operationId: sitemap
      summary: Extract sitemap URLs from a website
      description: 'Discover the public URLs of a site when you need to know which pages exist before fetching any of them — scoping a crawl, auditing coverage, or locating a section. With the default `type: "sitemap"` this reads the site''s own sitemap: `Sitemap:` directives in robots.txt first, then the conventional paths, following sitemap index files and gzipped sitemaps, and returns the URLs it declares for the requested host. A site that publishes no sitemap falls back to the links found on the given page. Use `type: "all"` to skip the sitemap and get every link on that one page instead, including links to other hosts. Neither mode fetches page content. Costs 1 credit per request.'
      tags:
      - Crawl
      x-codeSamples:
      - id: js
        lang: ts
        label: TypeScript SDK
        source: 'import { Search1API } from ''@search1api/client'';


          const client = new Search1API();

          const response = await client.sitemap(''https://example.com'');


          console.log(response);'
      - id: python
        lang: python
        label: Python SDK
        source: 'from search1api import Search1API


          client = Search1API()

          response = client.sitemap("https://example.com")


          print(response)'
      responses:
        '200':
          description: Successful response
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/SitemapResponse'
        '400':
          description: Bad Request
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
        '401':
          description: Unauthorized
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
        '402':
          description: Payment Required
          content:
            application/problem+json:
              schema:
                $ref: '#/components/schemas/ApiError'
        '422':
          description: Validation Error
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
        '429':
          description: Too Many Requests
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
        '500':
          description: Internal Server Error
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
        '502':
          description: Bad Gateway
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
      security:
      - bearerAuth: []
      x-payment-info:
        protocols:
        - mpp
        pricingMode: fixed
        price: '0.003'
      requestBody:
        required: true
        content:
          application/json:
            schema:
              type: object
              properties:
                url:
                  type: string
                  format: uri
                type:
                  type: string
                  enum:
                  - sitemap
                  - all
              required:
              - url
              additionalProperties: false
    servers:
    - url: https://api.search1api.com
  /extract:
    post:
      operationId: extract
      summary: Extract structured content from a URL
      description: Pull structured data out of a single webpage using a natural-language prompt and a JSON schema you supply. Use it when you need specific fields — prices, specifications, contact details — rather than the whole document; when you want the full text, use POST /crawl. Costs 10 credits per request.
      tags:
      - Crawl
      x-codeSamples:
      - id: js
        lang: ts
        label: TypeScript SDK
        source: "import { Search1API } from '@search1api/client';\n\nconst client = new Search1API();\nconst response = await client.extract('https://example.com', {\n  prompt: 'Extract the page title and description.',\n});\n\nconsole.log(response);"
      - id: python
        lang: python
        label: Python SDK
        source: "from search1api import Search1API\n\nclient = Search1API()\nresponse = client.extract(\n    \"https://example.com\",\n    prompt=\"Extract the page title and description.\",\n)\n\nprint(response)"
      responses:
        '200':
          description: Successful response
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ExtractResponse'
        '400':
          description: Bad Request
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
        '401':
          description: Unauthorized
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
        '402':
          description: Payment Required
          content:
            application/problem+json:
              schema:
                $ref: '#/components/schemas/ApiError'
        '422':
          description: Validation Error
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
        '429':
          description: Too Many Requests
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
        '500':
          description: Internal Server Error
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
        '502':
          description: Bad Gateway
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
      security:
      - bearerAuth: []
      x-payment-info:
        protocols:
        - mpp
        pricingMode: fixed
        price: '0.03'
      requestBody:
        required: true
        content:
          application/json:
            schema:
              type: object
              properties:
                url:
                  type: string
                  format: uri
                prompt:
                  type: string
                response_format:
                  type: object
                  additionalProperties: {}
              required:
              - url
              additionalProperties: false
    servers:
    - url: https://api.search1api.com
  /deepcrawl:
    post:
      operationId: deepcrawl
      summary: Deep crawl a website across multiple pages
      description: Start an asynchronous crawl of an entire site and package the pages as documents. Use it for whole-site ingestion; a single page is POST /crawl. The call returns a task id immediately rather than the result — poll GET /deepcrawl/status/{taskId} until the task reports completion. Costs 20 credits per request.
      tags:
      - Crawl
      x-codeSamples:
      - id: js
        lang: ts
        label: TypeScript SDK
        source: "import { Search1API } from '@search1api/client';\n\nconst client = new Search1API();\nconst task = await client.startDeepcrawl('https://example.com', {\n  type: 'all',\n});\n\nconsole.log(task.taskId);"
      - id: python
        lang: python
        label: Python SDK
        source: 'from search1api import Search1API


          client = Search1API()

          task = client.start_deepcrawl("https://example.com", type="all")


          print(task["taskId"])'
      responses:
        '202':
          description: Successful response
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/DeepcrawlAcceptedResponse'
        '400':
          description: Bad Request
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
        '401':
          description: Unauthorized
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
        '402':
          description: Payment Required
          content:
            application/problem+json:
              schema:
                $ref: '#/components/schemas/ApiError'
        '422':
          description: Validation Error
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
        '429':
          description: Too Many Requests
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
        '500':
          description: Internal Server Error
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
        '502':
          description: Bad Gateway
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
      security:
      - bearerAuth: []
      x-payment-info:
        protocols:
        - mpp
        pricingMode: fixed
        price: '0.06'
      requestBody:
        required: true
        content:
          application/json:
            schema:
              type: object
              properties:
                url:
                  type: string
                  format: uri
                type:
                  type: string
                  enum:
                  - sitemap
                  - all
              required:
              - url
              additionalProperties: false
    servers:
    - url: https://api.search1api.com
  /deepcrawl/status/{taskId}:
    get:
      operationId: deepcrawlStatus
      summary: Check deepcrawl task status
      description: Poll a deepcrawl task started by POST /deepcrawl. Returns the task's current state and, once it finishes, where to retrieve the packaged result. Safe to call repeatedly. Free to call.
      tags:
      - Crawl
      x-codeSamples:
      - id: js
        lang: ts
        label: TypeScript SDK
        source: 'import { Search1API } from ''@search1api/client'';


          const client = new Search1API();

          const status = await client.getDeepcrawlStatus(''task_id'');


          console.log(status);'
      - id: python
        lang: python
        label: Python SDK
        source: 'from search1api import Search1API


          client = Search1API()

          status = client.get_deepcrawl_status("task_id")


          print(status)'
      responses:
        '200':
          description: Successful response
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/DeepcrawlStatusResponse'
        '400':
          description: Bad Request
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
        '404':
          description: Not Found
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
        '500':
          description: Internal Server Error
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
        '502':
          description: Bad Gateway
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiError'
      security:
      - bearerAuth: []
      parameters:
      - in: path
        name: taskId
        required: true
        schema:
          type: string
          minLength: 1
    servers:
    - url: https://api.search1api.com
components:
  schemas:
    SitemapResponse:
      type: object
      required:
      - links
      properties:
        links:
          type: array
          items:
            type: string
            format: uri
    ExtractResponse:
      type: object
      required:
      - success
      - extractParameters
      - results
      properties:
        success:
          type: boolean
          enum:
          - true
        extractParameters:
          type: object
          required:
          - url
          properties:
            url:
              type: string
              format: uri
        results: {}
    DeepcrawlStatusResponse:
      type: object
      required:
      - taskId
      properties:
        taskId:
          type: string
        status:
          type: string
          enum:
          - queued
          - waiting
          - processing
          - completed
          - failed
          - not_found
        success:
          type: boolean
        message:
          type: string
        error:
          type:
          - string
          - 'null'
        r2Key:
          type: string
        zipUrl:
          type:
          - string
          - 'null'
          format: uri
      additionalProperties: true
    CrawlResult:
      type: object
      required:
      - title
      - link
      - content
      properties:
        title:
          type: string
        link:
          type: string
          format: uri
        content:
          type: string
        metadata:
          type: object
          additionalProperties: true
      additionalProperties: true
    CrawlResponse:
      type: object
      required:
      - crawlParameters
      - results
      properties:
        crawlParameters:
          type: object
          required:
          - url
          properties:
            url:
              type: string
              format: uri
        results:
          $ref: '#/components/schemas/CrawlResult'
    DeepcrawlAcceptedResponse:
      type: object
      required:
      - taskId
      - status
      properties:
        taskId:
          type: string
        status:
          type: string
          enum:
          - queued
    ApiError:
      type: object
      description: 'Search1API error. Every JSON error carries `ok: false`, `error` (a short status-derived label) and `message` (human-readable detail); validation failures add `errors`. Payment challenges may use RFC 9457 problem detail fields.'
      properties:
        ok:
          type: boolean
          enum:
          - false
        error:
          type: string
        message:
          type: string
        errors:
          type: array
          items:
            type: object
            properties:
              field:
                type: string
              message:
                type: string
              code:
                type: string
        type:
          type: string
          format: uri
        title:
          type: string
        status:
          type: integer
        detail:
          type: string
      additionalProperties: true
  securitySchemes:
    bearerAuth:
      type: http
      scheme: bearer
      bearerFormat: API Key
x-refined-from:
- search1api-openapi.json
- s1-dev-openapi.yml
x-discovery:
  ownershipProofs: []