Diffbot · OpenAPI Overlay 1.0.0

API Evangelist conversational phrasing for Diffbot Crawl API

6 actions 6 updates phrasing extends openapi/diffbot-crawl-api-openapi.yml
Generated by API Evangelist Written by API Evangelist tooling for Diffbot's API. It is a proposal applied on top of the contract, not a document Diffbot publishes.
View Overlay File View on GitHub Overlay Specification

What the actions change

x-apievangelist-phrasing

Targets 6

$.info
$.paths['/crawl'].get
$.paths['/crawl'].post
$.paths['/crawl/data'].get
$.paths['/v3/crawl'].get
$.paths['/v3/crawl/data'].get

OpenAPI Overlay

Raw ↑
# Generated by API Evangelist (build-phrasing.py). Our phrasing, not observed demand.
overlay: 1.0.0
info:
  title: API Evangelist conversational phrasing for Diffbot Crawl API
  version: 1.0.0
extends: openapi/diffbot-crawl-api-openapi.yml
actions:
- target: $.info
  update:
    x-apievangelist-phrasing:
      method: generated
      generated: '2026-09-26'
      generator: build-phrasing.py
      label: Generated by API Evangelist
      operations: 5
- target: $.paths['/crawl'].get
  update:
    x-apievangelist-phrasing:
      intent: Pause, restart, delete or check a crawl job
      effect: write
      questions:
      - How do I pause a crawl that's already spidering a site?
      - Can I manually kick off a new round of a repeating crawl?
      - What is the status of my crawl job right now?
      instructions:
      - text: Pause crawl job {name}.
        slots:
          name: query.name
      - text: Start a new crawl round for {name} now.
        slots:
          name: query.name
      - text: Delete crawl job {name} and everything it collected.
        slots:
          name: query.name
      method: generated
      generated: '2026-09-26'
- target: $.paths['/crawl'].post
  update:
    x-apievangelist-phrasing:
      intent: Create a crawl that spiders and extracts a site
      effect: write
      questions:
      - How do I spider an entire website and extract structured data from every page?
      - Can I limit a site crawl to URLs matching a pattern or a maximum depth?
      - Is there a way to have a crawl repeat automatically every few days?
      instructions:
      - text: Create crawl {name} starting at {seeds} and process pages with {apiUrl}.
        slots:
          name: requestBody.name
          seeds: requestBody.seeds
          apiUrl: requestBody.apiUrl
      - text: Crawl {seeds} as job {name} via {apiUrl}, only processing URLs containing {urlProcessPattern}.
        slots:
          seeds: requestBody.seeds
          name: requestBody.name
          apiUrl: requestBody.apiUrl
          urlProcessPattern: requestBody.urlProcessPattern
      - text: Start crawl {name} from {seeds} using {apiUrl} with a max of {maxToCrawl} pages.
        slots:
          name: requestBody.name
          seeds: requestBody.seeds
          apiUrl: requestBody.apiUrl
          maxToCrawl: requestBody.maxToCrawl
      method: generated
      generated: '2026-09-26'
- target: $.paths['/crawl/data'].get
  update:
    x-apievangelist-phrasing:
      intent: Download the results of a crawl job
      effect: read
      questions:
      - Where do I download the pages a crawl job extracted?
      - Can I export crawl results to CSV?
      - Is there a report of every URL a crawl visited?
      instructions:
      - text: Download the extracted data from crawl job {name}.
        slots:
          name: query.name
      - text: Export crawl {name} results in {format} format.
        slots:
          name: query.name
          format: query.format
      - text: Get the URL report for crawl {name}.
        slots:
          name: query.name
      method: generated
      generated: '2026-09-26'
- target: $.paths['/v3/crawl'].get
  update:
    x-apievangelist-phrasing:
      intent: Manage a crawl via the v3 token endpoint
      effect: write
      questions:
      - Can I create or check a crawl with a single GET request on the v3 path?
      - Is there a v3 crawl endpoint that takes seeds and a job name in the query string?
      instructions:
      - text: Using the v3 crawl endpoint, start job {name} with seeds {seeds}.
        slots:
          name: query.name
          seeds: query.seeds
      - text: Check crawl {name} through the v3 crawl endpoint.
        slots:
          name: query.name
      method: generated
      generated: '2026-09-26'
- target: $.paths['/v3/crawl/data'].get
  update:
    x-apievangelist-phrasing:
      intent: Fetch crawl results via the v3 endpoint
      effect: read
      questions:
      - Can I pull a crawl's extracted results from the v3 crawl data path?
      - Is there a v3 endpoint that returns crawl output by job name?
      instructions:
      - text: Fetch the v3 crawl data for job {name}.
        slots:
          name: query.name
      - text: Pull extraction results for crawl {name} from the v3 data endpoint.
        slots:
          name: query.name
      method: generated
      generated: '2026-09-26'