Diffbot · OpenAPI Overlay 1.0.0
API Evangelist conversational phrasing for Diffbot Crawl API
6 actions
6 updates
phrasing
extends
openapi/diffbot-crawl-api-openapi.yml
Generated by API Evangelist
Written by API Evangelist tooling for Diffbot's API. It is a proposal applied on top of the contract, not a document Diffbot publishes.
What the actions change
x-apievangelist-phrasing
Targets 6
$.info
$.paths['/crawl'].get
$.paths['/crawl'].post
$.paths['/crawl/data'].get
$.paths['/v3/crawl'].get
$.paths['/v3/crawl/data'].get
OpenAPI Overlay
# Generated by API Evangelist (build-phrasing.py). Our phrasing, not observed demand.
overlay: 1.0.0
info:
title: API Evangelist conversational phrasing for Diffbot Crawl API
version: 1.0.0
extends: openapi/diffbot-crawl-api-openapi.yml
actions:
- target: $.info
update:
x-apievangelist-phrasing:
method: generated
generated: '2026-09-26'
generator: build-phrasing.py
label: Generated by API Evangelist
operations: 5
- target: $.paths['/crawl'].get
update:
x-apievangelist-phrasing:
intent: Pause, restart, delete or check a crawl job
effect: write
questions:
- How do I pause a crawl that's already spidering a site?
- Can I manually kick off a new round of a repeating crawl?
- What is the status of my crawl job right now?
instructions:
- text: Pause crawl job {name}.
slots:
name: query.name
- text: Start a new crawl round for {name} now.
slots:
name: query.name
- text: Delete crawl job {name} and everything it collected.
slots:
name: query.name
method: generated
generated: '2026-09-26'
- target: $.paths['/crawl'].post
update:
x-apievangelist-phrasing:
intent: Create a crawl that spiders and extracts a site
effect: write
questions:
- How do I spider an entire website and extract structured data from every page?
- Can I limit a site crawl to URLs matching a pattern or a maximum depth?
- Is there a way to have a crawl repeat automatically every few days?
instructions:
- text: Create crawl {name} starting at {seeds} and process pages with {apiUrl}.
slots:
name: requestBody.name
seeds: requestBody.seeds
apiUrl: requestBody.apiUrl
- text: Crawl {seeds} as job {name} via {apiUrl}, only processing URLs containing {urlProcessPattern}.
slots:
seeds: requestBody.seeds
name: requestBody.name
apiUrl: requestBody.apiUrl
urlProcessPattern: requestBody.urlProcessPattern
- text: Start crawl {name} from {seeds} using {apiUrl} with a max of {maxToCrawl} pages.
slots:
name: requestBody.name
seeds: requestBody.seeds
apiUrl: requestBody.apiUrl
maxToCrawl: requestBody.maxToCrawl
method: generated
generated: '2026-09-26'
- target: $.paths['/crawl/data'].get
update:
x-apievangelist-phrasing:
intent: Download the results of a crawl job
effect: read
questions:
- Where do I download the pages a crawl job extracted?
- Can I export crawl results to CSV?
- Is there a report of every URL a crawl visited?
instructions:
- text: Download the extracted data from crawl job {name}.
slots:
name: query.name
- text: Export crawl {name} results in {format} format.
slots:
name: query.name
format: query.format
- text: Get the URL report for crawl {name}.
slots:
name: query.name
method: generated
generated: '2026-09-26'
- target: $.paths['/v3/crawl'].get
update:
x-apievangelist-phrasing:
intent: Manage a crawl via the v3 token endpoint
effect: write
questions:
- Can I create or check a crawl with a single GET request on the v3 path?
- Is there a v3 crawl endpoint that takes seeds and a job name in the query string?
instructions:
- text: Using the v3 crawl endpoint, start job {name} with seeds {seeds}.
slots:
name: query.name
seeds: query.seeds
- text: Check crawl {name} through the v3 crawl endpoint.
slots:
name: query.name
method: generated
generated: '2026-09-26'
- target: $.paths['/v3/crawl/data'].get
update:
x-apievangelist-phrasing:
intent: Fetch crawl results via the v3 endpoint
effect: read
questions:
- Can I pull a crawl's extracted results from the v3 crawl data path?
- Is there a v3 endpoint that returns crawl output by job name?
instructions:
- text: Fetch the v3 crawl data for job {name}.
slots:
name: query.name
- text: Pull extraction results for crawl {name} from the v3 data endpoint.
slots:
name: query.name
method: generated
generated: '2026-09-26'