Diffbot website screenshot

Diffbot

Diffbot is a company that provides AI-powered web scraping and data extraction services. Their technology allows businesses to automatically extract and organize data from any website, turning unstructured web content into structured data that can be easily analyzed and used for various purposes. Diffbot's solution is used by companies across industries to gather competitive intelligence, monitor market trends, track online mentions, and more.

Diffbot publishes 9 APIs on the APIs.io network, including Extract API, Crawl API, Bulk Extract API, and 6 more. Tagged areas include Extraction, Harvesting, Scraping, Web, and Knowledge Graph.

The Diffbot catalog on APIs.io includes 1 event-driven AsyncAPI specification.

Diffbot’s developer surface includes changelog, CLI, sandbox, developer console, API reference, getting-started guide, authentication, and 44 more developer resources.

71.4/100 exemplar ▬ flat Agent 42/100 agent ready saas Full breakdown ↓
scored 2026-09-08 · rubric v0.20.0
AccessFreemiumSelf serve⚡ Free to try
9 APIs 1 MCP Servers
ExtractionHarvestingScrapingWebKnowledge GraphCrawlingWeb SearchNatural LanguageEntity ResolutionAI

Kin Score

Kin Score Kin Score How this is scored →
scored 2026-09-08 · rubric v0.20.0
Create-or-Update Ergonomics applies to this provider. This API accepts writes, so it carries 10 points of the composite. It is scored from the published contracts themselves: whether a caller can create-or-update in one call, whether the write accepts a key the caller already holds, and whether the response says which branch ran. Without that, every write needs a search-and-branch in front of it, and the first time that check is skipped a duplicate record is created. Scored against the observed mean rather than raw — a provider at the catalog average is unchanged by this facet, not penalised by it.
Improve this rating by publishing the missing artifacts — every area above can be raised, and the full rubric is at apis.io/rating/. Every facet and dimension name above is a link: it opens that measurement's own page — what it means, the exact checks that feed it, how the whole catalog distributes on it, and the providers at the top of it. This rating is computed from github.com/api-evangelist/diffbot: open an issue to ask a question, or submit a pull request to add artifacts. Submit an artifact on GitHub — free → Manage your own listing — the Influence plan, $499/mo →

APIs 10

Individual APIs this provider publishes, each with its own machine-readable definition.

Diffbot Extract API

Diffbot Extract API is a powerful tool that allows users to automatically extract multiple types of data from web pages. This API is capable of extracting information such as ar...

Diffbot Crawl API

Diffbot Crawl API is a powerful tool that automates the process of extracting content and data from websites on a large scale. By using advanced machine learning algorithms, the...

Diffbot Bulk Extract API

Diffbot Bulk Extract API is a tool that allows users to extract data at scale from a variety of sources, including websites, documents, and social media platforms. This API util...

Diffbot DQL API

The Diffbot DQL API is a powerful tool that allows users to query and retrieve data from the web in a structured format. By using a simple query language, users can access a wea...

Diffbot Enhance API

Diffbot Enhance API enhances data by providing additional context and insights. By analyzing text and images, the API can identify and extract key information, such as entities,...

Diffbot Natural Language API

Diffbot Natural Language API allows users to extract and analyze textual content from websites. By utilizing advanced natural language processing algorithms, the API can automat...

Diffbot Web Search API

Search Diffbot's own web index — the largest independently crawled index outside Google and Bing, over 150TB. Candidates are retrieved and reranked by a cross-encoder trained to...

Diffbot Account API

Retrieve account details, plan, token metadata and usage activity for a Diffbot token. The only introspection surface Diffbot publishes: because the APIs return no rate-limit or...

Diffbot Knowledge Graph API

The Knowledge Graph itself — a linked graph of over 10 billion entities (organizations, people, articles, places, products, job posts and more) crawled and structured from the p...

Diffbot Crawl/Bulk Job API

The Diffbot Crawl/Bulk Job API is a powerful tool that allows users to automatically extract and organize large amounts of web data. It enables users to create custom scraping j...

Scroll for all 10

Open Collections 6

Open, tool-agnostic API collections (OpenAPI-derived and Bruno).

API Collection

OPEN COLLECTION

Diffbot Crawl API

OPEN COLLECTION

Diffbot Crawl Extract API

OPEN COLLECTION

Diffbot API

OPEN COLLECTION

MCP Servers 1

Model Context Protocol servers that expose these APIs to AI agents.

Diffbot MCP Server

Diffbot's first-party MCP server. Exposes the Extract, Web Search, Knowledge Graph (Enhance / DQL), Crawl and Natural Language surfaces as seven agent tools. Diffbot operates a ...

MCP SERVER

Pricing Plans 1

Published pricing tiers and plan structures.

Diffbot Plans Pricing

4 plans

PLANS

Rate Limits 1

Documented rate limits and quota policies.

Diffbot Rate Limits

12 limits

RATE LIMITS

FinOps 1

Cost, billing, and metering signals for API financial operations.

Event Specifications 1

AsyncAPI definitions for this provider's event-driven and streaming APIs.

Security Posture 2

Authentication, domain security, vulnerability disclosure, and trust-center signals.

Diffbot Authentication

apiKey/http · 3 schemes

SECURITY

Diffbot Domain Security

TLSv1.3 · HSTS · DMARC

SECURITY

Agentic Access 1

Recommended x-agentic-access execution contracts for AI agents.

Diffbot Agentic Access

14 operations · 1 acting

14 operations · 1 acting

AGENTIC

Resources

Get Started 6

Portal, sign-up, and the first successful call

Documentation 3

Reference material describing how the API behaves

Agent Surfaces 5

MCP servers, agent skills, and machine-readable catalogs

Design & Contract 6

Pagination, idempotency, versioning, errors, and events

Build 6

SDKs, sample code, and the tooling you integrate with

Access & Security 4

Authentication, authorization, and security posture

Learn 1

Tutorials, courses, talks, and written guidance

Operate 6

Status, limits, changes, and where to get help

Commercial 5

Pricing, plans, and the legal terms of use

Company 5

The organization behind the API

Other 4

Properties that don't map to a standard resource type

Source (apis.yml)

apis.yml Raw ↑
aid: diffbot
url: https://raw.githubusercontent.com/api-search/diffbot/refs/heads/main/apis.yml
apis:
- aid: diffbot:diffbot-extract-api
  name: Diffbot Extract API
  image: https://kinlane-images.s3.amazonaws.com/shared/apis-json/icons/diffbot.png
  baseURL: https://api.diffbot.com/v3
  humanURL: https://www.diffbot.com/docs/extract/
  tags:
  - Extraction
  - Harvesting
  - Scraping
  - Web
  properties:
  - type: OpenAPI
    url: openapi/_original/diffbot-extract-openapi.json
  - type: OpenAPI
    name: Refined one-per-tag split
    url: openapi/diffbot-extract-api-openapi.yml
  - type: Overlay
    url: overlays/diffbot-extract-overlay.yaml
  - type: Documentation
    url: https://www.diffbot.com/docs/extract/
  - type: APIReference
    url: https://www.diffbot.com/docs/extract/analyze
  description: Diffbot Extract API is a powerful tool that allows users to automatically extract multiple types of data from
    web pages. This API is capable of extracting information such as article text, author details, images, and even product
    prices from a variety of websites.
- aid: diffbot:diffbot-crawl-api
  name: Diffbot Crawl API
  image: https://kinlane-images.s3.amazonaws.com/shared/apis-json/icons/diffbot.png
  baseURL: https://api.diffbot.com/v3
  humanURL: https://www.diffbot.com/docs/crawl/
  tags:
  - Crawling
  - Extraction
  - Harvesting
  - Scraping
  properties:
  - type: OpenAPI
    url: openapi/_original/diffbot-crawl-openapi.json
  - type: OpenAPI
    name: Refined one-per-tag split
    url: openapi/diffbot-crawl-api-openapi.yml
  - type: Overlay
    url: overlays/diffbot-crawl-overlay.yaml
  - type: Documentation
    url: https://www.diffbot.com/docs/crawl/
  - type: APIReference
    url: https://www.diffbot.com/docs/crawl/create
  description: Diffbot Crawl API is a powerful tool that automates the process of extracting content and data from websites
    on a large scale. By using advanced machine learning algorithms, the API can analyze and extract information from web
    pages with speed and accuracy.
- aid: diffbot:diffbot-bulk-extract-api
  name: Diffbot Bulk Extract API
  image: https://kinlane-images.s3.amazonaws.com/shared/apis-json/icons/diffbot.png
  baseURL: https://api.diffbot.com/v3
  humanURL: https://www.diffbot.com/docs/bulk/
  tags:
  - Bulk
  - Extraction
  - Harvesting
  - Scraping
  properties:
  - type: OpenAPI
    url: openapi/_original/diffbot-bulk-openapi.json
  - type: Overlay
    url: overlays/diffbot-bulk-overlay.yaml
  - type: Documentation
    url: https://www.diffbot.com/docs/bulk/
  - type: APIReference
    url: https://www.diffbot.com/docs/bulk/create
  description: Diffbot Bulk Extract API is a tool that allows users to extract data at scale from a variety of sources, including
    websites, documents, and social media platforms. This API utilizes machine learning algorithms to automatically identify
    and extract relevant information from large amounts of unstructured data.
- aid: diffbot:diffbot-dql-api
  name: Diffbot DQL API
  image: https://kinlane-images.s3.amazonaws.com/shared/apis-json/icons/diffbot.png
  baseURL: https://kg.diffbot.com
  humanURL: https://www.diffbot.com/docs/dql/
  tags:
  - Extraction
  - Harvesting
  - Knowledge Graph
  - Query Language
  - Scraping
  properties:
  - type: OpenAPI
    url: openapi/_original/diffbot-dql-openapi.json
  - type: Overlay
    url: overlays/diffbot-dql-overlay.yaml
  - type: Documentation
    url: https://www.diffbot.com/docs/dql/
  - type: APIReference
    url: https://www.diffbot.com/docs/dql/get
  description: The Diffbot DQL API is a powerful tool that allows users to query and retrieve data from the web in a structured
    format. By using a simple query language, users can access a wealth of information from articles, images, videos, and
    more. This API is particularly useful for developers and data analysts who need to extract specific data from websites
    quickly and efficiently.
- aid: diffbot:diffbot-enhance-api
  name: Diffbot Enhance API
  image: https://kinlane-images.s3.amazonaws.com/shared/apis-json/icons/diffbot.png
  baseURL: https://kg.diffbot.com
  humanURL: https://www.diffbot.com/docs/enhance/
  tags:
  - Enhancements
  - Extraction
  - Harvesting
  - Knowledge Graph
  - Scraping
  properties:
  - type: OpenAPI
    url: openapi/_original/diffbot-enhance-openapi.json
  - type: Overlay
    url: overlays/diffbot-enhance-overlay.yaml
  - type: Documentation
    url: https://www.diffbot.com/docs/enhance/
  - type: APIReference
    url: https://www.diffbot.com/docs/enhance/get
  description: Diffbot Enhance API enhances data by providing additional context and insights. By analyzing text and images,
    the API can identify and extract key information, such as entities, topics, and sentiment, to enrich existing data sets.
- aid: diffbot:diffbot-natural-language-api
  name: Diffbot Natural Language API
  image: https://kinlane-images.s3.amazonaws.com/shared/apis-json/icons/diffbot.png
  baseURL: https://nl.diffbot.com
  humanURL: https://www.diffbot.com/docs/natural-language/
  tags:
  - Extraction
  - Harvesting
  - Language
  - Scraping
  properties:
  - type: OpenAPI
    url: openapi/_original/diffbot-natural-language-openapi.json
  - type: OpenAPI
    name: Refined one-per-tag split
    url: openapi/diffbot-natural-language-api-openapi.yml
  - type: Overlay
    url: overlays/diffbot-natural-language-overlay.yaml
  - type: Documentation
    url: https://www.diffbot.com/docs/natural-language/
  - type: APIReference
    url: https://www.diffbot.com/docs/natural-language/process-text
  description: Diffbot Natural Language API allows users to extract and analyze textual content from websites. By utilizing
    advanced natural language processing algorithms, the API can automatically identify and extract valuable information such
    as article titles, authors, publication dates, and key entities mentioned in the text.
- aid: diffbot:diffbot-web-search-api
  name: Diffbot Web Search API
  image: https://kinlane-images.s3.amazonaws.com/shared/apis-json/icons/diffbot.png
  baseURL: https://llm.diffbot.com/api/v1/web_search
  humanURL: https://www.diffbot.com/docs/web-search/
  tags:
  - Search
  - Web
  - Index
  - Retrieval
  properties:
  - type: OpenAPI
    url: openapi/_original/diffbot-web-search-openapi.json
  - type: Overlay
    url: overlays/diffbot-web-search-overlay.yaml
  - type: Documentation
    url: https://www.diffbot.com/docs/web-search/
  - type: APIReference
    url: https://www.diffbot.com/docs/web-search/get
  description: Search Diffbot's own web index — the largest independently crawled index outside Google and Bing, over 150TB.
    Candidates are retrieved and reranked by a cross-encoder trained to favour factual relevance over popularity and primary
    sources over secondary ones, and the most relevant chunk of each page is returned inline for citation. This is the one
    Diffbot surface that authenticates with an Authorization Bearer header rather than a token query parameter.
- aid: diffbot:diffbot-account-api
  name: Diffbot Account API
  image: https://kinlane-images.s3.amazonaws.com/shared/apis-json/icons/diffbot.png
  baseURL: https://api.diffbot.com/v4
  humanURL: https://www.diffbot.com/docs/account
  tags:
  - Account
  - Billing
  - Usage
  - Metering
  properties:
  - type: OpenAPI
    url: openapi/_original/diffbot-account-openapi.json
  - type: Overlay
    url: overlays/diffbot-account-overlay.yaml
  - type: Documentation
    url: https://www.diffbot.com/docs/account
  - type: APIReference
    url: https://www.diffbot.com/docs/account
  description: 'Retrieve account details, plan, token metadata and usage activity for a Diffbot token. The only introspection
    surface Diffbot publishes: because the APIs return no rate-limit or quota response headers, this endpoint is how a client
    discovers how much of its credit allotment it has spent.'
- aid: diffbot:diffbot-knowledge-graph-api
  name: Diffbot Knowledge Graph API
  image: https://kinlane-images.s3.amazonaws.com/shared/apis-json/icons/diffbot.png
  baseURL: https://kg.diffbot.com/kg/v3
  humanURL: https://www.diffbot.com/docs/ontology/
  tags:
  - Knowledge Graph
  - Entities
  - Ontology
  - Linked Data
  properties:
  - type: OpenAPI
    name: Refined one-per-tag split
    url: openapi/diffbot-knowledge-graph-api-openapi.yml
  - type: DataModel
    url: data-model/diffbot-data-model.yml
  - type: Documentation
    url: https://www.diffbot.com/docs/ontology/
  - type: APIReference
    url: https://www.diffbot.com/docs/ontology/all-entities
  description: The Knowledge Graph itself — a linked graph of over 10 billion entities (organizations, people, articles, places,
    products, job posts and more) crawled and structured from the public web. It is reached through two contracts, the DQL
    API for structured search and the Enhance API for single-entity enrichment, over a shared ontology. Entities carry first-class
    provenance and quality fields (confidence, salience, importance, origin, nbIncomingEdges, crawlTimestamp, sources) and
    are classified against NACE Rev 2.1, NAICS, ISO 3166 and IAB taxonomies. The entity shapes are documented as reference
    pages rather than declared as schemas in the published OpenAPI.
- aid: diffbot:diffbot-crawlbulk-job-api
  name: Diffbot Crawl/Bulk Job API
  image: https://kinlane-images.s3.amazonaws.com/shared/apis-json/icons/diffbot.png
  baseURL: https://api.diffbot.com/v3
  humanURL: https://www.diffbot.com/docs/crawl/search
  tags:
  - Bulk
  - Crawling
  - Extraction
  - Harvesting
  - Scraping
  properties:
  - type: Documentation
    url: https://www.diffbot.com/docs/crawl/search
  - type: APIReference
    url: https://www.diffbot.com/docs/crawl/search
  description: The Diffbot Crawl/Bulk Job API is a powerful tool that allows users to automatically extract and organize large
    amounts of web data. It enables users to create custom scraping jobs that can gather information from multiple websites
    in a structured and organized manner.
name: Diffbot
tags:
- Extraction
- Harvesting
- Scraping
- Web
- Knowledge Graph
- Crawling
- Web Search
- Natural Language
- Entity Resolution
- AI
type: Index
deliveryModel:
  model: saas
  open_source: false
  commercial: true
  callable_host: true
  label: Hosted service · you call their endpoint
  confidence: high
  source:
  - openapi
  - pricing
  generated: '2026-08-28'
  method: derived
accessModel:
  pricing: freemium
  onboarding: self-serve
  trial: false
  try_now: true
  public: false
  label: Freemium · Self-serve signup
  confidence: medium
  source:
  - plans
  - authentication
  - security
  generated: '2026-09-03'
  method: derived
image: https://kinlane-images.s3.amazonaws.com/shared/apis-json/icons/diffbot.png
access: 3rd-Party
common:
- type: OpenAPI
  name: Diffbot Extract APIs (first-party OpenAPI 3.1.0)
  url: openapi/_original/diffbot-extract-openapi.json
- type: Packages
  url: packages/diffbot-packages.yml
- type: SDKs
  name: Diffbot client libraries
  url: packages/diffbot-packages.yml
- type: MCPServer
  name: Diffbot MCP Server
  url: mcp/diffbot-mcp.yml
- type: ToolCrosswalk
  url: mcp/diffbot-tool-crosswalk.yml
- type: AgentSkill
  name: Diffbot Agent Skills (provider-published)
  url: skills/_index.yml
- type: Conventions
  url: conventions/diffbot-conventions.yml
- type: ErrorCatalog
  url: errors/diffbot-error-codes.yml
- type: Lifecycle
  url: lifecycle/diffbot-lifecycle.yml
- type: StatusPage
  name: Diffbot Status
  url: https://status.diffbot.com/
- type: Deprecation
  name: Migrating from the legacy DQL API
  url: https://www.diffbot.com/docs/dql/migrating-from-legacy-api
- type: ChangeLog
  name: Diffbot changelog (structured)
  url: changelog/diffbot-changelog.yml
- type: CLI
  name: db (Diffbot CLI)
  url: cli/diffbot-cli.yml
- type: Sandbox
  name: Diffbot Test Drive
  url: sandbox/diffbot-sandbox.yml
- type: Console
  name: Test Drive Extract
  url: https://www.diffbot.com/products/extract/testdrive
- type: Conformance
  url: conformance/diffbot-conformance.yml
- type: Compliance
  name: Diffbot GDPR and CCPA program
  url: conformance/diffbot-conformance.yml
- type: DataModel
  url: data-model/diffbot-data-model.yml
- type: Webhooks
  name: Crawl and Bulk notifyWebhook callbacks
  url: asyncapi/diffbot-webhooks.yml
- type: RateLimits
  url: rate-limits/diffbot-rate-limits.yml
- type: Plans
  url: plans/diffbot-plans-pricing.yml
- type: FinOps
  url: finops/diffbot-finops.yml
- type: Overlay
  name: API Evangelist overlay (Extract APIs)
  url: overlays/diffbot-extract-overlay.yaml
- type: Postman
  name: Diffbot on Postman
  url: https://www.postman.com/diffbotai
- type: DeveloperPortal
  name: Diffbot Developer Documentation
  url: https://www.diffbot.com/docs/
- type: APIReference
  name: Diffbot API reference
  url: https://www.diffbot.com/docs/interfaces/api
- type: GettingStarted
  name: Products overview
  url: https://www.diffbot.com/docs/products-overview
- type: Authentication
  name: Authenticating Diffbot API requests
  url: https://www.diffbot.com/docs/authentication
- type: SignUp
  name: Get started with a free Diffbot token
  url: https://app.diffbot.com/get-started
- type: Login
  name: Diffbot Dashboard
  url: https://app.diffbot.com/
- type: Support
  name: Diffbot support
  url: mailto:support@diffbot.com
- type: Careers
  url: https://www.diffbot.com/company/careers
- type: AgenticAccess
  url: agentic-access/diffbot-agentic-access.yml
- type: DomainSecurity
  url: security/diffbot-domain-security.yml
- type: Authentication
  url: authentication/diffbot-authentication.yml
- type: GitHubOrganization
  url: https://github.com/diffbot
- type: LinkedIn
  url: https://www.linkedin.com/company/diffbot
- name: Diffbot Changelog
  description: Dated product/API changelog back to 2015, with an RSS 2.0 feed.
  url: https://www.diffbot.com/changelog
  type: ChangeLog
- name: Diffbot - Knowledge Graph, AI Web Data Extraction and Crawling
  description: Diffbot's company homepage and product overview.
  url: https://www.diffbot.com/
  type: Website
- name: Diffbot - Plans and Pricing
  description: Published plan tiers, credit allotments, per-credit overage rates and a workload cost calculator.
  url: https://www.diffbot.com/pricing/
  type: Pricing
- name: Diffbot - Customer Stories
  description: Named customer case studies for the Extract, Crawl and Knowledge Graph products.
  url: https://www.diffbot.com/customer-stories/
  type: Customers
- name: Diffbot Documentation
  description: The Diffbot documentation site.
  url: https://www.diffbot.com/docs/
  type: Documentation
- name: Diffbot - Company News
  description: Company news and press releases.
  url: https://www.diffbot.com/company/news/
  type: News
- name: Diffblog - Intelligent applications start here
  description: The Diffbot engineering and product blog.
  url: https://blog.diffbot.com/
  type: Blog
- name: Knowledge Graph Glossary - Diffblog
  description: Glossary of Knowledge Graph and web-data terminology.
  url: https://blog.diffbot.com/knowledge-graph-glossary/
  type: Glossary
- name: Webinars - Diffblog
  description: Recorded product and technique webinars.
  url: https://blog.diffbot.com/webinars/
  type: Webinars
- name: Diffbot - Terms and Conditions
  description: Terms and conditions of service.
  url: https://www.diffbot.com/company/terms/
  type: TermsOfService
- name: Diffbot - Our Commitment to Your Privacy
  description: Privacy policy, with EEA and California-specific notices.
  url: https://www.diffbot.com/company/privacy/
  type: PrivacyPolicy
- name: Is Diffbot Compliant with GDPR/EU Data Laws?
  description: GDPR posture for Knowledge Graph and account data.
  url: https://www.diffbot.com/docs/account-billing/gdpr
  type: DataLicensing
- type: LlmsText
  name: Diffbot llms.txt
  url: llms/diffbot-llms.txt
- type: LLMsTxt
  name: Diffbot llms.txt (served)
  url: https://www.diffbot.com/llms.txt
created: '2024-11-13'
modified: '2026-09-06'
position: Consuming
description: Diffbot is a company that provides AI-powered web scraping and data extraction services. Their technology allows
  businesses to automatically extract and organize data from any website, turning unstructured web content into structured
  data that can be easily analyzed and used for various purposes. Diffbot's solution is used by companies across industries
  to gather competitive intelligence, monitor market trends, track online mentions, and more.
maintainers:
- FN: Kin Lane
  email: kin@apievangelist.com
specificationVersion: '0.23'
x-enrichment:
  date: '2026-09-06'
  status: enriched
  artifacts_added: 42
  pass: local-v3

Work with this as data

Every provider here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for providers

9 MCP tools reach this
  • find_providersBrowse and filter every provider in the catalog.
  • get_provider_artifactsEvery artifact this provider publishes, grouped by type.
  • get_provider_operationsEvery operation across all of their OpenAPIs — one call instead of parsing every spec.
  • get_provider_toolsEvery MCP tool they ship, with the operation each wraps.
  • get_provider_evidenceHow each part of their score was established. Free — the basis for a claim should not sit behind it.
  • get_provider_ratingPRO — composite, band, trend and facet scores.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools →

Call it yourself

curl for this page
This provider
curl "https://apis.io/api/v1/providers/diffbot"
All providers
curl "https://apis.io/api/v1/providers?limit=25"
Every operation they expose
curl "https://apis.io/api/v1/providers/diffbot/operations?limit=25"
How their score was established
curl "https://apis.io/api/v1/providers/diffbot/evidence"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no form to fill in. Signing in shares your email address with us — we store it to create your key and to recognise you if you sign in with another provider. See our Privacy Policy and Terms.

A second provider on the same verified email joins the account you already have.