Cognite Documents API

A document is a file that has been indexed by the document search engine. Every time a file is uploaded, updated or deleted in the Files API, it will also be scheduled for processing by the document search engine. After some processing, it will be possible to search for the file in the document search API. The document search engine is able to extract content from a variety of document types, and perform classification, contextualization and other operations on the file. This extracted and derived information is made available in the form of a `Document` object. The document structure consists of a selection of derived fields, such as the `title`, `author` and `language` of the document, plus some of the original fields from the raw file. The fields from the raw file can be found in the `sourceFile` structure. The derived fields are described in more detail below. ### Derived fields #### title Some document types (such as PDFs) contain additional metadata fields. If the document contains its title as part of this metadata, this field will be populated with that title. Note that we do not currently extract the title from the document content itself. If there is a need for this, we may consider adding such functionality in the future. #### author Similar to the `title` field, the author field is another field that can often be extracted from the document's metadata. #### producer The `producer` field also exists in the document metadata. It contains information about the software or the system that was used to create the document. #### createdTime The `createdTime` we assign to the document is not exactly the same as the one found in the Files API. We first try to extract the created time from the document metadata. If the document does not contain such a timestamp, we fall back to the time set in the Files API. #### mimeType If there is a mime type set in on the file in the Files API, this field will be set to the same mime type. If there is no mime type set on the file, we will try to auto-detect it. #### extension This field contains the extension of the file, derived from the file name. For instance, if the file name is `My Document.docx`, the `extension` field will contain `docx`. #### pageCount Contains the number of pages in the document, if possible to determine. #### type The `type` field contains a high level file type, derived from the mime type. Mime types are not that pleasant to look at, and not always easy to understand. That is why we map the mime types into more user-friendly types. Below is the list of types currently returned, but be aware that this list may be extended in the future. - `Document`: Document files from Microsoft Word or similar word processing software. - `PDF`: PDF files. - `Spreadsheet`: Files from Microsoft Excel or similar spreadsheet software. - `Presentation`: Slides from Microsoft Powerpoint or similar. - `Image`: Any kind of image such as PNG or JPG files. - `Video`: Any kind of video such as MOV or MP4 files. - `Tabular data`: Csv, tsv and other kinds of tabular data files. - `Plain text`: Plain text files. - `Compressed`: ZIP files and other kinds of compressed archive files. - `Script`: Program code such as python or matlab. - `Other`: Anything that doesn't fit in any of the above types. #### geoLocation If there is a geolocation set on the file in the Files API, then this field will contain the same geolocation. If there is no explicitly assigned geolocation, the document processing system will try to detect a location using two different techniques; 1. We will extract locations from files that contain embedded GPS locations. Photos and videos often have this kind of metadata. 2. We will look at related assets that have locations, and assign the same location(s) to the document. ### File type support We create a document for each uploaded file, but only derive data from certain files. The following file types are eligible for further data extraction & enrichment: - PDF files - Spreadsheets, documents, and presentations from the Microsoft, Libre Office and macOS office suites - Plain text files - Images

Operations 5

POST /documents/search Search for documents #
POST /documents/passages/search Semantic search for passages #
POST /documents/aggregate Aggregate documents #
POST /documents/list List documents #
POST /documents/content Retrieve document content #

Work with this as data

Every API here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for apis

7 MCP tools reach this
  • find_apisBrowse and filter every API in the catalog.
  • get_api_artifactsOne API's artifacts, grouped by type.
  • get_openapiThe primary OpenAPI for this API.
  • find_similar_apisAPIs that look like this one.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools →

Call it yourself

curl for this page
This API
curl "https://apis.io/api/v1/apis/cognite-documents-api"
All apis
curl "https://apis.io/api/v1/apis?limit=25"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no form to fill in. Signing in shares your email address with us — we store it to create your key and to recognise you if you sign in with another provider. See our Privacy Policy and Terms.

A second provider on the same verified email joins the account you already have.

OpenAPI Specification

cognite-documents-api-openapi.yml Raw ↑
openapi: 3.2.0
info:
  title: Cognite Documents API
  description: '# Introduction

    This is the reference documentation for the Cognite API with

    an overview of all the available methods.'
  version: v1
  contact:
    name: Cognite Support
    url: https://support.cognite.com
    email: support@cognite.com
servers:
- url: https://{cluster}.cognitedata.com/api/v1/projects/{project}
  description: The URL for the CDF cluster to connect to
  variables:
    cluster:
      enum:
      - api
      - az-tyo-gp-001
      - az-eastus-1
      - az-power-no-northeurope
      - westeurope-1
      - asia-northeast1-1
      - gc-dsm-gp-001
      default: api
      description: The CDF cluster to connect to
    project:
      default: publicdata
      description: The CDF project name.
security:
- oidc-token:
  - https://{cluster}.cognitedata.com/.default
- oauth2-client-credentials:
  - https://{cluster}.cognitedata.com/.default
- oauth2-open-industrial-data:
  - https://api.cognitedata.com/.default
- oauth2-auth-code:
  - https://{cluster}.cognitedata.com/.default
tags:


# --- truncated at 32 KB (72 KB total) ---
# Full source: https://raw.githubusercontent.com/api-evangelist/cognite/refs/heads/main/openapi/cognite-documents-api-openapi.yml