Cognite Document preview API

The document preview service is a utility API that can render most document types as an image or PDF. This can be very helpful if you want to display a preview of a file in a frontend, or for other tasks that require one of these formats. For both rendered formats there is a concept of a page. The actual meaning of a page depends on the source document. E.g. an image will always have exactly one page, while a spreadsheet will typically have one page representing each individual sheet. The document preview service can only generate preview for document sizes that do not exceed 150 MiB. Trying to preview a larger document will give an error. ### File type support Previews can be created for the following types of files: - PDF files - Spreadsheets, documents and presentations from the Microsoft and Libre Office office suites - Images

OpenAPI Specification

cognite-document-preview-api-openapi.yml Raw ↑
openapi: 3.1.0
info:
  title: Cognite 3D Asset Mapping Document preview API
  description: "# Introduction\nThis is the reference documentation for the Cognite API with\nan overview of all the available methods.\n\n# Postman\nSelect the **Download** button to download our OpenAPI specification to get started.\n\nTo import your data into Postman, select **Import**, and the Import modal opens.\nYou can import items by dragging or dropping files or folders. You can choose how to import your API and manage the import settings in **View Import Settings**.\n\nIn the Import Settings, set the **Folder organization** to **Tags**, select\n**Enable optional parameters** to turn off the settings, and select **Always inherit authentication** to turn on the settings. Select **Import**.\n\nSet the Authorization to **Oauth2.0**. By default, the settings are for Open Industrial Data. Navigate to [Cognite Hub](https://hub.cognite.com/open-industrial-data-211) to understand how to get the credentials for use in Postman.\n\nFor more information, see [Getting Started with Postman](https://developer.cognite.com/dev/guides/postman/).\n\n# Pagination\nMost resource types can be paginated, indicated by the field `nextCursor` in the response.\nBy passing the value of `nextCursor` as the cursor you will get the next page of `limit` results.\nNote that all parameters except `cursor` has to stay the same.\n\n# Parallel retrieval\nAs general guidance, Parallel Retrieval is a technique that should be used when due to query complexity, retrieval of data in a single request is significantly slower than it would otherwise be for a simple request.  Parallel retrieval does not act as a speed multiplier on optimally running queries.  By parallelizing such requests, data retrieval performance can be tuned to meet the client application needs. \n\nCDF supports parallel retrieval through the `partition` parameter, which has the format `m/n` where `n` is the amount of partitions you would like to split the entire data set into.\nIf you want to download the entire data set by splitting it into 10 partitions, do the following in parallel with `m` running from 1 to 10:\n  - Make a request to `/events` with `partition=m/10`.\n  - Paginate through the response by following the cursor as explained above. Note that the `partition` parameter needs to be passed to all subqueries.\n\nProcessing of parallel retrieval requests is subject to concurrency quota availability. The request returns the `429` response upon exceeding concurrency limits. See the Request throttling chapter below.\n\nTo prevent unexpected problems and to maximize read throughput, you should at most use 10 partitions. \nSome CDF resources will automatically enforce a maximum of 10 partitions.\nFor more specific and detailed information, please read the ```partition``` attribute documentation for the CDF resource you're using.  \n\n# Requests throttling\nCognite Data Fusion (CDF) returns the HTTP `429` (too many requests) response status code when project capacity exceeds the limit.\n\nThe throttling can happen:\n  - If a user or a project sends too many (more than allocated) concurrent requests.\n  - If a user or a project sends a too high (more than allocated) rate of requests in a given amount of time.\n\nCognite recommends using a retry strategy based on truncated exponential backoff to handle sessions with HTTP response codes 429.\n\nCognite recommends using a reasonable number (up to 10) of  `Parallel retrieval` partitions.\n\nFollowing these strategies lets you slow down the request frequency to maximize productivity without having to re-submit/retry failing requests.\n\nSee more [here](https://docs.cognite.com/dev/concepts/resource_throttling).\n\n# API versions\n## Version headers\nThis API uses calendar versioning, and version names follow the `YYYYMMDD` format.\nYou can find the versions currently available by using the version selector at the top of this page.\n\nTo use a specific API version, you can pass the `cdf-version: $version` header along with your requests to the API.\n\n## Beta versions\nThe beta versions provide a preview of what the stable version will look like in the future.\nBeta versions contain functionality that is reasonably mature, and highly likely to become a part of the stable API.\n\nBeta versions are indicated by a `-beta` suffix after the version name. For example, the beta version header for the\n2023-01-01 version is then `cdf-version: 20230101-beta`.\n\n## Alpha versions\nAlpha versions contain functionality that is new and experimental, and not guaranteed to ever become a part of the stable API.\nThis functionality presents no guarantee of service, so its use is subject to caution.\n\nAlpha versions are indicated by an `-alpha` suffix after the version name. For example, the alpha version header for\nthe 2023-01-01 version is then `cdf-version: 20230101-alpha`."
  version: v1
  contact:
    name: Cognite Support
    url: https://support.cognite.com
    email: support@cognite.com
servers:
- url: https://{cluster}.cognitedata.com/api/v1/projects/{project}
  description: The URL for the CDF cluster to connect to
  variables:
    cluster:
      enum:
      - api
      - az-tyo-gp-001
      - az-eastus-1
      - az-power-no-northeurope
      - westeurope-1
      - asia-northeast1-1
      - gc-dsm-gp-001
      default: api
      description: The CDF cluster to connect to
    project:
      default: publicdata
      description: The CDF project name.
security:
- oidc-token:
  - https://{cluster}.cognitedata.com/.default
- oauth2-client-credentials:
  - https://{cluster}.cognitedata.com/.default
- oauth2-open-industrial-data:
  - https://api.cognitedata.com/.default
- oauth2-auth-code:
  - https://{cluster}.cognitedata.com/.default
tags:
- name: Document preview
  x-displayName: Preview
  description: 'The document preview service is a utility API that can render most document types as an image or PDF.

    This can be very helpful if you want to display a preview of a file in a frontend, or for other

    tasks that require one of these formats.


    For both rendered formats there is a concept of a page. The actual meaning of a page depends on

    the source document. E.g. an image will always have exactly one page, while a spreadsheet

    will typically have one page representing each individual sheet.


    The document preview service can only generate preview for document sizes that do not

    exceed 150 MiB. Trying to preview a larger document will give an error.


    ### File type support

    Previews can be created for the following types of files:

    - PDF files

    - Spreadsheets, documents and presentations from the Microsoft and Libre Office office suites

    - Images'
paths:
  /documents/{documentId}/preview/image/pages/{pageNumber}:
    get:
      tags:
      - Document preview
      summary: Retrieve a image preview of a page from a document
      operationId: documentsPreviewImagePage
      description: '


        > **Required capabilities:** `filesAcl:READ`


        This endpoint returns a rendered image preview for a specific page of the specified document.


        The `accept` request header MUST be set to `image/png`. Other values will

        give an HTTP 406 error.


        The rendered image will be downsampled to a maximum of 2400x2400 pixels.

        Only PNG format is supported and only the first 10 pages can be rendered.


        Previews will be rendered if neccessary during the request. Be prepared

        for the request to take a few seconds to complete.'
      parameters:
      - in: path
        name: documentId
        schema:
          type: integer
        required: true
        description: Internal ID for document to preview
      - in: path
        name: pageNumber
        schema:
          type: integer
          minimum: 1
          maximum: 10
        required: true
        description: Page number to preview. Starting at 1 for first page
      responses:
        '200':
          description: OK
          content:
            image/png:
              schema:
                description: Rendered PNG image
                type: string
                format: binary
        '400':
          $ref: '#/components/responses/ErrorResponse'
        '401':
          $ref: '#/components/responses/ErrorResponse'
        '406':
          $ref: '#/components/responses/ErrorResponse'
        '422':
          $ref: '#/components/responses/ErrorResponse'
      x-capability:
      - filesAcl:READ
      x-code-samples:
      - lang: Python
        label: Python SDK
        source: 'client.documents.previews.download_page_as_png("previews", id=123, page_number=5)

          content = client.documents.previews.download_page_as_png_bytes(id=123, page_number=5)


          from IPython.display import Image

          binary_png = client.documents.previews.download_page_as_png_bytes(id=123, page_number=5)

          Image(binary_png)

          '
  /documents/{documentId}/preview/pdf:
    get:
      tags:
      - Document preview
      summary: Retrieve a PDF preview of a document
      operationId: documentsPreviewPdf
      description: '


        > **Required capabilities:** `filesAcl:READ`


        This endpoint returns a rendered PDF preview for a specified document.


        The `accept` request header MUST be set to `application/pdf`. Other values will

        give an HTTP 406 error.


        This endpoint is optimized for in-browser previews. We reserve the right

        to adjust the quality and other attributes of the output with this in mind.

        Please reach out to us if you have a different use case and requirements.


        Previews will be rendered if neccessary during the request. Be prepared

        for the request to take a few seconds to complete.'
      parameters:
      - $ref: '#/components/parameters/cdfversionheader'
      - in: path
        name: documentId
        schema:
          type: integer
        required: true
        description: Internal ID for document to preview
      responses:
        '200':
          description: OK
          content:
            application/pdf:
              schema:
                description: Rendered PDF document
                type: string
                format: binary
        '400':
          $ref: '#/components/responses/ErrorResponse'
        '401':
          $ref: '#/components/responses/ErrorResponse'
        '406':
          $ref: '#/components/responses/ErrorResponse'
        '422':
          $ref: '#/components/responses/ErrorResponse'
      x-capability:
      - filesAcl:READ
      x-code-samples:
      - lang: Python
        label: Python SDK
        source: 'client.documents.previews.download_document_as_pdf("previews", id=123)

          content = client.documents.previews.download_document_as_pdf_bytes(id=123)

          '
  /documents/{documentId}/preview/pdf/temporarylink:
    get:
      tags:
      - Document preview
      summary: Retrieve a temporary link to a PDF preview of a document
      operationId: documentsPreviewPdfTemporaryLink
      description: '


        > **Required capabilities:** `filesAcl:READ`


        This endpoint works similar as the normal preview endpoint except

        it returns a short-lived temporary link to download the rendered preview instead

        of returning the binary data.'
      parameters:
      - $ref: '#/components/parameters/cdfversionheader'
      - in: path
        name: documentId
        schema:
          type: integer
        required: true
        description: Internal ID for document to preview
      responses:
        '200':
          $ref: '#/components/responses/DocumentsPreviewTemporaryLinkResponse'
        '400':
          $ref: '#/components/responses/ErrorResponse'
        '401':
          $ref: '#/components/responses/ErrorResponse'
        '422':
          $ref: '#/components/responses/ErrorResponse'
      x-capability:
      - filesAcl:READ
      x-code-samples:
      - lang: Python
        label: Python SDK
        source: 'link = client.documents.previews.retrieve_pdf_link(id=123)

          '
components:
  responses:
    ErrorResponse:
      description: The response for a failed request.
      content:
        application/json:
          schema:
            type: object
            required:
            - error
            properties:
              error:
                $ref: '#/components/schemas/Error'
    DocumentsPreviewTemporaryLinkResponse:
      description: OK
      content:
        application/json:
          schema:
            description: 'A temporary link to download a preview of the document. The link is reachable without additional authentication details for a limited time.

              '
            type: object
            required:
            - temporaryLink
            - expirationTime
            properties:
              temporaryLink:
                type: string
              expirationTime:
                example: 1519862400000
                allOf:
                - $ref: '#/components/schemas/EpochTimestamp'
  parameters:
    cdfversionheader:
      in: header
      name: cdf-version
      description: cdf version header. Use this to specify the requested CDF release.
      schema:
        type: string
        example: alpha
  schemas:
    EpochTimestamp:
      description: The number of milliseconds since 00:00:00 Thursday, 1 January 1970, Coordinated Universal Time (UTC), minus leap seconds.
      type: integer
      minimum: 0
      format: int64
      example: 1730204346000
    Error:
      type: object
      required:
      - code
      - message
      description: Cognite API error.
      properties:
        code:
          type: integer
          description: HTTP status code.
          format: int32
          example: 401
        message:
          type: string
          description: Error message.
          example: Could not authenticate.
        missing:
          type: array
          description: List of lookup objects that do not match any results.
          items:
            type: object
            additionalProperties: true
        duplicated:
          type: array
          description: List of objects that are not unique.
          items:
            type: object
            additionalProperties: true
  securitySchemes:
    oidc-token:
      type: http
      scheme: bearer
      bearerFormat: OpenID Connect or OAuth2 token
      description: Access token issued by the CDF project's configured identity provider. Access token must be an OpenID Connect token, and the project must be configured to accept OpenID Connect tokens. Use a header key of 'Authorization' with a value of 'Bearer $accesstoken'. The token can be obtained through any flow supported by the identity provider.
    oauth2-client-credentials:
      type: oauth2
      description: Access token issued by the CDF project's configured identity provider. Access token must be an OpenID Connect token, and the project must be configured to accept OpenID Connect tokens. Use a header key of 'Authorization' with a value of 'Bearer $accesstoken'. The token can be obtained through any flow supported by the identity provider.
      flows:
        clientCredentials:
          tokenUrl: https://your-idps.token.url/
          scopes:
            default: https://{cluster}.cognitedata.com/.default
    oauth2-auth-code:
      type: oauth2
      description: Access token issued by the CDF project's configured identity provider. Access token must be an OpenID Connect token, and the project must be configured to accept OpenID Connect tokens. Use a header key of 'Authorization' with a value of 'Bearer $accesstoken'. The token can be obtained through any flow supported by the identity provider.
      flows:
        authorizationCode:
          authorizationUrl: https://your-idps.authorization.url/
          tokenUrl: https://your-idps.token.url/
          scopes:
            default: https://{cluster}.cognitedata.com/.default
    oauth2-open-industrial-data:
      type: oauth2
      description: Auth flow for Open Industrial Data. Get your client secret from https://hub.cognite.com/open-industrial-data-211.
      flows:
        clientCredentials:
          tokenUrl: https://login.microsoftonline.com/48d5043c-cf70-4c49-881c-c638f5796997/oauth2/v2.0/token
          scopes:
            default: https://api.cognitedata.com/.default
    org-oidc-token:
      type: openIdConnect
      openIdConnectUrl: https://auth.cognite.com/.well-known/openid-configuration
      description: 'Access token issued by the Cognite authorization server, and valid for the target organization. The token must

        be an OpenID Connect token, and it can be obtained by performing an OIDC login flow toward `auth.cognite.com`.

        This is a single URL for all CDF organizations.'
x-tagGroups:
- name: Changelog
  tags:
  - Changelog
- name: Organizations and projects
  tags:
  - Organizations
  - Projects
- name: Identity and access management
  tags:
  - Principals
  - Groups
  - Security categories
  - Sessions
  - Token
  - User profiles
  - Project Deletion Reporting
- name: Data modeling
  tags:
  - Data Modeling
  - Data models
  - Spaces
  - Views
  - Containers
  - Nodes
  - Instances
  - Statistics
  - Streams
  - Records
- name: Asset-centric data model
  tags:
  - Assets
  - Time series
  - Synthetic Time Series
  - Data point subscriptions
  - Events
  - Files
  - Sequences
  - Geospatial
  - Seismic
- name: 3D
  tags:
  - 3D Models
  - 3D Model Revisions
  - 3D Files
  - 3D Asset Mapping
  - 3D Contextualization
  - 3D Jobs
  - 3D Migration
  - 3D Scenes
- name: Contextualization
  tags:
  - Entity matching
  - Entity matching pipelines
  - Engineering diagrams
  - Vision
  - Advanced joins
- name: Cognite AI
  tags:
  - Agents
  - Skills
  - Chat Completions
  - Document AI
  - Models
- name: Documents
  tags:
  - Documents
  - Document preview
- name: Data ingestion
  tags:
  - Raw
  - Extraction Pipelines
  - Extraction Pipelines Runs
  - Extraction Pipelines Config
  - Extractors
- name: Data organization
  tags:
  - Data sets
  - Data domains
  - Data products
  - Rule sets
  - Labels
  - Relationships
  - Annotations
- name: Transformations
  tags:
  - Transformations
  - Transformation Jobs
  - Transformation Schedules
  - Transformation Notifications
  - Query
  - Schema
- name: Functions
  tags:
  - Functions
  - Function calls
  - Function schedules
- name: Hosted Extractors
  tags:
  - Sources
  - Jobs
  - Destinations
  - Mappings
- name: PostgreSQL Gateway
  tags:
  - Postgres Gateway Users
  - Postgres Gateway Tables
- name: SAP Writeback
  tags:
  - SAP Instances
  - SAP Endpoints
  - Schema Mappings
  - Writeback Requests
- name: Data workflows
  tags:
  - Workflows
  - Workflow versions
  - Workflow executions
  - Workflow triggers
  - Tasks
  - Workers
- name: Simulators
  tags:
  - Simulators
  - Simulator Integrations
  - Simulator Models
  - Simulator Routines
  - Simulation Runs
  - Simulator Logs
- name: Units
  tags:
  - Units
  - Unit Systems
- name: ''
  tags:
  - ''