Clarifeye Extraction Flows API

Manage extraction flows (auto-sync DAGs) — list, run, inspect statistics, update, and publish

Operations 5

GET /projects/{project_id}/extraction-flows/ List extraction flows #
PATCH /projects/{project_id}/extraction-flows/{flow_id}/ Update an extraction flow #
POST /projects/{project_id}/extraction-flows/{flow_id}/run-sync/ Run an extraction flow #
POST /projects/{project_id}/extraction-flows/{flow_id}/node-stats/ Get extraction flow statistics #
POST /projects/{project_id}/extraction-flows/{flow_id}/publish/ Publish an extraction flow #

Work with this as data

Every API here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for apis

7 MCP tools reach this
  • find_apisBrowse and filter every API in the catalog.
  • get_api_artifactsOne API's artifacts, grouped by type.
  • get_openapiThe primary OpenAPI for this API.
  • find_similar_apisAPIs that look like this one.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools →

Call it yourself

curl for this page
This API
curl "https://apis.io/api/v1/apis/clarifeye-extraction-flows-api"
All apis
curl "https://apis.io/api/v1/apis?limit=25"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no form to fill in. Signing in shares your email address with us — we store it to create your key and to recognise you if you sign in with another provider. See our Privacy Policy and Terms.

A second provider on the same verified email joins the account you already have.

OpenAPI Specification

clarifeye-extraction-flows-api-openapi.yml Raw ↑
openapi: 3.2.0
info:
  title: Clarifeye Platform Extraction Flows API
  description: REST API for the Clarifeye Platform - Document intelligence and AI-powered analysis.
  version: 1.0.0
  contact:
    name: Clarifeye Support
servers:
- url: https://eu.app.clarifeye.ai/api/v1
  description: EU
- url: https://us.app.clarifeye.ai/api/v1
  description: US
security:
- BearerAuth: []
- TokenAuth: []
tags:
- name: Extraction Flows
  description: Manage extraction flows (auto-sync DAGs) — list, run, inspect statistics, update, and publish
paths:
  /projects/{project_id}/extraction-flows/:
    get:
      tags:
      - Extraction Flows
      summary: List extraction flows
      description: List all extraction flows for the project, ordered by most recently updated.
      operationId: listExtractionFlows
      parameters:
      - $ref: '#/components/parameters/ProjectId'
      responses:
        '200':
          description: Successful response
          content:
            application/json:
              schema:
                type: array
                items:
                  $ref: '#/components/schemas/ExtractionFlow'
        '401':
          $ref: '#/components/responses/Unauthorized'
        '403':
          $ref: '#/components/responses/Forbidden'
        '404':
          $ref: '#/components/responses/NotFound'
  /projects/{project_id}/extraction-flows/{flow_id}/:
    patch:
      tags:
      - Extraction Flows
      summary: Update an extraction flow
      description: 'Partially update an extraction flow. The most common use is to switch its

        **publish mode** between `auto_publish` and `manual_publish`.


        **Publish modes:**

        - `auto_publish` — Each successful flow run automatically publishes the

        extracted data. No separate publish step

        is required.

        - `manual_publish` — Flow runs only compute and persist extracted data in

        the warehouse. Publishing to downstream indexes must be triggered

        explicitly via the `publish` action. Use this when you want to review

        the results of a run before exposing them to search / agents.'
      operationId: updateExtractionFlow
      parameters:
      - $ref: '#/components/parameters/ProjectId'
      - $ref: '#/components/parameters/FlowId'
      requestBody:
        required: true
        content:
          application/json:
            schema:
              type: object
              properties:
                name:
                  type: string
                  description: Human-readable name of the flow
                publish_mode:
                  type: string
                  enum:
                  - auto_publish
                  - manual_publish
                  description: 'Controls whether extracted data is pushed to downstream indexes

                    automatically after each run (`auto_publish`) or only on an

                    explicit `publish` call (`manual_publish`).

                    '
                dag:
                  type: object
                  description: DAG definition (`name` + `nodes`) for the extraction flow.
            example:
              publish_mode: manual_publish
      responses:
        '200':
          description: Flow updated successfully
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ExtractionFlow'
        '400':
          description: Invalid payload (e.g. malformed DAG or unknown publish_mode)
        '401':
          $ref: '#/components/responses/Unauthorized'
        '403':
          $ref: '#/components/responses/Forbidden'
        '404':
          $ref: '#/components/responses/NotFound'
  /projects/{project_id}/extraction-flows/{flow_id}/run-sync/:
    post:
      tags:
      - Extraction Flows
      summary: Run an extraction flow
      description: 'Queue a pipeline run that executes the flow''s DAG across the given

        documents (or all project documents if `document_ids` is omitted).


        Returns the `pipeline_run_id` you can poll to track progress.'
      operationId: runExtractionFlow
      parameters:
      - $ref: '#/components/parameters/ProjectId'
      - $ref: '#/components/parameters/FlowId'
      requestBody:
        required: false
        content:
          application/json:
            schema:
              type: object
              properties:
                document_ids:
                  type: array
                  items:
                    type: string
                    format: uuid
                  description: 'Optional list of document UUIDs to run the flow on. If omitted

                    or empty, the flow runs against all documents in the project.

                    '
            example:
              document_ids:
              - 550e8400-e29b-41d4-a716-446655440000
      responses:
        '200':
          description: Sync started
          content:
            application/json:
              schema:
                type: object
                properties:
                  message:
                    type: string
                    example: Sync started.
                  pipeline_run_id:
                    type: string
                    format: uuid
        '400':
          description: Flow DAG is empty, inconsistent, or document_ids are invalid
        '401':
          $ref: '#/components/responses/Unauthorized'
        '403':
          $ref: '#/components/responses/Forbidden'
        '404':
          $ref: '#/components/responses/NotFound'
  /projects/{project_id}/extraction-flows/{flow_id}/node-stats/:
    post:
      tags:
      - Extraction Flows
      summary: Get extraction flow statistics
      description: 'Dry-run cache simulation of the flow — returns, per DAG node, how many

        inputs would be reused from cache vs. recomputed, without actually

        executing any extraction. Useful to preview the cost/impact of a run.'
      operationId: getExtractionFlowStatistics
      parameters:
      - $ref: '#/components/parameters/ProjectId'
      - $ref: '#/components/parameters/FlowId'
      requestBody:
        required: false
        content:
          application/json:
            schema:
              type: object
              properties:
                document_ids:
                  type: array
                  items:
                    type: string
                    format: uuid
                  description: 'Optional list of document UUIDs to scope the simulation to.

                    When omitted, statistics are computed over all documents in

                    the project.

                    '
            example:
              document_ids:
              - 550e8400-e29b-41d4-a716-446655440000
      responses:
        '200':
          description: Per-node statistics
          content:
            application/json:
              schema:
                type: object
                description: 'Map of DAG node names to their cache/recompute statistics.

                  '
                additionalProperties:
                  type: object
        '400':
          description: Flow DAG is empty, inconsistent, or document_ids are invalid
        '401':
          $ref: '#/components/responses/Unauthorized'
        '403':
          $ref: '#/components/responses/Forbidden'
        '404':
          $ref: '#/components/responses/NotFound'
  /projects/{project_id}/extraction-flows/{flow_id}/publish/:
    post:
      tags:
      - Extraction Flows
      summary: Publish an extraction flow
      description: 'Queue a pipeline run that publishes previously-extracted data to the

        downstream indexes without re-running extraction.


        > **Only valid for flows whose `publish_mode` is `manual_publish`.**

        > For `auto_publish` flows, publishing happens automatically at the end

        > of each `run-sync`, so calling this endpoint is unnecessary.'
      operationId: publishExtractionFlow
      parameters:
      - $ref: '#/components/parameters/ProjectId'
      - $ref: '#/components/parameters/FlowId'
      requestBody:
        required: false
        content:
          application/json:
            schema:
              type: object
              properties:
                document_ids:
                  type: array
                  items:
                    type: string
                    format: uuid
                  description: 'Optional list of document UUIDs to scope publishing to. When

                    omitted, publishing covers all documents in the project.

                    '
            example:
              document_ids:
              - 550e8400-e29b-41d4-a716-446655440000
      responses:
        '200':
          description: Publish started
          content:
            application/json:
              schema:
                type: object
                properties:
                  message:
                    type: string
                    example: Publish started.
                  pipeline_run_id:
                    type: string
                    format: uuid
        '400':
          description: Flow DAG is empty or inconsistent
        '401':
          $ref: '#/components/responses/Unauthorized'
        '403':
          $ref: '#/components/responses/Forbidden'
        '404':
          $ref: '#/components/responses/NotFound'
components:
  parameters:
    FlowId:
      name: flow_id
      in: path
      required: true
      description: UUID of the extraction flow
      schema:
        type: string
        format: uuid
    ProjectId:
      name: project_id
      in: path
      required: true
      description: UUID of the project
      schema:
        type: string
        format: uuid
  responses:
    NotFound:
      description: Not found - resource does not exist
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
          example:
            error: Not found.
    Forbidden:
      description: Forbidden - insufficient permissions
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
          example:
            error: You do not have permission to perform this action.
    Unauthorized:
      description: Unauthorized - missing or invalid authentication
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
          example:
            error: Authentication credentials were not provided.
  schemas:
    ExtractionFlow:
      type: object
      properties:
        id:
          type: string
          format: uuid
        name:
          type: string
        publish_mode:
          $ref: '#/components/schemas/ExtractionFlowPublishMode'
        dag:
          type: object
          description: DAG definition (`name` + `nodes`) describing the extractor pipeline.
        tpuf_namespaces:
          type: object
          description: Per-namespace-type TurboPuffer namespace mapping populated on publish.
        last_pushed_to_tpuf:
          type:
          - string
          - 'null'
          format: date-time
        created_at:
          type: string
          format: date-time
        updated_at:
          type: string
          format: date-time
    ExtractionFlowPublishMode:
      type: string
      enum:
      - auto_publish
      - manual_publish
      description: '- `auto_publish` — runs automatically push extracted data to downstream indexes.

        - `manual_publish` — publishing is triggered explicitly via the `publish` action.

        '
    Error:
      type: object
      properties:
        error:
          type: string
          description: Error message
      example:
        error: User not found
  securitySchemes:
    BearerAuth:
      type: http
      scheme: bearer
      description: 'Use Authorization: Bearer <token>'
    TokenAuth:
      type: apiKey
      in: header
      name: Authorization
      description: 'Use Authorization: Token <token>'