Mindee · Arazzo Workflow

Mindee Custom Document Extraction

Version 1.0.0

Enqueue an arbitrary document against a custom model with raw text capture, poll until processed, then read fields and full text.

1 workflow 2 source APIs 1 provider
View Spec View on GitHub Document ParsingOCRIDPArtificial IntelligenceMachine-LearningInvoicesReceiptsIDSComputer-VisionArazzoWorkflows

Provider

mindee

Workflows

custom-document-extraction
Upload a custom document with raw text capture and read fields plus full text.
Sends a document to the extraction enqueue endpoint against a custom model with raw_text enabled, polls the job until processing finishes, and retrieves both the extracted fields and the raw document text.
3 steps inputs: authorization, file, filename, modelId, textContext outputs: fields, jobId, rawText
1
enqueueDocument
Send the document to the asynchronous extraction queue against the custom model with raw_text enabled so the full document text is returned.
2
pollJob
Poll the shared jobs endpoint until the custom extraction job reports Processed or Failed.
3
getResult
Retrieve the completed extraction inference and read the custom fields and the full raw text parsed from the document.

Source API Descriptions

Arazzo Workflow Specification

Raw ↑
arazzo: 1.0.1
info:
  title: Mindee Custom Document Extraction
  summary: Enqueue an arbitrary document against a custom model with raw text capture, poll until processed, then read fields and full text.
  description: >-
    Applies the Mindee asynchronous extraction pattern to a custom document
    type backed by a user-defined extraction model. The workflow uploads any
    document against the supplied model with the raw_text option enabled, polls
    the shared jobs endpoint until the job is Processed, and fetches the
    inference to read both the model-specific extracted fields and the complete
    raw text of the document. Every step spells out its request inline so the
    flow can be read and executed without opening the underlying OpenAPI
    description.
  version: 1.0.0
sourceDescriptions:
- name: extractionApi
  url: ../openapi/mindee-extraction-api-openapi.yml
  type: openapi
- name: jobsApi
  url: ../openapi/mindee-jobs-api-openapi.yml
  type: openapi
workflows:
- workflowId: custom-document-extraction
  summary: Upload a custom document with raw text capture and read fields plus full text.
  description: >-
    Sends a document to the extraction enqueue endpoint against a custom model
    with raw_text enabled, polls the job until processing finishes, and
    retrieves both the extracted fields and the raw document text.
  inputs:
    type: object
    required:
    - authorization
    - modelId
    - file
    properties:
      authorization:
        type: string
        description: Mindee API key sent in the Authorization header.
      modelId:
        type: string
        description: UUID of the custom extraction model to apply.
      file:
        type: string
        description: The document file to upload as binary form data.
      filename:
        type: string
        description: Optional filename to associate with the uploaded document.
      textContext:
        type: string
        description: Optional additional context passed to the model for this inference.
  steps:
  - stepId: enqueueDocument
    description: >-
      Send the document to the asynchronous extraction queue against the custom
      model with raw_text enabled so the full document text is returned.
    operationId: Enqueue_Extraction_Product_Inference_v2_products_extraction_enqueue_post
    parameters:
    - name: Authorization
      in: header
      value: $inputs.authorization
    requestBody:
      contentType: multipart/form-data
      payload:
        model_id: $inputs.modelId
        file: $inputs.file
        filename: $inputs.filename
        raw_text: true
        text_context: $inputs.textContext
    successCriteria:
    - condition: $statusCode == 202
    outputs:
      jobId: $response.body#/job/id
      status: $response.body#/job/status
  - stepId: pollJob
    description: >-
      Poll the shared jobs endpoint until the custom extraction job reports
      Processed or Failed.
    operationId: Get_Job_Status_v2_jobs__job_id__get
    parameters:
    - name: Authorization
      in: header
      value: $inputs.authorization
    - name: job_id
      in: path
      value: $steps.enqueueDocument.outputs.jobId
    - name: redirect
      in: query
      value: false
    successCriteria:
    - condition: $statusCode == 200
    outputs:
      status: $response.body#/job/status
    onSuccess:
    - name: jobProcessed
      type: goto
      stepId: getResult
      criteria:
      - context: $response.body
        condition: $.job.status == "Processed"
        type: jsonpath
    - name: jobPending
      type: goto
      stepId: pollJob
      criteria:
      - context: $response.body
        condition: $.job.status == "Processing"
        type: jsonpath
  - stepId: getResult
    description: >-
      Retrieve the completed extraction inference and read the custom fields
      and the full raw text parsed from the document.
    operationId: Get_Extraction_Product_Result_v2_products_extraction_results__inference_id__get
    parameters:
    - name: Authorization
      in: header
      value: $inputs.authorization
    - name: inference_id
      in: path
      value: $steps.enqueueDocument.outputs.jobId
    successCriteria:
    - condition: $statusCode == 200
    outputs:
      inferenceId: $response.body#/inference/id
      fields: $response.body#/inference/result/fields
      rawText: $response.body#/inference/result/raw_text
  outputs:
    jobId: $steps.enqueueDocument.outputs.jobId
    fields: $steps.getResult.outputs.fields
    rawText: $steps.getResult.outputs.rawText

Work with this as data

Every workflow here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for arazzo workflows

4 MCP tools reach this
  • find_arazzoBrowse and filter every workflow in the catalog.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools →

Call it yourself

curl for this page
This workflow
curl "https://apis.io/api/v1/arazzo/mindee-custom-document-extraction-workflow"
All arazzo workflows
curl "https://apis.io/api/v1/arazzo?limit=25"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no form to fill in. Signing in shares your email address with us — we store it to create your key and to recognise you if you sign in with another provider. See our Privacy Policy and Terms.

A second provider on the same verified email joins the account you already have.