Syntage Extractions API

Extractions are tasks that retrieve data for an entity from a datasource. They are used to collect invoices, tax returns, tax status, tax compliance checks, RPC records, RUG guarantees, Buró de Crédito reports, BIL reports, and other datasource-backed records. An extraction runs asynchronously. Creating one starts or queues the work, and the returned Extraction resource tracks its status, options, timing, errors, and the number of resources created or updated. ## How Extractions Work 1. Create or retrieve the entity you want to extract data for. 2. Create an extraction with the entity IRI, extractor name, and extractor options. 3. Store the returned extraction ID. 4. Retrieve the extraction until its `status` is `finished` or `failed`. 5. Read the extracted data from the related resource endpoints, such as invoices or tax returns. If the same extraction is already pending or running for the same entity, extractor, and options, Syntage returns a duplicate extraction response instead of starting the same work again. ## Options Options containing a period (`.`) in the name like `period.from` are a way to denote object paths. For example, `period.from` should actually be sent as: ```json { "options": { "period": { "from": "2020-01-01", "to": "2020-12-31" } } } ``` Extractor names and options are specific to the data you are extracting. For extraction-backed resources, use the resource overview page for the extractor name, required credentials, supported options, and the endpoint where the extracted records can be read. # Extraction status | Status | Description | | ------ | ----------- | | pending | The initial extraction status. The extraction is enqueued and waiting to be processed. | | waiting | The extraction does not meet the requirements to begin. | | running | The extraction process started and is currently running. The running time varies depending on the extractor type and the entity's transactional volume. You may find partial data available in our API endpoints during this time. | | finished | The extraction finished successfully. All data is available to be retrieved from our API endpoints. | | failed | The extraction couldn't start or failed during the process. Our internal retry policies weren't able to finish the extraction successfully. We may have partial data available in our API endpoints, but new extractions should be created to ensure all the entity's data is available. You can check the [extraction error code](#section/Extraction-error-codes) to understand why it failed and determine whether it can be retried or not. | | stopping | The extraction was requested to be stopped by the user. It is in the process of being stopped. | | stopped | The extraction was stopped by the user after it started running. This extraction is included in billing. | | cancelled | The extraction was stopped by the user before it started running. This extraction is not included in billing. | # Extraction error codes | Code | Description | Retryable | | ---- | ----------- | --------- | | invalid_credentials | The SAT [Credential](#tag/Credentials) is no longer valid. | No | | login_failed | We couldn't log in with the SAT credential. | Yes | | unrecoverable | The extraction process failed many times, and we reached a maximum number of retries. | Yes | | sat_unavailable | We detected that SAT itself is down or unresponsive. | Yes | | internal_error | We detected an internal error in our own infrastructure. | Yes | | undefined | We couldn't determine the error cause and our internal team will investigate it. | Yes |

Operations 4

POST /extractions Create an extraction #
GET /extractions List all extractions #
GET /extractions/{id} Retrieve an extraction #
DELETE /extractions/{id}/stop Request the extraction to be stopped #

Work with this as data

Every API here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for apis

7 MCP tools reach this
  • find_apisBrowse and filter every API in the catalog.
  • get_api_artifactsOne API's artifacts, grouped by type.
  • get_openapiThe primary OpenAPI for this API.
  • find_similar_apisAPIs that look like this one.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools →

Call it yourself

curl for this page
This API
curl "https://apis.io/api/v1/apis/syntage-extractions-api"
All apis
curl "https://apis.io/api/v1/apis?limit=25"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no form to fill in. Signing in shares your email address with us — we store it to create your key and to recognise you if you sign in with another provider. See our Privacy Policy and Terms.

A second provider on the same verified email joins the account you already have.

OpenAPI Specification

syntage-extractions-api-openapi.yml Raw ↑
openapi: 3.2.0
info:
  version: '2020-06-28'
  title: Syntage Extractions API
  contact:
    name: Email
    email: support@syntage.com
  description: '# Introduction


    The Syntage API is organized around REST.'
servers:
- url: https://api.syntage.com
  description: Production
- url: https://api.sandbox.syntage.com
  description: Sandbox
security:
- ApiKey: []
tags:
- name: Extractions
  description: Extractions are tasks that retrieve data for an entity from a datasource.
paths:
  /extractions:
    post:
      tags:
      - Extractions
      operationId: CreateExtraction
      summary: Create an extraction
      description: 'Create an extraction for a specific entity, extractor, and options.


        If an equivalent extraction is already `pending` or `running` for the same entity, extractor, and options, the API returns `409 Conflict` instead of starting a duplicate. Retrieve the in-progress extraction to track its status rather than retrying the request.'
      requestBody:
        $ref: '#/components/requestBodies/ExtractionCreate'
      responses:
        '202':
          $ref: '#/components/responses/Extraction'
        '400':
          $ref: '#/components/responses/ConstraintViolation'
        '401':
          $ref: '#/components/responses/Unauthorized'
        '409':
          $ref: '#/components/responses/DuplicateExtraction'
    get:
      tags:
      - Extractions
      operationId: ListExtraction
      summary: List all extractions
      description: List extraction tasks and filter them by entity, extractor, status, dates, or SAT rate limit state.
      responses:
        '200':
          $ref: '#/components/responses/ExtractionCollection'
        '401':
          $ref: '#/components/responses/Unauthorized'
      parameters:
      - name: taxpayer.id
        in: query
        description: Filter by RFC (Registro Federal de Contribuyentes, exact match)
        schema:
          type: string
          minLength: 12
          maxLength: 13
          example: PEIC211118IS0
      - name: extractor
        in: query
        description: Filter by extractor (partial match)
        example: invoice
        schema:
          $ref: '#/components/schemas/ExtractionExtractor'
      - name: status
        in: query
        description: Filter by status (exact match)
        schema:
          $ref: '#/components/schemas/ExtractionStatus'
      - name: startedAt[before]
        in: query
        description: Filter by start date (less than or equal `<=`)
        example: '2020-01-17 11:47:27'
        schema:
          type: string
          format: date-time
      - name: startedAt[strictly_before]
        in: query
        description: Filter by start date (less than `<`)
        example: '2020-01-17 11:47:27'
        schema:
          type: string
          format: date-time
      - name: startedAt[after]
        in: query
        description: Filter by start date (greater than or equal `>=`)
        example: '2020-01-17 11:47:27'
        schema:
          type: string
          format: date-time
      - name: startedAt[strictly_after]
        in: query
        description: Filter by start date (greater than `>`)
        example: '2020-01-17 11:47:27'
        schema:
          type: string
          format: date-time
      - name: finishedAt[before]
        in: query
        description: Filter by finished date (less than or equal `<=`)
        example: '2020-01-17 11:47:27'
        schema:
          type: string
          format: date-time
      - name: finishedAt[strictly_before]
        in: query
        description: Filter by finished date (less than `<`)
        example: '2020-01-17 11:47:27'
        schema:
          type: string
          format: date-time
      - name: finishedAt[after]
        in: query
        description: Filter by finished date (greater than or equal `>=`)
        example: '2020-01-17 11:47:27'
        schema:
          type: string
          format: date-time
      - name: finishedAt[strictly_after]
        in: query
        description: Filter by finished date (greater than `>`)
        example: '2020-01-17 11:47:27'
        schema:
          type: string
          format: date-time
      - name: exists[rateLimitedAt]
        in: query
        description: Filter extractions that was rate limited by SAT
        example: '1'
        schema:
          type: boolean
          example: true
      - $ref: '#/components/parameters/createdAtBefore'
      - $ref: '#/components/parameters/createdAtStrictlyBefore'
      - $ref: '#/components/parameters/createdAtAfter'
      - $ref: '#/components/parameters/createdAtStrictlyAfter'
      - name: order[startedAt]
        in: query
        description: Order by start date
        schema:
          $ref: '#/components/schemas/CollectionOrder'
      - name: order[finishedAt]
        in: query
        description: Order by finished date
        schema:
          $ref: '#/components/schemas/CollectionOrder'
      - $ref: '#/components/parameters/orderCreatedAt'
      - $ref: '#/components/parameters/orderUpdatedAt'
      - $ref: '#/components/parameters/collectionCursorNextPageParam'
      - $ref: '#/components/parameters/collectionCursorPreviousPageParam'
      - $ref: '#/components/parameters/collectionLimit'
  /extractions/{id}:
    get:
      tags:
      - Extractions
      operationId: GetExtraction
      summary: Retrieve an extraction
      description: Retrieve a single extraction, including its extractor, options, status, timing, error code, and data point counts.
      parameters:
      - $ref: '#/components/parameters/resourceId'
      responses:
        '200':
          $ref: '#/components/responses/Extraction'
        '401':
          $ref: '#/components/responses/Unauthorized'
        '404':
          $ref: '#/components/responses/NotFound'
  /extractions/{id}/stop:
    delete:
      tags:
      - Extractions
      operationId: StopExtraction
      summary: Request the extraction to be stopped
      description: Request cancellation for a pending or running extraction.
      parameters:
      - $ref: '#/components/parameters/resourceId'
      responses:
        '204':
          description: Extraction stopping process started
        '401':
          $ref: '#/components/responses/Unauthorized'
        '404':
          $ref: '#/components/responses/NotFound'
components:
  responses:
    DuplicateExtraction:
      description: 'One or more requested extractions are already `pending` or `running` for the same entity, extractor, and options.

        The response lists the extractors whose equivalent extraction is already in progress; remove them before retrying the request.

        '
      content:
        application/ld+json:
          schema:
            type: object
            required:
            - conflictingExtractors
            properties:
              conflictingExtractors:
                type: array
                description: Extractors whose equivalent extraction is already in progress.
                minItems: 1
                uniqueItems: true
                items:
                  $ref: '#/components/schemas/ExtractionExtractor'
    Extraction:
      description: Extraction resource response
      content:
        application/ld+json:
          schema:
            allOf:
            - type: object
              properties:
                '@context':
                  type: string
                  default: /contexts/Extraction
            - $ref: '#/components/schemas/Extraction'
    ConstraintViolation:
      description: Validation error
      content:
        application/ld+json:
          schema:
            type: object
            properties:
              '@context':
                type: string
                default: /contexts/ConstraintViolation
              '@type':
                type: string
                default: ConstraintViolation
              hydra:title:
                type: string
                default: An error occurred
              hydra:description:
                type: string
                description: Concatenated violation messages
                example: 'exampleField: This value should not be blank.'
              violations:
                type: array
                items:
                  type: object
                  properties:
                    propertyPath:
                      type: string
                      description: Property access path
                      example: exampleField
                    message:
                      type: string
                      description: Validaton error message
                      example: This value should not be blank.
    NotFound:
      description: Not found
      content:
        application/json:
          schema:
            type: object
            properties:
              message:
                type: string
    ExtractionCollection:
      description: Extraction collection response
      content:
        application/ld+json:
          schema:
            $ref: '#/components/schemas/ExtractionCollection'
    Unauthorized:
      description: Unauthorized
      content:
        application/json:
          schema:
            type: object
            properties:
              message:
                type: string
  schemas:
    Taxpayer:
      type: object
      properties:
        '@id':
          type: string
          format: iri-reference
          description: Taxpayer IRI reference
          example: /taxpayers/PEIC211118IS0
        '@type':
          type: string
          default: Taxpayer
        id:
          $ref: '#/components/schemas/TaxpayerID'
        personType:
          $ref: '#/components/schemas/TaxpayerPersonType'
        registrationDate:
          type: string
          format: date
          description: Taxpayer registration date
        name:
          $ref: '#/components/schemas/TaxpayerName'
    updatedAt:
      type: string
      description: Date and time the resource was last updated
      example: '2020-01-01T12:15:00.000Z'
    CollectionLimit:
      type: integer
      default: 20
      minimum: 1
      maximum: 1000
    ExtractionUpdatedDataPoints:
      type: integer
      description: 'When your taxpayer has resources by a previous extraction and new extraction found changes between our data and Sat data it would update this resource. For example, when a new extraction finds a Canceled invoice which you previously extracted as an Active invoice.

        '
      example: 2
    ExtractionStatus:
      type: string
      enum:
      - pending
      - running
      - finished
      - failed
      - stopping
      - stopped
    ExtractionErrorCode:
      type:
      - string
      - 'null'
      description: '[Extraction error code](#section/Extraction-error-codes)

        '
      example: null
      enum:
      - invalid_credentials
      - login_failed
      - unrecoverable
      - sat_unavailable
      - internal_error
      - undefined
    CursorCollection:
      type: object
      properties:
        '@context':
          type: string
        '@id':
          type: string
        '@type':
          type: string
          default: hydra:Collection
        hydra:member:
          type: array
          items:
            type: object
        hydra:view:
          type: object
          description: Pagination information
          properties:
            '@id':
              type: string
              format: iri-reference
              description: Current page IRI reference
            '@type':
              type: string
              default: hydra:PartialCollectionView
            hydra:next:
              type: string
              example: /entity/2a15f539-3251-48e1-aaeb-a154dc9c6edb/resource?id[lt]=9b8e5365-0b36-45f5-9c76-fbe439632367
              description: Next page IRI reference; omitted when there is no pagination
            hydra:last:
              type: string
              example: /entity/2a15f539-3251-48e1-aaeb-a154dc9c6edb/resource?id[gt]=9b8e5365-0b36-45f5-9c76-fbe439632367
              description: Last page IRI reference; omitted when there is no pagination
        hydra:search:
          type: object
          properties:
            '@type':
              type: string
            hydra:template:
              type: string
            hydra:variableRepresentation:
              type: string
            hydra:mapping:
              type: array
              items:
                type: object
                properties:
                  '@type':
                    type: string
                  variable:
                    type: string
                  property:
                    type: string
                  required:
                    type: boolean
    CollectionOrder:
      type: string
      enum:
      - asc
      - desc
      example: asc
    ExtractionRateLimitedAt:
      type: string
      format: date-time
      description: 'This field returns a date when the taxpayer reached the max amount of downloaded CFDI files allowed by SAT in a range of 24hrs. You should retry this extraction 24hrs later. If your taxpayer has not reached the max amount this value should be `NULL`

        '
      example: '2019-01-03T21:10:40.000Z'
    ExtractionCollection:
      allOf:
      - $ref: '#/components/schemas/CursorCollection'
      - type: object
        properties:
          '@context':
            default: /contexts/Extraction
          '@id':
            default: /extractions
          hydra:member:
            items:
              $ref: '#/components/schemas/Extraction'
    TaxpayerName:
      type: string
      minLength: 1
      maxLength: 254
      description: Taxpayer name
      example: Pedro Infante Cruz
    ExtractionOptions:
      type: object
      description: Extractor options; check the available options for each extractor in [Extractions](https://docs.syntage.com/api-reference/extractions/overview#extractors)
      example:
        types:
        - I
        - E
        - P
        period:
          from: '2020-01-01T00:00:00.000Z'
          to: '2020-03-31T23:59:59.000Z'
    ExtractionCreatedDataPoints:
      type: integer
      description: 'This is the number of created resources. For example, if 1 invoice was created with its XML and PDF file this value should be 3. And If you previously extracted all the resources and there are no more resources to create this value is 0.

        '
      example: 3
    ExtractionExtractor:
      type: string
      description: Extractor to run, which retrieves resources from the corresponding datasource
      enum:
      - invoice
      - monthly_tax_return
      - annual_tax_return
      - rif_tax_return
      - tax_status
      - tax_retention
      - tax_compliance
      - electronic_accounting
      - sat_certificates
      - rpc
      - buro_de_credito_report
      - bil
      - rug
      - background_check
    TaxpayerID:
      type: string
      minLength: 12
      maxLength: 13
      description: RFC (Registro Federal de Contribuyentes)
      example: PEIC211118IS0
    createdAt:
      type: string
      description: Date and time the resource was created
      example: '2020-01-01T12:15:00.000Z'
    Extraction:
      type: object
      properties:
        '@id':
          type: string
          format: iri-reference
          description: Extraction IRI reference
          example: /extractions/91106968-1abd-4d64-85c1-4e73d96fb997
        '@type':
          type: string
          description: JSON-LD resource type
          default: Extraction
        id:
          type: string
          format: uuid
          description: Unique extraction ID
          example: 91106968-1abd-4d64-85c1-4e73d96fb997
        taxpayer:
          $ref: '#/components/schemas/Taxpayer'
        extractor:
          $ref: '#/components/schemas/ExtractionExtractor'
        options:
          $ref: '#/components/schemas/ExtractionOptions'
        status:
          $ref: '#/components/schemas/ExtractionStatus'
        startedAt:
          type:
          - string
          - 'null'
          format: date-time
          description: Time when the _status_ changed from **pending** to **running**
        finishedAt:
          type:
          - string
          - 'null'
          format: date-time
          description: Time when the _status_ changed from **running** to **finished** or **failed**
        rateLimitedAt:
          $ref: '#/components/schemas/ExtractionRateLimitedAt'
        errorCode:
          $ref: '#/components/schemas/ExtractionErrorCode'
        createdDataPoints:
          $ref: '#/components/schemas/ExtractionCreatedDataPoints'
        updatedDataPoints:
          $ref: '#/components/schemas/ExtractionUpdatedDataPoints'
        createdAt:
          $ref: '#/components/schemas/createdAt'
        updatedAt:
          $ref: '#/components/schemas/updatedAt'
    TaxpayerPersonType:
      type: string
      enum:
      - physical
      - legal
      example: physical
  parameters:
    resourceId:
      name: id
      in: path
      required: true
      example: 91106968-1abd-4d64-85c1-4e73d96fb997
      schema:
        type: string
        format: uuid
    collectionLimit:
      name: itemsPerPage
      in: query
      required: false
      description: Number of items per page
      schema:
        $ref: '#/components/schemas/CollectionLimit'
    createdAtStrictlyAfter:
      name: createdAt[strictly_after]
      in: query
      description: Filter by resource creation date (greater than `>`)
      example: '2020-01-17 11:47:27'
      schema:
        type: string
        format: date-time
    createdAtStrictlyBefore:
      name: createdAt[strictly_before]
      in: query
      description: Filter by resource creation date (less than `<`)
      example: '2020-01-17 11:47:27'
      schema:
        type: string
        format: date-time
    orderCreatedAt:
      name: order[createdAt]
      in: query
      description: Order by resource creation date
      schema:
        $ref: '#/components/schemas/CollectionOrder'
    createdAtBefore:
      name: createdAt[before]
      in: query
      description: Filter by resource creation date (less than or equal `<=`)
      example: '2020-01-17 11:47:27'
      schema:
        type: string
        format: date-time
    collectionCursorNextPageParam:
      name: id[lt]
      in: query
      required: false
      example: 91106968-1abd-4d64-85c1-4e73d96fb997
      description: Collection cursor pointer to the next page
      schema:
        type: string
    orderUpdatedAt:
      name: order[updatedAt]
      in: query
      description: Order by resource update date
      schema:
        $ref: '#/components/schemas/CollectionOrder'
    createdAtAfter:
      name: createdAt[after]
      in: query
      description: Filter by resource creation date (greater than or equal `>=`)
      example: '2020-01-17 11:47:27'
      schema:
        type: string
        format: date-time
    collectionCursorPreviousPageParam:
      name: id[gt]
      in: query
      required: false
      example: 91106968-1abd-4d64-85c1-4e73d96fb997
      description: Collection cursor pointer to the previous page
      schema:
        type: string
  requestBodies:
    ExtractionCreate:
      required: true
      content:
        application/json:
          schema:
            oneOf:
            - type: object
              required:
              - entity
              - extractor
              properties:
                entity:
                  type: string
                  format: iri-reference
                  description: Entity IRI reference
                  example: /entities/7b3a25a9-a53a-4846-abe6-f2574c9c2d5d
                extractor:
                  $ref: '#/components/schemas/ExtractionExtractor'
                options:
                  $ref: '#/components/schemas/ExtractionOptions'
            - type: object
              deprecated: true
              required:
              - taxpayer
              - extractor
              properties:
                taxpayer:
                  type: string
                  format: iri-reference
                  description: Taxpayer IRI reference
                  example: /taxpayers/PEIC211118IS0
                extractor:
                  $ref: '#/components/schemas/ExtractionExtractor'
                options:
                  $ref: '#/components/schemas/ExtractionOptions'
  securitySchemes:
    ApiKey:
      type: apiKey
      in: header
      name: X-API-Key
      description: 'Your API key is available in the [Production](https://app.syntage.com/settings/api-keys) and [Sandbox](https://app.sandbox.syntage.com/settings/api-keys) dashboards.

        '
x-readme:
  explorer-enabled: true
  proxy-enabled: true
  samples-enabled: true