Amazon Comprehend · Schema

DatasetInputDataConfig

Specifies the format and location of the input data for the dataset.

Machine-LearningNatural Language ProcessingNLPText Analysis

Properties

Name Type Description
AugmentedManifests object
DataFormat object
DocumentClassifierInputDataConfig object
EntityRecognizerInputDataConfig object
View JSON Schema on GitHub

JSON Schema

openapi.yml-dataset-input-data-config-schema.json Raw ↑
{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "$id": "https://raw.githubusercontent.com/api-evangelist/amazon-comprehend/refs/heads/main/json-schema/openapi.yml-dataset-input-data-config-schema.json",
  "title": "DatasetInputDataConfig",
  "description": "Specifies the format and location of the input data for the dataset.",
  "type": "object",
  "properties": {
    "AugmentedManifests": {
      "allOf": [
        {
          "$ref": "#/components/schemas/DatasetAugmentedManifestsList"
        },
        {
          "description": "A list of augmented manifest files that provide training data for your custom model. An augmented manifest file is a labeled dataset that is produced by Amazon SageMaker Ground Truth. "
        }
      ]
    },
    "DataFormat": {
      "allOf": [
        {
          "$ref": "#/components/schemas/DatasetDataFormat"
        },
        {
          "description": "<p> <code>COMPREHEND_CSV</code>: The data format is a two-column CSV file, where the first column contains labels and the second column contains documents.</p> <p> <code>AUGMENTED_MANIFEST</code>: The data format </p>"
        }
      ]
    },
    "DocumentClassifierInputDataConfig": {
      "allOf": [
        {
          "$ref": "#/components/schemas/DatasetDocumentClassifierInputDataConfig"
        },
        {
          "description": "<p>The input properties for training a document classifier model. </p> <p>For more information on how the input file is formatted, see <a href=\"https://docs.aws.amazon.com/comprehend/latest/dg/prep-classifier-data.html\">Preparing training data</a> in the Comprehend Developer Guide. </p>"
        }
      ]
    },
    "EntityRecognizerInputDataConfig": {
      "allOf": [
        {
          "$ref": "#/components/schemas/DatasetEntityRecognizerInputDataConfig"
        },
        {
          "description": "The input properties for training an entity recognizer model."
        }
      ]
    }
  }
}

Work with this as data

Every JSON Schema here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for schemas

4 MCP tools reach this
  • find_json_schemasBrowse and filter every JSON Schema in the catalog.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools

Call it yourself

curl for this page
This JSON Schema
curl "https://apis.io/api/v1/json-schemas/openapi.yml-dataset-input-data-config"
All schemas
curl "https://apis.io/api/v1/json-schemas?limit=25"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no email required.

A second provider on the same verified email joins the account you already have.