Amazon Comprehend · Schema

DatasetEntityRecognizerDocuments

Describes the documents submitted with a dataset for an entity recognizer model.

Machine LearningNatural Language ProcessingNLPText Analysis

Properties

Name Type Description
S3Uri object
InputFormat object
View JSON Schema on GitHub

JSON Schema

openapi.yml-dataset-entity-recognizer-documents-schema.json Raw ↑
{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "$id": "https://raw.githubusercontent.com/api-evangelist/amazon-comprehend/refs/heads/main/json-schema/openapi.yml-dataset-entity-recognizer-documents-schema.json",
  "title": "DatasetEntityRecognizerDocuments",
  "description": "Describes the documents submitted with a dataset for an entity recognizer model.",
  "type": "object",
  "properties": {
    "S3Uri": {
      "allOf": [
        {
          "$ref": "#/components/schemas/S3Uri"
        },
        {
          "description": " Specifies the Amazon S3 location where the documents for the dataset are located. "
        }
      ]
    },
    "InputFormat": {
      "allOf": [
        {
          "$ref": "#/components/schemas/InputFormat"
        },
        {
          "description": " Specifies how the text in an input file should be processed. This is optional, and the default is ONE_DOC_PER_LINE. ONE_DOC_PER_FILE - Each file is considered a separate document. Use this option when you are processing large documents, such as newspaper articles or scientific papers. ONE_DOC_PER_LINE - Each line in a file is considered a separate document. Use this option when you are processing many short documents, such as text messages."
        }
      ]
    }
  },
  "required": [
    "S3Uri"
  ]
}