Apache Tika website screenshot

Apache Tika

Apache Tika is a toolkit for detecting and extracting metadata and structured text content from over 1,000 file formats including PDF, Microsoft Office (Word, Excel, PowerPoint), OpenDocument, HTML, XML, images, audio, video, and archive formats. Tika provides a REST API server, Java library, and command-line tool. It is used by Apache Solr, Apache Nutch, and many other systems for content extraction and indexing. It is maintained by the Apache Software Foundation.

Apache Tika publishes 12 APIs on the APIs.io network, including Apache Tika Server REST API API, Detect API, Detectors API, and 9 more. Tagged areas include Content Extraction, Document Processing, Metadata, Text Extraction, and Open Source.

Apache Tika’s developer surface includes documentation, developer portal, getting-started guide, release notes, support, and 6 more developer resources.

38.9/100 thin ▬ flat Agent 22/100 agent aware Full breakdown ↓
scored 2026-07-28 · rubric v0.6
AccessFreemium
13 APIs 7 Features 4 Use Cases
Content ExtractionDocument ProcessingMetadataText ExtractionOpen Source

Kin Score

Kin Score Kin Score How this is scored →
scored 2026-07-28 · rubric v0.6
Composite quality — 38.9/100 · thin
Contract Quality 8.5 / 25
Developer Ergonomics 7.8 / 20
Commercial Clarity 10.0 / 20
Operational Transparency 6.2 / 13
Governance 0.0 / 12
Discoverability 6.5 / 10
Agent readiness — 22/100 · agent aware
Machine-Readable Contract 18 / 18
Agentic Access Contract 10 / 10
MCP Server 0 / 12
Machine-Readable Auth 0 / 10
Idempotency 0 / 9
Stable Error Semantics 0 / 8
Request/Response Examples 0 / 7
Rate-Limit Signaling 7 / 7
Typed Event Surface 0 / 6
Agent Skills 0 / 5
Well-Known Catalog 0 / 4
Consent & Bot Identity 0 / 3
A2A Agent Card 0 / 8
Dry-Run / Simulate Mode 0 / 4
Improve this rating by publishing the missing artifacts — every area above can be raised, and the full rubric is at apis.io/rating/. This rating is computed from github.com/api-evangelist/apache-tika: open an issue to ask a question, or submit a pull request to add artifacts. Want it done for you? Prioritized profiling — $2,500 →

APIs 13

Individual APIs this provider publishes, each with its own machine-readable definition.

Apache Tika Java API

The Tika Java API provides the AutoDetectParser for automatic format detection and parsing, Metadata class for reading extracted metadata fields, ContentHandler for streaming SA...

Apache Tika Apache Tika Server REST API API

The Apache Tika Server REST API API from Apache Tika — 1 operation(s) for apache tika server rest api.

Apache Tika Detect API

The Detect API from Apache Tika — 1 operation(s) for detect.

Apache Tika Detectors API

The Detectors API from Apache Tika — 1 operation(s) for detectors.

Apache Tika Language API

The Language API from Apache Tika — 2 operation(s) for language.

Apache Tika Meta API

The Meta API from Apache Tika — 3 operation(s) for meta.

Apache Tika Mime Types API

The Mime Types API from Apache Tika — 1 operation(s) for mime types.

Apache Tika Parsers API

The Parsers API from Apache Tika — 2 operation(s) for parsers.

Apache Tika Rmeta API

The Rmeta API from Apache Tika — 5 operation(s) for rmeta.

Apache Tika Status API

The Status API from Apache Tika — 1 operation(s) for status.

Apache Tika Tika API

The Tika API from Apache Tika — 4 operation(s) for tika.

Apache Tika Translate API

The Translate API from Apache Tika — 2 operation(s) for translate.

Apache Tika Unpack API

The Unpack API from Apache Tika — 2 operation(s) for unpack.

Scroll for all 13

Open Collections 1

Open, tool-agnostic API collections (OpenAPI-derived and Bruno).

Pricing Plans 1

Published pricing tiers and plan structures.

Rate Limits 1

Documented rate limits and quota policies.

Apache Tika Rate Limits

5 limits

RATE LIMITS

FinOps 1

Cost, billing, and metering signals for API financial operations.

Features 7

Notable capabilities this provider offers.

1000+ Format Support

Detect and extract content from over 1,000 file formats using parser plugins.

Metadata Extraction

Extract document metadata including author, creation date, title, and format-specific properties.

Language Detection

Automatic language detection from extracted text content.

MIME Type Detection

Accurate MIME type detection based on file content (magic bytes) not just file extension.

REST Server

Standalone HTTP server for document processing without Java library dependency.

OCR Integration

Optional Tesseract OCR integration for text extraction from images and scanned PDFs.

Recursive Parsing

Recursive parsing of archive formats (ZIP, TAR, JAR) and embedded documents.

Scroll for all 7

Security Posture 2

Authentication, domain security, vulnerability disclosure, and trust-center signals.

Apache Tika Domain Security

TLSv1.3 · HSTS · DMARC

SECURITY

Apache Tika Vulnerability Disclosure

security.txt · contact published

SECURITY

Agentic Access 1

Recommended x-agentic-access execution contracts for AI agents.

Apache Tika Agentic Access

33 operations · 23 acting

33 operations · 23 acting

AGENTIC

Use Cases 4

What developers build with this provider.

Search Indexing

Extract text from documents for indexing in Apache Solr or Elasticsearch.

Document Intelligence

Automated metadata extraction and classification for document management systems.

Content Migration

Batch content extraction during digital archive migration and transformation.

E-Discovery

Legal e-discovery content extraction from diverse document collections.

Integrations 5

Pre-built integrations with other platforms and tools.

Apache Solr

Solr Cell uses Tika for extracting text from uploaded documents for indexing.

Apache Nutch

Nutch web crawler uses Tika for parsing fetched web page content.

Elasticsearch

Ingest attachment processor uses Tika for document content extraction.

Tesseract OCR

Optional Tesseract integration for OCR on images and scanned documents.

Apache NiFi

NiFi processor integration for automated document parsing workflows.

Resources

Get Started 2

Portal, sign-up, and the first successful call

Documentation 1

Reference material describing how the API behaves

Agent Surfaces 1

MCP servers, agent skills, and machine-readable catalogs

Build 2

SDKs, sample code, and the tooling you integrate with

Access & Security 2

Authentication, authorization, and security posture

Operate 2

Status, limits, changes, and where to get help

Commercial 1

Pricing, plans, and the legal terms of use

Source (apis.yml)

apis.yml Raw ↑
aid: apache-tika
name: Apache Tika
description: Apache Tika is a toolkit for detecting and extracting metadata and structured text content from over 1,000 file
  formats including PDF, Microsoft Office (Word, Excel, PowerPoint), OpenDocument, HTML, XML, images, audio, video, and archive
  formats. Tika provides a REST API server, Java library, and command-line tool. It is used by Apache Solr, Apache Nutch,
  and many other systems for content extraction and indexing. It is maintained by the Apache Software Foundation.
type: Index
accessModel:
  pricing: freemium
  onboarding: unknown
  trial: false
  try_now: false
  public: false
  label: Freemium
  confidence: medium
  source:
  - plans
  generated: '2026-07-22'
  method: derived
position: Consumer
access: 3rd-Party
image: https://kinlane-images.s3.amazonaws.com/shared/apis-json/icons/apache-tika.png
tags:
- Content Extraction
- Document Processing
- Metadata
- Text Extraction
- Open Source
created: '2026-03-16'
modified: '2026-05-19'
url: https://raw.githubusercontent.com/api-evangelist/apache-tika/refs/heads/main/apis.yml
specificationVersion: '0.19'
apis:
- aid: apache-tika:apache-tika-java-api
  name: Apache Tika Java API
  description: The Tika Java API provides the AutoDetectParser for automatic format detection and parsing, Metadata class
    for reading extracted metadata fields, ContentHandler for streaming SAX-based text extraction, and Detector for MIME type
    identification. The facade Tika class provides a simple one-line API for text extraction from any supported format.
  humanURL: https://tika.apache.org/
  tags:
  - Java
  - Content Extraction
  - Parser
  - Metadata
  properties:
  - type: Documentation
    url: https://tika.apache.org/
  - type: APIReference
    url: https://tika.apache.org/1.28/api/
  - type: SDKs
    url: https://search.maven.org/search?q=org.apache.tika
    title: Maven Java SDK
- aid: apache-tika:apache-tika-apache-tika-server-rest-api-api
  name: Apache Tika Apache Tika Server REST API API
  description: The Apache Tika Server REST API API from Apache Tika — 1 operation(s) for apache tika server rest api.
  humanURL: https://cwiki.apache.org/confluence/display/TIKA/TikaServer
  tags:
  - Apache Tika Server REST API
  properties:
  - type: OpenAPI
    url: openapi/apache-tika-apache-tika-server-rest-api-api-openapi.yml
  - type: Documentation
    url: https://cwiki.apache.org/confluence/display/TIKA/TikaServer
  - type: Documentation
    url: https://tika.apache.org/
- aid: apache-tika:apache-tika-detect-api
  name: Apache Tika Detect API
  description: The Detect API from Apache Tika — 1 operation(s) for detect.
  humanURL: https://cwiki.apache.org/confluence/display/TIKA/TikaServer
  tags:
  - Detect
  properties:
  - type: OpenAPI
    url: openapi/apache-tika-detect-api-openapi.yml
  - type: Documentation
    url: https://cwiki.apache.org/confluence/display/TIKA/TikaServer
  - type: Documentation
    url: https://tika.apache.org/
- aid: apache-tika:apache-tika-detectors-api
  name: Apache Tika Detectors API
  description: The Detectors API from Apache Tika — 1 operation(s) for detectors.
  humanURL: https://cwiki.apache.org/confluence/display/TIKA/TikaServer
  tags:
  - Detectors
  properties:
  - type: OpenAPI
    url: openapi/apache-tika-detectors-api-openapi.yml
  - type: Documentation
    url: https://cwiki.apache.org/confluence/display/TIKA/TikaServer
  - type: Documentation
    url: https://tika.apache.org/
- aid: apache-tika:apache-tika-language-api
  name: Apache Tika Language API
  description: The Language API from Apache Tika — 2 operation(s) for language.
  humanURL: https://cwiki.apache.org/confluence/display/TIKA/TikaServer
  tags:
  - Language
  properties:
  - type: OpenAPI
    url: openapi/apache-tika-language-api-openapi.yml
  - type: Documentation
    url: https://cwiki.apache.org/confluence/display/TIKA/TikaServer
  - type: Documentation
    url: https://tika.apache.org/
- aid: apache-tika:apache-tika-meta-api
  name: Apache Tika Meta API
  description: The Meta API from Apache Tika — 3 operation(s) for meta.
  humanURL: https://cwiki.apache.org/confluence/display/TIKA/TikaServer
  tags:
  - Meta
  properties:
  - type: OpenAPI
    url: openapi/apache-tika-meta-api-openapi.yml
  - type: Documentation
    url: https://cwiki.apache.org/confluence/display/TIKA/TikaServer
  - type: Documentation
    url: https://tika.apache.org/
- aid: apache-tika:apache-tika-mime-types-api
  name: Apache Tika Mime Types API
  description: The Mime Types API from Apache Tika — 1 operation(s) for mime types.
  humanURL: https://cwiki.apache.org/confluence/display/TIKA/TikaServer
  tags:
  - Mime Types
  properties:
  - type: OpenAPI
    url: openapi/apache-tika-mime-types-api-openapi.yml
  - type: Documentation
    url: https://cwiki.apache.org/confluence/display/TIKA/TikaServer
  - type: Documentation
    url: https://tika.apache.org/
- aid: apache-tika:apache-tika-parsers-api
  name: Apache Tika Parsers API
  description: The Parsers API from Apache Tika — 2 operation(s) for parsers.
  humanURL: https://cwiki.apache.org/confluence/display/TIKA/TikaServer
  tags:
  - Parsers
  properties:
  - type: OpenAPI
    url: openapi/apache-tika-parsers-api-openapi.yml
  - type: Documentation
    url: https://cwiki.apache.org/confluence/display/TIKA/TikaServer
  - type: Documentation
    url: https://tika.apache.org/
- aid: apache-tika:apache-tika-rmeta-api
  name: Apache Tika Rmeta API
  description: The Rmeta API from Apache Tika — 5 operation(s) for rmeta.
  humanURL: https://cwiki.apache.org/confluence/display/TIKA/TikaServer
  tags:
  - Rmeta
  properties:
  - type: OpenAPI
    url: openapi/apache-tika-rmeta-api-openapi.yml
  - type: Documentation
    url: https://cwiki.apache.org/confluence/display/TIKA/TikaServer
  - type: Documentation
    url: https://tika.apache.org/
- aid: apache-tika:apache-tika-status-api
  name: Apache Tika Status API
  description: The Status API from Apache Tika — 1 operation(s) for status.
  humanURL: https://cwiki.apache.org/confluence/display/TIKA/TikaServer
  tags:
  - Status
  properties:
  - type: OpenAPI
    url: openapi/apache-tika-status-api-openapi.yml
  - type: Documentation
    url: https://cwiki.apache.org/confluence/display/TIKA/TikaServer
  - type: Documentation
    url: https://tika.apache.org/
- aid: apache-tika:apache-tika-tika-api
  name: Apache Tika Tika API
  description: The Tika API from Apache Tika — 4 operation(s) for tika.
  humanURL: https://cwiki.apache.org/confluence/display/TIKA/TikaServer
  tags:
  - Tika
  properties:
  - type: OpenAPI
    url: openapi/apache-tika-tika-api-openapi.yml
  - type: Documentation
    url: https://cwiki.apache.org/confluence/display/TIKA/TikaServer
  - type: Documentation
    url: https://tika.apache.org/
- aid: apache-tika:apache-tika-translate-api
  name: Apache Tika Translate API
  description: The Translate API from Apache Tika — 2 operation(s) for translate.
  humanURL: https://cwiki.apache.org/confluence/display/TIKA/TikaServer
  tags:
  - Translate
  properties:
  - type: OpenAPI
    url: openapi/apache-tika-translate-api-openapi.yml
  - type: Documentation
    url: https://cwiki.apache.org/confluence/display/TIKA/TikaServer
  - type: Documentation
    url: https://tika.apache.org/
- aid: apache-tika:apache-tika-unpack-api
  name: Apache Tika Unpack API
  description: The Unpack API from Apache Tika — 2 operation(s) for unpack.
  humanURL: https://cwiki.apache.org/confluence/display/TIKA/TikaServer
  tags:
  - Unpack
  properties:
  - type: OpenAPI
    url: openapi/apache-tika-unpack-api-openapi.yml
  - type: Documentation
    url: https://cwiki.apache.org/confluence/display/TIKA/TikaServer
  - type: Documentation
    url: https://tika.apache.org/
common:
- type: AgenticAccess
  url: agentic-access/apache-tika-agentic-access.yml
- type: VulnerabilityDisclosure
  url: security/apache-tika-vulnerability-disclosure.yml
- type: DomainSecurity
  url: security/apache-tika-domain-security.yml
- type: GitHubRepository
  url: https://github.com/apache/tika
- type: Documentation
  url: https://tika.apache.org/
- type: Portal
  url: https://tika.apache.org/
- type: GettingStarted
  url: https://tika.apache.org/gettingstarted.html
- type: ReleaseNotes
  url: https://github.com/apache/tika/releases
- type: Support
  url: https://cwiki.apache.org/confluence/display/TIKA/MailingLists
- type: TermsOfService
  url: https://www.apache.org/licenses/
- type: SDKs
  url: https://pypi.org/project/tika/
  title: Python Tika Package
- type: Features
  data:
  - name: 1000+ Format Support
    description: Detect and extract content from over 1,000 file formats using parser plugins.
  - name: Metadata Extraction
    description: Extract document metadata including author, creation date, title, and format-specific properties.
  - name: Language Detection
    description: Automatic language detection from extracted text content.
  - name: MIME Type Detection
    description: Accurate MIME type detection based on file content (magic bytes) not just file extension.
  - name: REST Server
    description: Standalone HTTP server for document processing without Java library dependency.
  - name: OCR Integration
    description: Optional Tesseract OCR integration for text extraction from images and scanned PDFs.
  - name: Recursive Parsing
    description: Recursive parsing of archive formats (ZIP, TAR, JAR) and embedded documents.
- type: UseCases
  data:
  - name: Search Indexing
    description: Extract text from documents for indexing in Apache Solr or Elasticsearch.
  - name: Document Intelligence
    description: Automated metadata extraction and classification for document management systems.
  - name: Content Migration
    description: Batch content extraction during digital archive migration and transformation.
  - name: E-Discovery
    description: Legal e-discovery content extraction from diverse document collections.
- type: Integrations
  data:
  - name: Apache Solr
    description: Solr Cell uses Tika for extracting text from uploaded documents for indexing.
  - name: Apache Nutch
    description: Nutch web crawler uses Tika for parsing fetched web page content.
  - name: Elasticsearch
    description: Ingest attachment processor uses Tika for document content extraction.
  - name: Tesseract OCR
    description: Optional Tesseract integration for OCR on images and scanned documents.
  - name: Apache NiFi
    description: NiFi processor integration for automated document parsing workflows.
maintainers:
- FN: Kin Lane
  email: info@apievangelist.com