Apache Tika
Apache Tika is a toolkit for detecting and extracting metadata and structured text content from over 1,000 file formats including PDF, Microsoft Office (Word, Excel, PowerPoint), OpenDocument, HTML, XML, images, audio, video, and archive formats. Tika provides a REST API server, Java library, and command-line tool. It is used by Apache Solr, Apache Nutch, and many other systems for content extraction and indexing. It is maintained by the Apache Software Foundation.
Apache Tika publishes 12 APIs on the APIs.io network, including Apache Tika Server REST API API, Detect API, Detectors API, and 9 more. Tagged areas include Content Extraction, Document Processing, Metadata, Text Extraction, and Open Source.
Apache Tika’s developer surface includes documentation, developer portal, getting-started guide, release notes, support, and 6 more developer resources.
Kin Score
APIs 13
Individual APIs this provider publishes, each with its own machine-readable definition.
Apache Tika Java API
The Tika Java API provides the AutoDetectParser for automatic format detection and parsing, Metadata class for reading extracted metadata fields, ContentHandler for streaming SA...
Apache Tika Apache Tika Server REST API API
The Apache Tika Server REST API API from Apache Tika — 1 operation(s) for apache tika server rest api.
Apache Tika Detect API
The Detect API from Apache Tika — 1 operation(s) for detect.
Apache Tika Detectors API
The Detectors API from Apache Tika — 1 operation(s) for detectors.
Apache Tika Language API
The Language API from Apache Tika — 2 operation(s) for language.
Apache Tika Meta API
The Meta API from Apache Tika — 3 operation(s) for meta.
Apache Tika Mime Types API
The Mime Types API from Apache Tika — 1 operation(s) for mime types.
Apache Tika Parsers API
The Parsers API from Apache Tika — 2 operation(s) for parsers.
Apache Tika Rmeta API
The Rmeta API from Apache Tika — 5 operation(s) for rmeta.
Apache Tika Status API
The Status API from Apache Tika — 1 operation(s) for status.
Apache Tika Tika API
The Tika API from Apache Tika — 4 operation(s) for tika.
Apache Tika Translate API
The Translate API from Apache Tika — 2 operation(s) for translate.
Apache Tika Unpack API
The Unpack API from Apache Tika — 2 operation(s) for unpack.
Scroll for all 13
Open Collections 1
Open, tool-agnostic API collections (OpenAPI-derived and Bruno).
Apache Tika Server REST API
OPEN COLLECTIONPricing Plans 1
Published pricing tiers and plan structures.
Rate Limits 1
Documented rate limits and quota policies.
Apache Tika Rate Limits
RATE LIMITSFinOps 1
Cost, billing, and metering signals for API financial operations.
Apache Tika Finops
FINOPSFeatures 7
Notable capabilities this provider offers.
1000+ Format Support
Detect and extract content from over 1,000 file formats using parser plugins.
Metadata Extraction
Extract document metadata including author, creation date, title, and format-specific properties.
Language Detection
Automatic language detection from extracted text content.
MIME Type Detection
Accurate MIME type detection based on file content (magic bytes) not just file extension.
REST Server
Standalone HTTP server for document processing without Java library dependency.
OCR Integration
Optional Tesseract OCR integration for text extraction from images and scanned PDFs.
Recursive Parsing
Recursive parsing of archive formats (ZIP, TAR, JAR) and embedded documents.
Scroll for all 7
Security Posture 2
Authentication, domain security, vulnerability disclosure, and trust-center signals.
Agentic Access 1
Recommended x-agentic-access execution contracts for AI agents.
Use Cases 4
What developers build with this provider.
Search Indexing
Extract text from documents for indexing in Apache Solr or Elasticsearch.
Document Intelligence
Automated metadata extraction and classification for document management systems.
Content Migration
Batch content extraction during digital archive migration and transformation.
E-Discovery
Legal e-discovery content extraction from diverse document collections.
Integrations 5
Pre-built integrations with other platforms and tools.
Apache Solr
Solr Cell uses Tika for extracting text from uploaded documents for indexing.
Apache Nutch
Nutch web crawler uses Tika for parsing fetched web page content.
Elasticsearch
Ingest attachment processor uses Tika for document content extraction.
Tesseract OCR
Optional Tesseract integration for OCR on images and scanned documents.
Apache NiFi
NiFi processor integration for automated document parsing workflows.
Resources
Get Started 2
Portal, sign-up, and the first successful call
Documentation 1
Reference material describing how the API behaves
Agent Surfaces 1
MCP servers, agent skills, and machine-readable catalogs
Build 2
SDKs, sample code, and the tooling you integrate with
Access & Security 2
Authentication, authorization, and security posture
Operate 2
Status, limits, changes, and where to get help
Commercial 1
Pricing, plans, and the legal terms of use