Apache PDFBox
Apache PDFBox is an open-source Java library for working with PDF documents. It allows creation of new PDF documents, manipulation of existing documents, and the ability to extract content from documents with support for digital signatures.
Apache PDFBox publishes 5 APIs on the APIs.io network, including Documents API, Forms API, Operations API, and 2 more. Tagged areas include Document Processing, Java, PDF, Text Extraction, and Apache.
The Apache PDFBox catalog on APIs.io includes 1 JSON-LD context and 2 Spectral governance rulesets.
Apache PDFBox’s developer surface includes documentation, engineering blog, and 7 more developer resources.
Kin Score
APIs 5
Individual APIs this provider publishes, each with its own machine-readable definition.
Apache PDFBox Documents API
The Documents API from Apache PDFBox — 3 operation(s) for documents.
Apache PDFBox Forms API
The Forms API from Apache PDFBox — 1 operation(s) for forms.
Apache PDFBox Operations API
The Operations API from Apache PDFBox — 2 operation(s) for operations.
Apache PDFBox Pages API
The Pages API from Apache PDFBox — 1 operation(s) for pages.
Apache PDFBox Signatures API
The Signatures API from Apache PDFBox — 1 operation(s) for signatures.
Pricing Plans 1
Published pricing tiers and plan structures.
Rate Limits 1
Documented rate limits and quota policies.
Apache Pdfbox Rate Limits
RATE LIMITSFinOps 1
Cost, billing, and metering signals for API financial operations.
Apache Pdfbox Finops
FINOPSFeatures 7
Notable capabilities this provider offers.
PDF Text Extraction
Extract plain text and structured content from PDF documents
PDF Creation
Create new PDF documents programmatically with Java API
PDF Manipulation
Merge, split, rotate, and resize pages in existing PDFs
Digital Signatures
Apply and verify digital signatures for document authenticity
Form Filling
Read and fill interactive PDF forms (AcroForms)
PDF/A Validation
Validate and create PDF/A documents for archiving
Font Handling
Embed and extract fonts, handle Type 1, TrueType, and OpenType
Scroll for all 7
Semantic Vocabularies 1
JSON-LD contexts and semantic vocabularies used across these APIs.
Apache Pdfbox Context
JSON-LDSpectral Rules 2
Spectral governance rulesets for linting and validating these APIs.
Apache PDFBox API Rules
SPECTRALApache PDFBox API Rules
SPECTRALJSON Schema 11
Standalone JSON Schema definitions for this provider's data models.
CreateDocumentRequest
JSON SCHEMADocumentInfo
JSON SCHEMADocumentMetadata
JSON SCHEMAFormField
JSON SCHEMAFormFields
JSON SCHEMAMergeRequest
JSON SCHEMAPageInfo
JSON SCHEMAPageList
JSON SCHEMASignRequest
JSON SCHEMASplitRequest
JSON SCHEMATextExtractionResult
JSON SCHEMAScroll for all 11
JSON Structure 11
JSON Structure definitions describing this provider's data shapes.
Apache Pdfbox Create Document Request Structure
JSON STRUCTUREApache Pdfbox Document Info Structure
JSON STRUCTUREApache Pdfbox Document Metadata Structure
JSON STRUCTUREApache Pdfbox Form Field Structure
JSON STRUCTUREApache Pdfbox Form Fields Structure
JSON STRUCTUREApache Pdfbox Merge Request Structure
JSON STRUCTUREApache Pdfbox Page Info Structure
JSON STRUCTUREApache Pdfbox Page List Structure
JSON STRUCTUREApache Pdfbox Sign Request Structure
JSON STRUCTUREApache Pdfbox Split Request Structure
JSON STRUCTUREApache Pdfbox Text Extraction Result Structure
JSON STRUCTUREScroll for all 11
Examples 11
Example request and response payloads for these APIs.
Scroll for all 11
Security Posture 2
Authentication, domain security, vulnerability disclosure, and trust-center signals.
Agentic Access 1
Recommended x-agentic-access execution contracts for AI agents.
Use Cases 5
What developers build with this provider.
Invoice Processing
Extract data from PDF invoices for automated processing
Document Generation
Generate PDF reports, contracts, and certificates programmatically
Legal Document Management
Digitally sign and verify legal documents
Form Data Collection
Fill PDF forms and extract submitted data
Archive Management
Convert documents to PDF/A for long-term archiving
Integrations 4
Pre-built integrations with other platforms and tools.
Apache Tika
Content detection and text extraction integration
Spring Boot
Spring Boot starter for PDF processing in web applications
Maven Central
Available as org.apache.pdfbox on Maven Central
iText/OpenPDF
Complementary PDF library for advanced PDF generation
Resources
Documentation 1
Reference material describing how the API behaves
Agent Surfaces 1
MCP servers, agent skills, and machine-readable catalogs
Design & Contract 3
Pagination, idempotency, versioning, errors, and events
Build 1
SDKs, sample code, and the tooling you integrate with
Access & Security 2
Authentication, authorization, and security posture
Company 1
The organization behind the API