Apache PDFBox website screenshot

Apache PDFBox

Apache PDFBox is an open-source Java library for working with PDF documents. It allows creation of new PDF documents, manipulation of existing documents, and the ability to extract content from documents with support for digital signatures.

Apache PDFBox publishes 5 APIs on the APIs.io network, including Documents API, Forms API, Operations API, and 2 more. Tagged areas include Document Processing, Java, PDF, Text Extraction, and Apache.

The Apache PDFBox catalog on APIs.io includes 1 JSON-LD context and 2 Spectral governance rulesets.

Apache PDFBox’s developer surface includes documentation, engineering blog, and 7 more developer resources.

40.4/100 thin ▼ -9.1 Agent 22/100 agent aware Full breakdown ↓
scored 2026-07-28 · rubric v0.6
AccessFreemium
5 APIs 7 Features 5 Use Cases
Document ProcessingJavaPDFText ExtractionApacheOpen Source

Kin Score

Kin Score Kin Score How this is scored →
scored 2026-07-28 · rubric v0.6
Composite quality — 40.4/100 · thin
Contract Quality 10.8 / 25
Developer Ergonomics 2.2 / 20
Commercial Clarity 7.9 / 20
Operational Transparency 4.8 / 13
Governance 8.3 / 12
Discoverability 6.5 / 10
Agent readiness — 22/100 · agent aware
Machine-Readable Contract 18 / 18
Agentic Access Contract 10 / 10
MCP Server 0 / 12
Machine-Readable Auth 0 / 10
Idempotency 0 / 9
Stable Error Semantics 0 / 8
Request/Response Examples 0 / 7
Rate-Limit Signaling 7 / 7
Typed Event Surface 0 / 6
Agent Skills 0 / 5
Well-Known Catalog 0 / 4
Consent & Bot Identity 0 / 3
A2A Agent Card 0 / 8
Dry-Run / Simulate Mode 0 / 4
Improve this rating by publishing the missing artifacts — every area above can be raised, and the full rubric is at apis.io/rating/. This rating is computed from github.com/api-evangelist/apache-pdfbox: open an issue to ask a question, or submit a pull request to add artifacts. Want it done for you? Prioritized profiling — $2,500 →

APIs 5

Individual APIs this provider publishes, each with its own machine-readable definition.

Apache PDFBox Documents API

The Documents API from Apache PDFBox — 3 operation(s) for documents.

Apache PDFBox Forms API

The Forms API from Apache PDFBox — 1 operation(s) for forms.

Apache PDFBox Operations API

The Operations API from Apache PDFBox — 2 operation(s) for operations.

Apache PDFBox Pages API

The Pages API from Apache PDFBox — 1 operation(s) for pages.

Apache PDFBox Signatures API

The Signatures API from Apache PDFBox — 1 operation(s) for signatures.

Pricing Plans 1

Published pricing tiers and plan structures.

Rate Limits 1

Documented rate limits and quota policies.

Apache Pdfbox Rate Limits

5 limits

RATE LIMITS

FinOps 1

Cost, billing, and metering signals for API financial operations.

Features 7

Notable capabilities this provider offers.

PDF Text Extraction

Extract plain text and structured content from PDF documents

PDF Creation

Create new PDF documents programmatically with Java API

PDF Manipulation

Merge, split, rotate, and resize pages in existing PDFs

Digital Signatures

Apply and verify digital signatures for document authenticity

Form Filling

Read and fill interactive PDF forms (AcroForms)

PDF/A Validation

Validate and create PDF/A documents for archiving

Font Handling

Embed and extract fonts, handle Type 1, TrueType, and OpenType

Scroll for all 7

Semantic Vocabularies 1

JSON-LD contexts and semantic vocabularies used across these APIs.

Apache Pdfbox Context

11 classes · 32 properties

JSON-LD

Spectral Rules 2

Spectral governance rulesets for linting and validating these APIs.

Apache PDFBox API Rules

5 rules · 3 warnings 2 info

SPECTRAL

Apache PDFBox API Rules

12 rules · 5 errors 4 warnings 3 info

SPECTRAL

JSON Schema 11

Standalone JSON Schema definitions for this provider's data models.

CreateDocumentRequest

3 properties

JSON SCHEMA

DocumentInfo

5 properties

JSON SCHEMA

DocumentMetadata

7 properties

JSON SCHEMA

FormField

4 properties

JSON SCHEMA

FormFields

2 properties

JSON SCHEMA

MergeRequest

2 properties

JSON SCHEMA

PageInfo

4 properties

JSON SCHEMA

PageList

3 properties

JSON SCHEMA

SignRequest

4 properties

JSON SCHEMA

SplitRequest

1 properties

JSON SCHEMA

TextExtractionResult

4 properties

JSON SCHEMA

Scroll for all 11

JSON Structure 11

JSON Structure definitions describing this provider's data shapes.

Apache Pdfbox Document Info Structure

5 properties

JSON STRUCTURE

Apache Pdfbox Document Metadata Structure

7 properties

JSON STRUCTURE

Apache Pdfbox Form Field Structure

4 properties

JSON STRUCTURE

Apache Pdfbox Form Fields Structure

2 properties

JSON STRUCTURE

Apache Pdfbox Merge Request Structure

2 properties

JSON STRUCTURE

Apache Pdfbox Page Info Structure

4 properties

JSON STRUCTURE

Apache Pdfbox Page List Structure

3 properties

JSON STRUCTURE

Apache Pdfbox Sign Request Structure

4 properties

JSON STRUCTURE

Apache Pdfbox Split Request Structure

1 properties

JSON STRUCTURE

Scroll for all 11

Examples 11

Example request and response payloads for these APIs.

Scroll for all 11

Security Posture 2

Authentication, domain security, vulnerability disclosure, and trust-center signals.

Apache Pdfbox Domain Security

TLSv1.3 · HSTS · DMARC

SECURITY

Apache Pdfbox Vulnerability Disclosure

security.txt · contact published

SECURITY

Agentic Access 1

Recommended x-agentic-access execution contracts for AI agents.

Apache Pdfbox Agentic Access

8 operations · 4 acting

8 operations · 4 acting

AGENTIC

Use Cases 5

What developers build with this provider.

Invoice Processing

Extract data from PDF invoices for automated processing

Document Generation

Generate PDF reports, contracts, and certificates programmatically

Legal Document Management

Digitally sign and verify legal documents

Form Data Collection

Fill PDF forms and extract submitted data

Archive Management

Convert documents to PDF/A for long-term archiving

Integrations 4

Pre-built integrations with other platforms and tools.

Apache Tika

Content detection and text extraction integration

Spring Boot

Spring Boot starter for PDF processing in web applications

Maven Central

Available as org.apache.pdfbox on Maven Central

iText/OpenPDF

Complementary PDF library for advanced PDF generation

Resources

Documentation 1

Reference material describing how the API behaves

Agent Surfaces 1

MCP servers, agent skills, and machine-readable catalogs

Design & Contract 3

Pagination, idempotency, versioning, errors, and events

Build 1

SDKs, sample code, and the tooling you integrate with

Access & Security 2

Authentication, authorization, and security posture

Company 1

The organization behind the API

Source (apis.yml)

apis.yml Raw ↑
aid: apache-pdfbox
name: Apache PDFBox
description: Apache PDFBox is an open-source Java library for working with PDF documents. It allows creation of new PDF documents,
  manipulation of existing documents, and the ability to extract content from documents with support for digital signatures.
type: Index
accessModel:
  pricing: freemium
  onboarding: unknown
  trial: false
  try_now: false
  public: false
  label: Freemium
  confidence: medium
  source:
  - plans
  generated: '2026-07-22'
  method: derived
position: Consumer
access: 3rd-Party
image: https://kinlane-images.s3.amazonaws.com/shared/apis-json/icons/apache-pdfbox.png
tags:
- Document Processing
- Java
- PDF
- Text Extraction
- Apache
- Open Source
created: '2026-03-16'
modified: '2026-05-19'
url: https://raw.githubusercontent.com/api-evangelist/apache-pdfbox/refs/heads/main/apis.yml
specificationVersion: '0.19'
apis:
- aid: apache-pdfbox:apache-pdfbox-documents-api
  name: Apache PDFBox Documents API
  description: The Documents API from Apache PDFBox — 3 operation(s) for documents.
  humanURL: https://pdfbox.apache.org/2.0/getting-started.html
  tags:
  - Documents
  properties:
  - type: OpenAPI
    url: openapi/apache-pdfbox-documents-api-openapi.yml
  - type: Documentation
    url: https://pdfbox.apache.org/2.0/getting-started.html
- aid: apache-pdfbox:apache-pdfbox-forms-api
  name: Apache PDFBox Forms API
  description: The Forms API from Apache PDFBox — 1 operation(s) for forms.
  humanURL: https://pdfbox.apache.org/2.0/getting-started.html
  tags:
  - Forms
  properties:
  - type: OpenAPI
    url: openapi/apache-pdfbox-forms-api-openapi.yml
  - type: Documentation
    url: https://pdfbox.apache.org/2.0/getting-started.html
- aid: apache-pdfbox:apache-pdfbox-operations-api
  name: Apache PDFBox Operations API
  description: The Operations API from Apache PDFBox — 2 operation(s) for operations.
  humanURL: https://pdfbox.apache.org/2.0/getting-started.html
  tags:
  - Operations
  properties:
  - type: OpenAPI
    url: openapi/apache-pdfbox-operations-api-openapi.yml
  - type: Documentation
    url: https://pdfbox.apache.org/2.0/getting-started.html
- aid: apache-pdfbox:apache-pdfbox-pages-api
  name: Apache PDFBox Pages API
  description: The Pages API from Apache PDFBox — 1 operation(s) for pages.
  humanURL: https://pdfbox.apache.org/2.0/getting-started.html
  tags:
  - Pages
  properties:
  - type: OpenAPI
    url: openapi/apache-pdfbox-pages-api-openapi.yml
  - type: Documentation
    url: https://pdfbox.apache.org/2.0/getting-started.html
- aid: apache-pdfbox:apache-pdfbox-signatures-api
  name: Apache PDFBox Signatures API
  description: The Signatures API from Apache PDFBox — 1 operation(s) for signatures.
  humanURL: https://pdfbox.apache.org/2.0/getting-started.html
  tags:
  - Signatures
  properties:
  - type: OpenAPI
    url: openapi/apache-pdfbox-signatures-api-openapi.yml
  - type: Documentation
    url: https://pdfbox.apache.org/2.0/getting-started.html
maintainers:
- FN: Kin Lane
  email: info@apievangelist.com
common:
- type: AgenticAccess
  url: agentic-access/apache-pdfbox-agentic-access.yml
- type: VulnerabilityDisclosure
  url: security/apache-pdfbox-vulnerability-disclosure.yml
- type: DomainSecurity
  url: security/apache-pdfbox-domain-security.yml
- type: GitHubOrganization
  url: https://github.com/apache/pdfbox
- type: Documentation
  url: https://pdfbox.apache.org/
- type: SpectralRules
  url: rules/apache-pdfbox-spectral-rules.yml
- type: Vocabulary
  url: vocabulary/apache-pdfbox-vocabulary.yaml
- type: JSONLD
  url: json-ld/apache-pdfbox-context.jsonld
- type: Blog
  url: https://pdfbox.apache.org/blog/
- type: Features
  data:
  - name: PDF Text Extraction
    description: Extract plain text and structured content from PDF documents
  - name: PDF Creation
    description: Create new PDF documents programmatically with Java API
  - name: PDF Manipulation
    description: Merge, split, rotate, and resize pages in existing PDFs
  - name: Digital Signatures
    description: Apply and verify digital signatures for document authenticity
  - name: Form Filling
    description: Read and fill interactive PDF forms (AcroForms)
  - name: PDF/A Validation
    description: Validate and create PDF/A documents for archiving
  - name: Font Handling
    description: Embed and extract fonts, handle Type 1, TrueType, and OpenType
- type: UseCases
  data:
  - name: Invoice Processing
    description: Extract data from PDF invoices for automated processing
  - name: Document Generation
    description: Generate PDF reports, contracts, and certificates programmatically
  - name: Legal Document Management
    description: Digitally sign and verify legal documents
  - name: Form Data Collection
    description: Fill PDF forms and extract submitted data
  - name: Archive Management
    description: Convert documents to PDF/A for long-term archiving
- type: Integrations
  data:
  - name: Apache Tika
    description: Content detection and text extraction integration
  - name: Spring Boot
    description: Spring Boot starter for PDF processing in web applications
  - name: Maven Central
    description: Available as org.apache.pdfbox on Maven Central
  - name: iText/OpenPDF
    description: Complementary PDF library for advanced PDF generation