Apache ORC
Apache ORC is a self-describing, type-aware columnar file format designed for Hadoop workloads. It provides high compression ratios and fast read performance for large-scale data processing with support for complex data types.
Apache ORC publishes 3 APIs on the APIs.io network: Conversion API, Files API, and Operations API. Tagged areas include Big Data, Columnar Storage, Compression, File Format, and Hadoop.
The Apache ORC catalog on APIs.io includes 1 JSON-LD context and 2 Spectral governance rulesets.
Apache ORC’s developer surface includes documentation, engineering blog, and 7 more developer resources.
Kin Score
APIs 3
Individual APIs this provider publishes, each with its own machine-readable definition.
Apache ORC Conversion API
The Conversion API from Apache ORC — 1 operation(s) for conversion.
Apache ORC Files API
The Files API from Apache ORC — 4 operation(s) for files.
Apache ORC Operations API
The Operations API from Apache ORC — 1 operation(s) for operations.
Pricing Plans 1
Published pricing tiers and plan structures.
Rate Limits 1
Documented rate limits and quota policies.
Apache Orc Rate Limits
RATE LIMITSFinOps 1
Cost, billing, and metering signals for API financial operations.
Apache Orc Finops
FINOPSFeatures 6
Notable capabilities this provider offers.
Columnar Storage
Stores data by column for efficient compression and query performance
Predicate Pushdown
Skip reading data that does not match query predicates
Column Projection
Read only the columns needed for a query
ACID Support
Full ACID transactional support when used with Apache Hive
Schema Evolution
Add, rename, and remove columns while preserving backward compatibility
Compression
Supports ZLIB, Snappy, LZO, LZ4, and ZSTD compression codecs
Semantic Vocabularies 1
JSON-LD contexts and semantic vocabularies used across these APIs.
Apache Orc Context
JSON-LDSpectral Rules 2
Spectral governance rulesets for linting and validating these APIs.
Apache ORC API Rules
SPECTRALApache ORC API Rules
SPECTRALJSON Schema 12
Standalone JSON Schema definitions for this provider's data models.
ColumnStatisticsResponse
JSON SCHEMAColumnStatistics
JSON SCHEMAColumnType
JSON SCHEMAConversionRequest
JSON SCHEMAConversionResult
JSON SCHEMAFileInfo
JSON SCHEMAFileList
JSON SCHEMAFileMetadata
JSON SCHEMAMergeRequest
JSON SCHEMAOperationResult
JSON SCHEMAOrcSchema
JSON SCHEMAStripeInfo
JSON SCHEMAScroll for all 12
JSON Structure 12
JSON Structure definitions describing this provider's data shapes.
Apache Orc Column Statistics Response Structure
JSON STRUCTUREApache Orc Column Statistics Structure
JSON STRUCTUREApache Orc Column Type Structure
JSON STRUCTUREApache Orc Conversion Request Structure
JSON STRUCTUREApache Orc Conversion Result Structure
JSON STRUCTUREApache Orc File Info Structure
JSON STRUCTUREApache Orc File List Structure
JSON STRUCTUREApache Orc File Metadata Structure
JSON STRUCTUREApache Orc Merge Request Structure
JSON STRUCTUREApache Orc Operation Result Structure
JSON STRUCTUREApache Orc Orc Schema Structure
JSON STRUCTUREApache Orc Stripe Info Structure
JSON STRUCTUREScroll for all 12
Examples 12
Example request and response payloads for these APIs.
Apache Orc File Info Example
EXAMPLEApache Orc File List Example
EXAMPLEScroll for all 12
Security Posture 2
Authentication, domain security, vulnerability disclosure, and trust-center signals.
Agentic Access 1
Recommended x-agentic-access execution contracts for AI agents.
Use Cases 4
What developers build with this provider.
Hive Data Warehousing
Store Hive tables in highly efficient ORC format
Spark Analytics
Process large ORC datasets with Apache Spark SQL
Presto/Trino Queries
Fast analytical queries over ORC files with Presto or Trino
Data Lake Storage
Efficient columnar storage for data lake architectures
Integrations 5
Pre-built integrations with other platforms and tools.
Apache Hive
Native ORC support as default Hive storage format
Apache Spark
ORC data source support in Spark SQL
Presto/Trino
Fast ORC reading with native vectorized reader
Apache Flink
ORC file format support for batch and streaming
Apache Arrow
ORC to Arrow conversion for in-memory analytics
Resources
Documentation 1
Reference material describing how the API behaves
Agent Surfaces 1
MCP servers, agent skills, and machine-readable catalogs
Design & Contract 3
Pagination, idempotency, versioning, errors, and events
Build 1
SDKs, sample code, and the tooling you integrate with
Access & Security 2
Authentication, authorization, and security posture
Company 1
The organization behind the API