Kensho Extract API
Transforms unstructured PDF and image documents into machine-readable JSON, identifying titles, subtitles, paragraphs, tables, and footers in natural reading order. Optional OCR and Figure Extraction (FigEx). REST API at extract.kensho.com with asynchronous extractions and presigned upload/download URLs.
Documentation
Documentation
https://docs.kensho.com/extract
APIReference
https://docs.kensho.com/extract/api
GettingStarted
https://docs.kensho.com/extract/quickstart
Authentication
https://docs.kensho.com/authentication
Specifications
OpenAPI
https://raw.githubusercontent.com/api-evangelist/sp-global/refs/heads/main/openapi/sp-global-extractions-api-openapi.yml
OpenAPI
Harvested original (superseded by the per-tag refined specs above)
Other Resources
Overlay
https://raw.githubusercontent.com/api-evangelist/sp-global/refs/heads/main/overlays/sp-global-extract-overlay.yaml
Tutorials
https://docs.kensho.com/extract/toolkit
Signup
https://kensho.com/extract
JSONLD
https://raw.githubusercontent.com/api-evangelist/sp-global/refs/heads/main/json-ld/sp-global-context.jsonld