arXiv website screenshot

arXiv

arXiv is the open-access e-print repository operated by Cornell Tech, hosting more than two million preprints across physics, mathematics, computer science, quantitative biology, quantitative finance, statistics, electrical engineering, and economics. arXiv exposes two principal programmatic interfaces: a REST Query API that returns Atom 1.0 XML and an OAI-PMH v2.0 endpoint for bulk metadata harvesting, plus daily RSS feeds and Amazon S3 / Kaggle distributions for full-text corpora.

arXiv publishes 2 APIs on the APIs.io network: OAI-PMH API and Query API. Tagged areas include Science And Math, Scholarly Publishing, Preprints, Open Access, and Research.

The arXiv catalog on APIs.io includes 1 JSON-LD context and 2 Spectral governance rulesets.

arXiv’s developer surface includes documentation, engineering blog, support, changelog, tooling, and 24 more developer resources.

52.2/100 developing ▬ flat Agent 22/100 agent aware Full breakdown ↓
scored 2026-07-28 · rubric v0.6
AccessFreeOpen⚡ Free to try
4 APIs 9 Features 6 Use Cases
Science And MathScholarly PublishingPreprintsOpen AccessResearchOpen SourcePublic APIs

Kin Score

Kin Score Kin Score How this is scored →
scored 2026-07-28 · rubric v0.6
Composite quality — 52.2/100 · developing
Contract Quality 16.5 / 25
Developer Ergonomics 6.1 / 20
Commercial Clarity 8.4 / 20
Operational Transparency 4.8 / 13
Governance 8.3 / 12
Discoverability 8.2 / 10
Agent readiness — 22/100 · agent aware
Machine-Readable Contract 18 / 18
Agentic Access Contract 10 / 10
MCP Server 0 / 12
Machine-Readable Auth 0 / 10
Idempotency 0 / 9
Stable Error Semantics 0 / 8
Request/Response Examples 0 / 7
Rate-Limit Signaling 7 / 7
Typed Event Surface 0 / 6
Agent Skills 0 / 5
Well-Known Catalog 0 / 4
Consent & Bot Identity 0 / 3
A2A Agent Card 0 / 8
Dry-Run / Simulate Mode 0 / 4
Improve this rating by publishing the missing artifacts — every area above can be raised, and the full rubric is at apis.io/rating/. This rating is computed from github.com/api-evangelist/arxiv: open an issue to ask a question, or submit a pull request to add artifacts. Want it done for you? Prioritized profiling — $2,500 →

APIs 4

Individual APIs this provider publishes, each with its own machine-readable definition.

arXiv RSS Feeds

Daily RSS feeds of new arXiv submissions, organised by archive and subject category. Primarily intended for human consumption; the OAI-PMH and query APIs are recommended for mac...

arXiv Bulk Data

Full-text and source bulk distribution channels: an Amazon S3 Requester-Pays bucket containing every arXiv PDF and source archive, plus a periodically refreshed Kaggle dataset o...

arXiv OAI-PMH API

OAI-PMH v2.0 verbs for metadata harvesting.

arXiv Query API

Search and retrieve article metadata from arXiv.

Open Collections 2

Open, tool-agnostic API collections (OpenAPI-derived and Bruno).

arXiv OAI-PMH API

OPEN COLLECTION

arXiv Query API

OPEN COLLECTION

Pricing Plans 1

Published pricing tiers and plan structures.

Arxiv Plans Pricing

1 plans

PLANS

Rate Limits 1

Documented rate limits and quota policies.

Arxiv Rate Limits

0 limits

RATE LIMITS

Features 9

Notable capabilities this provider offers.

Field-Prefix Search

Targeted search across title, author, abstract, comment, journal reference, category, report number, and ID.

Boolean Query Composition

AND, OR, and ANDNOT operators with phrase grouping and parentheses.

Date-Range Filtering

submittedDate and lastUpdatedDate ranges in UTC.

Sort Control

Sort by relevance, lastUpdatedDate, or submittedDate, ascending or descending.

ID-Lookup Mode

Fetch metadata for an explicit comma-separated list of arXiv IDs.

OAI-PMH Bulk Harvest

Industry-standard metadata harvesting with resumption tokens and incremental from-date queries.

Three Metadata Formats

oai_dc, arXiv, and arXivRaw exposed via OAI-PMH.

Bulk Full-Text

Amazon S3 Requester-Pays buckets and periodic Kaggle dataset.

Open Source Stack

arXiv operates its services from a public GitHub organization (arXiv) with 50+ active repositories.

Scroll for all 9

Semantic Vocabularies 1

JSON-LD contexts and semantic vocabularies used across these APIs.

Arxiv Context

15 classes · 3 properties

JSON-LD

Spectral Rules 2

Spectral governance rulesets for linting and validating these APIs.

arXiv API Rules

5 rules · 4 warnings 1 info

SPECTRAL

arXiv API Rules

7 rules · 4 errors 3 warnings

SPECTRAL

JSON Schema 1

Standalone JSON Schema definitions for this provider's data models.

Article

14 properties

JSON SCHEMA

JSON Structure 1

JSON Structure definitions describing this provider's data shapes.

Arxiv Article Structure

14 properties

JSON STRUCTURE

Examples 2

Example request and response payloads for these APIs.

Security Posture 1

Authentication, domain security, vulnerability disclosure, and trust-center signals.

Arxiv Domain Security

TLSv1.3 · HSTS · DMARC

SECURITY

Agentic Access 1

Recommended x-agentic-access execution contracts for AI agents.

Arxiv Agentic Access

3 operations · 1 acting

3 operations · 1 acting

AGENTIC

Use Cases 6

What developers build with this provider.

Research Discovery Tools

Build search and recommendation interfaces over the arXiv corpus.

Citation And Bibliographic Apps

Pull metadata, DOIs, and journal references for reference managers.

AI Training And RAG

Build domain corpora for retrieval-augmented generation across scientific literature.

Topic Watching And Alerts

Schedule incremental harvests and notify users of new submissions in a category.

Bibliometrics And Trend Analysis

Aggregate metadata to study research trends, author networks, and category growth.

Academic Workflow Integration

Embed arXiv search into LaTeX editors, IDEs, note-taking tools, and chat assistants via MCP.

Integrations 6

Pre-built integrations with other platforms and tools.

Semantic Scholar

Citation graph and paper-similarity overlay used by community tooling.

NASA ADS

Cross-references and bibliography overlay used in arXiv-bib-overlay.

DOI / CrossRef

Articles surface DOIs once a publisher version of record exists.

Amazon S3

Bulk PDF and source distribution through Requester-Pays buckets.

Kaggle

Periodically refreshed full metadata dataset.

Model Context Protocol

Multiple community MCP servers expose arXiv search to AI assistants.

Solutions 4

Packaged solutions this provider offers.

Query API

Programmatic search and metadata retrieval.

OAI-PMH Harvest

Bulk metadata sync for downstream indexes.

RSS Feeds

Daily new-submission feeds per archive or subject.

Bulk Full-Text

S3 and Kaggle distributions for corpus-scale work.

Resources

Get Started 1

Portal, sign-up, and the first successful call

Documentation 1

Reference material describing how the API behaves

Agent Surfaces 1

MCP servers, agent skills, and machine-readable catalogs

Design & Contract 3

Pagination, idempotency, versioning, errors, and events

Build 12

SDKs, sample code, and the tooling you integrate with

Scroll for all 12

Access & Security 1

Authentication, authorization, and security posture

Operate 4

Status, limits, changes, and where to get help

Commercial 3

Pricing, plans, and the legal terms of use

Company 2

The organization behind the API

Other 1

Properties that don't map to a standard resource type

Source (apis.yml)

apis.yml Raw ↑
aid: arxiv
accessModel:
  pricing: free
  onboarding: open
  trial: false
  try_now: true
  public: true
  label: Free · Open access
  confidence: medium
  source:
  - plans
  generated: '2026-07-22'
  method: derived
image: https://kinlane-images.s3.amazonaws.com/shared/apis-json/icons/arxiv.png
name: arXiv
description: 'arXiv is the open-access e-print repository operated by Cornell Tech, hosting more than two million preprints
  across physics, mathematics, computer science, quantitative biology, quantitative finance, statistics, electrical engineering,
  and economics. arXiv exposes two principal programmatic interfaces: a REST Query API that returns Atom 1.0 XML and an OAI-PMH
  v2.0 endpoint for bulk metadata harvesting, plus daily RSS feeds and Amazon S3 / Kaggle distributions for full-text corpora.'
url: https://info.arxiv.org/help/api/index.html
specificationVersion: '0.20'
created: '2026-05-28'
modified: '2026-05-29'
x-source: public-apis/public-apis
x-type: opensource
x-category: Science & Math
x-tier: 1
x-tier-reason: Cornell-operated open scholarship infrastructure with multiple long-lived public APIs.
tags:
- Science And Math
- Scholarly Publishing
- Preprints
- Open Access
- Research
- Open Source
- Public APIs
apis:
- name: arXiv RSS Feeds
  description: Daily RSS feeds of new arXiv submissions, organised by archive and subject category. Primarily intended for
    human consumption; the OAI-PMH and query APIs are recommended for machine integration.
  humanURL: https://info.arxiv.org/help/rss.html
  baseURL: https://rss.arxiv.org
  tags:
  - Scholarly Publishing
  - Feeds
  properties:
  - type: Documentation
    url: https://info.arxiv.org/help/rss.html
- name: arXiv Bulk Data
  description: 'Full-text and source bulk distribution channels: an Amazon S3 Requester-Pays bucket containing every arXiv
    PDF and source archive, plus a periodically refreshed Kaggle dataset of the complete metadata corpus.'
  humanURL: https://info.arxiv.org/help/bulk_data.html
  baseURL: https://info.arxiv.org/help/bulk_data_s3.html
  tags:
  - Bulk Data
  - Open Data
  properties:
  - type: Documentation
    url: https://info.arxiv.org/help/bulk_data.html
  - type: Resources
    url: https://info.arxiv.org/help/bulk_data_s3.html
    title: Amazon S3 Bulk Buckets
  - type: Resources
    url: https://www.kaggle.com/datasets/Cornell-University/arxiv
    title: Kaggle arXiv Dataset
- aid: arxiv:arxiv-oai-pmh-api
  name: arXiv OAI-PMH API
  description: OAI-PMH v2.0 verbs for metadata harvesting.
  humanURL: https://info.arxiv.org/help/api/user-manual.html
  baseURL: https://export.arxiv.org/api/query
  tags:
  - OAI-PMH
  properties:
  - type: OpenAPI
    url: openapi/arxiv-oai-pmh-api-openapi.yml
  - type: Documentation
    url: https://info.arxiv.org/help/api/user-manual.html
  - type: APIReference
    url: https://info.arxiv.org/help/api/user-manual.html#_calling_the_api
  - type: GettingStarted
    url: https://info.arxiv.org/help/api/basics.html
  - type: JSONSchema
    url: json-schema/arxiv-article-schema.json
  - type: JSONStructure
    url: json-structure/arxiv-article-structure.json
  - type: Examples
    url: examples/arxiv-query-articles-example.json
  - type: SDKs
    url: https://pypi.org/project/arxiv/
  - type: Documentation
    url: https://info.arxiv.org/help/oa/index.html
  - type: Examples
    url: examples/arxiv-oaipmh-listrecords-example.json
- aid: arxiv:arxiv-query-api
  name: arXiv Query API
  description: Search and retrieve article metadata from arXiv.
  humanURL: https://info.arxiv.org/help/api/user-manual.html
  baseURL: https://export.arxiv.org/api/query
  tags:
  - Query
  properties:
  - type: OpenAPI
    url: openapi/arxiv-query-api-openapi.yml
  - type: Documentation
    url: https://info.arxiv.org/help/api/user-manual.html
  - type: APIReference
    url: https://info.arxiv.org/help/api/user-manual.html#_calling_the_api
  - type: GettingStarted
    url: https://info.arxiv.org/help/api/basics.html
  - type: JSONSchema
    url: json-schema/arxiv-article-schema.json
  - type: JSONStructure
    url: json-structure/arxiv-article-structure.json
  - type: Examples
    url: examples/arxiv-query-articles-example.json
  - type: SDKs
    url: https://pypi.org/project/arxiv/
  - type: Documentation
    url: https://info.arxiv.org/help/oa/index.html
  - type: Examples
    url: examples/arxiv-oaipmh-listrecords-example.json
common:
- type: AgenticAccess
  url: agentic-access/arxiv-agentic-access.yml
- type: DomainSecurity
  url: security/arxiv-domain-security.yml
- type: Website
  url: https://arxiv.org
- type: DeveloperPortal
  url: https://info.arxiv.org/help/api/index.html
- type: Documentation
  url: https://info.arxiv.org/help/api/user-manual.html
- type: TermsOfService
  url: https://info.arxiv.org/help/api/tou.html
- type: PrivacyPolicy
  url: https://info.arxiv.org/help/policies/privacy_policy.html
- type: StatusPage
  url: https://status.arxiv.org/
- type: Blog
  url: https://blog.arxiv.org/
- type: Support
  url: https://info.arxiv.org/help/contact.html
- type: GitHubOrganization
  url: https://github.com/arXiv
- type: ChangeLog
  url: https://github.com/arXiv/arxiv-docs/commits/develop
- type: Plans
  url: plans/arxiv-plans-pricing.yml
- type: RateLimits
  url: rate-limits/arxiv-rate-limits.yml
- type: SpectralRules
  url: rules/arxiv-rules.yml
- type: Vocabulary
  url: vocabulary/arxiv-vocabulary.yml
- type: JSONLD
  url: json-ld/arxiv-context.jsonld
  title: arXiv JSON-LD Context
- type: PublicAPIsListing
  url: https://github.com/public-apis/public-apis
- type: Tools
  url: https://github.com/blazickjp/arxiv-mcp-server
  title: arXiv MCP Server (blazickjp)
- type: Tools
  url: https://github.com/shoumikdc/arXiv-mcp
  title: arXiv MCP (shoumikdc)
- type: Tools
  url: https://github.com/Tejas242/arxiv-mcp
  title: arXiv MCP (Tejas242)
- type: Tools
  url: https://github.com/glaforge/arxiv-mcp-server
  title: arXiv MCP Server in Java (glaforge)
- type: Tools
  url: https://github.com/kelvingao/arxiv-mcp
  title: arXiv MCP (kelvingao)
- type: SDKs
  url: https://pypi.org/project/arxiv/
  title: arxiv Python wrapper (lukasschwab/arxiv.py)
- type: SDKs
  url: https://github.com/titipata/arxivpy
  title: arxivpy Python client (titipata/arxivpy)
- type: GitHubRepository
  url: https://github.com/arXiv/arxiv-search
  title: arxiv-search (Search UI and APIs)
- type: GitHubRepository
  url: https://github.com/arXiv/oaipmh
  title: oaipmh (OAI-PMH service)
- type: GitHubRepository
  url: https://github.com/arXiv/arxiv-feed
  title: arxiv-feed (Atom and RSS service)
- type: GitHubRepository
  url: https://github.com/arXiv/arxiv-canonical
  title: arxiv-canonical (JSON schema for arXiv metadata)
- type: Features
  data:
  - name: Field-Prefix Search
    description: Targeted search across title, author, abstract, comment, journal reference, category, report number, and
      ID.
  - name: Boolean Query Composition
    description: AND, OR, and ANDNOT operators with phrase grouping and parentheses.
  - name: Date-Range Filtering
    description: submittedDate and lastUpdatedDate ranges in UTC.
  - name: Sort Control
    description: Sort by relevance, lastUpdatedDate, or submittedDate, ascending or descending.
  - name: ID-Lookup Mode
    description: Fetch metadata for an explicit comma-separated list of arXiv IDs.
  - name: OAI-PMH Bulk Harvest
    description: Industry-standard metadata harvesting with resumption tokens and incremental from-date queries.
  - name: Three Metadata Formats
    description: oai_dc, arXiv, and arXivRaw exposed via OAI-PMH.
  - name: Bulk Full-Text
    description: Amazon S3 Requester-Pays buckets and periodic Kaggle dataset.
  - name: Open Source Stack
    description: arXiv operates its services from a public GitHub organization (arXiv) with 50+ active repositories.
- type: UseCases
  data:
  - name: Research Discovery Tools
    description: Build search and recommendation interfaces over the arXiv corpus.
  - name: Citation And Bibliographic Apps
    description: Pull metadata, DOIs, and journal references for reference managers.
  - name: AI Training And RAG
    description: Build domain corpora for retrieval-augmented generation across scientific literature.
  - name: Topic Watching And Alerts
    description: Schedule incremental harvests and notify users of new submissions in a category.
  - name: Bibliometrics And Trend Analysis
    description: Aggregate metadata to study research trends, author networks, and category growth.
  - name: Academic Workflow Integration
    description: Embed arXiv search into LaTeX editors, IDEs, note-taking tools, and chat assistants via MCP.
- type: Integrations
  data:
  - name: Semantic Scholar
    description: Citation graph and paper-similarity overlay used by community tooling.
  - name: NASA ADS
    description: Cross-references and bibliography overlay used in arXiv-bib-overlay.
  - name: DOI / CrossRef
    description: Articles surface DOIs once a publisher version of record exists.
  - name: Amazon S3
    description: Bulk PDF and source distribution through Requester-Pays buckets.
  - name: Kaggle
    description: Periodically refreshed full metadata dataset.
  - name: Model Context Protocol
    description: Multiple community MCP servers expose arXiv search to AI assistants.
- type: Solutions
  data:
  - name: Query API
    description: Programmatic search and metadata retrieval.
  - name: OAI-PMH Harvest
    description: Bulk metadata sync for downstream indexes.
  - name: RSS Feeds
    description: Daily new-submission feeds per archive or subject.
  - name: Bulk Full-Text
    description: S3 and Kaggle distributions for corpus-scale work.
maintainers:
- FN: Kin Lane
  email: kin@apievangelist.com