Apache PySpark website screenshot

Apache PySpark

Python API for Apache Spark - A unified analytics engine for large-scale data processing supporting batch processing, streaming, machine learning, and graph computing.

Apache PySpark publishes 5 APIs on the APIs.io network. Tagged areas include Big Data, Data Processing, Distributed Computing, Machine Learning, and Python.

Apache PySpark’s developer surface includes getting-started guide, release notes, and 10 more developer resources.

24.7/100 emerging ▬ flat Agent 3/100 human only Full breakdown ↓
scored 2026-08-05 · rubric v0.9.1
AccessFreemium
5 APIs
Big DataData ProcessingDistributed ComputingMachine LearningPythonStreaming

Kin Score

Kin Score Kin Score How this is scored →
scored 2026-08-05 · rubric v0.9.1
Composite quality — 24.7/100 · emerging
Contract Quality 0.0 / 25
Developer Ergonomics 3.0 / 20
Commercial Clarity 7.9 / 20
Operational Transparency 8.2 / 13
Governance 0.0 / 12
Discoverability 5.6 / 10
Agent readiness — 3/100 · human only
Machine-Readable Contract 0 / 18
Agentic Access Contract 0 / 10
MCP Server 0 / 12
Machine-Readable Auth 0 / 10
Idempotency 0 / 9
Stable Error Semantics 0 / 8
Request/Response Examples 0 / 7
Rate-Limit Signaling 7 / 7
Typed Event Surface 0 / 6
Agent Skills 0 / 5
Well-Known Catalog 0 / 4
Consent & Bot Identity 0 / 3
A2A Agent Card 0 / 8
Dry-Run / Simulate Mode 0 / 4
Improve this rating by publishing the missing artifacts — every area above can be raised, and the full rubric is at apis.io/rating/. This rating is computed from github.com/api-evangelist/pyspark: open an issue to ask a question, or submit a pull request to add artifacts. Want it done for you? Prioritized profiling — $2,500 →

APIs 5

Individual APIs this provider publishes, each with its own machine-readable definition.

PySpark Core API

Core Spark functionality including RDDs, SparkContext, and basic operations.

PySpark SQL

Structured data processing with DataFrame and SQL operations.

PySpark Streaming

Real-time stream processing capabilities using DStreams and Structured Streaming.

PySpark MLlib

Machine learning library with scalable algorithms for classification, regression, clustering, and more.

PySpark ML (DataFrame-based)

DataFrame-based machine learning API with pipelines and feature transformers.

Pricing Plans 1

Published pricing tiers and plan structures.

Pyspark Plans Pricing

3 plans

PLANS

Rate Limits 1

Documented rate limits and quota policies.

Pyspark Rate Limits

5 limits

RATE LIMITS

FinOps 1

Cost, billing, and metering signals for API financial operations.

Security Posture 2

Authentication, domain security, vulnerability disclosure, and trust-center signals.

Pyspark Domain Security

TLSv1.3 · HSTS · DMARC

SECURITY

Pyspark Vulnerability Disclosure

security.txt · contact published

SECURITY

Resources

Get Started 2

Portal, sign-up, and the first successful call

Build 1

SDKs, sample code, and the tooling you integrate with

Access & Security 3

Authentication, authorization, and security posture

Operate 3

Status, limits, changes, and where to get help

Company 2

The organization behind the API

Other 1

Properties that don't map to a standard resource type

Source (apis.yml)

apis.yml Raw ↑
aid: pyspark
name: Apache PySpark
description: Python API for Apache Spark - A unified analytics engine for large-scale data processing supporting batch processing,
  streaming, machine learning, and graph computing.
type: Index
accessModel:
  pricing: freemium
  onboarding: unknown
  trial: false
  try_now: false
  public: false
  label: Freemium
  confidence: medium
  source:
  - plans
  generated: '2026-07-22'
  method: derived
image: https://kinlane-images.s3.amazonaws.com/shared/apis-json/icons/pyspark.png
tags:
- Big Data
- Data Processing
- Distributed Computing
- Machine Learning
- Python
- Streaming
url: https://raw.githubusercontent.com/api-evangelist/pyspark/refs/heads/main/apis.yml
created: '2024-01-01'
modified: '2026-04-28'
specificationVersion: '0.19'
apis:
- aid: pyspark:pyspark-core-api
  name: PySpark Core API
  description: Core Spark functionality including RDDs, SparkContext, and basic operations.
  humanURL: https://spark.apache.org/docs/latest/api/python/reference/pyspark.html
  tags:
  - RDD
  - Spark Context
  properties:
  - type: Documentation
    url: https://spark.apache.org/docs/latest/api/python/reference/pyspark.html
  - type: APIReference
    url: https://spark.apache.org/docs/latest/api/python/reference/api/pyspark.SparkContext.html
- aid: pyspark:pyspark-sql
  name: PySpark SQL
  description: Structured data processing with DataFrame and SQL operations.
  humanURL: https://spark.apache.org/docs/latest/sql-programming-guide.html
  tags:
  - DataFrame
  - SQL
  properties:
  - type: Documentation
    url: https://spark.apache.org/docs/latest/api/python/reference/pyspark.sql/index.html
  - type: GettingStarted
    url: https://spark.apache.org/docs/latest/sql-getting-started.html
- aid: pyspark:pyspark-streaming
  name: PySpark Streaming
  description: Real-time stream processing capabilities using DStreams and Structured Streaming.
  humanURL: https://spark.apache.org/docs/latest/streaming-programming-guide.html
  tags:
  - Streaming
  - Real-Time
  properties:
  - type: Documentation
    url: https://spark.apache.org/docs/latest/api/python/reference/pyspark.streaming/index.html
  - type: ProgrammingGuide
    url: https://spark.apache.org/docs/latest/streaming-programming-guide.html
- aid: pyspark:pyspark-mllib
  name: PySpark MLlib
  description: Machine learning library with scalable algorithms for classification, regression, clustering, and more.
  humanURL: https://spark.apache.org/docs/latest/ml-guide.html
  tags:
  - Machine Learning
  - MLlib
  properties:
  - type: Documentation
    url: https://spark.apache.org/docs/latest/api/python/reference/pyspark.ml.html
  - type: MLGuide
    url: https://spark.apache.org/docs/latest/ml-guide.html
- aid: pyspark:pyspark-ml
  name: PySpark ML (DataFrame-based)
  description: DataFrame-based machine learning API with pipelines and feature transformers.
  humanURL: https://spark.apache.org/docs/latest/ml-pipeline.html
  tags:
  - Machine Learning
  - Pipeline
  properties:
  - type: Documentation
    url: https://spark.apache.org/docs/latest/api/python/reference/pyspark.ml.html
  - type: PipelineGuide
    url: https://spark.apache.org/docs/latest/ml-pipeline.html
common:
- type: VulnerabilityDisclosure
  url: security/pyspark-vulnerability-disclosure.yml
- type: DomainSecurity
  url: security/pyspark-domain-security.yml
- type: LinkedIn
  url: https://www.linkedin.com/company/apachespark
- type: Website
  url: https://spark.apache.org/
- type: GitHubOrganization
  url: https://github.com/apache/spark
- type: GettingStarted
  url: https://spark.apache.org/docs/latest/api/python/getting_started/install.html
- type: GettingStarted
  url: https://spark.apache.org/docs/latest/quick-start.html
- type: Downloads
  url: https://spark.apache.org/downloads.html
- type: Community
  url: https://spark.apache.org/community.html
- type: IssueTracker
  url: https://issues.apache.org/jira/projects/SPARK
- type: ReleaseNotes
  url: https://spark.apache.org/releases/
- type: Security
  url: https://spark.apache.org/security.html
maintainers:
- FN: Kin Lane
  email: kin@apievangelist.com