Amazon EMR · Arazzo Workflow

Amazon EMR Run a Spark ETL Job

Version 1.0.0

Launch a Spark cluster and queue an ETL processing step in one call.

1 workflow 1 source API 1 provider
View Spec View on GitHub Amazon Web ServicesAnalyticsApache SparkBig DataData ProcessingHadoopArazzoWorkflows

Provider

amazon-emr

Workflows

run-spark-etl-job
Run a Spark cluster with ETL processing steps queued.
Creates and starts a new EMR cluster with Spark installed and queues the supplied ETL processing steps to run once the cluster is provisioned, returning the identifier of the newly created cluster.
1 step inputs: instances, name, releaseLabel, steps outputs: jobFlowId
1
runSparkEtl
Create and start a new EMR cluster with Spark installed and queue the supplied ETL processing steps to run once the cluster is provisioned.

Source API Descriptions

Arazzo Workflow Specification

Raw ↑
arazzo: 1.0.1
info:
  title: Amazon EMR Run a Spark ETL Job
  summary: Launch a Spark cluster and queue an ETL processing step in one call.
  description: >-
    Launches a managed Amazon EMR cluster with Apache Spark installed and
    submits the supplied ETL processing steps in the same RunJobFlow call, so an
    extract-transform-load workload begins as soon as the cluster is
    provisioned. The workflow passes through the caller supplied name, instance
    configuration, release label, and steps, requests the Spark application, and
    returns the new cluster's JobFlowId. Every step spells out its request
    inline, including the AWS JSON protocol X-Amz-Target header, so the flow can
    be read and executed without opening the underlying OpenAPI description.
  version: 1.0.0
sourceDescriptions:
- name: clustersApi
  url: ../openapi/amazon-emr-clusters-api-openapi.yml
  type: openapi
workflows:
- workflowId: run-spark-etl-job
  summary: Run a Spark cluster with ETL processing steps queued.
  description: >-
    Creates and starts a new EMR cluster with Spark installed and queues the
    supplied ETL processing steps to run once the cluster is provisioned,
    returning the identifier of the newly created cluster.
  inputs:
    type: object
    required:
    - name
    - instances
    - releaseLabel
    - steps
    properties:
      name:
        type: string
        description: The name of the cluster to create.
      instances:
        type: object
        description: The instance configuration for the cluster.
      releaseLabel:
        type: string
        description: The Amazon EMR release label (e.g. emr-6.10.0).
      steps:
        type: array
        description: The ordered list of ETL processing steps to run after cluster creation.
        items:
          type: object
  steps:
  - stepId: runSparkEtl
    description: >-
      Create and start a new EMR cluster with Spark installed and queue the
      supplied ETL processing steps to run once the cluster is provisioned.
    operationId: RunJobFlow
    parameters:
    - name: X-Amz-Target
      in: header
      value: ElasticMapReduce.RunJobFlow
    requestBody:
      contentType: application/json
      payload:
        Name: $inputs.name
        Instances: $inputs.instances
        ReleaseLabel: $inputs.releaseLabel
        Applications:
        - Name: Spark
        Steps: $inputs.steps
    successCriteria:
    - condition: $statusCode == 200
    outputs:
      jobFlowId: $response.body#/JobFlowId
  outputs:
    jobFlowId: $steps.runSparkEtl.outputs.jobFlowId

Work with this as data

Every workflow here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for arazzo workflows

4 MCP tools reach this
  • find_arazzoBrowse and filter every workflow in the catalog.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools →

Call it yourself

curl for this page
This workflow
curl "https://apis.io/api/v1/arazzo/amazon-emr-run-spark-etl-job-workflow"
All arazzo workflows
curl "https://apis.io/api/v1/arazzo?limit=25"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no form to fill in. Signing in shares your email address with us — we store it to create your key and to recognise you if you sign in with another provider. See our Privacy Policy and Terms.

A second provider on the same verified email joins the account you already have.