ElevenLabs Sound Generation API

Generate sound effects and non-speech audio from a text prompt.

Operations 1

Each operation below carries the questions people ask an LLM about it and the instructions they give an agent to run it. Generated by API Evangelist overlay

POST /v1/sound-generation Generate a sound effect from text · Sound Generation #
Ask an LLM
“Can I create a sound effect just by describing it?”
“How long can a generated sound effect be, and can it loop?”
Tell an agent
Generate a sound effect of {text}.
Make a {duration_seconds}-second sound effect of {text}.

Documentation

Specifications

Other Resources

Work with this as data

Every API here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for apis

7 MCP tools reach this
  • find_apisBrowse and filter every API in the catalog.
  • get_api_artifactsOne API's artifacts, grouped by type.
  • get_openapiThe primary OpenAPI for this API.
  • find_similar_apisAPIs that look like this one.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools →

Call it yourself

curl for this page
This API
curl "https://apis.io/api/v1/apis/elevenlabs-sound-generation-api"
All apis
curl "https://apis.io/api/v1/apis?limit=25"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no form to fill in. Signing in shares your email address with us — we store it to create your key and to recognise you if you sign in with another provider. See our Privacy Policy and Terms.

A second provider on the same verified email joins the account you already have.

OpenAPI Specification

elevenlabs-sound-generation-api-openapi.yml Raw ↑
openapi: 3.2.0
info:
  title: ElevenLabs API Documentation Sound Generation API
  description: This is the documentation for the ElevenLabs API. You can use this API to use our service programmatically, this is done by using your API key. You can find your API key in the dashboard at https://elevenlabs.io/app/settings/api-keys.
  version: '1.0'
tags:
- name: sound-generation
  description: Generate sound effects and non-speech audio from a text prompt.
paths:
  /v1/sound-generation:
    post:
      tags:
      - sound-generation
      summary: Sound Generation
      description: Turn text into sound effects for your videos, voice-overs or video games using the most advanced sound effects models in the world.
      operationId: sound_generation
      parameters:
      - name: output_format
        in: query
        required: false
        schema:
          $ref: '#/components/schemas/AllowedOutputFormats'
          title: Output format of the generated audio.
          description: Output format of the generated audio. Formatted as codec_sample_rate_bitrate. So an mp3 with 22.05kHz sample rate at 32kbs is represented as mp3_22050_32. MP3 with 192kbps bitrate requires you to be subscribed to Creator tier or above. PCM with 44.1kHz sample rate requires you to be subscribed to Pro tier or above. Note that the μ-law format (sometimes written mu-law, often approximated as u-law) is commonly used for Twilio audio inputs.
          enum:
          - mp3_22050_32
          - mp3_24000_48
          - mp3_44100_32
          - mp3_44100_64
          - mp3_44100_96
          - mp3_44100_128
          - mp3_44100_192
          - pcm_8000
          - pcm_16000
          - pcm_22050
          - pcm_24000
          - pcm_32000
          - pcm_44100
          - pcm_48000
          - ulaw_8000
          - alaw_8000
          - opus_48000_32
          - opus_48000_64
          - opus_48000_96
          - opus_48000_128
          - opus_48000_192
          default: mp3_44100_128
        description: Output format of the generated audio. Formatted as codec_sample_rate_bitrate. So an mp3 with 22.05kHz sample rate at 32kbs is represented as mp3_22050_32. MP3 with 192kbps bitrate requires you to be subscribed to Creator tier or above. PCM with 44.1kHz sample rate requires you to be subscribed to Pro tier or above. Note that the μ-law format (sometimes written mu-law, often approximated as u-law) is commonly used for Twilio audio inputs.
      - name: xi-api-key
        in: header
        required: false
        schema:
          anyOf:
          - type: string
          - type: 'null'
          description: Your API key. This is required by most endpoints to access our API programmatically. You can view your xi-api-key using the 'Profile' tab on the website.
          title: Xi-Api-Key
        description: Your API key. This is required by most endpoints to access our API programmatically. You can view your xi-api-key using the 'Profile' tab on the website.
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/Body_Sound_Generation_v1_sound_generation_post'
      responses:
        '200':
          description: The generated sound effect as an MP3 file
          content:
            audio/mpeg:
              schema:
                type: string
                format: binary
          headers:
            character-cost:
              description: The number of characters used for billing
              schema:
                type: string
        '422':
          description: Validation Error
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/HTTPValidationError'
      x-fern-sdk-group-name: text_to_sound_effects
      x-fern-sdk-method-name: convert
components:
  schemas:
    HTTPValidationError:
      properties:
        detail:
          items:
            $ref: '#/components/schemas/ValidationError'
          type: array
          title: Detail
      type: object
      title: HTTPValidationError
    Body_Sound_Generation_v1_sound_generation_post:
      properties:
        text:
          type: string
          title: Text
          description: The text that will get converted into a sound effect.
          examples:
          - A large, ancient wooden door slowly opening in an eerie, abandoned castle..
        loop:
          type: boolean
          title: Loop
          description: Whether to create a sound effect that loops smoothly. Only available for the 'eleven_text_to_sound_v2 model'.
          default: false
        duration_seconds:
          anyOf:
          - type: number
          - type: 'null'
          title: Duration Seconds
          description: The duration of the sound which will be generated in seconds. Must be at least 0.5 and at most 30. If set to None we will guess the optimal duration using the prompt. Defaults to None.
        prompt_influence:
          anyOf:
          - type: number
          - type: 'null'
          title: Prompt Influence
          description: A higher prompt influence makes your generation follow the prompt more closely while also making generations less variable. Must be a value between 0 and 1. Defaults to 0.3.
          default: 0.3
        model_id:
          $ref: '#/components/schemas/SFXModelId'
          description: The model ID to use for the sound generation.
          default: eleven_text_to_sound_v2
          examples:
          - eleven_text_to_sound_v2
      type: object
      required:
      - text
      title: Body_Sound_Generation_v1_sound_generation_post
    AllowedOutputFormats:
      type: string
      enum:
      - mp3_22050_32
      - mp3_24000_48
      - mp3_44100_32
      - mp3_44100_64
      - mp3_44100_96
      - mp3_44100_128
      - mp3_44100_192
      - pcm_8000
      - pcm_16000
      - pcm_22050
      - pcm_24000
      - pcm_32000
      - pcm_44100
      - pcm_48000
      - ulaw_8000
      - alaw_8000
      - opus_48000_32
      - opus_48000_64
      - opus_48000_96
      - opus_48000_128
      - opus_48000_192
    SFXModelId:
      type: string
      enum:
      - eleven_text_to_sound_v2
      title: SFXModelId
      default: eleven_text_to_sound_v2
    ValidationError:
      properties:
        loc:
          items:
            anyOf:
            - type: string
            - type: integer
          type: array
          title: Location
        msg:
          type: string
          title: Message
        type:
          type: string
          title: Error Type
      type: object
      required:
      - loc
      - msg
      - type
      title: ValidationError