Gradium TTS API

Text-to-Speech endpoints for converting text to audio

Operations 2

GET /speech/tts TTS WebSocket Stream #
POST /post/speech/tts TTS POST Endpoint #

Work with this as data

Every API here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for apis

7 MCP tools reach this
  • find_apisBrowse and filter every API in the catalog.
  • get_api_artifactsOne API's artifacts, grouped by type.
  • get_openapiThe primary OpenAPI for this API.
  • find_similar_apisAPIs that look like this one.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools →

Call it yourself

curl for this page
This API
curl "https://apis.io/api/v1/apis/gradium-tts-api"
All apis
curl "https://apis.io/api/v1/apis?limit=25"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no form to fill in. Signing in shares your email address with us — we store it to create your key and to recognise you if you sign in with another provider. See our Privacy Policy and Terms.

A second provider on the same verified email joins the account you already have.

OpenAPI Specification

gradium-tts-api-openapi.yml Raw ↑
openapi: 3.2.0
info:
  title: Gradium TTS API
  description: 'This documentation covers the Gradium API.


    This API exposes our Text-To-Speech and Speech-To-Text models, which offers low-latency, high-quality & natural sounding output and best in class accuracy.


    For issues, questions, or feature requests, please contact us at support@gradium.ai'
  version: 0.1.0
servers:
- url: https://api.gradium.ai/api
  description: Gradium API
tags:
- name: TTS
  description: Text-to-Speech endpoints for converting text to audio
paths:
  /speech/tts:
    get:
      tags:
      - TTS
      summary: TTS WebSocket Stream
      description: Connect to this endpoint via WebSocket for real-time text-to-speech conversion with low latency audio streaming.
      parameters:
      - name: x-api-key
        in: header
        required: true
        schema:
          type: string
        description: Your Gradium API key
      responses:
        '101':
          description: WebSocket connection established
      x-codeSamples:
      - lang: cURL
        source: "wscat -c \"wss://api.gradium.ai/api/speech/tts\" \\\n  -H \"x-api-key: your_api_key\"\n# After connection, paste:\n# {\"type\":\"setup\",\"voice_id\":\"YTpq7expH9539ERJ\",\"model_name\":\"default\",\"output_format\":\"wav\"}\n# {\"type\":\"text\",\"text\":\"Hello, world!\"}\n# {\"type\":\"end_of_stream\"}\n"
      - lang: Python
        source: "import asyncio\nimport base64\nimport json\n\nimport websockets\n\n\nasync def synthesise(api_key: str, voice_id: str, text: str) -> bytes:\n    setup = {\n        \"type\": \"setup\",\n        \"voice_id\": voice_id,\n        \"model_name\": \"default\",\n        \"output_format\": \"wav\",\n    }\n    audio_chunks = []\n\n    async with websockets.connect(\n        \"wss://api.gradium.ai/api/speech/tts\",\n        additional_headers={\"x-api-key\": api_key},\n    ) as ws:\n        await ws.send(json.dumps(setup))\n        ready = json.loads(await ws.recv())\n        assert ready[\"type\"] == \"ready\"\n\n        await ws.send(json.dumps({\"type\": \"text\", \"text\": text}))\n        await ws.send(json.dumps({\"type\": \"end_of_stream\"}))\n\n        while True:\n            msg = json.loads(await ws.recv())\n            if msg[\"type\"] == \"audio\":\n                audio_chunks.append(base64.b64decode(msg[\"audio\"]))\n            elif msg[\"type\"] == \"end_of_stream\":\n                break\n            elif msg[\"type\"] == \"error\":\n                raise RuntimeError(msg[\"message\"])\n\n    return b\"\".join(audio_chunks)\n\n\naudio = asyncio.run(synthesise(\"your_api_key\", \"YTpq7expH9539ERJ\", \"Hello, world!\"))\nwith open(\"output.wav\", \"wb\") as f:\n    f.write(audio)\n"
      operationId: getSpeechTts
      x-operation-id-source: derived
  /post/speech/tts:
    post:
      tags:
      - TTS
      summary: TTS POST Endpoint
      description: 'Use this HTTP POST endpoint for simple, text-to-speech conversion. The audio

        data is sent back in a streaming way.


        **Endpoint URL:**


        ```

        https://api.gradium.ai/api/post/speech/tts

        ```


        **Authentication:**

        Include your API key in the request header:

        - Header: `x-api-key: your_api_key`


        ---


        ## Quick Example


        ```bash

        curl -L -X POST https://api.gradium.ai/api/post/speech/tts \

        -H "x-api-key: your_api_key" \

        -H "Content-Type: application/json" \

        -d ''{"text": "Hello, this is a test of the text to speech system.", "voice_id": "YTpq7expH9539ERJ", "output_format": "wav", "only_audio": true}'' \

        > output.wav

        ```


        ---


        ## Request Format


        **Method:** POST

        **Content-Type:** application/json


        **Request Body:**

        ```json

        {

        "text": "Hello, this is a test of the text to speech system.",

        "voice_id": "YTpq7expH9539ERJ",

        "output_format": "wav",

        "json_config": "{}",

        "only_audio": true

        }

        ```


        **Fields:**

        - `text` (string, required): The text to be converted to speech

        - `voice_id` (string, required): Voice ID from the library (e.g.,

        "YTpq7expH9539ERJ") or a custom voice ID

        - `output_format` (string, required): Audio format - "wav", "pcm", or "opus"

        (ogg wrapped opus data).

        - `json_config` (string, optional): Additional configuration in JSON string format (e.g., `{"padding_bonus": -1.2}`)

        - `model_name` (string, optional): The TTS model to use (default: "default")

        - `only_audio` (boolean, optional): When `true`, returns only the raw audio

        bytes. When `false` or omitted, returns a stream of JSON messages containing

        the audio and metadata. The format is the same as with the websocket endpoint.


        ---


        ## Response Format


        ### When `only_audio` is `true`


        The response body contains the raw audio bytes in the requested format. Save directly to a file:


        ```bash

        curl ... > output.wav

        ```


        **Content-Type:** Depends on the output format:

        - `audio/wav` for WAV format

        - `audio/ogg` for Ogg wrapped Opus format

        - `audio/pcm` for PCM format


        ### When `only_audio` is `false` or omitted


        The response is a stream of JSON messages using the same format as the

        WebSocket endpoint. Read the body line-by-line until it closes — the

        body closing signals that synthesis is complete.


        ## Error Handling


        If the request fails before the response stream has started, the server

        responds with `HTTP 500` and a plain-text body. Two body shapes occur:


        - **Upstream errors** (with a numeric code) such as authentication

        failures or worker-level rejections:


        ```

        error from server :

        ```


        For example, a revoked or expired API key returns

        `error from server 1008: API key is revoked or expired`.


        - **Proxy-level rejections** (e.g. unsupported `Content-Type`, malformed

        request body) come back as raw error strings without the `error from

        server` prefix.


        In both cases the body is plain text (not JSON). Errors that occur

        after the response stream has started (when `only_audio` is `false`)

        are surfaced as `{"type": "error", ...}` JSON messages within the

        stream rather than as a different HTTP status.


        ---


        ## When to Use POST vs WebSocket


        The POST endpoint is ideal for simple, text-to-speech generations.

        The main difference with the WebSocket endpoint is that the input is not

        handled in a streaming way; the entire text is sent in one request. The audio is

        still streamed back to the client, allowing for efficient handling of large

        audio outputs and lower latency.


        So if your use case involves sending complete text blocks and receiving audio

        responses, the POST endpoint is a straightforward choice. For more interactive

        or real-time applications where text input is streamed, the WebSocket endpoint

        is more suitable.'
      parameters:
      - name: x-api-key
        in: header
        required: true
        schema:
          type: string
        description: Your Gradium API key
      requestBody:
        required: true
        content:
          application/json:
            schema:
              type: object
              required:
              - text
              - voice_id
              - output_format
              properties:
                text:
                  type: string
                  description: The text to convert to speech
                voice_id:
                  type: string
                  description: Voice ID from the library or custom voice ID
                output_format:
                  type: string
                  enum:
                  - wav
                  - pcm
                  - opus
                  - ulaw_8000
                  - mulaw_8000
                  - alaw_8000
                  - pcm_8000
                  - pcm_16000
                  - pcm_22050
                  - pcm_24000
                  - pcm_44100
                  - pcm_48000
                  description: Audio output format
                only_audio:
                  type: boolean
                  description: When true, returns raw audio bytes instead of JSON
      responses:
        '200':
          description: Audio data returned successfully
        '500':
          description: 'Pre-stream error. Body is plain text. Upstream errors (authentication, worker rejections) are formatted as `error from server <code>: <reason>`; proxy-level rejections (e.g. malformed request body) come back as raw error strings.'
      x-codeSamples:
      - lang: cURL
        source: "curl -L -X POST https://api.gradium.ai/api/post/speech/tts \\\n  -H \"x-api-key: your_api_key\" \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\"text\": \"Hello, world!\", \"voice_id\": \"YTpq7expH9539ERJ\", \"output_format\": \"wav\", \"only_audio\": true}' \\\n  > output.wav\n"
      - lang: Python
        source: "import requests\n\nresp = requests.post(\n    \"https://api.gradium.ai/api/post/speech/tts\",\n    json={\n        \"text\": \"Hello, world!\",\n        \"voice_id\": \"YTpq7expH9539ERJ\",\n        \"output_format\": \"wav\",\n        \"only_audio\": True,\n    },\n    headers={\"x-api-key\": \"your_api_key\"},\n)\nresp.raise_for_status()\nwith open(\"output.wav\", \"wb\") as f:\n    f.write(resp.content)\n"
      operationId: postPostSpeechTts
      x-operation-id-source: derived
x-tagGroups:
- name: Documentation
  tags:
  - Documentation
  - FAQ
  - Release notes
- name: API Reference
  tags:
  - TTS
  - STT
  - Voices
  - Pronunciations
  - Credits