Gradium TTS API
Text-to-Speech endpoints for converting text to audio
Text-to-Speech endpoints for converting text to audio
Every API here is available over the APIs.io API and to AI agents over MCP.
One button, every client — Claude, Cursor, VS Code and the rest.
https://apis.io/mcp
find_apisBrowse and filter every API in the catalog.get_api_artifactsOne API's artifacts, grouped by type.get_openapiThe primary OpenAPI for this API.find_similar_apisAPIs that look like this one.apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.resolveTurn a domain, URL or GitHub org into the provider it belongs to.find_cohortsEvery scored population of providers in the catalog.curl "https://apis.io/api/v1/apis/gradium-tts-api"
curl "https://apis.io/api/v1/apis?limit=25"
Discovery needs no key. Ratings and market analysis are Pro.
Free tier, no form to fill in. Signing in shares your email address with us — we store it to create your key and to recognise you if you sign in with another provider. See our Privacy Policy and Terms.
A second provider on the same verified email joins the account you already have.
openapi: 3.2.0
info:
title: Gradium TTS API
description: 'This documentation covers the Gradium API.
This API exposes our Text-To-Speech and Speech-To-Text models, which offers low-latency, high-quality & natural sounding output and best in class accuracy.
For issues, questions, or feature requests, please contact us at support@gradium.ai'
version: 0.1.0
servers:
- url: https://api.gradium.ai/api
description: Gradium API
tags:
- name: TTS
description: Text-to-Speech endpoints for converting text to audio
paths:
/speech/tts:
get:
tags:
- TTS
summary: TTS WebSocket Stream
description: Connect to this endpoint via WebSocket for real-time text-to-speech conversion with low latency audio streaming.
parameters:
- name: x-api-key
in: header
required: true
schema:
type: string
description: Your Gradium API key
responses:
'101':
description: WebSocket connection established
x-codeSamples:
- lang: cURL
source: "wscat -c \"wss://api.gradium.ai/api/speech/tts\" \\\n -H \"x-api-key: your_api_key\"\n# After connection, paste:\n# {\"type\":\"setup\",\"voice_id\":\"YTpq7expH9539ERJ\",\"model_name\":\"default\",\"output_format\":\"wav\"}\n# {\"type\":\"text\",\"text\":\"Hello, world!\"}\n# {\"type\":\"end_of_stream\"}\n"
- lang: Python
source: "import asyncio\nimport base64\nimport json\n\nimport websockets\n\n\nasync def synthesise(api_key: str, voice_id: str, text: str) -> bytes:\n setup = {\n \"type\": \"setup\",\n \"voice_id\": voice_id,\n \"model_name\": \"default\",\n \"output_format\": \"wav\",\n }\n audio_chunks = []\n\n async with websockets.connect(\n \"wss://api.gradium.ai/api/speech/tts\",\n additional_headers={\"x-api-key\": api_key},\n ) as ws:\n await ws.send(json.dumps(setup))\n ready = json.loads(await ws.recv())\n assert ready[\"type\"] == \"ready\"\n\n await ws.send(json.dumps({\"type\": \"text\", \"text\": text}))\n await ws.send(json.dumps({\"type\": \"end_of_stream\"}))\n\n while True:\n msg = json.loads(await ws.recv())\n if msg[\"type\"] == \"audio\":\n audio_chunks.append(base64.b64decode(msg[\"audio\"]))\n elif msg[\"type\"] == \"end_of_stream\":\n break\n elif msg[\"type\"] == \"error\":\n raise RuntimeError(msg[\"message\"])\n\n return b\"\".join(audio_chunks)\n\n\naudio = asyncio.run(synthesise(\"your_api_key\", \"YTpq7expH9539ERJ\", \"Hello, world!\"))\nwith open(\"output.wav\", \"wb\") as f:\n f.write(audio)\n"
operationId: getSpeechTts
x-operation-id-source: derived
/post/speech/tts:
post:
tags:
- TTS
summary: TTS POST Endpoint
description: 'Use this HTTP POST endpoint for simple, text-to-speech conversion. The audio
data is sent back in a streaming way.
**Endpoint URL:**
```
https://api.gradium.ai/api/post/speech/tts
```
**Authentication:**
Include your API key in the request header:
- Header: `x-api-key: your_api_key`
---
## Quick Example
```bash
curl -L -X POST https://api.gradium.ai/api/post/speech/tts \
-H "x-api-key: your_api_key" \
-H "Content-Type: application/json" \
-d ''{"text": "Hello, this is a test of the text to speech system.", "voice_id": "YTpq7expH9539ERJ", "output_format": "wav", "only_audio": true}'' \
> output.wav
```
---
## Request Format
**Method:** POST
**Content-Type:** application/json
**Request Body:**
```json
{
"text": "Hello, this is a test of the text to speech system.",
"voice_id": "YTpq7expH9539ERJ",
"output_format": "wav",
"json_config": "{}",
"only_audio": true
}
```
**Fields:**
- `text` (string, required): The text to be converted to speech
- `voice_id` (string, required): Voice ID from the library (e.g.,
"YTpq7expH9539ERJ") or a custom voice ID
- `output_format` (string, required): Audio format - "wav", "pcm", or "opus"
(ogg wrapped opus data).
- `json_config` (string, optional): Additional configuration in JSON string format (e.g., `{"padding_bonus": -1.2}`)
- `model_name` (string, optional): The TTS model to use (default: "default")
- `only_audio` (boolean, optional): When `true`, returns only the raw audio
bytes. When `false` or omitted, returns a stream of JSON messages containing
the audio and metadata. The format is the same as with the websocket endpoint.
---
## Response Format
### When `only_audio` is `true`
The response body contains the raw audio bytes in the requested format. Save directly to a file:
```bash
curl ... > output.wav
```
**Content-Type:** Depends on the output format:
- `audio/wav` for WAV format
- `audio/ogg` for Ogg wrapped Opus format
- `audio/pcm` for PCM format
### When `only_audio` is `false` or omitted
The response is a stream of JSON messages using the same format as the
WebSocket endpoint. Read the body line-by-line until it closes — the
body closing signals that synthesis is complete.
## Error Handling
If the request fails before the response stream has started, the server
responds with `HTTP 500` and a plain-text body. Two body shapes occur:
- **Upstream errors** (with a numeric code) such as authentication
failures or worker-level rejections:
```
error from server :
```
For example, a revoked or expired API key returns
`error from server 1008: API key is revoked or expired`.
- **Proxy-level rejections** (e.g. unsupported `Content-Type`, malformed
request body) come back as raw error strings without the `error from
server` prefix.
In both cases the body is plain text (not JSON). Errors that occur
after the response stream has started (when `only_audio` is `false`)
are surfaced as `{"type": "error", ...}` JSON messages within the
stream rather than as a different HTTP status.
---
## When to Use POST vs WebSocket
The POST endpoint is ideal for simple, text-to-speech generations.
The main difference with the WebSocket endpoint is that the input is not
handled in a streaming way; the entire text is sent in one request. The audio is
still streamed back to the client, allowing for efficient handling of large
audio outputs and lower latency.
So if your use case involves sending complete text blocks and receiving audio
responses, the POST endpoint is a straightforward choice. For more interactive
or real-time applications where text input is streamed, the WebSocket endpoint
is more suitable.'
parameters:
- name: x-api-key
in: header
required: true
schema:
type: string
description: Your Gradium API key
requestBody:
required: true
content:
application/json:
schema:
type: object
required:
- text
- voice_id
- output_format
properties:
text:
type: string
description: The text to convert to speech
voice_id:
type: string
description: Voice ID from the library or custom voice ID
output_format:
type: string
enum:
- wav
- pcm
- opus
- ulaw_8000
- mulaw_8000
- alaw_8000
- pcm_8000
- pcm_16000
- pcm_22050
- pcm_24000
- pcm_44100
- pcm_48000
description: Audio output format
only_audio:
type: boolean
description: When true, returns raw audio bytes instead of JSON
responses:
'200':
description: Audio data returned successfully
'500':
description: 'Pre-stream error. Body is plain text. Upstream errors (authentication, worker rejections) are formatted as `error from server <code>: <reason>`; proxy-level rejections (e.g. malformed request body) come back as raw error strings.'
x-codeSamples:
- lang: cURL
source: "curl -L -X POST https://api.gradium.ai/api/post/speech/tts \\\n -H \"x-api-key: your_api_key\" \\\n -H \"Content-Type: application/json\" \\\n -d '{\"text\": \"Hello, world!\", \"voice_id\": \"YTpq7expH9539ERJ\", \"output_format\": \"wav\", \"only_audio\": true}' \\\n > output.wav\n"
- lang: Python
source: "import requests\n\nresp = requests.post(\n \"https://api.gradium.ai/api/post/speech/tts\",\n json={\n \"text\": \"Hello, world!\",\n \"voice_id\": \"YTpq7expH9539ERJ\",\n \"output_format\": \"wav\",\n \"only_audio\": True,\n },\n headers={\"x-api-key\": \"your_api_key\"},\n)\nresp.raise_for_status()\nwith open(\"output.wav\", \"wb\") as f:\n f.write(resp.content)\n"
operationId: postPostSpeechTts
x-operation-id-source: derived
x-tagGroups:
- name: Documentation
tags:
- Documentation
- FAQ
- Release notes
- name: API Reference
tags:
- TTS
- STT
- Voices
- Pronunciations
- Credits