Gradium STT API

Speech-to-Text endpoints for converting audio to text

Operations 2

GET /speech/asr STT WebSocket Stream #
POST /post/speech/asr STT POST Endpoint #

Work with this as data

Every API here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for apis

7 MCP tools reach this
  • find_apisBrowse and filter every API in the catalog.
  • get_api_artifactsOne API's artifacts, grouped by type.
  • get_openapiThe primary OpenAPI for this API.
  • find_similar_apisAPIs that look like this one.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools →

Call it yourself

curl for this page
This API
curl "https://apis.io/api/v1/apis/gradium-stt-api"
All apis
curl "https://apis.io/api/v1/apis?limit=25"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no form to fill in. Signing in shares your email address with us — we store it to create your key and to recognise you if you sign in with another provider. See our Privacy Policy and Terms.

A second provider on the same verified email joins the account you already have.

OpenAPI Specification

gradium-stt-api-openapi.yml Raw ↑
openapi: 3.2.0
info:
  title: Gradium STT API
  description: 'This documentation covers the Gradium API.


    This API exposes our Text-To-Speech and Speech-To-Text models, which offers low-latency, high-quality & natural sounding output and best in class accuracy.


    For issues, questions, or feature requests, please contact us at support@gradium.ai'
  version: 0.1.0
servers:
- url: https://api.gradium.ai/api
  description: Gradium API
tags:
- name: STT
  description: Speech-to-Text endpoints for converting audio to text
paths:
  /speech/asr:
    get:
      tags:
      - STT
      summary: STT WebSocket Stream
      description: Connect to this endpoint via WebSocket for real-time speech-to-text conversion with streaming audio input.
      parameters:
      - name: x-api-key
        in: header
        required: true
        schema:
          type: string
        description: Your Gradium API key
      responses:
        '101':
          description: WebSocket connection established
      x-codeSamples:
      - lang: cURL
        source: "wscat -c \"wss://api.gradium.ai/api/speech/asr\" \\\n  -H \"x-api-key: your_api_key\"\n# After connection, paste:\n# {\"type\":\"setup\",\"model_name\":\"default\",\"input_format\":\"pcm\",\"json_config\":{\"language\":\"en\",\"delay_in_frames\":16}}\n"
      - lang: Python
        source: "import asyncio\nimport base64\nimport json\n\nimport websockets\n\nCHUNK_BYTES = 1920 * 2  # 80 ms at 24 kHz, 16-bit mono.\n\n\nasync def transcribe(api_key: str, pcm_audio: bytes):\n    setup = {\n        \"type\": \"setup\",\n        \"model_name\": \"default\",\n        \"input_format\": \"pcm\",\n        \"json_config\": {\n            \"language\": \"en\",\n            \"delay_in_frames\": 16,\n        },\n    }\n\n    async with websockets.connect(\n        \"wss://api.gradium.ai/api/speech/asr\",\n        additional_headers={\"x-api-key\": api_key},\n    ) as ws:\n        await ws.send(json.dumps(setup))\n        ready = json.loads(await ws.recv())\n        assert ready[\"type\"] == \"ready\"\n\n        async def producer():\n            for off in range(0, len(pcm_audio), CHUNK_BYTES):\n                chunk = pcm_audio[off : off + CHUNK_BYTES]\n                await ws.send(json.dumps({\n                    \"type\": \"audio\",\n                    \"audio\": base64.b64encode(chunk).decode(),\n                }))\n            await ws.send(json.dumps({\"type\": \"end_of_stream\"}))\n\n        async def consumer():\n            while True:\n                msg = json.loads(await ws.recv())\n                if msg[\"type\"] == \"text\":\n                    print(msg[\"text\"])\n                elif msg[\"type\"] == \"end_of_stream\":\n                    return\n                elif msg[\"type\"] == \"error\":\n                    raise RuntimeError(msg[\"message\"])\n\n        await asyncio.gather(producer(), consumer())\n\n\nasyncio.run(transcribe(\"your_api_key\", open(\"input.pcm\", \"rb\").read()))\n"
      operationId: getSpeechAsr
      x-operation-id-source: derived
  /post/speech/asr:
    post:
      tags:
      - STT
      summary: STT POST Endpoint
      description: 'Use this HTTP POST endpoint for simple, one-shot speech-to-text

        transcription.'
      parameters:
      - name: x-api-key
        in: header
        required: true
        schema:
          type: string
        description: Your Gradium API key
      - name: Content-Type
        in: header
        required: false
        schema:
          type: string
          enum:
          - audio/wav
          - audio/pcm
          - audio/ogg
          - audio/opus
        description: Format of the audio in the request body. Defaults to audio/wav when omitted.
      - name: model
        in: query
        required: false
        schema:
          type: string
          default: default
        description: Speech-to-Text model name.
      - name: input_format
        in: query
        required: false
        schema:
          type: string
          enum:
          - wav
          - pcm
          - opus
        description: Overrides the audio format detected from Content-Type.
      - name: json_config
        in: query
        required: false
        schema:
          type: string
        description: 'JSON-encoded model configuration. Example: {"language": "en"}'
      requestBody:
        required: true
        content:
          audio/wav:
            schema:
              type: string
              format: binary
              description: WAV audio file.
          audio/pcm:
            schema:
              type: string
              format: binary
              description: 'Raw PCM audio: 24 kHz, 16-bit signed little-endian, mono.'
          audio/ogg:
            schema:
              type: string
              format: binary
              description: Ogg-wrapped Opus audio.
      responses:
        '200':
          description: NDJSON stream of transcription messages.
          content:
            application/x-ndjson:
              schema:
                type: string
                description: 'Newline-delimited JSON messages: text, end_text, or error. The body closes when transcription is complete.'
        '500':
          description: 'Pre-stream error. Body is plain text. Upstream errors (authentication, worker rejections) are formatted as `error from server <code>: <reason>`; proxy-level rejections (e.g. unsupported Content-Type) come back as raw error strings.'
      x-codeSamples:
      - lang: cURL
        source: "curl -L -X POST https://api.gradium.ai/api/post/speech/asr \\\n  -H \"x-api-key: your_api_key\" \\\n  -H \"Content-Type: audio/wav\" \\\n  --data-binary @input.wav\n"
      - lang: Python
        source: "import json\n\nimport requests\n\nwith open(\"input.wav\", \"rb\") as f:\n    audio = f.read()\n\nwith requests.post(\n    \"https://api.gradium.ai/api/post/speech/asr\",\n    data=audio,\n    headers={\n        \"x-api-key\": \"your_api_key\",\n        \"Content-Type\": \"audio/wav\",\n    },\n    stream=True,\n) as resp:\n    resp.raise_for_status()\n    for line in resp.iter_lines(decode_unicode=True):\n        if not line:\n            continue\n        msg = json.loads(line)\n        if msg[\"type\"] == \"text\":\n            print(msg[\"text\"])\n"
      operationId: postPostSpeechAsr
      x-operation-id-source: derived
x-tagGroups:
- name: Documentation
  tags:
  - Documentation
  - FAQ
  - Release notes
- name: API Reference
  tags:
  - TTS
  - STT
  - Voices
  - Pronunciations
  - Credits