Gladia Transcription V2 API
The Transcription V2 API from Gladia — 3 operation(s) for transcription v2.
The Transcription V2 API from Gladia — 3 operation(s) for transcription v2.
Every API here is available over the APIs.io API and to AI agents over MCP.
One button, every client — Claude, Cursor, VS Code and the rest.
https://apis.io/mcp
find_apisBrowse and filter every API in the catalog.get_api_artifactsOne API's artifacts, grouped by type.get_openapiThe primary OpenAPI for this API.find_similar_apisAPIs that look like this one.apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.resolveTurn a domain, URL or GitHub org into the provider it belongs to.find_cohortsEvery scored population of providers in the catalog.curl "https://apis.io/api/v1/apis/gladia-transcription-v2-api"
curl "https://apis.io/api/v1/apis?limit=25"
Discovery needs no key. Ratings and market analysis are Pro.
Free tier, no form to fill in. Signing in shares your email address with us — we store it to create your key and to recognise you if you sign in with another provider. See our Privacy Policy and Terms.
A second provider on the same verified email joins the account you already have.
openapi: 3.2.0
info:
title: Gladia Control AudioToText Transcription V2 API
description: Gladia AI audio infrastructure API for speech-to-text transcription via REST and WebSocket. Supports asynchronous pre-recorded audio processing and real-time live transcription with speaker diarization, automatic language detection across 100+ languages, and audio intelligence features.
version: '1.0'
contact: {}
servers:
- url: https://api.gladia.io/
description: Gladia API production URL
tags:
- name: Transcription V2
paths:
/v2/transcription:
post:
operationId: TranscriptionController_initPreRecordedJob_v2
parameters: []
requestBody:
required: true
content:
application/json:
schema:
$ref: '#/components/schemas/InitTranscriptionRequest'
responses:
'201':
description: The transcription job has been initiated
content:
application/json:
schema:
$ref: '#/components/schemas/InitPreRecordedTranscriptionResponse'
'400':
description: Something is wrong with the request
content:
application/json:
schema:
$ref: '#/components/schemas/BadRequestErrorResponse'
'401':
description: You don't have the permissions to initiate a new transcription job
content:
application/json:
schema:
$ref: '#/components/schemas/UnauthorizedErrorResponse'
'422':
description: The parameters you gave are incorrect
content:
application/json:
schema:
$ref: '#/components/schemas/UnprocessableEntityErrorResponse'
security:
- x_gladia_key: []
summary: Initiate a new transcription job
tags:
- Transcription V2
get:
operationId: TranscriptionController_list_v2
parameters:
- name: offset
required: false
in: query
description: The starting point for pagination. A value of 0 starts from the first item.
schema:
minimum: 0
default: 0
type: integer
- name: limit
required: false
in: query
description: The maximum number of items to return. Useful for pagination and controlling data payload size.
schema:
minimum: 1
default: 20
type: integer
- name: date
required: false
in: query
description: Filter items relevant to a specific date in ISO format (YYYY-MM-DD).
schema:
format: date-time
example: '2026-06-12'
type: string
- name: before_date
required: false
in: query
description: Include items that occurred before the specified date in ISO format.
schema:
format: date-time
example: '2026-06-12T21:00:09.947Z'
type: string
- name: after_date
required: false
in: query
description: Filter for items after the specified date. Use with `before_date` for a range. Date in ISO format.
schema:
format: date-time
example: '2026-06-12T21:00:09.947Z'
type: string
- name: status
required: false
in: query
description: Filter the list based on item status. Accepts multiple values from the predefined list.
schema:
example:
- done
type: array
items:
type: string
enum:
- queued
- processing
- done
- error
- name: custom_metadata
required: false
in: query
schema:
additionalProperties: true
example:
user: John Doe
type: object
- name: kind
required: false
in: query
description: Filter the list based on the item type. Supports multiple values from the predefined list.
schema:
example:
- pre-recorded
type: array
items:
type: string
enum:
- pre-recorded
- live
responses:
'200':
description: A list of transcription jobs matching the parameters.
content:
application/json:
schema:
$ref: '#/components/schemas/ListTranscriptionResponse'
'401':
description: You don't have the permissions to access transcription jobs
content:
application/json:
schema:
$ref: '#/components/schemas/UnauthorizedErrorResponse'
security:
- x_gladia_key: []
summary: Get transcription jobs based on query parameters
tags:
- Transcription V2
/v2/transcription/{id}:
get:
operationId: TranscriptionController_getTranscript_v2
parameters:
- name: id
required: true
in: path
description: Id of the transcription job
schema:
example: 45463597-20b7-4af7-b3b3-f5fb778203ab
type: string
responses:
'200':
description: The transcription job's metadata
content:
application/json:
schema:
example:
id: 45463597-20b7-4af7-b3b3-f5fb778203ab
request_id: G-45463597
version: 2
kind: pre-recorded
created_at: '2023-12-28T09:04:17.210Z'
status: queued
file:
id: f0dcZE10-23d8-47f0-a25d-74a6eed88721
filename: split_infinity.wav
source: http://files.gladia.io/example/audio-transcription/split_infinity.wav
audio_duration: 20
number_of_channels: 1
request_params:
audio_url: http://files.gladia.io/example/audio-transcription/split_infinity.wav
subtitles: false
diarization: false
translation: false
summarization: false
sentences: false
moderation: false
named_entity_recognition: false
name_consistency: false
custom_spelling: false
structured_data_extraction: false
chapterization: false
sentiment_analysis: false
display_mode: false
audio_enhancer: false
language_config:
code_switching: false
languages:
- fr
- en
accurate_words_timestamps: false
diarization_enhanced: false
punctuation_enhanced: false
completed_at: null
custom_metadata: null
error_code: null
result: null
oneOf:
- $ref: '#/components/schemas/PreRecordedResponse'
- $ref: '#/components/schemas/StreamingResponse'
discriminator:
propertyName: kind
mapping:
pre-recorded: '#/components/schemas/PreRecordedResponse'
live: '#/components/schemas/StreamingResponse'
'401':
description: You don't have the permissions to access the transcription job
content:
application/json:
schema:
$ref: '#/components/schemas/UnauthorizedErrorResponse'
'404':
description: The transcription job doesn't exist or has been deleted
content:
application/json:
schema:
$ref: '#/components/schemas/NotFoundErrorResponse'
security:
- x_gladia_key: []
summary: Get the transcription job's metadata
tags:
- Transcription V2
delete:
operationId: TranscriptionController_deleteTranscript_v2
parameters:
- name: id
required: true
in: path
description: Id of the transcription job
schema:
example: 45463597-20b7-4af7-b3b3-f5fb778203ab
type: string
responses:
'202':
description: The transcription job has been successfully deleted
'401':
description: You don't have the permissions to delete this transcription job
content:
application/json:
schema:
$ref: '#/components/schemas/UnauthorizedErrorResponse'
'403':
description: The transcription job is not in a deletable state
content:
application/json:
schema:
$ref: '#/components/schemas/ForbiddenErrorResponse'
'404':
description: The transcription job doesn't exist or has been deleted
content:
application/json:
schema:
$ref: '#/components/schemas/NotFoundErrorResponse'
security:
- x_gladia_key: []
summary: Delete the transcription job
tags:
- Transcription V2
/v2/transcription/{id}/file:
get:
operationId: TranscriptionController_getAudio_v2
parameters:
- name: id
required: true
in: path
description: Id of the transcription job
schema:
example: 45463597-20b7-4af7-b3b3-f5fb778203ab
type: string
responses:
'200':
description: The audio file used for this transcription job
content:
application/octet-stream:
schema:
type: string
format: binary
example: <binary>
'401':
description: You don't have the permissions to access this transcription job or its audio file
content:
application/json:
schema:
$ref: '#/components/schemas/UnauthorizedErrorResponse'
'404':
description: The transcription job or its audio file doesn't exist or has been deleted
content:
application/json:
schema:
$ref: '#/components/schemas/NotFoundErrorResponse'
security:
- x_gladia_key: []
summary: Download the audio file used for this transcription job
tags:
- Transcription V2
components:
schemas:
TranslationResultDTO:
type: object
properties:
error:
description: Contains the error details of the failed addon
allOf:
- $ref: '#/components/schemas/AddonErrorDTO'
full_transcript:
type: string
description: All transcription on text format without any other information
languages:
type: array
description: All the detected languages in the audio sorted from the most detected to the less detected
example:
- en
items:
$ref: '#/components/schemas/TranslationLanguageCodeEnum'
sentences:
description: If `sentences` has been enabled, sentences results for this translation
type: array
items:
$ref: '#/components/schemas/SentencesDTO'
subtitles:
description: If `subtitles` has been enabled, subtitles results for this translation
type: array
items:
$ref: '#/components/schemas/SubtitleDTO'
utterances:
description: Transcribed speech utterances present in the audio
type: array
items:
$ref: '#/components/schemas/UtteranceDTO'
required:
- error
- full_transcript
- languages
- utterances
PreProcessingConfig:
type: object
properties:
audio_enhancer:
type: boolean
description: If true, apply pre-processing to the audio stream to enhance the quality.
default: false
speech_threshold:
type: number
description: Sensitivity configuration for Speech Threshold. A value close to 1 will apply stricter thresholds, making it less likely to detect background sounds as speech.
default: 0.6
minimum: 0
maximum: 1
UtteranceDTO:
type: object
properties:
start:
type: number
description: Start timestamp in seconds of this utterance
end:
type: number
description: End timestamp in seconds of this utterance
confidence:
type: number
description: Confidence on the transcribed utterance (1 = 100% confident)
channel:
type: integer
description: Audio channel of where this utterance has been transcribed from
minimum: 0
speaker:
type: integer
description: If `diarization` enabled, speaker identification number
minimum: 0
words:
description: List of words of the utterance, split by timestamp
type: array
items:
$ref: '#/components/schemas/WordDTO'
text:
type: string
description: Transcription for this utterance
language:
description: Spoken language in this utterance
example: en
allOf:
- $ref: '#/components/schemas/TranscriptionLanguageCodeEnum'
required:
- start
- end
- confidence
- channel
- words
- text
- language
CustomVocabularyEntryDTO:
type: object
properties:
value:
type: string
description: The text used to replace in the transcription.
example: Gladia
intensity:
type: number
description: The global intensity of the feature.
example: 0.5
minimum: 0
maximum: 1
pronunciations:
description: The pronunciations used in the transcription.
type: array
items:
type: string
language:
description: Specify the language in which it will be pronounced when sound comparison occurs. Default to transcription language.
example: en
allOf:
- $ref: '#/components/schemas/TranscriptionLanguageCodeEnum'
required:
- value
TranscriptionResultDTO:
type: object
properties:
metadata:
description: Metadata for the given transcription & audio file
allOf:
- $ref: '#/components/schemas/TranscriptionMetadataDTO'
transcription:
description: Transcription of the audio speech
allOf:
- $ref: '#/components/schemas/TranscriptionDTO'
translation:
description: If `translation` has been enabled, translation of the audio speech transcription
allOf:
- $ref: '#/components/schemas/TranslationDTO'
summarization:
description: If `summarization` has been enabled, summarization of the audio speech transcription
allOf:
- $ref: '#/components/schemas/SummarizationDTO'
moderation:
description: If `moderation` has been enabled, moderation of the audio speech transcription
allOf:
- $ref: '#/components/schemas/ModerationDTO'
named_entity_recognition:
description: If `named_entity_recognition` has been enabled, the detected entities
allOf:
- $ref: '#/components/schemas/NamedEntityRecognitionDTO'
name_consistency:
description: If `name_consistency` has been enabled, Gladia will improve consistency of the names accross the transcription
allOf:
- $ref: '#/components/schemas/NamesConsistencyDTO'
structured_data_extraction:
description: If `structured_data_extraction` has been enabled, structured data extraction results
allOf:
- $ref: '#/components/schemas/StructuredDataExtractionDTO'
sentiment_analysis:
description: If `sentiment_analysis` has been enabled, sentiment analysis of the audio speech transcription
allOf:
- $ref: '#/components/schemas/SentimentAnalysisDTO'
audio_to_llm:
description: If `audio_to_llm` has been enabled, audio to llm results of the audio speech transcription
allOf:
- $ref: '#/components/schemas/AudioToLlmListDTO'
sentences:
description: 'If `sentences` has been enabled, sentences of the audio speech transcription. Deprecated: content will move to the `transcription` object.'
deprecated: true
allOf:
- $ref: '#/components/schemas/SentencesDTO'
display_mode:
description: If `display_mode` has been enabled, the output will be reordered, creating new utterances when speakers overlapped
allOf:
- $ref: '#/components/schemas/DisplayModeDTO'
chapterization:
description: If `chapterization` has been enabled, will generate chapters name for different parts of the given audio.
allOf:
- $ref: '#/components/schemas/ChapterizationDTO'
diarization:
description: If `diarization` has been requested and an error has occurred, the result will appear here
allOf:
- $ref: '#/components/schemas/DiarizationDTO'
required:
- metadata
DiarizationConfigDTO:
type: object
properties:
number_of_speakers:
type: integer
description: Exact number of speakers in the audio
example: 3
minimum: 1
min_speakers:
type: integer
description: Minimum number of speakers in the audio
example: 1
minimum: 0
max_speakers:
type: integer
description: Maximum number of speakers in the audio
example: 2
minimum: 0
ListTranscriptionResponse:
type: object
properties:
first:
type: string
description: URL to fetch the first page
format: uri
example: https://api.gladia.io/v2/transcription?status=done&offset=0&limit=20
current:
type: string
description: URL to fetch the current page
format: uri
example: https://api.gladia.io/v2/transcription?status=done&offset=0&limit=20
next:
type:
- string
- 'null'
description: URL to fetch the next page
format: uri
example: https://api.gladia.io/v2/transcription?status=done&offset=20&limit=20
items:
description: List of transcriptions
discriminator:
propertyName: kind
mapping:
pre-recorded: '#/components/schemas/PreRecordedResponse'
live: '#/components/schemas/StreamingResponse'
type: array
items:
oneOf:
- $ref: '#/components/schemas/PreRecordedResponse'
- $ref: '#/components/schemas/StreamingResponse'
required:
- first
- current
- next
- items
SentencesDTO:
type: object
properties:
success:
type: boolean
description: The audio intelligence model succeeded to get a valid output
is_empty:
type: boolean
description: The audio intelligence model returned an empty value
exec_time:
type: number
description: Time audio intelligence model took to complete the task
error:
description: '`null` if `success` is `true`. Contains the error details of the failed model'
allOf:
- $ref: '#/components/schemas/AddonErrorDTO'
results:
description: If `sentences` has been enabled, transcription as sentences.
type:
- array
- 'null'
items:
type: string
required:
- success
- is_empty
- exec_time
- error
- results
SentimentAnalysisDTO:
type: object
properties:
success:
type: boolean
description: The audio intelligence model succeeded to get a valid output
is_empty:
type: boolean
description: The audio intelligence model returned an empty value
exec_time:
type: number
description: Time audio intelligence model took to complete the task
error:
description: '`null` if `success` is `true`. Contains the error details of the failed model'
allOf:
- $ref: '#/components/schemas/AddonErrorDTO'
results:
type: string
description: If `sentiment_analysis` has been enabled, Gladia will analyze the sentiments and emotions of the audio
required:
- success
- is_empty
- exec_time
- error
- results
SummarizationConfigDTO:
type: object
properties:
type:
description: The type of summarization to apply
default: general
allOf:
- $ref: '#/components/schemas/SummaryTypesEnum'
FileResponse:
type: object
properties:
id:
type: string
description: The file id
filename:
type:
- string
- 'null'
description: The name of the uploaded file
source:
type:
- string
- 'null'
description: The link used to download the file if audio_url was used
audio_duration:
type:
- number
- 'null'
description: Duration of the audio file
example: 3600
number_of_channels:
type:
- integer
- 'null'
description: Number of channels in the audio file
minimum: 1
example: 1
required:
- id
- filename
- source
- audio_duration
- number_of_channels
BadRequestErrorResponse:
type: object
properties:
timestamp:
type: string
description: Date of when the error occurred
example: '2023-12-28T09:04:17.210Z'
path:
type: string
description: Path to the API endpoint
example: /v2/transcription/45463597-20b7-4af7-b3b3-f5fb778203ab
request_id:
type: string
description: Debug id
example: G-821fe9df
statusCode:
type: number
description: HTTP status code of the error
example: 400
message:
type: string
description: Error message
example: Content-Type is missing Multipart Boundary.
validation_errors:
description: List of validation errors, if any
example:
- Field "language" must be a string
- Field "min_speakers" must be a number
type: array
items:
type: string
required:
- timestamp
- path
- request_id
- statusCode
- message
SubtitleDTO:
type: object
properties:
format:
description: Format of the current subtitle
example: srt
allOf:
- $ref: '#/components/schemas/SubtitlesFormatEnum'
subtitles:
type: string
description: Transcription on the asked subtitle format
required:
- format
- subtitles
CallbackMethodEnum:
type: string
enum:
- POST
- PUT
description: 'The HTTP method to be used. Allowed values are `POST` or `PUT` (default: `POST`)'
InitPreRecordedTranscriptionResponse:
type: object
properties:
id:
type: string
description: Id of the job
format: uuid
example: 45463597-20b7-4af7-b3b3-f5fb778203ab
result_url:
type: string
description: Prebuilt URL with your transcription `id` to fetch the result
example: https://api.gladia.io/v2/transcription/45463597-20b7-4af7-b3b3-f5fb778203ab
format: uri
required:
- id
- result_url
CallbackConfig:
type: object
properties:
url:
type: string
description: URL on which we will do a `POST` request with configured messages
example: https://callback.example
format: uri
receive_partial_transcripts:
type: boolean
description: If true, partial transcript will be sent to the defined callback.
default: false
receive_final_transcripts:
type: boolean
description: If true, final transcript will be sent to the defined callback.
default: true
receive_speech_events:
type: boolean
description: If true, begin and end speech events will be sent to the defined callback.
default: false
receive_pre_processing_events:
type: boolean
description: If true, pre-processing events will be sent to the defined callback.
default: true
receive_realtime_processing_events:
type: boolean
description: If true, realtime processing events will be sent to the defined callback.
default: true
receive_post_processing_events:
type: boolean
description: If true, post-processing events will be sent to the defined callback.
default: true
receive_acknowledgments:
type: boolean
description: If true, acknowledgments will be sent to the defined callback.
default: false
receive_errors:
type: boolean
description: If true, errors will be sent to the defined callback.
default: false
receive_lifecycle_events:
type: boolean
description: If true, lifecycle events will be sent to the defined callback.
default: true
AudioToLlmListConfigDTO:
type: object
properties:
prompts:
description: The list of prompts applied on the audio transcription
example:
- Extract the key points from the transcription
minItems: 1
type: array
items:
type: array
model:
type: string
description: The model to use for the prompt execution. You can find the list of supported models [here](https://openrouter.ai/models).
default: openai/gpt-5.4-nano
required:
- prompts
StreamingSupportedSampleRateEnum:
type: number
enum:
- 8000
- 16000
- 32000
- 44100
- 48000
description: The sample rate of the audio stream
NamedEntityRecognitionDTO:
type: object
properties:
success:
type: boolean
description: The audio intelligence model succeeded to get a valid output
is_empty:
type: boolean
description: The audio intelligence model returned an empty value
exec_time:
type: number
description: Time audio intelligence model took to complete the task
error:
description: '`null` if `success` is `true`. Contains the error details of the failed model'
allOf:
- $ref: '#/components/schemas/AddonErrorDTO'
results:
description: If `named_entity_recognition` has been enabled, the detected entities.
type:
- array
- 'null'
items:
$ref: '#/components/schemas/NamedEntityRecognitionResult'
required:
- success
- is_empty
- exec_time
- error
- results
TranscriptionLanguageCodeEnum:
type: string
enum:
- af
- am
- ar
- as
- az
- ba
- be
- bg
- bn
- bo
- br
- bs
- ca
- cs
- cy
- da
- de
- el
- en
- es
- et
- eu
- fa
- fi
- fo
- fr
- gl
- gu
- ha
- haw
- he
- hi
- hr
- ht
- hu
- hy
- id
- is
- it
- ja
- jw
- ka
- kk
- km
- kn
- ko
- la
- lb
- ln
- lo
- lt
- lv
- mg
- mi
- mk
- ml
- mn
- mr
- ms
- mt
- my
- ne
- nl
- nn
- 'no'
- oc
- pa
- pl
- ps
- pt
- ro
- ru
- sa
- sd
- si
- sk
- sl
- sn
- so
- sq
- sr
- su
- sv
- sw
- ta
- te
- tg
- th
- tk
- tl
- tr
- tt
- uk
- ur
- uz
- vi
- yi
- yo
- zh
description: Specify the language in which it will be pronounced when sound comparison occurs. Default to transcription language.
TranscriptionMetadataDTO:
type: object
properties:
audio_duration:
type: number
description: Duration of the transcribed audio file
example: 3600
number_of_distinct_channels:
type: integer
description: Number of distinct channels in the transcribed audio file
minimum: 1
example: 1
billing_time:
type: number
description: Billed duration in seconds (audio_duration * number_of_distinct_channels)
example: 3600
transcription_time:
type: number
description: Duration of the transcription in seconds
example: 20
required:
- audio_duration
- number_of_distinct_channels
- billing_time
- transcription_time
DisplayModeDTO:
type: object
properties:
success:
type: boolean
description: The audio intelligence model succeeded to get a valid output
is_empty:
type: boolean
description: The audio intelligence model returned an empty value
exec_time:
type: number
description: Time audio intelligence model took to complete the task
error:
description: '`null` if `success` is `true`. Contains the error details of the failed model'
allOf:
- $ref: '#/components/schemas/AddonErrorDTO'
results:
description: If `display_mode` has been enabled, proposes an alternative display output.
type:
- array
- 'null'
items:
type: string
required:
- success
- is_empty
- exec_time
- error
- results
PiiRedactionConfigDTO:
type: object
properties:
entity_types:
description: The entity types to redact
example:
- GDPR
- HEALTH_INFORMATION
- HIPAA_SAFE_HARBOR
- QUEBEC_PRIVACY_ACT
- EMAIL_ADDRESS
- NAME
- PHONE_NUMBER
allOf:
- $ref: '#/components/schemas/PiiRedactionEntityTypeEnum'
processed_text_type:
type: string
description: The type of processed text to return (marker or mask)
enum:
- MARKER
- MASK
example: MARKER
StreamingSupportedBitDepthEnum:
type: number
enum:
- 8
- 16
- 24
- 32
description: The bit depth of the audio stream
SummarizationDTO:
type: object
properties:
success:
type: boolean
description: The audio intelligence model succeeded to get a valid output
is_empty:
type: boolean
description: The audio intelligence model returned an empty value
exec_time:
type: number
description: Time audio intelligence model took to complete the task
error:
description: '`null` if `success` is `true`. Contains the error details of the failed model'
allOf:
- $ref: '#/components/schemas/AddonErrorDTO'
results:
type:
- string
- 'null'
description: If `summarization` has been enabled, summary of the transcription
required:
- success
- is_empty
- exec_time
- error
- results
TranslationDTO:
type: object
properties:
success:
type: boolean
d
# --- truncated at 32 KB (76 KB total) ---
# Full source: https://raw.githubusercontent.com/api-evangelist/gladia/refs/heads/main/openapi/gladia-transcription-v2-api-openapi.yml