Modulate
Modulate is a voice AI company based in Cambridge, Massachusetts, building audio-native voice intelligence for trust, safety, and conversation understanding. Its Velma-2 platform exposes a suite of REST (batch) and WebSocket (streaming) model APIs — multilingual and English speech-to-text transcription with speaker diarization, deepfake (synthetic voice) detection, emotion and accent detection, PII/PHI tagging and redaction, language detection, and music/speech and AI-music detection — alongside Velma conversation analysis (behaviors, topics, sentiment, participant roles) and the ToxMod voice-safety product for gaming and social platforms. Authentication is via an X-API-Key header (an api_key query parameter for WebSocket streams), billed per hour of audio processed. Backed by Sierra Ventures.
Modulate publishes 11 APIs on the APIs.io network, including Velma 2 Accent Batch API, Velma 2 Ai Music Detection Batch API, Velma 2 Batch API, and 8 more. Tagged areas include Company, Ai, Voice AI, Speech to Text, and Transcription.
The Modulate catalog on APIs.io includes 7 event-driven AsyncAPI specifications.
Modulate’s developer surface includes authentication, documentation, API reference, getting-started guide, support, engineering blog, pricing, and 21 more developer resources.
Kin Score
APIs 11
Individual APIs this provider publishes, each with its own machine-readable definition.
Modulate Velma 2 Accent Batch API
The Velma 2 Accent Batch API from Modulate — 1 operation(s) for velma 2 accent batch.
Modulate Velma 2 Ai Music Detection Batch API
The Velma 2 Ai Music Detection Batch API from Modulate — 1 operation(s) for velma 2 ai music detection batch.
Modulate Velma 2 Batch API
The Velma 2 Batch API from Modulate — 2 operation(s) for velma 2 batch.
Modulate Velma 2 Emotion Batch API
The Velma 2 Emotion Batch API from Modulate — 1 operation(s) for velma 2 emotion batch.
Modulate Velma 2 Language Detection Batch API
The Velma 2 Language Detection Batch API from Modulate — 1 operation(s) for velma 2 language detection batch.
Modulate Velma 2 Music Detection Batch API
The Velma 2 Music Detection Batch API from Modulate — 1 operation(s) for velma 2 music detection batch.
Modulate Velma 2 Pii Phi Redaction Batch API
The Velma 2 Pii Phi Redaction Batch API from Modulate — 1 operation(s) for velma 2 pii phi redaction batch.
Modulate Velma 2 Stt Batch API
The Velma 2 Stt Batch API from Modulate — 1 operation(s) for velma 2 stt batch.
Modulate Velma 2 Stt Batch English Vfast API
The Velma 2 Stt Batch English Vfast API from Modulate — 1 operation(s) for velma 2 stt batch english vfast.
Modulate Velma 2 Stt Batch Multilingual Vfast API
The Velma 2 Stt Batch Multilingual Vfast API from Modulate — 1 operation(s) for velma 2 stt batch multilingual vfast.
Modulate Velma 2 Synthetic Voice Detection Batch API
The Velma 2 Synthetic Voice Detection Batch API from Modulate — 1 operation(s) for velma 2 synthetic voice detection batch.
Scroll for all 11
MCP Servers 1
Model Context Protocol servers that expose these APIs to AI agents.
modulate-mcp.yml
MCP SERVERRate Limits 1
Documented rate limits and quota policies.
Modulate Rate Limits
RATE LIMITSEvent Specifications 7
AsyncAPI definitions for this provider's event-driven and streaming APIs.
Velma 2 AI Music Detection Streaming API
Real-time AI music detection over WebSocket. The client streams audio and receives per-window vocal AI verdicts as they become available, followed by a final clip-level summary ...
ASYNCAPIVelma 2 Music Detection Streaming API
Real-time frame-level music and speech classification over WebSocket. The client streams audio and receives per-frame probabilities as they become available, followed by a final...
ASYNCAPIVelma 2 PII/PHI Redaction Streaming API
Real-time speech-to-text with PII/PHI redaction over WebSocket. Provides live transcription with automatic language detection, PII/PHI detection and text redaction, and a redact...
ASYNCAPIVelma 2 STT Streaming API
Real-time speech-to-text over WebSocket. Provides live multilingual transcription with automatic per-utterance language detection, delivering each utterance as it is completed. ...
ASYNCAPIVelma 2 STT Streaming English v2 API
Low-latency English speech-to-text over WebSocket. Pure transcription only - no diarization, emotion detection, accent detection, or PII/PHI tagging. Streams interim partial tra...
ASYNCAPIVelma 2 Synthetic Voice Detection Streaming API
Real-time synthetic voice detection for single-speaker audio over WebSocket. The client streams audio and receives per-frame verdicts as they become available.
ASYNCAPIModulate Velma-2 Streaming Server
Streaming velma-2 over WebSocket. Transcribes and analyzes a live audio stream against a client-supplied analysis configuration, emitting JSON events that describe clips and ana...
ASYNCAPIScroll for all 7
Security Posture 3
Authentication, domain security, vulnerability disclosure, and trust-center signals.
Resources
Get Started 5
Portal, sign-up, and the first successful call
Documentation 3
Reference material describing how the API behaves
Agent Surfaces 3
MCP servers, agent skills, and machine-readable catalogs
Design & Contract 5
Pagination, idempotency, versioning, errors, and events
Build 1
SDKs, sample code, and the tooling you integrate with
Access & Security 4
Authentication, authorization, and security posture
Operate 2
Status, limits, changes, and where to get help
Commercial 3
Pricing, plans, and the legal terms of use
Company 2
The organization behind the API