Inworld AI Realtime Voice Agent
Inworld AI is a research lab and inference provider focused on realtime AI models for consumer-facing applications. We build voice AI that feels as human as it sounds. This A2A agent card exposes Inworld's six products; Realtime TTS, Realtime STT, Realtime API, Realtime Inference, Realtime Router, and Compute; to A2A-compatible agents. The voice that makes AI agents human. Realtime AI for consumer-facing applications. Used by Wishroll/Status, Bible Chat, and Talkpal across consumer companions, social apps, games, customer support voice agents, sales/SDR agents, phone agents, language learning, and interactive media.
AgentCard object. It is an agent card in spirit rather than in schema — useful for
discovery, but an A2A client cannot assume it will parse.
no-protocolVersionno-defaultInputModesno-defaultOutputModesno-preferredTransport
An A2A client finds this agent by fetching the well-known path on the provider's own host. That is what makes an agent card different from every other agent artifact in this catalog: it is provider-published by construction — it cannot be derived, generated, or reconstructed on a provider's behalf.
Realtime TTS realtime-tts
Voices that sound human enough that users stay on the call and come back. Streaming text-to-speech with word, phoneme, and viseme timestamps. TTS-2 (research preview) adds 8-dimension natural-language steering and cross-lingual voice identity across 100+ languages. Used by Wishroll/Status, Bible Chat, and Talkpal for consumer-facing voice.
Realtime STT realtime-stt
Captures what users said, including how they said it, so the agent responds with context. Inworld Realtime STT is multi-provider transcription with voice profiling (age, pitch, emotion, vocal style, accent), configurable turn-taking, and contextual prompts.
Realtime Router realtime-router
Pick the right model for each user, scenario, and price point and switch without rewiring. Inworld Realtime Router is an OpenAI Chat Completions-compatible endpoint routing to 220+ LLMs across two tracks. 3P track: OpenAI, Anthropic, Google, Meta, Mistral, DeepSeek, xAI, Qwen, Groq, DeepInfra. 1P track: Realtime Inference. Used by Wishroll for consumer-facing applications.
Realtime Inference realtime-inference
Run open-source models fast enough for live voice and cheap enough for consumer-scale free tiers. Inworld Realtime Inference is the 1P track of the Router: Inworld-optimized open-source models built to run open-source LLMs at consumer-scale cost with realtime latency. Confirmed 1P models: Gemma 4, DeepSeek V3.2 / V4 family, GLM-5.1/5.2. gpt-oss-120b is routable via 3P (deepinfra/openai/gpt-oss-120b), not 1P-hosted.
Compute compute
Dedicated capacity for traffic-heavy customers; predictable latency when shared inference no longer fits. Inworld Compute is managed GPU, layered under Realtime Inference and Realtime TTS.
Realtime API realtime-api
One integrated voice loop instead of stitching three vendors; ships in days, fails in fewer places. Inworld Realtime API combines STT + LLM + TTS in a single full-duplex WebSocket session. OpenAI Realtime protocol compatible. Includes Inworld Silero VAD + Smart Turn detector.
2026-07-28 from https://inworld.ai/.well-known/agent-card.json,
HTTP 200. The body was saved verbatim and is the sole source for
everything on this page — no field is inferred.
View the captured card.
Providers change what they serve; if this card has moved or changed shape,
the provider profile carries the
current state as of the last build.