Inworld AI Realtime Voice Agent
Inworld AI is a research lab and inference provider focused on realtime AI models for consumer-facing applications. We build voice AI that feels as human as it sounds. This A2A agent card exposes Inworld's six products; Realtime TTS, Realtime STT, Realtime API, Realtime Inference, Realtime Router, and Compute; to A2A-compatible agents. The voice that makes AI agents human. Realtime AI for consumer-facing applications. Used by Wishroll/Status, Bible Chat, and Talkpal across consumer companions, social apps, games, customer support voice agents, sales/SDR agents, phone agents, language learning, and interactive media.
AgentCard object. It is an agent card in spirit rather than in schema — useful for
discovery, but an A2A client cannot assume it will parse.
no-protocolVersionno-defaultInputModesno-defaultOutputModesno-preferredTransport
An A2A client finds this agent by fetching the well-known path on the provider's own host. That is what makes an agent card different from every other agent artifact in this catalog: it is provider-published by construction — it cannot be derived, generated, or reconstructed on a provider's behalf.
Realtime TTS realtime-tts
Voices that sound human enough that users stay on the call and come back. Streaming text-to-speech with word, phoneme, and viseme timestamps. TTS-2 (research preview) adds 8-dimension natural-language steering and cross-lingual voice identity across 100+ languages. Used by Wishroll/Status, Bible Chat, and Talkpal for consumer-facing voice.
Realtime STT realtime-stt
Captures what users said, including how they said it, so the agent responds with context. Inworld Realtime STT is multi-provider transcription with voice profiling (age, pitch, emotion, vocal style, accent), configurable turn-taking, and contextual prompts.
Realtime Router realtime-router
Pick the right model for each user, scenario, and price point and switch without rewiring. Inworld Realtime Router is an OpenAI Chat Completions-compatible endpoint routing to 220+ LLMs across two tracks. 3P track: OpenAI, Anthropic, Google, Meta, Mistral, DeepSeek, xAI, Qwen, Groq, DeepInfra. 1P track: Realtime Inference. Used by Wishroll for consumer-facing applications.
Realtime Inference realtime-inference
Run open-source models fast enough for live voice and cheap enough for consumer-scale free tiers. Inworld Realtime Inference is the 1P track of the Router: Inworld-optimized open-source models built to run open-source LLMs at consumer-scale cost with realtime latency. Confirmed 1P models: Gemma 4, DeepSeek V3.2 / V4 family, GLM-5.1/5.2. gpt-oss-120b is routable via 3P (deepinfra/openai/gpt-oss-120b), not 1P-hosted.
Compute compute
Dedicated capacity for traffic-heavy customers; predictable latency when shared inference no longer fits. Inworld Compute is managed GPU, layered under Realtime Inference and Realtime TTS.
Realtime API realtime-api
One integrated voice loop instead of stitching three vendors; ships in days, fails in fewer places. Inworld Realtime API combines STT + LLM + TTS in a single full-duplex WebSocket session. OpenAI Realtime protocol compatible. Includes Inworld Silero VAD + Smart Turn detector.
2026-07-28 from https://inworld.ai/.well-known/agent-card.json,
HTTP 200. The body was saved verbatim and is the sole source for
everything on this page — no field is inferred.
View the captured card.
Providers change what they serve; if this card has moved or changed shape,
the provider profile carries the
current state as of the last build.
Work with this as data
Every agent card here is available over the APIs.io API and to AI agents over MCP. A2A Agent Cards is not yet its own endpoint on the v1 API. Reach it through catalog search and the tag graph, or the MCP server.