Vonage has published a measurement that anyone paying for structured output should read. In The Hidden Token Tax on JSON Schemas, Nimrod Taiblum of Vonage’s AI Center of Excellence starts from the mechanism: “JSON schemas increase input token usage because the schema is included in every request sent to the model.” The team adopted structured output across sentiment analysis, entity extraction, and interaction scoring, then measured what it cost. Five schemas, from a single field at 88 characters to a deeply nested one at 1,730, were sent with the same roughly 100-token prompt to Gemini models through Vertex AI and to Claude Haiku 4.5 through Bedrock, at temperature zero, with token counts read from each provider’s usage API.
The figures are Vonage’s own, from a disclosed method, which is the right kind. With the medium schema and a 100-token prompt, Claude Haiku 4.5 counted 415 input tokens against 121 without the schema, an overhead of 71%: “You’re paying nearly 3.5x more in input tokens than your actual content warrants.” Gemini 2.5 Flash added 24 tokens for the same schema. Over a million calls the bill ranged from $37.20 to $396.00 depending on the model. “The overhead is a fixed cost per call. As your input payload grows, it fades,” from 37% at 500 tokens to 3% at 10,000. The advice is specific: strip unused fields and flatten nesting, since “a schema shrunk from 1,730 characters to 555 saves you roughly 650 tokens per call on Claude,” move descriptions and examples out of the schema and into the prompt, skip the schema for single-label classification, and batch.
That last piece of advice is where the catalog has something to add, because the schema an LLM reads at inference time and the schema an API publishes in its contract are two different artifacts with opposite economics. In a published contract, descriptions and examples are what a human or an agent uses to understand the surface before calling it, and the catalog’s agent readiness rubric lights a dimension when they are present. In a structured-output call, the same text is paid for on every request. Vonage’s advice is right for the second and would be wrong for the first. The Vonage provider page lists 9 API pages, including the Vonage Messages API and the Vonage Verify API, and the agentic access profile maps 25 operations, 18 of them acting and 3 flagged human-in-the-loop.
The Kin Score is 62.2, strong band, carried by developer ergonomics at 76.8, discoverability at 75.0, and contract quality at 68.0, with contract governance at 31.8. The Agent Readiness score is 24.8, agent-aware. OpenAPI examples and error semantics are both unlit, and that is the tension worth stating plainly. Vonage has just shown, with its own numbers, where examples and descriptions cost money and should be trimmed. On its own public contract, the place where they cost nothing and buy understanding, the catalog cannot find them. The token tax is real, and it is a tax on the wrong document.