Nutrient accepts seven JSON Schema keywords, and publishes no contract of its own

Nutrient accepts seven JSON Schema keywords, and publishes no contract of its own

Nutrient has written an unusually candid post about a constraint in its own API. In AI Schema Generator: Auto-draft a JSON Schema for extraction, Marija Trpkovic explains that the Data Extraction API’s extract endpoint “accepts seven JSON Schema keywords, and not the dozens standard JSON Schema allows.” A schema that reaches for $ref to reuse a nested object or oneOf to model two invoice layouts “comes back rejected.” The root must be an object, every object is processed as closed, and setting additionalProperties explicitly is “rejected, not silently ignored.” The new AI Schema Generator in Nutrient Studio exists to draft a first schema that stays inside that dialect, from up to five example documents, a document type, and a plain-language description of the fields and rules that matter.

There are no figures, and the post’s sharpest point is about descriptions rather than types. “The extraction model reads the description on each field as an instruction when deciding which value on the page belongs in that field.” A field named total with no description “invites ambiguity on an invoice that shows a subtotal, a tax line, a discount, and a grand total,” and the generated descriptions “tend to carry the disambiguation and the negative constraints that hand-written first drafts omit.” The post is also honest that prompting a general model for a schema does not avoid the problem, because the model reaches for the same forbidden keywords a developer would. The review loop is the right one: generate, read the descriptions, run against a document that was not an example, and let the citations and confidence signals show where the schema is guessing.

The catalog has to report a gap the post does not mention. The Nutrient.io provider page carries four API pages, the Nutrient Viewer API, the Nutrient Processor API, the Nutrient Accessibility API and the Nutrient AI API. The Data Extraction API this post is about is not among them, and spec presence is unlit on the record, so the catalog has found no machine-readable contract for any of Nutrient’s APIs. Nutrient’s own site navigation lists an MCP server, and that dimension is unlit too. Rate limit signal is the only lit dimension of the eighteen.

The Kin Score is 24.5, emerging band, and the shape is a documentation site without a contract behind it. Discoverability is 69.6 and access clarity 40.8, while contract quality is 0.0 and contract governance 0.0, with operational transparency at 18.4. The Agent Readiness score is 2.5, human-only. A provider that has written a careful post about which JSON Schema keywords its endpoint accepts has, on the public record the catalog can reach, not published the schema of the endpoint itself. The seven-keyword dialect is documented in prose. The contract that would let an agent discover it is not published at all.

← Buf says generate OpenAPI from Protobuf, and its own API ships as Protobuf only
Teradata benchmarks Tera against Claude Code, and its record has no delegated identity →