LangChain puts a System One model in the loop to block a tool call before it runs

LangChain puts a System One model in the loop to block a tool call before it runs

LangChain has published a short guide to a different kind of model in the agent loop. In Building a Harness with Jev, Sydney Runkle and Hunter Lovell start from the cost of the loop itself: “Agents run in a loop: an LLM decides what to do, a tool executes, a model evaluates the results, and then continues in that loop until the task is complete.” Every step is another model call. Jev, from TypeSafe AI, is what that company calls a System One model: it does not generate text. “A System One model evaluates a state and returns typed answers and probabilities,” trained with reinforcement learning for calibrated decisions, and “System One models evaluate every question in a request in parallel.” The post wires it into LangGraph two ways: a routing middleware that sends simple turns to a cheap model and hard ones to a capable one, and an AutoModeMiddleware that “uses Jev to check tool calls for risky decisions it may take, and block calls before the tool executes.”

The numbers belong to TypeSafe, reported by LangChain, and the post says so: “The company reports up to 200x faster inference and 400x lower cost than comparable LLMs on classification tasks.” Those are vendor claims on an unnamed comparison and the post does not re-measure them. What survives is the design, and the honesty about its limits: “Jev isn’t a drop-in replacement for an LLM.” Use the LLM for open-ended reasoning and generation, and the classifier for the fast, structured decisions in between, which is where most of the loop’s cost and most of its risk sit. The guardrail use is the one with teeth. A model that returns a calibrated probability that a tool call is risky, before the tool runs, is the dry-run an agent harness has been missing.

The catalog has to say where this story lives, because it is not on the API it scores. The LangChain provider page lists 66 API pages, and they are the hosted LangSmith platform, tracing, evaluation, deployments, access policies and audit logs. The harness in the post is built in the open-source LangChain and LangGraph libraries, with the TypeSafe integration shipped as a package, so none of it passes through an operation the catalog measures. The agentic access profile maps 506 operations, 309 of them acting and 8 flagged human-in-the-loop, which is the surface a harness like this one would govern if it were pointed at LangSmith itself.

The Kin Score is 43.4, developing band. Discoverability carries it at 67.9 and access clarity at 63.2, with contract quality at 49.2. Contract governance is 0.0 and operational transparency is 21.1. The Agent Readiness score is 33.0, agent-ready, with error semantics, reversibility, agent skills, and OpenAPI examples lit. The dimensions still dark are the ones this post is about from the other direction: dry-run mode is unlit, and so is the MCP server and every identity dimension. LangChain has just shown how to put a probability in front of every risky tool call. Its own hosted API does not yet describe a way to rehearse one.

← APIs.io Insights reads 343,299 job postings for the APIs 984 companies name, and joins them to the catalog
Netlify adds a model that answers typed questions, with no keys to create →