Controllable LLM Agents with Jev: A 10-Step Blueprint for Cost and Drift Control
A 10-step architectural blueprint for placing Jev ahead of LLMs like Claude and Codex to cut token costs, route by confidence, and prevent agent drift.
Diogo Almeida, founder of Jev and TypeSafe, has published a concise 12-page guide detailing how to pair lightweight decision primitives with large language models to build faster, cheaper, and strictly controllable AI agent systems. Summarized by AI practitioner @0xwhrrari, this 10-step architectural blueprint demonstrates how to wrap generative models such as Claude, Codex, and Grok with semantic decision boundaries, confidence routing, and permanent audit receipts.

Image source: @0xwhrrari / X
Conventional agent tutorials frequently advocate expanding prompt instructions and passing cumulative conversation histories on every interaction. In production environments, this unconstrained approach leads to exploding token bills, unacceptable latency, and unpredictable agent drift. By delegating discrete semantic assessments to specialized System 1 decision models while leaving authority in deterministic code, engineering teams can achieve resilient system governance without sacrificing generative flexibility.
1. Responsibility Separation and State Formulation (Steps 1 to 3)
Establishing a controllable agent architecture begins with strict separation between open-ended generation, bounded semantic decision-making, and deterministic business logic.
- Step 1: Split the responsibilities
- Allocate tasks according to engine strengths: large language models handle natural language generation, lightweight decision models like Jev handle bounded semantic decisions, and deterministic code retains ultimate execution authority.
- High-consequence side effects and authorization rules must remain anchored in deterministic control structures rather than probabilistic text generation.
- Step 2: Build the state
- Avoid forwarding entire multi-turn conversation logs to the decision engine.
- Isolate and package only the active request, directly relevant evidence, governing system policy, and the proposed action into a minimal state payload.
- Step 3: Choose the right primitive
- Match each judgment with its corresponding Jev evaluation primitive:
- Choice: Selects a single route or workflow branch from a predefined set of mutual options.
- Score: Evaluates candidate items against an ordered rubric or graded scale.
- Noul: Computes calibrated probabilities (0.0 to 1.0) indicating whether a specific assertion or guard condition holds true.
2. Pre- and Post-Generative Guardrails with Atomic Decisions (Steps 4 to 7)
Positioning lightweight evaluations before and after generative inference dramatically curbs redundant token consumption and catches operational errors early.
- Step 4: Replace giant evaluation prompts with atomic questions
- Decompose monolithic evaluative prompts into distinct, typed semantic queries.
- Separate intent classification, task urgency, evidence sufficiency, operational risk, and execution scope into discrete evaluation steps.
- Step 5: Put Jev before the LLM
- Evaluate incoming state prior to dispatching expensive API requests to frontier generative models.
- Filter required context, select active tools, route to the appropriate model provider, and establish the target workflow branch before consuming generative tokens.
- Step 6: Give the LLM a bounded job
- Once the upstream decision layer resolves the route, supply the generative model with only the instructions, reference files, and tool definitions required for that specific branch.
- Restricting the execution surface mitigates context contamination, reduces latency, and prevents off-topic tool invocation.
- Step 7: Put Jev after the LLM
- Run verification immediately upon receiving generative output to confirm whether the completion answers the prompt, relies on sufficient evidence, and remains within permitted operational scope.
- Divergent or non-compliant outputs can be quarantined or redirected prior to reaching downstream consumers.
3. Confidence Routing, Batch Evaluation, and Decision Receipts (Steps 8 to 10)
Production reliability requires transparent calibration: decisions should branch dynamically based on probabilistic certainty and preserve end-to-end auditability.
- Step 8: Route by confidence
- Establish automated execution paths for high-confidence, low-risk evaluations.
- Direct ambiguous or low-confidence states to prompt users for supplemental context, while routing high-impact or destructive actions to human review queues.
- Step 9: Batch independent decisions
- Rather than dispatching serial model queries for each discrete criterion, batch multiple Choice, Score, and Noul evaluations across a unified state payload.
- Parallelizing independent judgments lowers cumulative network latency and eliminates redundant data serialization overhead.
- Step 10: Record the complete decision receipt
- Log an immutable decision receipt for every automated judgment, capturing the state version, query text, probability distribution, selected route, model identifier, execution latency, final outcome, and any subsequent human override.
- Persistent decision telemetry creates a reproducible audit log essential for evaluating model drift, conducting regressions, and optimizing control parameters over time.
While many AI courses emphasize composing ever-larger prompts, real-world deployment challenges demand a rigorous control harness constructed around every generative call. Implementing this 10-step discipline enables teams to deploy fast, budget-conscious, and auditable AI agent systems.
Original source
- X (Twitter) Thread: @0xwhrrari Original Post
- Diogo Almeida Guide PDF: Google Drive Document (12-Page Guide)