What Is Jev Engineering? A Guide for Agent Builders
Updated 8 skills8 stepsOrigin: Diogo Almeida
Created by Diogo Almeida - https://madewithjev.com/jev-engineering
Overview
Jev Engineering is a way to build AI agents: an LLM writes, Jev decides, and code acts. According to that page, every point where an agent picks a worker, scores a source or approves a tool call goes to Jev, which TypeSafe AI calls its System One model, instead of going to a model that generates text. The split gives each kind of work to the component built for it. Language models produce briefings, drafts and code. Jev handles bounded decisions such as routing, scoring, approving and escalating. Deterministic code executes those decisions and enforces exact rules. The code that holds these pieces together is an agent harness.
flowchart LR
S[Shared state] --> L[LLM writes]
S --> J[Jev decides]
subgraph H[Agent harness]
L
J
C[Code acts]
end
L --> C
J --> C
C --> V[Verify result]
V --> S
The decision component behaves differently from a chat model. You send it text and typed questions, and it returns a typed answer, a probability for every option and a confidence score, in 70 to 500 milliseconds. One practitioner thread describes it as answering several questions in a single pass instead of generating text token by token. It cannot generate text. That limit is deliberate. The method keeps a larger model in the loop for anything that has to be written and treats Jev as a specialized decision layer, not a replacement for generative models.
The term is very new. The model launched on September 15, 2026, when TypeSafe AI came out of stealth with $40M led by DCVC. Its creator, Diogo Almeida, co-invented RLHF during his time at OpenAI. The earliest identifiable use of the term is the creator-associated page What is Jev Engineering?, dated September 19, 2026. No earlier book, paper or talk turns up. One thread reports that the phrase started trending the week the model launched. Since then the material has grown from a three-part slogan into an implementation method covering typed questions, parallel decisions, calibrated probabilities, confidence thresholds and deterministic execution. The creator's reading list was still being updated at the time of writing, so expect the vocabulary to shift.
Most of the published evidence comes from the vendor's orbit or from individual practitioners. On a 400-item reference set, one report compared zero-shot Jev with two classic baselines:
| Approach | Accuracy, 400-item set | Source |
|---|---|---|
| Jev, zero-shot | 95.9% | OrcaRouter report |
| Hand-written keyword rules | 77.2% | OrcaRouter report |
| Supervised TF-IDF baseline | 66.0% | OrcaRouter report |
The reports disagree on calibration. One benchmark found a ten-bin expected calibration error of 0.0313 on a 1,200-item MMLU sample. The OrcaRouter tests found the error rising to 0.107, against a 0.024 noise floor, on 900 synthetic support tickets whose deciding policy was not in the text. Treat the headline speed claims with care. TypeSafe reportedly claimed significant performance improvements, and a practitioner write-up notes that the widely quoted figures came from a founder's talk, not from an independently replicated benchmark.
Several alternatives cover similar ground. Comparison pages set Jev beside zero-shot and fine-tuned classifiers, named-entity recognition models, and language models asked for structured output. One roundup names managed structured-output features and libraries such as BAML, Instructor and Outlines as the closest alternatives. Open clones exist as well. One architecture comparison describes Laya as a ModernBERT backbone with a candidate-scoring head and Kev-0.5B as Qwen2.5-0.5B with a custom decision head. The same comparison says none of these clones show equivalence to Jev's base model, training data or recipe. The practical difference from schema-constrained generation is that decisions sit in a separate component that returns probabilities your code can threshold.
No independent adoption survey or peer-reviewed evaluation turns up in current coverage (Knolli's overview). Treat Jev Engineering as a pattern to test in shadow mode before you trust it. The skill pages break the method into parts, starting with separating generation from decision-making and benchmarking and observing agent loops. Teams planning a migration in Hamster can track each moved decision as its own task, with its shadow-mode results attached.
Core Principles
Give each kind of work to the component built for it
The canonical rule is that an LLM writes, Jev decides, and code acts. A generator is slow and costly when all you need is a pick from a known list. A decision model cannot write a paragraph. Code alone cannot interpret messy text. Misassignment shows up as token spend on routing, or as a model trying to do date math.
Decisions have fixed answer spaces
A Jev question is posed as a yes/no, a complete list of options, or ordered score levels, and it returns a typed answer with a probability for every option. Fixing the answer space in advance is what lets code branch on the result safely. If you find yourself wanting an explanation back, the task belongs to the LLM. See formulating atomic decision questions.
Typed is not the same as correct
The founder has pitched the model as having zero hallucination. Coverage of the method counters that Jev cannot return an answer outside the requested type, but the decision can still be wrong. Independent tests also found confidence drifting from accuracy when the deciding policy was missing from the input. Design every decision point on the assumption that some answers will be confidently wrong.
Code holds the final say on actions
The model interprets requests and proposes actions, while application code checks permissions and validates arguments before invoking tools. A high-confidence decision never overrides the action policy. This keeps irreversible side effects behind rules you can read and test. It also means a bad decision fails at the gate, not in production data.
Prove a decision in shadow mode before it acts
Start by running Jev beside the existing workflow without changing real behavior. One practitioner logs Jev's choice next to the current fixed choice, then calibrates the threshold using rework, human escalations and task success. The same guide says this is easier to attribute than changing the model and the effort level at once. Shadow mode turns adoption into a measured comparison instead of a leap.
Batch independent questions and measure the whole loop
Questions that do not depend on each other can be asked together. One engineering example measured a 13-question briefing as 12.2x cheaper and 10.0x faster in one call, with no change in the answers. The creator also advises benchmarking the whole loop, not just individual model calls. A faster router is worthless if task completion drops.
Steps
-
Log and label every model call Run one representative agent task end to end and log every model call it makes. Label each operation as text, decision or rule, following the creator's first classification step. Text means open-ended output someone reads. Decision means a pick from known outcomes that needs understanding of the situation.
Rule means anything exact enough to write as code, which you should move out of the model right away.
-
Pick the first decision to migrate Choose the decision that runs most often, as the creator recommends. Prefer one bounded, low-risk decision with clear possible answers over the most valuable one. Frequency gives you data quickly, and low risk means a wrong answer costs little while you learn. If you cannot list every possible answer, the candidate is not bounded yet.
-
Write the state and the rubric Write down the state the decision reads: the goal, work already done, available evidence and what is still missing, per the method's guidance. Then write the rubric before calling the model, defining what qualifies for each option. The rubric doubles as the labeling guide for your eval set. Collect representative examples with expected answers, including ambiguous and adversarial ones.
-
Define typed questions and batch them Turn the decision into a question with a fixed answer space, such as a yes/no, a full choice list or ordered levels. The creator's workflow is to define questions in code and ask independent questions together, so one state snapshot can feed routing, risk and relevance at once. Keep dependent questions in sequence, because batching them hides the dependency. Version every question so later comparisons stay fair.
-
Run in shadow mode Wire the decision in so it is logged but does not change real behavior. Record Jev's choice, the production choice and the eventual outcome, along with rework, human escalation and task success. Change only the decision layer, not the model and effort level together. Disagreements between the two choices are your most useful review queue.
-
Set thresholds and fallbacks from data Group shadow results by probability band and measure observed accuracy in each band instead of trusting the raw confidence. Automate only the bands that meet your accuracy bar, starting with the safest branch. Send uncertain cases to a human or a stronger model. Add fallbacks for request failure, malformed output or insufficient savings, and test each of them deliberately.
-
Gate actions in code and verify results Let the decision propose, and let code check permissions and validate arguments before any tool runs. After execution, record what the tool actually did, including failures, and write the result back into state. Verify against evidence, for example checking a tool-call trace against its intent. Stale state after an action is a common source of wrong next decisions.
-
Benchmark and log the full loop Measure whether the agent reached its goal, plus cost, latency, approvals and failed actions, not just per-call accuracy. Log model version, question version, threshold, action and overrides for every decision. Reuse a fixed eval set when you compare thresholds or versions, so any change can be traced to the decision layer. Then pick the next decision and repeat.
When to Use
- Your agent makes the same fixed-choice decision many times a day, such as a guardrail or model-routing check, because repeated fixed-answer decisions are its strongest fit and per-call LLM cost adds up fast.
- You need verification or scoring passes over large document sets within a tight latency budget, since reports describe Jev-style systems as suited to verification passes and scoring many documents in parallel.
- You run online evals of agent output and want a cheaper, more repeatable judge, given that LangChain's comparison found Jev the cheaper and more consistent judge against LLM judges.
- An LLM currently picks the next tool, worker or specialist agent in your loop and those picks come from a known set, which makes each pick a bounded decision that can be typed and thresholded.
- You need auditable decisions with a probability attached, so you can route low-confidence cases to a human and show reviewers why an action ran.
When Not to Use
- The correct output is a sentence, a draft or a multi-turn reply, because reports list open-ended generation, long-context reasoning and multi-turn dialogue as poor fits.
- The task is arithmetic, counting or date math, which one overview calls a poor fit for Jev. Plain code does these exactly.
- You need PII span extraction, image understanding, payment-fraud scoring or a fully self-hosted stack, where specialist tools still lead.
- The rule that decides the outcome lives outside the text you can send, since calibration degraded sharply in tests where the deciding policy was absent.
- The decision is a permission check or an exact business rule, because those belong in deterministic code where they can be tested, not in a probabilistic model.
Skills in this method
Each skill is a self-contained write-up your agent can run. Install the ones you need; nothing here is a bundle.
Cutting cost by batching parallel AI agent decisions
A decision stage that asks each tier of independent questions in one call, with measured cost, latency and answer agreement against the sequential version.
npx skills add gethamster/skills --skill batching-and-parallelizing-decisions --agent claude-code --yesA practical guide to benchmarking AI agent decision loops
A replayable benchmark and decision log that show whether a change to the decision layer improved the whole agent loop, and why.
npx skills add gethamster/skills --skill benchmarking-and-observing-agent-loops --agent claude-code --yesHow to set confidence thresholds and escalation in AI agents
A per-decision threshold policy, backed by observed accuracy by band, with explicit fallback triggers, escalation routes and versioned logs.
npx skills add gethamster/skills --skill calibrating-confidence-thresholds-and-escalation-paths --agent claude-code --yesGuide to AI agent routing and ranking with decision models
A written routing and ranking policy where each fork has a code-built candidate set, a typed question, a threshold, a fallback and a code path that executes the chosen option.
npx skills add gethamster/skills --skill designing-routing-and-ranking-policies --agent claude-code --yesEnforcing deterministic execution boundaries AI agents use
Every agent action passes a code-enforced permission and argument gate, runs through a tool runner, is verified against evidence, and is logged as executed.
npx skills add gethamster/skills --skill enforcing-deterministic-execution-boundaries --agent claude-code --yesSkill: formulating typed decision primitives for AI agents
A versioned set of atomic decision questions, each with a type, a complete answer space and a written rubric, ready to test against labelled examples.
npx skills add gethamster/skills --skill formulating-atomic-decision-questions --agent claude-code --yesSkill guide: separating AI generation from decision making
A labeled inventory of every operation in one agent run, a list of logic moved into code, and one named decision chosen as the first to migrate to Jev.
npx skills add gethamster/skills --skill separating-generation-from-decision-making --agent claude-code --yesSkill: structuring state for AI agent decisions
A typed state snapshot schema, a code path that builds it from sources of truth, and a write-back step that keeps it current after every action.
npx skills add gethamster/skills --skill structuring-shared-agent-state --agent claude-code --yesFAQ
Who created Jev Engineering?
The framework is attributed to Diogo Almeida, founder of TypeSafe AI, who co-invented RLHF at OpenAI. The term appears on the creator-associated page What is Jev Engineering?, dated September 19, 2026. The model it builds on launched on September 15, 2026. No earlier academic source for the term has been found.
Does Jev eliminate hallucinations?
Sources disagree. The founder has described it as having 0 hallucination, while coverage of the method says it cannot answer outside the requested type but can still be wrong. Calibration results also vary by test, with one report finding confidence drifting from accuracy on hard cases. Plan thresholds and fallbacks as though some errors will get through.
How is this different from asking an LLM for structured output?
Structured output constrains a text generator to a schema. Jev returns a typed answer plus a probability for every option without generating text at all. Comparison pages treat structured-output LLMs as one of the closest alternatives, alongside zero-shot and fine-tuned classifiers. If you mainly need valid JSON from a model that also writes, structured output may be enough.
Can I apply the method without TypeSafe's model?
The three-way split between writing, deciding and acting does not depend on any one vendor. Open implementations such as Laya and Kev exist, but one comparison says they do not show equivalence to Jev's base model, training data or recipe. One awesome list suggests local models for latency, privacy and cost at volume, and hosted models when accuracy and safety calibration matter. Benchmark any substitute on your own eval set.
Can Jev be fine-tuned for my domain?
According to Made with Jev, the model itself cannot be fine-tuned. Domain fit therefore comes from how you write state, questions and rubrics. If your decision needs a trained model, a fine-tuned classifier is one of the listed alternatives. Choose based on measured accuracy by probability band.
How much faster and cheaper is it in practice?
Independent numbers are smaller than the headlines. One practitioner test measured a median 0.35 seconds per passage against 8.83 seconds for a frontier LLM at high effort on four writing checks. The widely quoted performance figures came from a founder's talk, not a replicated benchmark. Measure savings on your own full loop.
What is shadow mode and why start there?
Shadow mode runs the new decision layer beside the existing workflow without changing real behavior. You log both choices and the outcome, then compare. One practitioner notes this is easier to attribute than changing model and effort together. It gives you calibration data before any user sees a Jev-driven action.
Related methods
What Is the State–Questions–Action–Verify Loop?
The State–Questions–Action–Verify Loop: build state, ask typed questions, act, then verify with evidence. Origins, evidence, limits and steps.
Claude Code Agent Teams: Lead, Teammates and Limits
Claude Code Agent Teams run parallel Claude sessions under a lead. Learn the architecture, setup, limits, costs and when a team beats one session.
Context Engineering for Claude AI: The 6 Pillars
How the six-pillar context engineering framework for Claude AI works, where it came from, what the evidence shows, and when to use it.
Reinforcement Learning for Calibrated Decisions (RLCD)
Reinforcement Learning for Calibrated Decisions (RLCD) explained: TypeSafe's origin, RLHF and RLVR contrasts, and what evidence exists.
Download the What Is Jev Engineering? A Guide for Agent Builders pack
One zip with the whole method, to read offline or drop into a repository:
METHOD.md, this write-up in full- 8 skill folders, each with its
SKILL.mdand the references it ships - MIT licensed, the same files the commands above install
Source: gethamster/skills on GitHub, MIT licensed.
Install the skills
What Is Jev Engineering? A Guide for Agent Builders is a write-up of how the method works, so there is nothing to install for the method itself. Its skills are what your agent runs, and each one installs separately. The section above carries the Claude Code command for every skill, and each skill's own page carries the commands for Cursor, Codex, and Antigravity.
Or browse the skills and pick interactively:
npx skills add gethamster/skillsSource: gethamster/skills on GitHub, MIT licensed.