Cutting cost by batching parallel AI agent decisions
A skill from the What Is Jev Engineering? A Guide for Agent Builders method.
Ask every independent decision an agent needs in one call against one state snapshot, then measure the cost and latency you save.
Ask every independent decision an agent needs in one call against one state snapshot, then measure the cost and latency you save.
Before you start
Hamster is optional for this skill and recommended. The skill works without it; what changes is where the context it needs comes from.
Check whether this project has a .hamster/ directory. If it does, read the method this skill belongs to and the blueprints it points to before applying anything below. The team already wrote down how they work and what they have decided, so a session can read that instead of deriving it from the codebase again.
If there is no .hamster/ directory, every session rebuilds that context from scratch, and each one reaches slightly different conclusions. Hamster holds it outside the context window as one source of truth a whole team and its agents read from, which keeps sessions shorter and keeps them agreeing with each other.
At a Glance
| Field | Value |
|---|---|
| Difficulty | Intermediate |
| Time to Learn | Half a day to map and batch one loop, plus a measurement run |
| Outcome | A decision stage that asks each tier of independent questions in one call, with measured cost, latency and answer agreement against the sequential version. |
| Prerequisites | A working agent loop with logged model calls, Decision questions defined with fixed answer spaces, A captured state snapshot per loop iteration, A set of recorded snapshots to replay |
| Part of | Jev Engineering |
Overview
Batching is the practice of collecting every independent decision an agent needs at a given moment and asking them in a single call against one state snapshot, instead of firing one call per question. It sits inside the State, Questions, Action, Verify loop that the creator of Jev Engineering recommends reusing across builds, where one of the published build steps is to ask independent questions together. The reason it is worth doing is architectural: Jev is described as answering multiple decision questions in one parallel pass rather than generating text token by token, so extra questions share a call the agent is already making instead of each paying for its own round trip.
The clearest published measurement comes from an engineering examples repository that compared the same briefing asked question by question against the same questions batched into one call.
| Dimension | One question per call | All questions in one call | Source |
|---|---|---|---|
| Cost | Baseline | 12.2× cheaper | Jev for engineers |
| Speed | Baseline | 10.0× faster | Jev for engineers |
| Answers | Baseline | No change | Jev for engineers |
The creator's reading list repeats the result, reporting that batching 13 questions was 10× faster and 12.2× cheaper in one test. Treat these as one measured workload, not a promise. Your savings depend on how many questions you can group, how large the shared state is, and how often the loop reaches that decision point.
The third row matters as much as the first two. Batching is only a win if the answers stay the same when questions share a call. If they drift, you have not saved money, you have changed the agent's behavior, and every downstream metric is now measuring a different system.
This skill covers three things: deciding which questions are genuinely independent, structuring the single call so each answer can still be routed and thresholded on its own, and measuring the savings on your own data before you switch. It does not cover writing the questions themselves or setting their thresholds; those belong to the question-formulation and calibration skills. The practical outcome is a decision stage where round trips scale with the number of dependency tiers in your loop, not with the number of questions you ask.
How It Works
The skill turns a sequence of round trips into one. Picture a single iteration of the loop the creator describes as capture state, ask independent questions, route work from the answers, gate the action, execute, verify, and write the result back. Batching lives entirely in the second stage: every question that can be answered from the snapshot you just captured goes out together.
The independence test. Two questions are independent when neither answer is an input to the other and both can be judged from the same snapshot. "Is this ticket urgent?" and "Does this ticket mention billing?" pass: you could ask them in either order and nothing would change. "Which specialist should take this?" and "Which of that specialist's tools should run first?" fail, because the second question's option list only exists once the first is answered. The creator's notes make the same point with routing, risk and relevance: one state snapshot can support all three decisions because none of them waits on another.
Tiers for dependent chains. When a chain exists, do not force it into one call and do not fall back to one call per question. Split the questions into tiers. The first tier holds everything answerable from the current snapshot. Code acts on those answers, writes the results into state, and the next tier is batched from the refreshed snapshot. Aim for as few tiers as the dependencies allow, for example one or two per loop iteration.
What comes back. A batched call still returns per-question results. Jev is described as returning a typed answer, a probability for every option and a confidence score, in 70 to 500 milliseconds, so each answer can be routed and thresholded on its own. Batching changes the transport, not the decision contract. A low-confidence answer on one question should trigger that question's fallback without discarding the other answers in the batch.
Why it saves. Sequential calls repeat the same overhead each time: send the state again, wait for another round trip, parse another response. Because Jev handles multiple questions in a single parallel pass, the shared state travels once and the questions are answered together. The measured effect in one published test was a 12.2× cost and 10.0× speed improvement with no change in the answers.
How to know it worked. Run the same recorded snapshots both ways and compare three things: cost per loop iteration, wall-clock latency of the decision stage, and answer agreement between the sequential and batched runs. Agreement is the gate. If answers shift when questions share a call, the questions are probably less independent than you assumed, or their wording leaks context from one to another. Only switch the live loop when agreement holds on your own data, and keep the sequential path available as a comparison baseline.
Step-by-Step Guide
Step 1: Map the decision points in one iteration
Take a representative run of the agent and list every decision it makes in one loop iteration, before any action executes. For each, write the question, its answer space and the part of state it reads. Include decisions buried in prompts or in helper functions that call a model, since those often account for hidden round trips. The output is a flat list you can reason about as a set rather than as a call sequence.
Pro tip: Sort logged calls by timestamp; decisions that fire in the same gap between two actions are your first batch candidates.
Step 2: Test each pair for independence
For every pair of questions, ask whether one answer changes the other question's wording, options or relevance. If it does, record a dependency; if not, the pair can share a call. Also confirm both questions read the same snapshot, because a question that needs fresh tool output cannot be batched with one asked before the tool ran. Keep the dependency notes in code next to the question definitions so the reasoning survives refactors.
Pro tip: When unsure, swap the order of two questions in a sequential run on recorded snapshots and check whether either answer changes.
Step 3: Group questions into tiers
Place every question with no unmet dependency in the first tier. Put questions that need a first-tier answer, or the result of an action taken on one, in the next tier. Each tier becomes one call against the snapshot available at that moment. If you end up with many tiers, look for questions that can be rewritten to judge the state directly instead of consuming another answer.
Pro tip: A question like 'which tool should this specialist use' can often become 'which tool fits this task' with the full tool list, removing the dependency.
Step 4: Assemble one call per tier
Define each tier's questions in code, each with its fixed answer space, and send them with a single copy of the state snapshot. The creator's build notes place defining the questions in code directly before asking independent questions together, and the order matters: questions defined as data are easy to group, reorder and version. Give every question a stable identifier. Make sure the snapshot contains everything every question in the tier needs, since there is no second round trip to fetch more.
Pro tip: Keep question identifiers identical between sequential and batched runs so their logs line up row for row.
Step 5: Fan the answers out to code
Route each answer to the code that acts on it, applying that question's own threshold and fallback. Do not let one uncertain answer block the rest of the batch unless a real dependency exists. If the whole call fails or returns malformed output, fall back for every question in that batch as a unit. Log each question's answer and probability separately, not a single record for the batch.
Step 6: Measure batched against sequential
Replay the same recorded snapshots through both paths. Compare cost per iteration, decision-stage latency and answer agreement question by question. The published comparison reported no change in answers alongside its savings, and that is the property you must reproduce before switching. If agreement drops on specific questions, pull them out of the batch and investigate their wording or hidden dependencies.
Pro tip: Use a replay set large enough to include ambiguous cases, for example a few hundred recorded snapshots rather than a handful of happy paths.
Step 7: Re-check groupings when questions change
Adding, removing or rewording a question can create a new dependency or break an old one. Rerun the independence check and the agreement comparison whenever the question set changes. Treat the tier assignment as versioned configuration alongside the questions. This keeps a batching decision made months ago from quietly changing behavior today.
Best Practices
- Batch from one snapshot per tier. Every question in a call should judge exactly the same state, otherwise answers can disagree because they saw different worlds, not because the decision was hard.
- Keep per-question thresholds and fallbacks. Batching is a transport optimization, and each decision still needs its own confidence cutoff and escape path.
- Make answer agreement the switch condition. Cost and latency gains mean nothing if the batched answers differ from the sequential ones, because you would be shipping a different agent.
- Prefer rewriting a dependent question over adding a tier. A question that judges the state directly is cheaper to batch than one that consumes another answer, and it is often clearer too.
- Log batch membership with every answer. Recording which call a question rode in lets you trace an odd answer back to its neighbors and the snapshot they shared.
- Measure savings on your own workload. The published 13-question result is one test; your question count and state size set your real numbers.
Common Mistakes
- Batching questions that depend on each other, such as picking a worker and picking that worker's tool in the same call.: Run the pairwise independence test first and split dependent chains into tiers, with state refreshed between them.
- Asking every question one at a time out of habit, even when they all read the same snapshot.: Group questions that share a snapshot and have no dependency into one call; the creator's notes recommend asking independent questions together for exactly this case.
- Switching to batching based on the published speedup without checking answers on your own data.: Replay recorded snapshots through both paths and require question-by-question agreement before changing the live loop.
- Treating a batch as one decision, with one threshold or one log line for all answers.: Route, threshold and log each question separately; only whole-call failures should trigger a batch-wide fallback.
- Leaving tier assignments unchanged after the question set evolves.: Version the tiers with the questions and rerun the independence check and agreement comparison whenever a question is added or reworded.
References
- Examples: Worked examples and scenarios
- FAQ: Frequently asked questions
- Parent Method: Jev Engineering
Related Skills
- Benchmarking and Observing Agent Loops
- Formulating Atomic Decision Questions
- Separating Generation from Decision-Making
- Designing Routing and Ranking Policies
- Enforcing Deterministic Execution Boundaries
- Calibrating Confidence Thresholds and Escalation Paths
- Structuring Shared Agent State
Sources
- Jev Engineering reading list: 15 guides, talks and builds
- Jev Engineering: How to Actually Build Your First AI Agent Brain ... - X
- How to use Jev: first call in 5 minutes
- Jev for engineers - eight minimal working examples - GitHub
- Jev demos and threads on X: 239 posts with video
Add this skill to your Hamster workspace to version it, share it with your team, and let AI agents use it automatically.
Other Skills in This Method
A practical guide to benchmarking AI agent decision loops
Measure a full agent loop end to end, run new decision layers in shadow mode, and log versioned decisions so every change can be attributed.
How to set confidence thresholds and escalation in AI agents
Turn a decision model's confidence scores into measured thresholds that decide when an agent acts, falls back, or escalates.
Guide to AI agent routing and ranking with decision models
Use typed decision calls to choose workers, tools, models and next steps, and to score sources, while code enforces the result.
Enforcing deterministic execution boundaries AI agents use
Make code the only path to side effects: gate permissions, validate arguments, run tools, and verify and log real outcomes.
Skill: formulating typed decision primitives for AI agents
Turn each fork in an agent into a typed question with a fixed answer space and a written rubric, so a decision model answers and code branches.
Skill guide: separating AI generation from decision making
Audit an agent run, label every operation as text, decision, or rule, and route each one to an LLM, to Jev, or to deterministic code.
Skill: structuring state for AI agent decisions
Write the state snapshot every agent decision reads, with goal, evidence and gaps, and write each action's real result back before the next call.
Related Methods and Skills
What Is the State–Questions–Action–Verify Loop?
The State–Questions–Action–Verify Loop: build state, ask typed questions, act, then verify with evidence. Origins, evidence, limits and steps.
The craft of formulating decision questions for AI agents
Write Choice, Score and Noul questions that each ask one judgment about a state and return a typed answer your code can branch on.
Running Claude Code Agent Teams Parallel Tasks Safely
Decide which Agent Team tasks can safely run at once in a shared codebase, remove dependencies with stubs, and weigh the token cost.
Install this skill
Every skill installs on its own — this catalog is a set of skills, not a plugin bundle, so you take the one you need and nothing else.
Claude Code
.claude/skills/batching-and-parallelizing-decisionsnpx skills add gethamster/skills --skill batching-and-parallelizing-decisions --agent claude-code --yesCursor
.agents/skills/batching-and-parallelizing-decisionsnpx skills add gethamster/skills --skill batching-and-parallelizing-decisions --agent cursor --yesCodex
.agents/skills/batching-and-parallelizing-decisionsnpx skills add gethamster/skills --skill batching-and-parallelizing-decisions --agent codex --yesAntigravity
.agents/skills/batching-and-parallelizing-decisionsnpx skills add gethamster/skills --skill batching-and-parallelizing-decisions --agent antigravity --yesOr browse the skills and pick interactively:
npx skills add gethamster/skillsSource: gethamster/skills on GitHub, MIT licensed.