What Is the State–Questions–Action–Verify Loop?
Updated 7 skills7 stepsOrigin: TypeSafe AI
Created by TypeSafe AI - https://typesafe.ai
Overview
The State-Questions-Action-Verify Loop is a workflow pattern for AI agents and AI-assisted features. Each cycle builds an explicit state, asks narrow typed questions about it, and takes one bounded action chosen from the answers. It then checks the result with fresh evidence before deciding to continue, retry or stop. The name does not belong to an established framework with a documented creator. The four parts combine vocabulary from TypeSafe AI's documentation on state with the generic agent loop described in research and practitioner writing. Treat the name as a convenient label for a set of practices, not as a standard you can cite for authority.
The state and questions half comes from TypeSafe. Its docs define state as the content a System One model evaluates, such as a support message, a passage of text, or the current state of your application. Each request evaluates one state against one or more questions. Its primitives come in pairs: a question defines one judgment to make about the state, and the answer is a typed value returned to your code. The System One page describes the rest of the workflow. The application builds state, asks independent questions, combines the answers with deterministic checks in code, and routes the case for action or review. TypeSafe also documents verification as a separate use of the same primitives. Its SDE cascade extracts fields with a cheap model, checks each field with a yes/no question, and escalates to a more expensive model when a verifier fires. None of these pages uses the four-part name.
The action and verify half comes from the agent-loop literature. A 2026 review of agent systems describes the minimal loop as retrieve context, plan, act via tools, verify, update memory and repeat, and ties it to the ReAct and reflection patterns. A practitioner guide to agent architecture gives a longer version: observe, reason, choose action, act, observe result, update state, verify, then continue or terminate. This loop changes two things. It replaces open-ended reasoning with typed questions whose answers code can branch on, and it treats state as an artifact you build and persist, not as whatever sits in the context window.
flowchart TD
A[Build state] --> B[Ask typed questions]
B --> C[Combine answers in code]
C --> D[Take bounded action]
D --> E[Verify with fresh evidence]
E -->|pass, work remains| A
E -->|fail| F[Record cause]
F -->|retry| A
F -->|stuck| G[Hand to human]
E -->|done| H[Stop]
The evidence that verification loops help is suggestive, not conclusive. A 2026 practitioner article reports higher OSWorld success for self-correcting loops than for sequential or single-step approaches. That article does not state the study design, sample size, benchmark version, model setup or cost. Its self-correcting loops are also not this exact pattern. Read the figures as a direction, not as proof that this loop causes the gain.
| Approach | Shape | Reported OSWorld result | Caveat |
|---|---|---|---|
| Single-step | One prompt, one action | About 30% per this article | Method details not published |
| Multi-step sequential | Fixed chain, no checking | About 45% per this article | Method details not published |
| Self-correcting loop | Act, check, revise | 58.2% per this article | Not this exact pattern |
| ReAct-style loop | Plan, act, verify, update memory | Not reported | Verification often free-form |
| This loop | Typed questions on explicit state | Not reported | No published benchmarks |
If you adopt the loop, test it against baselines. A loop-engineering roadmap recommends comparing recurring loops with the current human workflow, a single-agent run and a bounded fixed-retry policy. It also recommends reporting failed runs, costs, human interventions and state diffs. Two limits deserve attention. The first is verifier reliability. When static matching is not possible, StateFlow uses an LLM to judge whether the problem is solved, so your checker can be wrong in the same ways as your agent. The second is the gap between partial progress and success. Attaching a check to every step lets you report which checks passed and which remain, instead of calling the task done. State format matters too. The TypeSafe API accepts state as a string, object or array. Structured state lets deterministic code check fields directly. Plain text leaves more of the judgment to the model.
In Hamster Studio, teams can keep the loop's state, questions and verification evidence in one shared workspace so agents and people read from the same record. The skill pages cover each step in detail: defining verifiable success criteria, representing current agent state, formulating typed decision questions, selecting and executing bounded actions, checking post-action results, recovering from verification failures and applying loop termination and continuation rules.
Core Principles
State is an artifact, not a memory
Build the state from sources you control and keep it outside the model. Loop-engineering guidance lists progress files, database checkpoints, traces and issue comments as the places state survives between runs. TypeSafe's definition of state makes the same point from the other side: state is the concrete content being judged. If you cannot print the state a decision was made against, you cannot audit or reproduce that decision.
One judgment per question
Each question should ask for exactly one decision. TypeSafe's primitives are designed this way: a question defines one judgment about a state, and its answer is a typed value. Compound questions blur which condition drove the answer and make thresholds impossible to tune. When a question needs the word and, split it into two.
Code decides, the model judges
Typed answers feed deterministic logic. The model does not free-write the next move. The System One workflow combines independent answers with deterministic checks in code before routing a case for action or review. This keeps policy in reviewable code and keeps the model's job narrow.
If business rules live only in a prompt, you have lost this property.
Every consequential action has an expected outcome
Before acting, write down what should be different afterward. The generic agent loop places observe result and verify directly after act, and that only works if you know what you are verifying against. Checking that an action ran is not the same as checking that it worked. A missing expectation is the most common reason verification quietly becomes a formality.
Evidence beats the agent's claim
Verification should rest on observable evidence. A practitioner guide to loop engineering names tests, builds, diffs, working links, screenshots and written acceptance criteria as acceptable evidence, and warns against the agent's unsupported judgment that it is done. Where only a model can judge, remember that StateFlow falls back to an LLM evaluator in that situation. Your verifier then needs its own testing.
Partial progress is not success
Track each required check separately so the loop can report exactly what passed and what remains. The pattern of pairing every step with a verify check makes partial completion visible instead of rounding it up to done. This matters most when a loop stops on budget. A clear list of remaining checks lets a person pick up the work without redoing it.
Exits and approvals are designed before the run
Decide in advance what done means, when to stop, and which actions need human approval. The pre-start questions in loop-engineering guidance cover scope, per-iteration work, verification, where progress is recorded and approval points. A loop-engineering roadmap adds that a mature loop should show why each action was allowed and when control returned to a person. Improvised stop conditions tend to show up only after the budget is gone.
Steps
-
Define done and how it will be checked Before anything runs, write the acceptance criteria as observable checks with pass thresholds. List the evidence each check will use, such as a test result, a field value or a reviewer sign-off. Decide what counts as partial progress and what counts as full success. Also record which actions will need human approval.
The detailed method is on the defining verifiable success criteria page.
-
Build the state Assemble the content the questions will be judged against: the input message, relevant records, the applicable policy and any progress from earlier iterations. Prefer structured fields over a text blob when code will need to check values directly. Keep it small enough that every item is relevant to at least one question. Persist it somewhere outside the model so the next iteration and any human reviewer see the same thing.
-
Ask typed questions Write one question per judgment and choose its answer type: a pick from options, a score on a described scale, or a yes/no probability. Batch questions that are independent of each other against the same state. Define in advance what each answer value means for the workflow. If you cannot say which branch a given answer triggers, the question is not ready.
The formulating typed decision questions page covers phrasing.
-
Choose one bounded action in code Combine the typed answers with deterministic checks to select a single action from a fixed, validated set. Include a no-match or blocked outcome so the loop is never forced into an invalid choice. Use answer confidence to decide between acting, asking for confirmation and routing to a person. Validate arguments and scope before execution.
-
Execute and observe fresh results Run the action, then gather new evidence about its effect instead of reusing observations taken before it. Compare that evidence with the expected outcome you wrote for the action. A successful API call or a finished command is not evidence that the intended change happened. Record the raw result alongside the state.
The checking post-action results page covers picking the right signal.
-
Handle a failed check When verification fails, write the failure evidence and its likely cause into the state before trying again. Retry only within a bounded count, and change something between attempts rather than repeating the same action. Escalate to a stronger model or hand control to a person when retries stop producing progress. Preserve the full state at handoff so the person does not start from scratch.
-
Persist and decide whether to continue After each iteration, save the updated state and the evidence, then apply the exit rules you defined in the first step. Stop on success, continue if required checks remain and progress is being made, and stop with a handoff on stall, budget exhaustion or a pending approval. Report partial progress as a list of passed and outstanding checks. A loop that can only exit on success is incomplete.
When to Use
- Recurring triage or routing decisions, such as support tickets or refund requests, where each case can be described as a bounded state and the next step is one of a known set of actions.
- Agent tasks with a cheap, objective check available after each action, such as tests, a build or a schema validation, because the loop's value depends on verification that can say yes or no.
- Extraction or classification pipelines where a cheap model does most of the work and you want per-field checks that escalate only suspicious items to a stronger model or a reviewer.
- Workflows that span several runs or sessions, where progress must survive restarts and a person may need to take over midway with a clear record of what passed.
- Situations where you need an audit trail of why an action was allowed, since typed answers, explicit state and stored evidence make each decision reviewable later.
When Not to Use
- Open-ended creative or exploratory work, such as brainstorming, where no observable check can separate a good result from a bad one and verification would be theater.
- One-shot requests with no follow-up action, where the overhead of building state, typed questions and exit rules costs more than simply reviewing the single output.
- Tasks where the only available verifier is the same model grading its own work with no independent signal, because the loop then adds cost without adding reliability.
- High-stakes irreversible actions where no confidence level justifies automation, since the correct design there is human approval on every case rather than a loop.
Skills in this method
Each skill is a self-contained write-up your agent can run. Install the ones you need; nothing here is a bundle.
Applying agent loop termination and continuation logic
A written set of exit and continuation rules that produces one recorded verdict per iteration and never leaves a loop running on a dead end.
npx skills add gethamster/skills --skill applying-loop-termination-and-continuation-rules --agent claude-code --yesPractical steps for verifying AI agent action results
Every consequential agent action ends in a passed, failed or could-not-verify verdict backed by fresh, inspectable evidence stored in state.
npx skills add gethamster/skills --skill checking-post-action-results --agent claude-code --yesSkill: defining success criteria for AI agent verification
A testable acceptance specification: observable checks, pass thresholds, failure conditions, evidence requirements and review conditions.
npx skills add gethamster/skills --skill defining-verifiable-success-criteria --agent claude-code --yesThe craft of formulating decision questions for AI agents
A named map of typed questions, each with defined options or a rubric and a written rule for how code acts on its answer.
npx skills add gethamster/skills --skill formulating-typed-decision-questions --agent claude-code --yesHow to do agent error recovery and failure handling
A written failure policy that routes every failed check to a bounded retry, an escalation, or a human handoff, with the cause and state preserved.
npx skills add gethamster/skills --skill recovering-from-verification-failures --agent claude-code --yesSkill: representing agent state in AI workflows
A state schema and persistence routine that every question, action and verification step reads from and writes back to.
npx skills add gethamster/skills --skill representing-current-agent-state --agent claude-code --yesSkill guide: bounded action selection in agent systems
A constrained action layer where every model decision maps to a validated, confidence-gated operation or an explicit handoff.
npx skills add gethamster/skills --skill selecting-and-executing-bounded-actions --agent claude-code --yesFAQ
Is the State-Questions-Action-Verify Loop an established framework?
No. The research found no source that uses this exact four-part name or credits it to a creator. The state and typed-question mechanics come from TypeSafe AI's documentation, and the act and verify cycle comes from general agent-loop writing. Use the name as shorthand for the combined practice, not as an authority.
How is it different from ReAct?
ReAct-style loops interleave reasoning and tool use, and a 2026 review places verification and memory updates inside the minimal agent loop. This pattern keeps that cycle but narrows the reasoning step into typed questions whose answers deterministic code branches on. It also treats state as an explicit, persisted artifact. You trade some flexibility for decisions that are easier to test and audit.
Does adding verification actually improve agent success rates?
There is suggestive evidence. A practitioner article reports 58.2% OSWorld success for self-correcting loops against about 45% for sequential and about 30% (source) for single-step approaches. The article does not publish the study design, model setup or cost, and it does not test this exact loop. Run your own comparison against a single-agent baseline and a fixed-retry policy, as a loop-engineering roadmap recommends.
What if my verifier has to be a model too?
That is common when no static check exists, and StateFlow uses an LLM to judge whether a problem is solved in exactly that case. The risk is that the verifier shares the agent's blind spots. Keep verifier questions narrow and per-field, as in TypeSafe's SDE cascade, and sample verifier decisions for human review. A survey of robot-policy verifiers describes KnowNo, which requests human help when several actions remain plausible, as one way to contain this risk.
Should state be plain text or structured data?
Both work. The TypeSafe API accepts state as a string, object or array. Structured state lets code check specific fields deterministically and makes diffs between iterations readable. Plain text suits messages and documents where the judgment is inherently linguistic. Many loops mix the two, with structured records plus a text field for the raw input.
How should the loop report partial progress?
Track each acceptance check separately and report which have passed and which remain. The pattern of attaching a verify check to every step makes this natural. Never collapse partial completion into done because the budget ran out. The outstanding checks are the handoff note for whoever continues.
Do I need TypeSafe to use this loop?
No. The pattern is tool-agnostic: any system that can hold explicit state, return typed answers and run checks after an action can implement it. TypeSafe's System One workflow is simply the clearest documented example of the state and questions half. The action and verification half appears across general agent-loop guidance.
Download the What Is the State–Questions–Action–Verify Loop? pack
One zip with the whole method, to read offline or drop into a repository:
METHOD.md, this write-up in full- 7 skill folders, each with its
SKILL.mdand the references it ships - MIT licensed, the same files the commands above install
Source: gethamster/skills on GitHub, MIT licensed.
Install the skills
What Is the State–Questions–Action–Verify Loop? is a write-up of how the method works, so there is nothing to install for the method itself. Its skills are what your agent runs, and each one installs separately. The section above carries the Claude Code command for every skill, and each skill's own page carries the commands for Cursor, Codex, and Antigravity.
Or browse the skills and pick interactively:
npx skills add gethamster/skillsSource: gethamster/skills on GitHub, MIT licensed.