Practical outcome based reinforcement learning reward design
A skill from the Reinforcement Learning for Calibrated Decisions (RLCD) method.
Build rewards that score a model's stated probabilities against realized outcomes so training pushes confidence toward observed correctness.
Build rewards that score a model's stated probabilities against realized outcomes so training pushes confidence toward observed correctness.
Before you start
Hamster is optional for this skill and recommended. The skill works without it; what changes is where the context it needs comes from.
Check whether this project has a .hamster/ directory. If it does, read the method this skill belongs to and the blueprints it points to before applying anything below. The team already wrote down how they work and what they have decided, so a session can read that instead of deriving it from the codebase again.
If there is no .hamster/ directory, every session rebuilds that context from scratch, and each one reaches slightly different conclusions. Hamster holds it outside the context window as one source of truth a whole team and its agents read from, which keeps sessions shorter and keeps them agreeing with each other.
At a Glance
| Field | Value |
|---|---|
| Difficulty | Advanced |
| Time to Learn | 2-4 weeks for a first working reward and calibration pass |
| Outcome | A reward signal and post-hoc calibration step that make a decision model's confidence values track how often its decisions turn out to be correct. |
| Prerequisites | Working knowledge of reinforcement learning and policy-gradient methods, Familiarity with softmax outputs and probability distributions, A decision task with a defined answer space and resolvable ground truth, Held-out data for post-training calibration |
| Part of | Reinforcement Learning for Calibrated Decisions (RLCD) |
Overview
Outcome-based reward design is the work of deciding what number a decision model receives after it commits to a probabilistic answer and the real outcome becomes known. In the calibrated-decision setting, the RLCD glossary definition ties the reward to whether the model's stated probability matches how often its answer turns out to be correct, rather than to human preference ratings or automatic verification. For the definition and history of the approach, see the RLCD method page; this page covers building the reward itself.
This is a design problem, not a configuration step, because TypeSafe has not published the reward it uses. Independent commentary notes that the RLCD reward function and training procedure remain undisclosed. Practitioners therefore work from open reimplementations. The OpenJev repository trains with a composite of strictly proper scoring rules, cross-entropy plus a multi-class Brier score, and a separate write-up suggests an implementation can apply the Brier score during training and again as a held-out objective for post-hoc calibration. Treat these as reference patterns you can test, not as the vendor's recipe.
The inputs are concrete. Each training example carries a state, a structured question, a correct answer and a defined output schema, as described in this RLCD term explainer. The model emits a probability distribution over the permitted answers, and the realized outcome arrives later. The outputs are a scalar reward per decision, a policy update, and optionally a calibration layer fitted after training.
The central decision is which scoring rule to use. A strictly proper rule gives its best expected score when the reported probability equals the true chance of the outcome, so a model cannot improve its reward by hedging toward the middle or by exaggerating certainty. A reward that only checks whether the top choice was right does not have this property: it pays the same for a correct answer at low and high confidence, so it gives no pressure toward honest probabilities.
You can tell the design went wrong when accuracy improves but confidence stops meaning anything, when one task type looks well calibrated and another drifts, or when your ground truth is itself shaky. Commentary on RLCD raises exactly this last risk, warning about benchmark bias when the correct answer is defined by an average of generative-AI outputs. A reward can only be as honest as the outcome it is scored against.
How It Works
The reward loop has four stages that repeat for every batch: the model produces a distribution over candidate decisions, the true outcome is observed, a proper scoring rule converts the gap between the two into a reward, and a policy-gradient step shifts the model toward higher expected reward.
flowchart LR
A[Structured example] --> B[Decision distribution]
B --> C[Observe ground truth]
C --> D[Proper scoring rule]
D --> E[Reward]
E --> F[Policy gradient update]
F --> B
F --> G[Held-out calibration]
G --> H[Calibrated probabilities]
The RLCD term explainer describes this shape: the model produces a probabilistic decision, a reward function evaluates how well its probability distribution matches the ground truth, and policy-gradient updates adjust the model to increase expected reward on future decisions. In practice the reward is the negative of a loss, so lowering the scoring-rule penalty and raising the reward are the same move.
The three components used in open implementations play different roles and run at different times.
| Component | Role | When applied |
|---|---|---|
| Cross-entropy | Penalizes low probability on the true outcome, part of OpenJev's composite loss | During training |
| Brier score | Squared probability error across all candidates, weighted 1.0 in the OpenJev loss | Training, and as a held-out objective |
| Temperature scaling | Rescales softmax outputs toward empirical accuracy via L-BFGS, per OpenJev | After training, on held-out data |
Cross-entropy looks only at the probability assigned to the outcome that happened. It punishes confident misses very hard, which gives strong gradients early in training but can make the model sensitive to mislabeled examples. The Brier score, which the OpenJev repository describes as penalizing squared probability discrepancies across all candidate outcomes, is bounded and looks at the whole distribution, so it cares whether the probability mass on wrong answers is sensibly spread. OpenJev adds the two with an equal weight of 1.0 in its published loss; that weight is one project's choice and a reasonable starting point to tune, not a standard.
Temperature scaling sits outside the reward loop. After the policy is trained, you fit a single temperature on held-out data so that softmax probabilities line up with how often the model is actually right, which is how OpenJev applies post-hoc L-BFGS temperature scaling. Using the Brier score as the held-out objective, as one RLCD write-up suggests, keeps the calibration target consistent with the training signal. Because temperature scaling preserves the ranking of candidates, it changes confidence without changing which decision the model picks.
The design intent throughout matches the glossary description: the reward tracks whether stated probabilities match correctness rates, not whether a rater liked the answer or a program accepted it.
Step-by-Step Guide
Step 1: Define the outcome and when it resolves
Write down, for each decision type, what counts as correct and when that becomes known. Some outcomes resolve instantly from a label, while others arrive days later from a downstream system. The reward cannot be computed until the outcome exists, so your training data pipeline has to join decisions to outcomes reliably. If the definition of correct is ambiguous, the scoring rule will faithfully reward that ambiguity.
Document exclusions, such as cases that never resolve, instead of silently treating them as wrong.
Pro tip: Keep a small set of hand-reviewed outcomes to spot-check the automated labels before they feed any reward.
Step 2: Build structured training examples
Package each example with a state, a structured question, the correct answer and a defined output schema, the anatomy described in the RLCD term explainer. The schema fixes the candidate set the model distributes probability over, which is what makes a scoring rule computable. Without a closed answer space you cannot score probability mass on wrong answers. Schema design itself is covered in designing schema-constrained decisions.
Step 3: Choose a strictly proper scoring rule
Pick cross-entropy, the Brier score, or a composite; the OpenJev implementation combines both. Cross-entropy gives sharp gradients but is harsh on confident mistakes and noisy labels. The Brier score is bounded and considers the full distribution. Avoid rewards based only on whether the top choice matched, because they carry no information about whether the stated confidence was right.
Record the rule and any weights so later calibration results can be traced back to them.
Pro tip: Start with the composite and an equal weight, then compare variants on held-out calibration rather than training loss.
Step 4: Turn the score into a reward and update the policy
Define the per-decision reward as the negative scoring-rule loss, so better-calibrated distributions earn more. Apply policy-gradient updates that raise expected reward on future decisions, the step the RLCD explainer describes. Normalize or baseline the reward to keep gradient variance manageable. Watch both average reward and the spread of predicted probabilities, since a collapse toward uniform or toward extreme values signals a problem.
Pro tip: Log the entropy of predicted distributions per batch; a steady slide toward zero usually means growing overconfidence.
Step 5: Fit post-hoc temperature scaling on held-out data
After training, reserve data the policy never saw and fit a single temperature that aligns softmax probabilities with empirical accuracy, as OpenJev does with L-BFGS. Use a proper scoring rule as the fitting objective; one RLCD write-up suggests reusing the Brier score. Temperature scaling does not change which decision is chosen, only how confident it looks. Refit whenever the input distribution shifts meaningfully.
Pro tip: Never fit the temperature on training data; it will look perfect and fix nothing.
Step 6: Audit the reward against calibration by task type
Check whether the reward actually produced calibrated probabilities by binning predictions and comparing stated confidence to observed accuracy. Do this separately for each decision type, because the RLCD practitioner guide warns that calibration can vary by task. If one type is badly off, inspect its outcome labels and its share of the training data before changing the scoring rule. The measurement mechanics are covered in evaluating probabilistic calibration.
Pro tip: Set a review cadence, for example after every retraining run, so drift is caught before software starts acting on stale confidence.
Best Practices
- Score the whole distribution, not only the chosen answer. A reward that ignores the probability on wrong candidates cannot teach the model how to express uncertainty between close alternatives.
- Keep the scoring rule strictly proper, as the OpenJev composite does. Proper rules make honest probabilities the best expected strategy, so the model gains nothing by gaming its confidence.
- Separate training from calibration data. Fitting temperature or evaluating the Brier score on held-out data, as suggested in one RLCD write-up, is the only way to see calibration the model has not memorized.
- Treat outcome quality as part of the reward. Commentary flags bias when ground truth is an average of generative-AI outputs, so prefer real resolved outcomes over proxy consensus labels.
- Label reference implementations for what they are. TypeSafe has not disclosed its reward, per independent analysis, so document that your loss weights come from open projects and your own tuning.
- Report calibration per decision type alongside the reward curve. A rising average reward can hide one task type drifting while another improves.
Common Mistakes
- Rewarding only whether the top decision was correct.: Use a proper scoring rule over the full distribution. The RLCD framing ties reward to whether stated probability matches correctness rates, which an accuracy-only reward cannot measure.
- Copying the OpenJev loss weight as if it were a standard.: The equal weighting in OpenJev's composite loss is one project's choice. Tune the balance against held-out calibration for your own task mix.
- Fitting temperature scaling on the training set.: Fit it on data the policy never saw, as post-hoc calibration intends. Training-set fits look well calibrated and fail in production.
- Assuming a public reimplementation reproduces TypeSafe's method.: Treat open code as a plausible pattern only. Analysts note the reward function and training procedure are unpublished, so no public loss can be claimed as the original.
- Pooling all decision types into one calibration check after training.: Audit calibration separately per task, because the practitioner guide notes it can vary by question type. A pooled number can mask a badly calibrated minority task.
References
- Examples: Worked examples and scenarios
- FAQ: Frequently asked questions
- Parent Method: Reinforcement Learning for Calibrated Decisions (RLCD)
Related Skills
- Calibrating Confidence to Outcomes
- Structuring Machine-to-Machine Decision Outputs
- Designing Schema-Constrained Decisions
- Handling Abstention and Uncertainty
- Estimating Decision Confidence
- Evaluating Probabilistic Calibration
Sources
- RLCD (Reinforcement Learning for Calibrated Decisions)
- RLCD explained: Reinforcement Learning for Calibrated Decisions
- TypeSafe AI「Jev」モデルの訓練ロジックとRLCDアルゴリズム
- Jev: The Language Model That Won't Talk - Anthony Maio
- Term: Reinforcement Learning for Calibrated Decisions (RLCD)
- GitHub - Heman10x-NGU/Verdict-open-jev: Non-autoregressive
- What Is RLCD? The Secret Behind Jev | Di Zhang
Add this skill to your Hamster workspace to version it, share it with your team, and let AI agents use it automatically.
Other Skills in This Method
Steps for calibrating AI confidence to real world accuracy
Line up a decision model's stated confidence with the success rates you actually observe, per task type, using logged predictions and outcomes.
Guide to schema constrained decision outputs for AI models
Define a bounded answer space and typed schema so a model returns decisions and probabilities software can validate and act on.
Confidence Estimation for AI Decisions, Step by Step
Attach a probability to every structured decision and check it against real outcomes, so software knows when to act and when to escalate.
Steps for evaluating probabilistic calibration of AI models
Check whether a model's stated probabilities match real outcome rates with probability bins, reliability plots and expected calibration error.
AI Abstention and Uncertainty Handling for Decisions
Set per-task confidence thresholds from logged outcomes so software acts on reliable decisions and escalates uncertain ones to people.
Mastering machine readable AI decision outputs integration
Treat a decision model's typed output and confidence value as a versioned software contract that code validates, routes and acts on.
Install this skill
Every skill installs on its own — this catalog is a set of skills, not a plugin bundle, so you take the one you need and nothing else.
Claude Code
.claude/skills/designing-outcome-based-reward-signalsnpx skills add gethamster/skills --skill designing-outcome-based-reward-signals --agent claude-code --yesCursor
.agents/skills/designing-outcome-based-reward-signalsnpx skills add gethamster/skills --skill designing-outcome-based-reward-signals --agent cursor --yesCodex
.agents/skills/designing-outcome-based-reward-signalsnpx skills add gethamster/skills --skill designing-outcome-based-reward-signals --agent codex --yesAntigravity
.agents/skills/designing-outcome-based-reward-signalsnpx skills add gethamster/skills --skill designing-outcome-based-reward-signals --agent antigravity --yesOr browse the skills and pick interactively:
npx skills add gethamster/skillsSource: gethamster/skills on GitHub, MIT licensed.