The three questions are not substitutes

The official primitive guide separates Choice, Score, and Noul because downstream code needs different meanings from each. Choice selects one defined option. Score assesses an ordered level of quality. Noul returns the probability that a yes/no proposition is true. Similar names do not make their outputs interchangeable.

Use Choice for a closed branch such as “which existing handler should receive this request?” Its response contains the selected value, probabilities for every candidate, and confidence, making competition between alternatives visible. Use Score when a retrieval candidate needs a comparable level such as none / weak / useful / direct for how well it answers a question. Use Noul for an independent condition such as “does this message contain a refund request?” If several tags can all be true, do not force them into an oversized Choice; ask a separate Noul for each condition.

  1. 1One closed branch
  2. 2Choice
  3. 3distribution across candidates
  1. 1Ordered quality
  2. 2Score
  3. 3distribution across levels
  1. 1Independent condition
  2. 2Noul
  3. 3probability of yes
  1. 1Distribution + business rules + execution authority
  2. 2automate, confirm, or refuse
Consider the sequence and each role.

Reading probability and confidence

A probability says how strongly an answer is supported for this question and state. For Choice and Score, confidence summarizes how concentrated the distribution is. It is not the probability that the workflow is correct, that the user gave permission, or that an action is safe in the real world. As the confidence documentation explains, two candidates may both be reasonable and split the distribution, yielding low confidence. That is acceptable when choosing a harmless display treatment. Conversely, a sharply concentrated answer cannot justify ordering from a stale price list passed in as state.

For Noul, a value near 0.5 means the yes and no answers are competing; it does not mean “moderately urgent.” Before sorting on is_urgent = 0.51, define urgency and decide whether missed urgency or excess alerts carries the larger loss. Using probabilities in decisions requires calibration and loss analysis on validation data.

Design safe gates

First divide actions into three levels. A is a reversible, low-impact suggestion such as retrieval ranking or a draft tag. B is an operation a person reviews before it is sent. C is high impact: payment, external publication, deletion, or physical execution. For A, representative examples and confidence can inform the design. For B, route to review when confidence is low, state is incomplete, or the target is new. For C, always require explicit authority, identity verification, deterministic policy checks, reconfirmation, and an audit record, regardless of the model value.

For example, a Choice that finds refund in a ticket must not call a refund API. Code must verify (a) the actor’s authority, (b) that the order ID exists, (c) the amount calculation, (d) that it was not already refunded, and (e) a human confirmation token before executing. Jev helps find the ticket; it does not delegate authority.

An experiment for a threshold

Do not declare “0.8 is safe.” Prepare a holdout set close to actual use and split it by time. Attach an expected human label and impact level to every case; record the model answer, all candidate probabilities, confidence, and latency. For thresholds from 0.5 through 0.95, tabulate coverage (the proportion automated), incorrect-action rate, review volume, and severe-error count. In a costly workflow, severe errors matter more than average accuracy.

You also need controls: the current rules, the existing non-random human flow, and a model-free retrieval ranking should all run on the same inputs. Numbers only observed after introduction cannot separate model effects from seasonality or a change in operators. Change one of question wording, criteria, candidate generation, or model version at a time. Never evaluate only high-confidence cases: that hides the real review burden. Report both the whole set and the automated subset.

Exercise: build a review queue

Add a needs_review Choice candidate to the memo classifier from the previous chapter. In a read-only application, keep an item in review rather than fixing its tag when confidence is low or needs_review has the greatest probability. Then build two reports: one for maximum candidate probability and one for confidence, without mixing them. Have a human label 20 holdout items and retain the reason each entered review. Finally, trace one case that was wrong despite high confidence. Explain whether the cause was state, candidate coverage, criteria, or the real-world label. That counterexample is the start of safe design.

An unexecuted gate example

The following short TypeScript example does not call an API. It shows how to handle an already acquired Choice response. Confirm the numeric range and field names for confidence in the adopted SDK’s official response types. The 0.85 threshold is a placeholder to replace after evaluation, not a universal value.

type Decision = { value: string; confidence: number; probabilities: Record<string, number> };
function route(d: Decision, permissions: { canAutoTag: boolean }) {
  if (!Number.isFinite(d.confidence) || d.confidence < 0 || d.confidence > 1) throw new Error("invalid confidence");
  const labels = ["research", "implementation", "idea", "needs_review"] as const;
  if (!(labels as readonly string[]).includes(d.value)) throw new Error("invalid Choice label");
  const unclear = d.value === "needs_review" || d.confidence < 0.85;
  if (!permissions.canAutoTag || unclear) return { kind: "review", reason: "authority-or-uncertainty" };
  return { kind: "display_tag", value: d.value }; // no external side effect
}

canAutoTag must not be produced by the model. Derive it deterministically from the signed-in user, the target record, and organizational policy. For an external send, return a confirmation screen instead of display_tag, and reload the target at confirmation time to ensure it has not changed. Repeated high-confidence answers never justify skipping an authorization check.

Experiment table and failure diagnosis

At minimum, an evaluation table should contain case_id, split, predicted, p_max, confidence, action, human_label, harm_class, latency_ms, question_version. action records the decision made by the workflow—display, automatic tag, review, or refusal—rather than only the model answer. It distinguishes “the classification was right but authority stopped it” from “the classification was wrong but review prevented harm.” Define severity before evaluation begins; do not rewrite it after seeing a convenient outcome.

Common failures include missing rare dangerous cases despite high accuracy on a balanced test, reviewers being influenced by the displayed model answer, and repeatedly fitting a threshold to the test set. Counter them by reporting dangerous cases as their own stratum, collecting blinded human decisions, and using the final set only once for a decision. Check whether probabilities are calibrated with a reliability diagram and observed accuracy per bin; do not over-trust bins with few samples.

Confidence is not a fourth primitive

Current System One documentation describes text-only Jev with Choice and Score confidence in the 0–1 range; Noul has no separate confidence. For a Choice with n candidates, the documented concentration calculation is (pmax − 1/n) / (1 − 1/n): pmax=.7 produces .4 with two candidates and .6 with four. This is arithmetic, not a measured reliability claim. A fixture’s .value is an internal teaching representation, not an SDK field name.

Validate every fixture before a gate: label must be in the current candidate tuple and confidence must be finite and within 0–1. Then calibrate probabilities on held-out, drifted, and adversarial cases. Brier score is the mean of (p − y)^2; report bin counts with each calibration point. For a toy loss where a wrong automatic tag costs 9 and review costs 1, automate only when a validated event probability exceeds 8/9. Do not substitute Choice confidence for that probability.

The 2026-09-22 BetterOffline discussion questions harness interpretation. It is community commentary: a deterministic process exit code may need no model, and no thread proves a model benefit. The dated September source boundary is the System One early-access post; vendor evals assume their harness is correct and are not human truth.

Calibrate against change, not one convenient split

Make temporal holdouts from newer documents and a drift set from changed vocabulary, policies, and candidate counts. For each probability bin, publish the count, mean predicted probability, observed outcome rate, and interval; a nearly perfect-looking bin of three cases is not calibration. Compute Brier score on the same locked cases, then inspect it by harm class and by candidate count. A confidence threshold has no universal meaning because the base rate, candidate set, policy, and annotator disagreement all change.

Loss is an explicit toy assumption, not a hidden policy. If a wrong automatic tag costs 9 and review costs 1, 8/9 follows only when a validated event probability describes that exact loss model. Review is also fallible: sample reviewer reversals, measure queue delay, and retain a separate escalation path. A probability may help decide a reversible display tag; it never replaces a signed authority check for a consequential action.

MENTAL MODEL / VERIFICATION COST

The value of a decision depends on downstream work.

Verify all sequentially
12 s
Verify all in parallel
4 s
Judge, then verify half
9 s

Assumptions: one second for the judgment, half of the candidates retained, and equal verification time. Full parallelism needs enough compute and concurrency. Compare success rate and total cost, including wrong judgments and retries. These figures are estimates, not measurements.

SOURCES

01
TypeSafe primitives ↗docs.typesafe.ai · unknown
02
Confidence ↗docs.typesafe.ai · unknown
03
Confidence-gated routing ↗docs.typesafe.ai · unknown
04
TypeSafe AI: Introducing System One Models and Jev ↗typesafe.ai · 2026-09-15
05
OpenAI: Introducing Structured Outputs ↗openai.com · 2024-08-06
06
RouteLLM ↗arxiv.org · 2024-06-26
07
SalesRLAgent ↗arxiv.org · 2025-03-30
08
Confidence routing paper ↗arxiv.org · 2025-09-23
09
TypeSafe Evals ↗evals.typesafe.ai · unknown
10
OpenJev model card ↗huggingface.co · unknown
11
Reddit r/LocalLLaMA: Jev architecture discussion ↗www.reddit.com · 2026-09-18
12
Reddit r/LocalLLaMA: OpenJev discussion ↗www.reddit.com · 2026-09-16
13
Reddit r/BetterOffline: Jev hype and harness discussion ↗www.reddit.com · 2026-09-22

YOUR NOTES