The finished form is a small pipeline, not a clever function
The most dangerous integration joins text input directly to an action function. The programming model in the official build guide keeps state, candidates, and execution in code; the judgment component resolves one narrow ambiguity. This exercise builds a non-sending workflow: select a candidate answer from published internal FAQs and display a draft only.
- 1Question
- 2deterministic top-k retrieval
- 3freeze candidate IDs and text in state
- 1State
- 2relevance Score and answerable Noul
- 3select or abstain
- 1Selection
- 2draft with cited passage
- 3person decides whether to send
- 1Log
- 2E2E evaluation
- 3improve question, candidates, and threshold one at a time
Implementation skeleton
Retrieval comes first. Split documents into chunks and use BM25 or an existing vector search to obtain top-k. Never admit documents the user cannot access. Next, ask a Score for each candidate with explicit criteria: “answers directly,” “related but insufficient,” or “irrelevant.” At the same time, ask a Noul whether the candidate set contains an answer. If Noul is negative, abstain instead of inventing an answer even when one candidate has the best Score. Finally, ordinary code selects only after checking the minimum level, the gap to alternatives, confidence, and that the cited source still exists.
abstain is a normal output that exposes a knowledge boundary, not a failure. When candidates are missing or outdated, policies conflict, or a question needs individual judgment, say that verified documentation is insufficient and hand it to the responsible team. Optimizing only for conversational fluency pressures teams to reduce this answer, but its role in preventing high-impact hallucination must be measured.
Function calling has the same structure. The official cookbook uses typed questions to select a function name and closed arguments, while ordinary typed functions perform execution. Include none among function candidates, enumerate argument candidates from APIs or databases, and authorize every execution. A selected function name is not permission.
Evaluate one layer beyond classification accuracy
Split offline evaluation into four layers. First, candidate-generation recall: was the right FAQ present in top-k? If not, the judgment component cannot recover it. Second, judgment quality: do per-candidate Scores and answerable Noul agree with human labels? Third, decision quality: was select, abstain, or review appropriate? Fourth, E2E behavior: does the displayed citation exist, stay within access rights, and let the user complete their goal? Mixing the layers misattributes a retrieval failure to the model.
Separate development, validation, and locked test data, ensuring paraphrases of one document do not leak into separate splits. After publication, use documents that first appeared in the future as a holdout. Record correct answers, incorrect answers, no-answer cases, dangerous wrong answers, review routing, latency, and cost. Compare the same retrieval with and without the model, the existing rules, and a human-only process. Read the amount and kinds of failure, not just a persuasive single example.
Debug failures in order
When an answer is poor, reconstruct the original state and candidate IDs from the execution log. Then ask whether the correct answer was a candidate. If not, investigate retrieval or access controls. If it was, read the question criteria: are boundary cases stated, and was the candidate text truncated? If the answer was reasonable but execution was bad, investigate the threshold or code composition. Observe network failures, rate limits, and schema errors separately from semantic mistakes, then return an explicit failure with its cause; do not switch backend or provider.
Do not publish personal experiment logs or customer data as teaching material. Use synthetic or permitted public data for the exercise and remove secrets, contact information, and identifiers from stored examples. Keep evaluation sets outside the confidentiality boundary of the production system.
Finishing exercise
Prepare 50 public FAQs and 80 fictional questions, deliberately making 20 unanswerable. Use top-1 retrieval text as the baseline; compare it with Score + Noul + abstain. Have people judge citation validity and whether unanswerable questions were held. Fix the threshold, evaluate the final 20 items once, and do not repeatedly inspect that final set after changes. Select the most expensive error and write which defense—candidates, question, authority, or review—should have caught it. That determines the next implementation task.
Test execution order without executing a model
An implementation’s E2E tests do not have to invoke a networked model on every run. Inject fixed simulated responses and exercise branches for no candidates, a negative Noul, tied Scores, missing authority, and a deleted citation. The conceptual example below accepts simulated output; it is not executed.
function decide(input: { candidates: string[]; answerableYes: number; best: string | null; canRead: boolean }) {
if (!Number.isFinite(input.answerableYes)) throw new Error("invalid answerable probability");
if (!input.canRead) return { kind: "deny" as const };
if (input.candidates.length === 0 || input.answerableYes < 0.7) return { kind: "abstain" as const };
if (!input.best) return { kind: "review" as const };
if (!input.candidates.includes(input.best)) throw new Error("best citation is not a candidate");
return { kind: "draft", citationId: input.best } as const;
}Unit tests for this function alone do not establish safety. Integration tests must verify that retrieval results do not leak before the access filter, a citation ID still identifies the same current text, and the draft screen cannot automatically trigger an external send. Run model-inclusive evaluation separately with input, question, model version, and retrieval date fixed; do not confuse network failure with semantic error.
A small experiment plan
Split 80 questions into train-design 40 / validation 20 / locked-test 20. Train-design does not train the model; it designs criteria and candidate count. Store each question as JSONL with answerable, correct_doc_id, acceptable_doc_ids, danger_if_wrong, and created_at. When documents change, also retain the document snapshot ID used at evaluation time so updated text does not make an old question accidentally correct. Report not only accuracy but appropriate abstention on no-answer cases, dangerous-wrong-answer rate, review rate, and p95 latency. This compares “answers often” and “stops safely” on the same page.
2024–2026 evaluation boundary: compare a decision interface with the job
SLS-style typed interfaces are not sufficient evidence: schema success and a valid citation ID do not prove the passage answers the user. SalesRLAgent (submitted 2025-03-30) concerns specialized sales prediction; it does not establish a general typed-instruction interface. The confidence-routing paper (submitted 2025-09-23) is useful history for abstention experiments, not a deployment alternate-provider policy.
Run an unexecuted offline lab with frozen candidates. First reject an unknown citation ID, non-finite answer probability, or best ID outside the candidate set. Next compare: deterministic top-1, typed select/abstain, and blinded human review. Break results out by candidate count, document age, access denial, prompt injection in a candidate, and no-answer cases. Re-run the locked set only after a named change and retain the old result; never convert an upstream outage into a second model/provider path.
The 2026-09-16 OpenJev thread contains community claims about cross-encoders, which are not proof about Jev. Its model card has a dated “Image Decisions serving update” subsection (2026-09-28), but that date does not date every checkpoint; it also discloses contamination and security weaknesses. Treat both as research leads, not Kumyu measurements. No paid call was made.
Freeze the version tuple before comparing runs
Every row needs a manifest: model and SDK version, question/criteria hash, candidate generator and index snapshot, document revision, policy version, region, input-token count, candidate count, and collection time. Without it, a score change cannot be assigned to retrieval, model behavior, or a moving document. Keep training-contamination checks separate: hash near-duplicates across design, validation, and locked sets; do not call a post-release document a clean holdout merely because its date is newer.
Test temporal holdouts, candidate-order permutations, and inputs that try to override the retrieval policy or request an unlisted action. Report p50 and p95 latency, input/output tokens, candidate count, region, explicit failure rate, and abstention/review rate alongside quality. A timeout, malformed response, access denial, or injection attempt has its own terminal state and cause. It must not silently become a different backend, provider, or answer path.
MENTAL MODEL / VERIFICATION COST
The value of a decision depends on downstream work.
Assumptions: one second for the judgment, half of the candidates retained, and equal verification time. Full parallelism needs enough compute and concurrency. Compare success rate and total cost, including wrong judgments and retries. These figures are estimates, not measurements.
SOURCES
01YOUR NOTES