Original conceptual diagram. Arrows show the dependency or decision sequence, not measured performance.

  1. 1Public terminology and fictional text
  2. 2evidence retrieval
  3. 3preserve negation, units and time
  4. 4expert error review
Consider the sequence and each role.

Before you begin

Prerequisite: Adapt to Japanese: vocabulary, notation and tasks. Estimated study time: 50 minutes.

This chapter is an editorial guide for terminology and document-processing prototypes. It is not clinical validation. The exercise is UNRUN (not-run) and uses fictional examples or public terminology with clear reuse conditions.

What you will learn

  • Create evaluation criteria for a medical terminology RAG system.
  • Build a prototype without using patient data.

Limit the initial task to terminology and document processing

The initial uses in this chapter are searching public terminology, suggesting possible meanings of abbreviations, extracting terms from fictional documents, and organizing content faithfully to the original. This course does not demonstrate that diagnosis, treatment recommendations or patient-specific judgments are safe. A readable output is not necessarily clinically correct.

For each glossary entry, record the name, synonyms, abbreviations, source, version and usage conditions. When an abbreviation means different things in different fields, show the candidates and supporting evidence. Do not settle on one meaning if context is insufficient. Even for public material, check reuse conditions and whether it is up to date.

Keep patient information out of the first experiment

The exercises do not upload patient information. Use fictional examples or public terminology with clear reuse conditions. Before moving to real data, consult the organization's responsible staff, legal, ethical and information-management procedures, and the rules applicable in that region. Local execution can reduce some avenues of leakage; it does not prove legal compliance or safety.

Removing a name does not necessarily make data anonymous. Dates, locations, rare events and combinations of free-text details may identify someone. The US HHS de-identification guidance is a reference for thinking about re-identification risk, but it does not replace the legal requirements of Japan or other regions. Information can also remain in models, adapters, caches and logs, so design the retention scope and access controls.

The HHS guidance describes Expert Determination and Safe Harbor under US HIPAA and residual identification risk; it does not establish Japanese compliance. The WHO guidance on large multi-modal models in health is a governance reference for the use boundary, not clinical evidence that this prototype works.

Failures that overall accuracy can miss

Score separately whether the system preserves these Japanese distinctions: 所見あり versus 所見なし (a finding being present versus absent), 疑い versus 確定 (suspected versus confirmed), 既往 versus 現在 (historical versus current), and 患者本人 versus 家族 (the patient versus a family member). Numbers, decimal points, orders of magnitude, units, times, left versus right, and the scope of negation can matter more than fluent prose.

For example, use this exact fictional Japanese input: 検査値Aは3.0 mg/L。前回は2.0 mg/L。症状Bは認めない. Its meaning is that test value A is 3.0 mg/L, the previous value was 2.0 mg/L, and symptom B is not present. Keep the Japanese input unchanged for the evaluation. The expected output preserves the values and their time points, does not reverse the negation, and adds no disease name or severity. This is not an example of a medical reference range or a diagnosis.

Connect expert review to the intended use

A terminology expert should define evaluation criteria and review critical errors separately. Check agreement between the source and output, abstention for terms absent from the collection, notation variations, and performance when the institution or time period of the data changes. Even a good average score does not justify an unrestricted use if serious negation reversals remain.

After RAG or additional training, do not automatically connect a model's assertions to medical decisions. Provide a correction interface, links to original text, the model version and the minimum records needed for audit. Any proposed clinical use needs separate validation and approval appropriate to that use.

Roles and input → process → output

Role Responsibility
Learner Design the hypothesis, data and evaluation.
Model Attempt the specified transformation.
Application Enforce limits, validation and permissions.
Stage What it contains
Input Terminology material with clear usage conditions and fictional examples.
Process Retrieve evidence → suggest candidates or structure the text → obtain expert evaluation.
Output Terminology candidates with sources, explicit uncertainty and an error classification.

Workflow

  1. Fix the scope so no patient information is used.
  2. Register the sources and versions of the terminology material.
  3. Create evaluation examples for negation, time points, units and the subject of a statement.
  4. Have an expert define critical errors.
  5. Check that the system does not generate judgments outside the intended task.

Exercise: extract structure from fictional text

Status: UNRUN (not-run). Create an evaluation that extracts values, units, time points, negation and supporting source sentences from fictional text.

Deliverable: structured examples containing no diagnosis, plus a scoring sheet for expert review.

Completion check: record negation reversal, missing units and unsupported additions as separate error types.

Quality checklist

  • Can you create evaluation criteria for medical terminology RAG?
  • Does the prototype avoid patient data?
  • Have you recorded negation reversal, missing units and unsupported additions separately?

Pitfalls and failure diagnosis

Do not use a medical benchmark score as a guarantee of diagnostic ability. Do not begin sharing or training merely because data is described as anonymized. Do not send patient data to a public Hub, external logs or a paid API.

Caveat What to check
Treating a medical benchmark score as a guarantee of diagnostic ability. Keep the exercise limited to public terms and fictional documents; clinical use needs separate validation and approval.
Starting sharing or training based only on a claim that data was anonymized. Review permission, provenance and the data-sharing or training scope before proceeding.
Sending patient data to a public Hub, external logs or a paid API. Use public or fictional inputs for this exercise and check what outputs and logs retain.

Use the experiment worksheet to plan or record this experiment.

MENTAL MODEL / MEMORY

Separate model weights from KV cache.

Weights and KV cache grow independently. These values are planning estimates.

4.5 GBWeights 4.0 GB + KV 0.5 GB

GB uses 10⁹ bytes. KV assumes 32 layers, 8 KV heads, head dimension 128, FP16, and batch size 1. Quantization metadata, runtime buffers, the OS, and model-specific structure need additional memory. For MoE, distinguish total from active parameters.

Sources

Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.

01
HHS de-identification guidance (US reference) ↗www.hhs.govPublished: Unknown · Accessed: 2026-10-03
02
WHO ethics and governance of large multimodal models in health ↗www.who.intPublished: Unknown · Accessed: 2026-10-03
03
RAG original paper ↗arxiv.orgPublished: Unknown · Accessed: 2026-10-03
Saved in this browser only.