Original conceptual diagram. Arrows show the dependency or decision sequence, not measured performance.

  1. 1Readable documents and versions
  2. 2permission-filtered retrieval
  3. 3cited evidence
  4. 4generation
  5. 5citation and abstention checks
Consider the sequence and each role.

Before you begin

Prerequisite: Choose GPU platforms and costs. Estimated study time: 50 minutes.

This chapter is an editorial guide. The exercise is UNRUN (not-run); it describes a plan rather than measured results.

What you will learn

  • Distinguish retrieval failures from generation failures.
  • Design a small RAG system with access control and traceable sources.

RAG changes the input

Retrieval-augmented generation, or RAG, retrieves relevant material and includes it in an LLM's input to produce an answer. Updating an ordinary search index does not train the LLM's weights. When specifications or glossaries change frequently, the ability to update the material without retraining is an advantage.

Adding RAG does not guarantee a correct answer. The system may miss the necessary paragraph, retrieve an old version, separate a table from its units, or attach a citation that does not support the claim. Evaluate the retriever and the generator separately.

The original RAG paper combines parametric generation with non-parametric retrieval. Updating an application’s external documents is separate from the original paper’s training method. The access-control and citation checks in this chapter are application design requirements, not guarantees supplied by RAG.

From documents to an answer

Obtain the original text and attach a document ID, version, publication date and access permissions. Split it into units that preserve headings and tables as far as possible, then build an embedding or keyword search index. Retrieve candidates for the question, reorder them with a reranker if needed, and fit the evidence into the available context. Require the output to cite its sources and to abstain when the evidence is insufficient.

Embeddings help represent similarity in meaning. Keyword search can be useful for model numbers, abbreviations and numerical values; hybrid search combines both approaches. Chunk size and top-k are not magic constants. Choose them by checking whether the retriever returns the passages that contain the answer.

Include permissions and prompt injection in the design

Exclude documents the user cannot read before they become retrieval candidates. Hiding them after generation is too late: the information has already been passed to the LLM. Treat instructions found inside documents as source material, not as authorization to change application permissions or send information elsewhere.

Evaluate the retrieval rate for the correct evidence, citation accuracy, the proportion of answers supported by evidence, the reflection of document updates, and abstention on questions outside the collection. Before improving RAG, check that the correct source material exists and that the user is allowed to read it.

Roles and input → process → output

Role Responsibility
Learner Design the hypothesis, data and evaluation.
Model Attempt the specified transformation.
Application Enforce limits, validation and permissions.
Stage What it contains
Input Documents with verified usage rights and access permissions, plus questions.
Process Split → retrieve → rerank → generate with evidence.
Output Answers, sources, abstentions and retrieval logs.

Workflow

  1. Prepare ten public manuals for fictional products.
  2. Assign an ID and version to each document.
  3. Write questions with known supporting evidence, as well as questions the collection cannot answer.
  4. Score the retrieval results on their own first.
  5. Then score the correspondence between answers and citations.

Exercise: test an updated specification

Status: UNRUN (not-run). Create a specification with an old version and a new version. Test whether the system answers using only evidence from the latest version.

Deliverable: a classification table separating retrieval failures, generation failures and missing source material.

Completion check: include a test showing that the system does not follow instructions embedded in retrieved documents.

Quality checklist

  • Can you distinguish retrieval failures from generation failures?
  • Does the small RAG design include access control and sources?
  • Does the evaluation test resistance to instructions inside retrieved documents?

Pitfalls and failure diagnosis

Do not treat a search result as immediate proof of a fact. Embeddings and indexes that contain personal information also require protection.

Caveat What to check
Treating something found by search as verified evidence. Check the retrieved passage’s source, revision and support for the answer.
Failing to protect embeddings or indexes that contain personal information. Apply the source documents’ access and retention requirements to the index and embeddings too.

Use the experiment worksheet to plan or record this experiment.

MENTAL MODEL / MEMORY

Separate model weights from KV cache.

Weights and KV cache grow independently. These values are planning estimates.

4.5 GBWeights 4.0 GB + KV 0.5 GB

GB uses 10⁹ bytes. KV assumes 32 layers, 8 KV heads, head dimension 128, FP16, and batch size 1. Quantization metadata, runtime buffers, the OS, and model-specific structure need additional memory. For MoE, distinguish total from active parameters.

Sources

Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.

01
RAG original paper ↗arxiv.orgPublished: Unknown · Accessed: 2026-10-03
Saved in this browser only.