Allow approximately 35 minutes. Exercise status: not run (not-run). This chapter is an editorial learning guide, not a record of measured operation or quality.
Original conceptual diagram. Arrows show the dependency or decision sequence, not measured performance.
- 1Service task
- 2rules, retrieval or generation
- 3validation in code
- 4authorized action
How to use this course
This course is for developers building services that use AI. No mathematics or machine-learning background is required. Experience running Python makes the exercises easier.
Chapters 1–5 establish a basis for decisions. Observe input and output with the small model in chapter 6. Add retrieval in chapter 8; change weights only when needed from chapter 9 onward. Read Mac inference and CUDA training as separate execution paths.
No models have been executed, software installed, models downloaded or paid APIs called to create this textbook. Code examples are unrun. Readers review their contents, licenses and capacity requirements, then manually execute only the selected steps.
Learning objectives
- Explain how AI, machine learning, deep learning and generative AI relate to each other.
- Separate prediction from generation by their inputs and outputs.
Responsibilities
- Learner: design hypotheses, data and evaluation.
- Model: try the specified transformation.
- Application: enforce limits, validation and permissions.
Input → process → output
| Input | Process | Output |
|---|---|---|
| The task, existing data and unacceptable failures | Decompose the task into rules, retrieval, classification and generation | The scope assigned to AI and the scope guaranteed by ordinary code |
AI describes a goal; machine learning is one way to build it
AI is a broad name for technologies that produce intelligent behavior. Machine learning adjusts relationships between inputs and desired outputs using data; not all AI learns. Fixed rule matching, search, statistical models and neural networks can coexist within the same service.
In an invoicing service, for example, ordinary code can validate file formats, OCR can read text, a classifier can suggest account categories, and an LLM can draft explanations. Dividing the work makes failures easier to locate and inspect than assigning every step to an LLM.
Read history through what each approach made easier
Early symbolic AI combined knowledge and rules written explicitly by people. Statistical machine learning estimated patterns from features people designed. Deep learning expanded the ability to learn the representations themselves through multiple layers of computation. This is not a history in which each new approach completely replaced the previous one. Tree-based models, for example, remain useful comparison baselines for tabular data.
The 2017 Transformer paper presented a sequence-transformation architecture centered on attention. Methods for adapting large pretrained models to specific uses subsequently spread. Remember RAG in 2020, LoRA in 2021 and QLoRA in 2023 as landmarks that distinguish referencing knowledge from outside the model from reducing the burden of additional training.
The landmarks above refer to the original Transformer paper, RAG paper, LoRA paper and QLoRA paper. Their publication years identify different mechanisms, rather than a ranking of current products.
What a service developer must decide first
“Add AI” is not a requirement. Decide what the input is, who uses the output and which failures are unacceptable. Some mistakes in ranking search candidates can be corrected. Finalizing money amounts or making medical judgments requires separating validation and human judgment from generation. Generating a convincing answer and completing a business process correctly are different evaluation dimensions.
Follow the workflow
- Choose one service you are building.
- Write down its inputs, transformations, outputs and users.
- Create a rules-only baseline and an alternative that uses AI.
- Decide where the workflow can return to a person after a failure.
Review quality
- Explain how AI, machine learning, deep learning and generative AI relate to each other.
- Separate prediction from generation by their inputs and outputs.
- Are permission to draft a response and permission to send it externally separate?
Diagnose failures
| Caveat | What to check |
|---|---|
| Do not choose an approach only because its LLM is new. | Write a rules-only baseline and an AI alternative for the same input, output and user. |
| Do not collapse factual accuracy, format, speed and safety into one word: accuracy. | Score factual support, format compliance, latency and safety separately. |
Exercise: Map AI, machine learning and LLMs
Status: not run (not-run).
Decompose customer support into retrieval, classification, answer generation and sending. Choose the steps that need an LLM.
Deliverable: A responsibility table for the four stages and the conditions that require human review.
Completion check: Are permission to draft a response and permission to send it externally separate?
Use the experiment worksheet to plan conditions and record observations. Leave unknown measurements as null; record formulas and assumptions for estimates.
Glossary
| Term | Meaning |
|---|---|
| parameter / weight | An internal model value adjusted during training. Weights account for most of the model file size. |
| token | A unit of text handled by a tokenizer. It differs from a character or word. |
| embedding | A numerical-vector representation of a token, document or other item. |
| context | The input, history and generation that one inference can refer to. |
| KV cache | Memory that reuses keys and values used in attention for past tokens. |
| quantization | Representing weights or other values with fewer bits, trading capacity against quality and support. |
| checkpoint | A saved set of weights and any required state at a stage of training. |
| adapter | A small set of additionally trained weights combined with a base model. |
| RAG | A system that retrieves external material and adds it to the generation input. |
| SFT | Supervised additional training using pairs of input and desired output. |
| LoRA | A parameter-efficient adaptation method using low-rank matrices. |
| TTFT | Time from the start of a request until the first output token becomes available. |
| GPU-hour | The computing-resource unit of one GPU used for one hour. It differs from token billing. |
| hallucination | Generating unsupported or incorrect content in a plausible form. |
| data leakage | Information that should be unknown during evaluation enters training or setting selection. |
MENTAL MODEL / MEMORY
Separate model weights from KV cache.
Weights and KV cache grow independently. These values are planning estimates.
GB uses 10⁹ bytes. KV assumes 32 layers, 8 KV heads, head dimension 128, FP16, and batch size 1. Quantization metadata, runtime buffers, the OS, and model-specific structure need additional memory. For MoE, distinguish total from active parameters.
Sources
Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.
01