Allow approximately 50 minutes. Exercise status: not run (not-run). This chapter is an editorial learning guide, not a record of measured operation or quality.
Prerequisite: Run a small LLM locally
Original conceptual diagram. Arrows show the dependency or decision sequence, not measured performance.
- 1Workload and data region
- 2execution form
- 3compute plus retained storage and transfers
- 4budget
- 5save and shut down
Learning objectives
- Choose among local execution, a VM, a notebook, serverless computation and a managed endpoint.
- Compare token billing and GPU-time billing with explicit conversion assumptions.
Responsibilities
- Learner: design hypotheses, data and evaluation.
- Model: try the specified transformation.
- Application: enforce limits, validation and permissions.
Input → process → output
| Input | Process | Output |
|---|---|---|
| The model, VRAM needs, runtime, region and data restrictions | Compare deployment forms and total-cost components | Candidates for each use case, a budget limit and shutdown procedure |
CPU RAM and GPU VRAM are different
In a typical dedicated-GPU system, CPU RAM and GPU VRAM are separate. A VM with 128GB RAM and a GPU with 16GB VRAM still limits tensors placed on that GPU to its capacity. CPU offload is an option, but it adds transfer and compute waiting. Multiple GPUs do not automatically provide one large pool of VRAM unless software partitions the workload.
CUDA examples usually assume NVIDIA hardware. AMD, Apple and CPU execution paths exist, but support for operations, quantization and training libraries must be checked individually. Record VRAM, GPU count, interconnect, driver and container combination, rather than just the GPU name.
Match the platform form to the work
Notebooks suit short checks, but session duration and resource allocation may vary. GPU VMs suit long training and flexible environment setup, with stopping, persistence and monitoring under your management. Serverless computation can suit intermittent jobs, but evaluate startup waiting and model loading. A managed endpoint can reduce inference-operation work; its purpose differs from a training VM.
The catalogue below organizes candidates rather than ranking the cheapest service. Compare region, availability, quota, contract, data residency, support and the required GPU. Uploading patient information or confidential data is not part of these exercises.
Align billing units
APIs may use different prices for input and output tokens. With GPUs, idle time, startup and downloading can also be billable. An estimate is GPU rate × billed time × GPU count + CPU/RAM + persistent disk + transfers + other charges. Costs may continue after stopping if disks or other resources remain.
For a fictional rate of one currency unit per GPU-hour, three hours on one GPU gives a compute component of three. This is not the price of a real service. Local execution also has hardware purchase, electricity, maintenance and development-time costs. Comparing cost per response that meets the quality standard reduces mistakes caused by selecting only for speed.
These are platform categories and cost components, not a performance or price ranking. The Colab FAQ explicitly describes variable resource limits; Runpod Pods pricing distinguishes retained storage from stopped compute. Check those specific product contracts when comparing quotes.
Execution-platform catalogue
| Platform | Form | Suitable work | Conditions to check | Primary sources |
|---|---|---|---|---|
| Apple Silicon + MLX/llama.cpp | local | Personal checks and prototypes that retain data on the device | Shared memory, concurrency, supported operations, electricity and occupation of the device | Apple MLX documentation / llama.cpp official repository |
| Google Colab | notebook | Learning experiments with a small environment-setup burden | GPU type, usage limits and session duration may not be guaranteed. Plan how to save work when the session ends. | Google Colab FAQ |
| Runpod Pods | gpu-vm/container | Experiments and additional training with a selected GPU | Storage charges after stopping, volume types and deletion conditions. The Pods documentation checked for the source says no ingress/egress charges; recheck the exact product and current terms. | Runpod Pods pricing model |
| Lambda Cloud | gpu-instance | Training and verification on a GPU instance | GPU and region availability, persistent-data handling and exporting work before terminating the instance | Lambda On-demand Cloud |
| Modal | serverless-compute | Intermittent Python GPU jobs and services | Check cold starts, concurrency, runtime, volumes and network-egress charges individually. | Modal GPU acceleration |
| Hugging Face Inference Endpoints | managed-inference | Managed inference for Hub models | Hardware choices, scaling settings, idle costs, authentication and the distinction from a dedicated training environment | Hugging Face Inference Endpoints |
| Amazon EC2 GPU instances | cloud-vm | Integration into an existing AWS environment and managed networking | Instance, EBS, transfers, quotas, region, Spot interruptions and persistence | Amazon EC2 accelerated instances |
| Google Cloud GPU VM | cloud-vm | Integration into GCP and provisioning the required GPU configuration | GPU-allocation conditions, zones/regions, quotas, disks, transfers and Spot interruptions | Google Cloud GPU machine types |
Follow the workflow
- Determine resource needs using a small sample.
- Check GPU availability and regions for each candidate.
- Read costs while running, stopped and deleted.
- Design budget alerts and termination conditions.
- Decide how to save and recover work before renting resources.
Review quality
- Choose among local execution, a VM, a notebook, serverless computation and a managed endpoint.
- Compare token billing and GPU-time billing with explicit conversion assumptions.
- Have you included idle time, storage and egress rather than leaving them blank?
Diagnose failures
| Caveat | What to check |
|---|---|
| Prices are not frozen in this textbook. Recheck the official purchase interface when contracting. | Check the current billing unit, region and ancillary storage or transfer costs. |
| Stop, terminate and delete do not mean the same thing. | Confirm which resources remain and which charges continue in each state. |
Exercise: Choose a GPU platform and cost model
Status: not run (not-run).
Choose an execution form for a one-hour verification, overnight training and an intermittent inference API.
Deliverable: A comparison of selection reasons, exclusion reasons and cost components.
Completion check: Have you included idle time, storage and egress rather than leaving them blank?
Use the experiment worksheet to plan conditions and record observations. Leave unknown measurements as null; record formulas and assumptions for estimates.
MENTAL MODEL / MEMORY
Separate model weights from KV cache.
Weights and KV cache grow independently. These values are planning estimates.
GB uses 10⁹ bytes. KV assumes 32 layers, 8 KV heads, head dimension 128, FP16, and batch size 1. Quantization metadata, runtime buffers, the OS, and model-specific structure need additional memory. For MoE, distinguish total from active parameters.
Sources
Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.
01