Allow approximately 50 minutes. Exercise status: not run (not-run). This chapter is an editorial learning guide, not a record of measured operation or quality.

Prerequisite: Run a small LLM locally

Original conceptual diagram. Arrows show the dependency or decision sequence, not measured performance.

  1. 1Workload and data region
  2. 2execution form
  3. 3compute plus retained storage and transfers
  4. 4budget
  5. 5save and shut down
Consider the sequence and each role.

Learning objectives

  • Choose among local execution, a VM, a notebook, serverless computation and a managed endpoint.
  • Compare token billing and GPU-time billing with explicit conversion assumptions.

Responsibilities

  • Learner: design hypotheses, data and evaluation.
  • Model: try the specified transformation.
  • Application: enforce limits, validation and permissions.

Input → process → output

Input Process Output
The model, VRAM needs, runtime, region and data restrictions Compare deployment forms and total-cost components Candidates for each use case, a budget limit and shutdown procedure

CPU RAM and GPU VRAM are different

In a typical dedicated-GPU system, CPU RAM and GPU VRAM are separate. A VM with 128GB RAM and a GPU with 16GB VRAM still limits tensors placed on that GPU to its capacity. CPU offload is an option, but it adds transfer and compute waiting. Multiple GPUs do not automatically provide one large pool of VRAM unless software partitions the workload.

CUDA examples usually assume NVIDIA hardware. AMD, Apple and CPU execution paths exist, but support for operations, quantization and training libraries must be checked individually. Record VRAM, GPU count, interconnect, driver and container combination, rather than just the GPU name.

Match the platform form to the work

Notebooks suit short checks, but session duration and resource allocation may vary. GPU VMs suit long training and flexible environment setup, with stopping, persistence and monitoring under your management. Serverless computation can suit intermittent jobs, but evaluate startup waiting and model loading. A managed endpoint can reduce inference-operation work; its purpose differs from a training VM.

The catalogue below organizes candidates rather than ranking the cheapest service. Compare region, availability, quota, contract, data residency, support and the required GPU. Uploading patient information or confidential data is not part of these exercises.

Align billing units

APIs may use different prices for input and output tokens. With GPUs, idle time, startup and downloading can also be billable. An estimate is GPU rate × billed time × GPU count + CPU/RAM + persistent disk + transfers + other charges. Costs may continue after stopping if disks or other resources remain.

For a fictional rate of one currency unit per GPU-hour, three hours on one GPU gives a compute component of three. This is not the price of a real service. Local execution also has hardware purchase, electricity, maintenance and development-time costs. Comparing cost per response that meets the quality standard reduces mistakes caused by selecting only for speed.

These are platform categories and cost components, not a performance or price ranking. The Colab FAQ explicitly describes variable resource limits; Runpod Pods pricing distinguishes retained storage from stopped compute. Check those specific product contracts when comparing quotes.

Execution-platform catalogue

Platform Form Suitable work Conditions to check Primary sources
Apple Silicon + MLX/llama.cpp local Personal checks and prototypes that retain data on the device Shared memory, concurrency, supported operations, electricity and occupation of the device Apple MLX documentation / llama.cpp official repository
Google Colab notebook Learning experiments with a small environment-setup burden GPU type, usage limits and session duration may not be guaranteed. Plan how to save work when the session ends. Google Colab FAQ
Runpod Pods gpu-vm/container Experiments and additional training with a selected GPU Storage charges after stopping, volume types and deletion conditions. The Pods documentation checked for the source says no ingress/egress charges; recheck the exact product and current terms. Runpod Pods pricing model
Lambda Cloud gpu-instance Training and verification on a GPU instance GPU and region availability, persistent-data handling and exporting work before terminating the instance Lambda On-demand Cloud
Modal serverless-compute Intermittent Python GPU jobs and services Check cold starts, concurrency, runtime, volumes and network-egress charges individually. Modal GPU acceleration
Hugging Face Inference Endpoints managed-inference Managed inference for Hub models Hardware choices, scaling settings, idle costs, authentication and the distinction from a dedicated training environment Hugging Face Inference Endpoints
Amazon EC2 GPU instances cloud-vm Integration into an existing AWS environment and managed networking Instance, EBS, transfers, quotas, region, Spot interruptions and persistence Amazon EC2 accelerated instances
Google Cloud GPU VM cloud-vm Integration into GCP and provisioning the required GPU configuration GPU-allocation conditions, zones/regions, quotas, disks, transfers and Spot interruptions Google Cloud GPU machine types

Follow the workflow

  1. Determine resource needs using a small sample.
  2. Check GPU availability and regions for each candidate.
  3. Read costs while running, stopped and deleted.
  4. Design budget alerts and termination conditions.
  5. Decide how to save and recover work before renting resources.

Review quality

  • Choose among local execution, a VM, a notebook, serverless computation and a managed endpoint.
  • Compare token billing and GPU-time billing with explicit conversion assumptions.
  • Have you included idle time, storage and egress rather than leaving them blank?

Diagnose failures

Caveat What to check
Prices are not frozen in this textbook. Recheck the official purchase interface when contracting. Check the current billing unit, region and ancillary storage or transfer costs.
Stop, terminate and delete do not mean the same thing. Confirm which resources remain and which charges continue in each state.

Exercise: Choose a GPU platform and cost model

Status: not run (not-run).

Choose an execution form for a one-hour verification, overnight training and an intermittent inference API.

Deliverable: A comparison of selection reasons, exclusion reasons and cost components.

Completion check: Have you included idle time, storage and egress rather than leaving them blank?

Use the experiment worksheet to plan conditions and record observations. Leave unknown measurements as null; record formulas and assumptions for estimates.

MENTAL MODEL / MEMORY

Separate model weights from KV cache.

Weights and KV cache grow independently. These values are planning estimates.

4.5 GBWeights 4.0 GB + KV 0.5 GB

GB uses 10⁹ bytes. KV assumes 32 layers, 8 KV heads, head dimension 128, FP16, and batch size 1. Quantization metadata, runtime buffers, the OS, and model-specific structure need additional memory. For MoE, distinguish total from active parameters.

Sources

Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.

01
Runpod Pods pricing model ↗docs.runpod.ioPublished: Unknown · Accessed: 2026-10-03
02
Modal GPU acceleration ↗modal.comPublished: Unknown · Accessed: 2026-10-03
03
Lambda on-demand cloud ↗docs.lambda.aiPublished: Unknown · Accessed: 2026-10-03
04
Google Colab FAQ ↗research.google.comPublished: Unknown · Accessed: 2026-10-03
05
Hugging Face Inference Endpoints ↗huggingface.coPublished: Unknown · Accessed: 2026-10-03
06
Amazon EC2 accelerated instances ↗docs.aws.amazon.comPublished: Unknown · Accessed: 2026-10-03
07
Google Cloud GPU machine types ↗docs.cloud.google.comPublished: Unknown · Accessed: 2026-10-03
08
Apple MLX documentationPublished: Unknown · Accessed: 2026-10-03
09
llama.cpp official repositoryPublished: Unknown · Accessed: 2026-10-03
Saved in this browser only.