Memory stack

Bandwidth is a rate, not a performance verdict

An accelerator can execute operations only if operands arrive in time. HBM is wide, nearby DRAM; it is not a synonym for fast AI. Ask how many bytes a kernel moves for every useful operation. That ratio is arithmetic intensity. A low-intensity operation that repeatedly reads a large vector can be limited by memory bandwidth. A high-intensity matrix multiply may be limited by compute. Capacity is separate: a tensor that does not fit cannot be made fast merely by adding bandwidth.

  1. 1weights + activations
  2. 2HBM reads/writes
  3. 3arithmetic units
  4. 4result
  1. 1bytes per operation
  2. 2arithmetic intensity
  3. 3bandwidth or compute ceiling
  1. 1capacity failure
  2. 2recompute/offload
  3. 3different data-movement cost
Consider the sequence and each role.

Micron specifies HBM3E with 1,024 I/O pins and more than 9.2 Gb/s per pin; multiplying then dividing bits by eight gives a little over 1.17 TB/s before protocol details, consistent with its stated more-than-1.2-TB/s placement bandwidth. Micron's product page is a vendor specification, not a workload result. NVIDIA lists 4.8 TB/s and 141 GB for H200 SXM, while its eight-GPU node is 1.1 TB total; the table shows why per-GPU and node memory must not be mixed.

A deliberately small roofline calculation

This is a toy calculation, not a benchmark. Suppose a kernel performs 240 billion operations while moving 120 GB from HBM: intensity is 2 operations/byte. On a hypothetical device with 4 TB/s sustainable HBM traffic, its bandwidth ceiling is 8 tera-operations/s. If compute capability is 40 tera-operations/s at the selected precision, memory traffic is the first ceiling. Doubling compute changes nothing; halving bytes moved through better reuse changes the ceiling to 16 TOP/s.

Now make the toy kernel 2,400 billion operations over the same 120 GB: intensity is 20 operations/byte and the bandwidth ceiling is 80 TOP/s. The 40 TOP/s compute ceiling matters. This is why peak bandwidth is not a universal speed multiplier. Real kernels add cache behavior, synchronization, layout, instruction overhead, precision rules, and contention. Sustainable bandwidth must be measured for the access pattern.

Capacity is equally concrete. If weights, activations, workspace, and KV cache total 150 GB, a 141-GB device cannot keep that configuration resident. Sharding, quantization, batching changes, checkpointing, and host offload alter both footprint and data path. Calling this a bandwidth problem loses the choice.

Claims, limits, and a useful test

NVIDIA's 2023 H200 announcement is a specification and product statement, not a promise for a specific model’s latency. The announcement compares generations, but a fair experiment preserves model, precision, context, batch policy, software, and interconnect. Micron’s numbers describe its part under stated conditions; neither source proves competitor quality, customer yield, market share, or financial outcome. A measured low-intensity workload may value local bandwidth; qualification, shipment, price, and yield are later links that can fail.

Exercise. Run one fixed step after warm-up. Record model revision, precision, batch, sequence length, device count, elapsed time, HBM traffic, and achieved compute. Change only one variable—sequence length, batch, or layout. State a falsifier: if achieved bandwidth and compute utilization are both low, investigate launch gaps or input stalls before telling a bandwidth story. Add a residency report so capacity and bandwidth remain separate.

Inspect a timeline as well as a peak counter. Short traffic bursts can coexist with idle gaps, and aggregate bytes can hide device imbalance. In a distributed job, measure collective communication separately: faster local HBM may leave the critical path unchanged when an all-reduce or host barrier dominates. The useful result is a trace-linked hypothesis and a next experiment, not an unqualified memory-bound label.

ROOFLINE / HYPOTHETICAL INPUTS

Does memory feed the compute?

512 TFLOP/sBounded by memory. Compute ceiling: 1,000 TFLOP/s.

Upper bound = min(compute ceiling, bandwidth × arithmetic intensity). Decimal TB = 10¹² bytes. Cache effects, access patterns, communication and actual utilization are omitted; this is not a device benchmark. More capacity does not necessarily increase bandwidth.

SOURCES

01
Micron HBM3E product page ↗www.micron.com · unknown
02
NVIDIA HGX AI Factory components ↗docs.nvidia.com · unknown
03
NVIDIA announces H200 ↗nvidianews.nvidia.com · 2023-11-13

YOUR NOTES