Memory stack

A hierarchy is a data-movement map

“AI needs memory” hides the decision that matters: which byte is waiting where. A register is closest to execution but tiny. Caches trade capacity for lower access latency. HBM and conventional DRAM hold the active tensors, model weights, and key/value (KV) state needed while an accelerator is executing. SSD holds checkpoints, training shards, indexes, and datasets that are not active at that instant. Networked storage and the network fabric can add another queue before a byte reaches a host. These layers are complements. Replacing a slow SSD with a faster one cannot make an attention kernel that is waiting on HBM complete faster; adding HBM cannot repair a training input pipeline starved by remote storage.

  1. 1SSD/object store
  2. 2host DRAM
  3. 3accelerator HBM
  4. 4cache/registers
  5. 5arithmetic unit
  1. 1cold checkpoint active batch active tensors immediate reuse
  1. 1measure queue time
  2. 2measure transfer
  3. 3measure bandwidth
  4. 4measure compute
Consider the sequence and each role.

The right first question is therefore not “is HBM better than NAND?” It is “at which edge does the application stop making progress?” During model loading, measure storage read time, decompression, host-to-device transfer, and allocator time separately. During inference, separate time to first token (prompt ingestion and KV-cache construction) from decode tokens per second. During training, separate dataloader stalls, collective communication, and kernel execution. A profiler trace and a fixed workload are stronger evidence than an advertised peak bandwidth number.

Work through a bounded example

Suppose a retrieval-augmented service has a 32k-token prompt, a fixed batch of four requests, and a latency budget. Record the model revision, precision, prompt tokens, output limit, device count, and retrieval payload. First run it with a local, already-warm index. Next, make only the index remote. If time to first token rises while decode speed is unchanged, the experiment has identified a storage or network boundary, not an HBM boundary. Next hold the input path constant and change only context length. If memory allocation fails or tail latency rises before decode, inspect KV-cache residency and fragmentation. Do not call either result “an AI memory shortage” without the trace.

This exercise also prevents a common financial error. An increase in accelerator deployments may increase demand for several layers, but it does not tell us which supplier captures the economics. HBM content per accelerator, NAND capacity per server, controller content per SSD, packaging yield, contract pricing, and capital spending are different variables. A higher byte count can coincide with lower realized revenue if price falls, supply ramps, or a design uses a competing component.

What issuer disclosures can and cannot establish

Micron’s fiscal Q3 2026 material describes its own products and financial results; its HBM white paper explains the company’s view of AI data-center memory. SK hynix’s 2025 HBM4 page describes its development and preparation claims. These are useful primary sources for what the issuers said on their dates. They are not independent measurements of a customer workload, market share, or a competitor’s yield. The Micron material includes product and business-unit information; it should be reconciled with the accompanying filing before using a non-GAAP or cash-flow figure in a model.

A practical worksheet has four columns: measured system boundary; component that could affect it; company evidence; disconfirming evidence. For example, “time to first token is dominated by device allocation” might make HBM capacity relevant, but it is not proof that a particular HBM supplier wins the design. The disconfirming evidence could be a lower-capacity model configuration, a different packaging choice, or a software change that reduces KV memory. Keep these as conditional branches rather than price targets.

Exercise: make an honest bottleneck statement

  1. Run a fixed prompt twice after warm-up and retain the trace.
  2. Change one variable—context length, batch, storage locality, or precision.
  3. Name the slowest boundary and one metric that would falsify the diagnosis.
  4. Write a supplier thesis only as: “If this boundary persists and this product is qualified, then revenue exposure may change.”
  5. Add an invalidation condition: lower utilization, a competing package, lower ASP, or a software reduction in active memory.

This keeps system engineering, issuer disclosure, and financial inference in their proper layers.

ROOFLINE / HYPOTHETICAL INPUTS

Does memory feed the compute?

512 TFLOP/sBounded by memory. Compute ceiling: 1,000 TFLOP/s.

Upper bound = min(compute ceiling, bandwidth × arithmetic intensity). Decimal TB = 10¹² bytes. Cache effects, access patterns, communication and actual utilization are omitted; this is not a device benchmark. More capacity does not necessarily increase bandwidth.

SOURCES

01
Micron HBM white paper ↗www.micron.com · unknown
02
Micron Q3 FY2026 quarterly results ↗investors.micron.com · 2026-06-24
03
SK hynix HBM4 development ↗news.skhynix.com · 2025-09-12

YOUR NOTES