Memory stack

Storage is a pipeline, not a capacity label

An SSD does useful work for AI when it supplies data at the time and granularity the pipeline needs. Training may read immutable shards, decode and augment them on CPUs, stage batches in host DRAM, then transfer them to accelerators. Inference may load weights at startup, retrieve documents or vectors, write logs, and checkpoint state. Capacity matters, but a “fast SSD” does not by itself describe latency, concurrent behavior, host CPU cost, network path, or how much of the workload is already cached.

  1. 1NAND pages
  2. 2SSD controller/FTL
  3. 3NVMe/PCIe
  4. 4host DRAM
  5. 5decode
  6. 6GPU HBM
  1. 1read queue mapping/GC transport staging CPU/GPU execute
Consider the sequence and each role.

NAND pages are programmed and read as groups while erases operate on larger blocks. SSD firmware’s flash translation layer maps logical addresses to physical locations, manages garbage collection, reserves spare blocks, and wear-levels writes. This is why sequential, random, read-heavy, and sustained write behavior differ. A drive can expose high burst throughput from cache and a different steady state after enough writes trigger reclamation. Endurance is a workload property as well as a drive rating: small random writes and frequent rewrites can create internal write amplification.

Measure the complete path

Consider a toy training loader that must deliver 8 GB/s of decoded samples to keep one accelerator busy. Four SSDs each report 3 GB/s sequential read, suggesting 12 GB/s at the storage layer. The loader still misses its target if compressed records expand slowly, CPU workers are insufficient, PCIe is shared, a remote filesystem serializes requests, or batches wait for a straggler. Conversely, if the application needs only 2 GB/s after caching, buying for 12 GB/s does not move training time. These numbers are illustrative; record actual file format, compression, block size, worker count, queue depth, and cache state.

Solidigm’s 2025 article describes a test that compares a direct GPU-storage DMA path with a conventional SSD → CPU/RAM → GPU path, and publishes its server, drive, software, block-size, and queue-depth setup. Read its methodology before interpreting its results. It is vendor-authored testing, not a portable performance law. The company’s AI storage page also argues that different pipeline stages need different drives; treat that as proposed product positioning. NVIDIA’s architecture table shows that local HBM bandwidth is orders of magnitude higher than a single SSD stream, reinforcing why staging, locality, and overlap matter. The table is a platform specification, not an end-to-end storage benchmark.

Failure modes worth naming

First, benchmark a warm page cache and accidentally attribute DRAM speed to SSD. Second, measure only large sequential reads while production makes small random requests. Third, count device throughput but omit decompression and data augmentation. Fourth, fill a drive and then discover garbage-collection tail latency. Fifth, equate a manufacturer’s capacity or power claim with a cluster result. Finally, treat direct storage-to-GPU movement as automatically faster: alignment, registration, I/O size, software support, topology, and concurrency decide whether the path helps.

The business boundary is equally important. More training data or longer retention can increase storage requirements, but shipment volume, capacity mix, controller cost, endurance qualification, price, inventory, and customer architecture determine revenue. A product announcement cannot establish any of those outcomes by itself.

Exercise: make a repeatable feeding test

Choose a fixed shard set and run three trials: cold cache, warm cache, and a deliberately constrained number of loader workers. Collect median and tail batch wait, GPU utilization, storage read bytes, host CPU time, and bytes transferred to the device. Change one variable at a time—record size, queue depth, or worker count—and preserve the same model and batch. If a proposed SSD upgrade helps only cold start while steady-state utilization is unchanged, document that narrow result. Add a sustained-write trial if checkpointing is relevant, and label every untested failure mode as unknown.

ROOFLINE / HYPOTHETICAL INPUTS

Does memory feed the compute?

512 TFLOP/sBounded by memory. Compute ceiling: 1,000 TFLOP/s.

Upper bound = min(compute ceiling, bandwidth × arithmetic intensity). Decimal TB = 10¹² bytes. Cache effects, access patterns, communication and actual utilization are omitted; this is not a device benchmark. More capacity does not necessarily increase bandwidth.

SOURCES

01
Solidigm: accelerating AI with high-performance storage ↗www.solidigm.com · 2025-09-02
02
Solidigm: AI storage solutions ↗www.solidigm.com · unknown
03
NVIDIA HGX AI Factory components ↗docs.nvidia.com · unknown

YOUR NOTES