First decide which devices must communicate
A job can communicate inside one accelerator server, across a rack-scale GPU domain, or between many servers. A storage or control network can have another purpose entirely. NVLink is a GPU communication fabric described by NVIDIA for scale-up domains; InfiniBand and Ethernet provide different scale-out choices and software/operations contracts. These names do not identify the same layer. The physical link can still involve copper or optics. A choice between protocol families does not directly specify which optical-module supplier captures the purchase.
- 1GPU group and local fabric
- 2host/network adapter
- 3rack switch
- 4inter-rack fabric
- 1Application collective
- 2transport/library
- 3adapter queues
- 4actual physical links
Compare software and physical boundaries separately
The NVIDIA NVLink page distinguishes generations and rack-scale domains. It also marks the displayed specifications as preliminary and subject to change. Treat its bandwidth values as vendor specifications for the stated platform, not as observed throughput for your workload. Aggregated GPU bandwidth is not the throughput of one useful collective operation. A comparison must state direction, units, GPU count, topology and the operation measured. An undated product page establishes what the vendor currently presents at access, not the date a product first shipped.
Linux's userspace-verbs documentation explains how software accesses InfiniBand hardware, including resource management, memory pinning and the distinction between slower setup and fast-path operations. That boundary matters: low-copy transport still needs registered memory, compatible drivers and correct resource ownership. RDMA is not a promise that the application avoids every copy, has no CPU overhead or achieves line rate. Kernel documentation describes an interface, not a benchmark for a particular NIC.
Work a communication-budget example
Use a hypothetical eight-device ring with a one-gigabyte payload per participating device. A simplified all-reduce traffic factor is 2×(n−1)/n, or 1.75 for n=8. At an assumed useful 50 gigabytes per second, the bandwidth-only time is approximately 1.75/50 seconds, or 35 milliseconds. This model omits latency, routing, overlap, protocol overhead and algorithm details. It is a lower-level planning calculation, not measured NVLink, InfiniBand or Ethernet performance. Convert gigabits to gigabytes before entering a port rate.
Now suppose each iteration computes for 40 milliseconds and communicates for 35 without overlap. Total time is 75 milliseconds. Doubling the useful communication rate reduces that part to 17.5 milliseconds and total to 57.5: about 1.30 times faster, not twice as fast. If compute and communication overlap, a different model is needed. The most expensive link is valuable only to the extent that it removes the application's actual waiting time.
Tradeoffs include operations and failure behavior
Compare topology, routing, congestion control, isolation, failure diagnosis, monitoring and the team's operating skills. A nominally fast fabric can perform badly when paths are oversubscribed or queues are incorrectly configured. A familiar Ethernet deployment can still need careful loss/congestion settings and a supported software stack for RDMA. Do not import configuration advice from a different protocol or generation without checking its contract. A network decision is not just a port-rate decision.
Draw two hypothetical racks with four devices each. Label local communication, cross-rack communication and storage traffic. Identify which transfers cross the same uplink and what happens if it fails. Keep protocol, connector, media, reach and replacement procedure in separate columns. Then choose a single collective benchmark and sweep payload size and participant count while keeping other settings fixed. Record median and tail time, retransmissions or error counters, effective throughput and accelerator idle time. This is a proposed experiment, not a claim that it has been run.
Translate the result into a business question
An increase in network spending can reach switching silicon, adapters, modules, lasers, cables and software support in different proportions. The topology determines port count; physical reach and serviceability determine media choices; qualification determines supplier access. Nothing in this chapter establishes AAOI's share of a fabric purchase. Combine the actual design with the issuer's disclosed product/customer and financial evidence before building that thesis.
NETWORK / HYPOTHETICAL INPUTS
Port labels and useful throughput differ.
An 800 Gb/s port multiplied by the selected useful fraction. This toy fraction combines idle time and overhead; it is not a measured link or a protocol model. Tail latency, topology, retries and collective algorithms need separate measurements. It cannot predict a supplier’s sales.
SOURCES
01YOUR NOTES