GPU Cluster Networking
Training workers wait when they cannot exchange results quickly enough. Compare communication costs and workload placement in a modeled GPU cluster.
Technical evidence
How to read this chart. Start with the left chart: lower lines mean less time to finish a training step. A value of 1 means the same time as running inside one compute hall. Solid lines keep communicating workers grouped together; dashed lines interleave them between halls. The right chart shows how adding more participants to a communication ring can increase the delay in the modeled case.





