Requires all-reduce sync per transformer layer ($< 2\,\mu\text{s}$ ceiling). Strict NVLink / intra-chassis only.
Activations passed across layer boundaries with micro-batching. WAN delay masked by bubble scheduling.
Asynchronous token routing across expert clusters with speculative caching.
Distributed AI Training & Inference Parallelism Across Network Boundaries
Mathematical limits governing communication overhead across NVLink, RoCE v2 InfiniBand, and Continental WAN.
| Parallelism Paradigm | Network Boundary | Latency Ceiling | Payload Size per Step | WAN Feasibility |
|---|---|---|---|---|
| Tensor Parallelism (TP) | Intra-Node / NVLink Fabric | ≤ 2.0 μs | All-Reduce per Attention Head (~4-16 MB) | Never Feasible (100% Stall) |
| Context Parallelism (CP) | Intra-Cluster / InfiniBand | ≤ 15.0 μs | Ring-Attention KV Chunks (~2-8 MB) | Prohibitive (>90% Bubble) |
| Pipeline Parallelism (PP) | Cross-Campus / Terrestrial WAN | ≤ 40.0 ms | Boundary Layer Activations (~1 MB/microbatch) | Feasible with 1F1B Scheduling |
| Data Parallelism (FSDP / ZeRO-3) | Metro-Area Low-Latency WAN | ≤ 5.0 ms | Gradient Shards & Weight Broadcast | Viable within Metro Rings |
| Federated / Agent Swarms | Global Continental WAN / Subsea | ≤ 250.0 ms | Semantic AST Diff & JSON State Graph | 100% Native Production Fit |
Mastering Global Interconnect & Physical Infrastructure
Explore the foundational engineering literature on optical dispersion, WAN pipelining, and confidential compute.