chokepoints.ai
SUBSCRIBE
10 layers580 nodes2,376 dependencies9 chokepoints112 bottlenecks6,500+ companiesnode size = companies identified

Scale-up fabric (intra-cluster GPU-to-GPU interconnect)

BOTTLENECK

NVLink and NVSwitch are the only fabric in volume production that ties hundreds of GPUs into one coherent memory space, and no open alternative ships today.

NVLink and NVSwitch fabrics tie GPUs within one training domain into shared memory space. Larger coherent clusters let bigger models train faster. NVIDIA captures full-stack margin; the UALink consortium is pursuing an open alternative.

Why the concentration exists

Scale-up fabric provides extremely high bandwidth between GPUs within a single training domain, significantly greater than what scale-out interconnects can deliver. This bandwidth handles heavy communication demands between accelerators efficiently. Poor network bandwidth or high latency causes idle GPUs and significantly increases training time and operational costs. Using high-speed, low-latency interconnects maximizes resource utilization and overall training performance.[3][6]

NVIDIA first introduced NVLink in 2016 to enable faster GPU-to-GPU communication beyond PCIe limitations. NVLink 6 delivers 3.6 TB/s of bidirectional GPU-to-GPU bandwidth per GPU, doubling the 1.8 TB/s of Blackwell. Fifth-generation NVLink, released in 2024, supports 72 GPUs all-to-all communication at 1,800 GB/s, providing 130 TB/s of aggregate bandwidth. Hopper's NVLink 4.0, launched in 2022, offers 900 GB/s bidirectional bandwidth per GPU.[13][14][21]

NVLink's proprietary nature means you cannot plug an NVIDIA GPU into an AMD Infinity cluster. The maturity of NVLink and its software support including CUDA, CUDA-aware MPI, and NCCL give it a huge first-mover advantage in AI and HPC. NVLink has been production-deployed at scale for nearly a decade with millions of Nvidia chips deployed across NVLink-capable systems. This proprietary lock-in and software ecosystem concentration creates barriers for alternatives.[2][16]

What the evidence shows

Nvidia’s NVLink is the predominant scale-up technology; 2026 is the first year credible alternatives reach market.

chipstrat.com

NVLink’s proprietary nature prevents plugging an NVIDIA GPU into an AMD Infinity cluster.

intuitionlabs.ai
RESCORED JUL 2026near-monopolyscaling8 companies

Who supplies it

NVIDIA's NVLink is the predominant scale-up technology. NVLink has been production-deployed at scale for nearly a decade. AWS is adopting NVIDIA NVLink Fusion to combine NVLink scale-up performance with Trainium4 chips in semi-custom rack infrastructure. Cadence adopted NVIDIA NVLink Fusion, enabling custom silicon scale-up for demanding workloads including model training and agentic AI inference.[1][12][13][16]

AMD offers Infinity Fabric as an alternative to NVLink for linking GPUs into a single domain. AMD's Mega Pod scale-up system is expected to pack up to 256 accelerators. Google's TPU uses ICI (inter-chip interconnect) for chip-to-chip communication within a pod. Google's ICI scale-up network supports a maximum world size of 9,216 TPUs for TPUv7 Ironwood.[5][9][10][20]

Who controls it

NVIDIA+7 more tracked
13.9% CAGRsource

What it depends on, and what depends on it

Within a node, NVLink provides 900 GB/s bandwidth between GPUs. Between nodes, InfiniBand or RoCE networks typically provide 400-800 Gb/s per node. On HGX H800, every GPU has two NVLinks to each of the four NVSwitches. On HGX B200/B300, every GPU has nine NVLinks to each of the two NVSwitches. Poor network configuration can bottleneck GPU utilization to 40-50 percent.[15][18]

GPU clustering becomes necessary when memory ceilings, long training times, or throughput limits restrict experimentation, deployment speed, or architectural flexibility in production-scale AI systems. A single GPU's memory ceiling is up to approximately 80 GB VRAM such as the NVIDIA H100 SXM, while a GPU cluster scales linearly with 8× H100s providing approximately 640 GB aggregate VRAM. A 70B parameter LLM in BF16 requires approximately 140 GB VRAM, so it does not fit on a single GPU and needs a minimum of 2× H100 80 GB GPUs.[19]

NVLink fabrics and switches are used in the NVIDIA Blackwell platform to improve computational lithography and device simulation for advanced chip manufacturing. Vera Rubin NVL72 combines Vera CPUs, Rubin GPUs, NVIDIA NVLink networking and security features into a unified platform. NVIDIA NVLink-C2C delivers 1.8 TB/s of coherent bandwidth between CPUs and GPUs, which is 7x the bandwidth of PCIe Gen6. NVLink Fusion connects XPUs to the NVIDIA AI infrastructure platform.[11][12][16]

Where it sits in the stack

Takes in: Switch ASIC wafers, copper/optical cables, backplanes, NVLink connector assemblies

Sends on: Unified GPU domain enabling model parallelism at scale-up cluster level

view in atlas

What would break it

2026 is the first year credible alternative fabrics and switches reach the market. Huawei's UB-Mesh, presented at Hot Chips 2025, aims to unify all interconnects including PCIe, NVLink, and TCP/IP into one massive mesh fabric supporting up to 10 Tbps per chip. If widely adopted, such standards could eventually supersede vendor-specific links. However, these alternatives are nascent and not yet mainstream.[1][2]

UALink is an open industry standardized interconnect purpose-built for GPU-to-GPU communication. Astera Labs will support a complete portfolio of UALink scale-up connectivity solutions as GPUs and AI accelerators integrate this interconnect option. UALink is driven by hyperscalers that aim to deploy scale-up GPU clusters based on open standards. This open alternative could reduce vendor lock-in if adoption grows.[4][7]

Nvidia's Kyber NVL144 architecture was designed to connect 144 Rubin Ultra GPUs using a copper-based NVLink 7 scale-up fabric. Its complex PCB midplane caused a delay of more than a year from 2027 to 2028. This demonstrates that pushing scale-up fabric technology to larger domains carries engineering risks. NVLink's proprietary nature prevents plugging an NVIDIA GPU into an AMD Infinity cluster, which limits customer flexibility.[2][9]

What to watch

2026 is the first year credible alternative fabrics and switches reach the market. Upscale AI's Spectrum X Ethernet scale-out systems will hit the market later in 2026 for AI data centers building diverse, multi-vendor infrastructure. Upscale AI will join the NVIDIA Partner Network to strengthen its position as a pureplay provider of NVIDIA AI-native networking infrastructure.[1][17]

Nvidia's Feynman generation, expected to start shipping in mid-to-late 2028, will be available with either copper or co-packaged optical NVLink interconnects. Nvidia's Kyber NVL144 architecture is now expected in 2028 after being delayed more than a year from 2027. Nvidia's Vera Rubin NVL576 will use a combination of copper and optical interconnects, with the first layer using copper in the rack and the second spine layer using pluggable modules.[8][9]

AMD's Mega Pod scale-up system is expected to pack up to 256 accelerators. Google's TPU 8i can provide roughly 1,024 to 1,152 accelerators within one low-latency domain. Google's TPU 8t can reach 9,600 chip packages per domain. Google supports configuration of TPUs into slice sizes ranging from 4 TPUs up to 2,048 TPUs.[9][10]

Related nodes

InfiniBand switches and routersEthernet switches and switch siliconOptical transceivers and pluggable modulesOptical components — lasers and EMLsPhotonic integrated circuits (PICs)DSPs, retimers, and gearbox chips

Sources

  1. chipstrat.com · 2025-12-20T19:05:18
  2. intuitionlabs.ai · 2026-03-01T06:15:47
  3. blog.apnic.net · 2025-06-03T00:00:00
  4. asteralabs.com · 2025-09-18T23:56:22
  5. viksnewsletter.com · 2025-09-08T06:47:30
  6. dell.com · 2025-09-24T20:29:00
  7. asteralabs.com · 2026-01-27T00:17:20
  8. theregister.com · sun 5 Apr 2026
  9. tomshardware.com
  10. newsletter.semianalysis.com
  11. nvidianews.nvidia.com
  12. blogs.nvidia.com
  13. developer.nvidia.com · 2024
  14. developer.nvidia.com
  15. docs.nvidia.com
  16. gamesbeat.com
  17. aimagazine.com · March 12, 2026
  18. together.ai
  19. hyperstack.cloud
  20. aleksagordic.com
  21. thundercompute.com · 2020

Full scorecard, owner shares, supply edges and the full tracked roster are in the desk letter.

GET THE BRIEFING