InfiniBand switches and routers
BOTTLENECKNVIDIA controls the full InfiniBand stack post-Mellanox and ships nearly every switch, making a qualified second source hard to stand up.
InfiniBand and Ethernet fabric switches for AI clusters with native RDMA support. These determine scale-out cluster size and training job scheduling efficiency. NVIDIA holds near-monopoly post-Mellanox, controlling NIC, switch, cable, and management software.
High-performance lossless fabric switches for AI scale-out cluster networking; RDMA-native; NVIDIA controls the full stack post-Mellanox acquisition.
Why the concentration exists
InfiniBand switches use a switched fabric architecture that enables point-to-point connections with minimal latency. The switches employ cut-through forwarding mode, which fetches only the header information of a data packet and immediately initiates forwarding once the destination port is determined. This hardware-based approach avoids putting the CPU under pressure to make forwarding decisions, which increases performance and throughput compared to software-dependent routing.[4][6]
The architecture supports SHARP (Scalable Hierarchical Aggregation and Reduction Protocol), which offloads collective operations like reductions to the switch hardware itself. These collective operations are common in MPI and AI training workloads, so offloading them reduces the computational burden on processors. NVIDIA Quantum InfiniBand switch systems include self-healing network capabilities that enable network recovery 5,000 times faster than software-based solutions.[1][9]
InfiniBand bandwidth has evolved through defined generations: 10 Gb/s in 2002 (SDR), 40 Gb/s in 2008 (QDR), 56 Gb/s in 2011 (FDR), 100 Gb/s in 2015 (EDR), 200 Gb/s in 2018 (HDR), and 400 Gb/s in 2021 (NDR). The architecture supports up to 48,000 nodes in a single subnet and unlimited scaling across subnets using routers. InfiniBand has historically offered an order of magnitude better latency than Ethernet switches, though the gap has narrowed to about two times with the latest high-performance Ethernet switches from vendors such as Cisco and Juniper Networks.[7][18]
What the evidence shows
Broadcom's G200 switch ASIC launched June 2023 to take on InfiniBand and started gaining traction in scale-out AI networks by May 2025.
nextplatform.comInfiniBand's SHARP offloads collective operations like reductions common in MPI and AI training to the switch hardware itself.
network-switch.comWho supplies it
North America commands a market share of around 35% in the InfiniBand market. The Asia Pacific region holds about 15% market share but is expanding at the highest rate among all regions. InfiniBand switch sales in AI back-end networks surged in the second quarter of 2025, driven by strong demand for 800 Gbps InfiniBand switches from NVIDIA's Blackwell Ultra platform.[15][16]
Who controls it
No named supplier is publicly confirmed for this node yet.
What it depends on, and what depends on it
NVIDIA Quantum InfiniBand fixed-configuration switches provide up to 64 ports of 400Gb/s non-blocking bandwidth in a 1U form factor. The modular switch configurations scale up to 2,048 ports of 400Gb/s non-blocking bandwidth in a single enclosure. NVIDIA Quantum HDR switches offer up to 40 ports with speeds of 200 Gbps, while Quantum-2 NDR switches offer up to 64 ports with speeds of 400 Gbps.[5][8]
A fully loaded Quantum-2 air-cooled switch typically draws 800 to 1,200 watts, while the Quantum-X800 can draw 1,200 to 1,800 watts depending on configuration. A single 42U rack with four Quantum-2 switches and fully populated optics can add over 5,000 watts of heat load. Edge switches based on the Mellanox InfiniBand architecture carry an aggregated bi-directional throughput of 51.2Tb/s with a capacity of more than 66.5 billion packets per second.[12][13]
For a 400-node cluster, InfiniBand requires only 15 NVIDIA Quantum 8000 series switches and 400 cables, while competing Omni-Path technology requires 24 switches and 876 cables for 384 nodes. The HDR CS8500 switch provides a maximum of 800 HDR 200Gb/s ports, and each 200 GB port can be split into 2X100G to support 1,600 HDR100 100Gb/s ports. The NVIDIA SB7880 InfiniBand router enables isolation and connectivity between up to six different InfiniBand subnets with 36 100Gb/s ports.[6][8]
Where it sits in the stack
Takes in: Switch ASICs (Quantum-3/X), transceivers, management processors, chassis
Sends on: Lossless AI training fabric for scale-out cluster communication
What would break it
Ethernet has emerged as a competitive threat to InfiniBand in AI back-end network deployments. Dell'Oro Group reports that Ethernet now leads AI back-end network deployments in 2025, driven by cost advantages, multi-vendor ecosystems, and operational familiarity. For a 512-GPU training cluster, the 3-year total cost of ownership for InfiniBand NDR totals $4.61 million versus Ethernet 400G/800G RoCE at $2.37 million, a gap of $2.24 million.[17]
Meta's documented conclusion on Llama 3 infrastructure was that both RoCE and InfiniBand provide equivalent performance when properly tuned for AI training. This finding undermines the performance justification for InfiniBand's hardware premium. InfiniBand's all-to-all optimization provides no benefit over standard high-speed Ethernet for inference clusters, which means over-specifying fabric for inference is a common and expensive mistake in AI infrastructure planning.[10][17]
Broadcom's G200 switch ASIC launched in June 2023 to take on InfiniBand in scale-out networks for AI and HPC clusters. The G200 started gaining traction in May 2025 as a more scalable and cheaper alternative to InfiniBand. NVIDIA continues selling substantial volumes of InfiniBand scale-out networks and has generated significant revenue from NVSwitch interconnects for back-end scale-up networks that link GPU memories.[3]
What to watch
NVIDIA's Quantum-X Photonics InfiniBand switches reach production in early 2026. Spectrum-X Photonics Ethernet switches follow in the second half of 2026. These photonics-integrated products represent NVIDIA's response to the thermal management and power consumption challenges that co-packaged optics technology faces.[11][2]
The InfiniBand Architecture Specifications Volume 1 and Volume 2, Release 2.0, published on July 31, 2025, introduce support for XDR speeds up to 200Gb/s per lane. This enables total link bandwidths of 800Gb/s using QSFP connectors and 1.6Tb/s using QSFP-DD and OSFP connectors. At IBTA Plugfest 42 in April 2025, over 400 devices were registered with 14 servers built including four new PCIe Gen5 servers.[14]
Related nodes
Sources
- network-switch.com · 2025-09-21T17:00:09
- fibermall.com · 2025-12-01T06:54:58
- nextplatform.com · 2026-02-19T20:00:11
- nebius.com · 2024-06-18T00:00:00
- weka.io
- fibermall.com
- github.com
- advancedhpc.com
- nvidia.com
- rack2cloud.com
- datagravity.dev · 2026
- fibermall.com
- nvidianews.nvidia.com
- infinibandta.org · 2025-07-31
- delloro.com · 2Q 2025
- virtuemarketresearch.com
- vitextech.com · 2025
- techtarget.com
Full scorecard, owner shares, supply edges and the full tracked roster are in the desk letter.
GET THE BRIEFING