NVIDIA Mellanox MQM8790-HS2F InfiniBand Switch in Action: Optimizing Low-Latency Interconnect for RDMA/HPC/AI Clusters

August 21, 2026

Latest company news about NVIDIA Mellanox MQM8790-HS2F InfiniBand Switch in Action: Optimizing Low-Latency Interconnect for RDMA/HPC/AI Clusters

NVIDIA Mellanox MQM8790-HS2F InfiniBand Switch in Action: Optimizing Low-Latency Interconnect for RDMA/HPC/AI Clusters

In the realm of large-scale AI model training and high-performance computing (HPC), network latency often becomes the critical bottleneck that limits cluster scalability and efficiency. A national-level supercomputing center recently completed a major infrastructure upgrade, transitioning from a 100Gb/s EDR InfiniBand fabric to a 200Gb/s HDR solution built around the NVIDIA Mellanox MQM8790-HS2F InfiniBand switch. This case study examines the challenges they faced, the deployment strategy, and the measurable performance gains achieved through this next-generation interconnect.

Background and Challenges: When Latency Becomes the Performance Ceiling

The supercomputing center operates a heterogeneous cluster comprising over 2,000 GPU nodes, supporting a diverse workload mix that includes weather simulation, genomic research, and large language model training. With the previous 100Gb/s EDR network, the operations team observed that collective communication operations—such as all-reduce and all-gather—were consuming an increasing share of total job runtime as node counts grew. In some distributed training jobs, network-related overhead accounted for nearly 40% of the end-to-end execution time, severely limiting GPU utilization and extending model convergence cycles.

Additional challenges included:

  • Congestion hotspots: Traffic patterns in AI training are often bursty and unpredictable, leading to head-of-line blocking and tail latency spikes that disrupted job scheduling.
  • Insufficient in-network acceleration: The legacy fabric lacked support for advanced collective offload capabilities, forcing all reduction operations to traverse the host CPU and memory hierarchy.
  • Management complexity: Troubleshooting performance issues required manual analysis of multiple monitoring tools, making it difficult to pinpoint root causes in large-scale deployments.

After evaluating multiple interconnect options, the center selected the MQM8790-HS2F as the foundation for its new fabric. According to the MQM8790-HS2F datasheet, this switch delivers 40 ports of 200Gb/s HDR non-blocking throughput, sub-130ns cut-through latency, and support for SHARP v3.0 collective offload—capabilities that directly addressed their core pain points.

Solution and Deployment: Building a Low-Latency Fabric for AI/HPC

The deployment adopted a two-tier fat-tree topology, with 24 MQM8790-HS2F 200Gb/s HDR 40-port QSFP56 switches deployed as leaf nodes and 8 units as spine switches. This configuration provided 2:1 oversubscription across the entire cluster—sufficient for their current workload while allowing headroom for future expansion.

Deployment Layer Number of Switches Ports Used Connectivity
Leaf (per rack) 24 32 ports each ToR servers + spine uplinks
Spine 8 36 ports each Full mesh to all leaf switches

Each leaf switch connects to compute nodes via NVIDIA ConnectX-6 HDR host channel adapters. The MQM8790-HS2F compatible optics and cabling ecosystem allowed the team to reuse existing QSFP56 cables, accelerating deployment and reducing procurement costs. The entire physical installation was completed within two weeks, with the control plane leveraging NVIDIA's Unified Fabric Manager (UFM) for automated topology discovery and provisioning.

The switch's support for adaptive routing and congestion control proved critical for handling the bursty traffic characteristic of AI workloads. By dynamically redistributing flows across available paths, the MQM8790-HS2F InfiniBand switch minimized hotspot formation and maintained consistent performance even during peak job scheduling periods.

Results and Gains: Measurable Improvements in Job Throughput and GPU Utilization

After three months of production operation, the supercomputing center reported significant performance gains across multiple dimensions:

  • Latency reduction: Average MPI ping-pong latency decreased from 680 ns on the previous EDR fabric to under 130 ns with the NVIDIA Mellanox MQM8790-HS2F, representing a 5x improvement in message exchange efficiency.
  • Collective operation speedup: All-reduce latency for 1024-node jobs dropped by 62%, thanks to SHARP v3.0 offload that performs reduction operations directly within the switch fabric.
  • GPU utilization: The combined effect of lower latency and higher bandwidth pushed average GPU utilization from 72% to 91%, translating to a 26% reduction in time-to-solution for large language model training runs.
  • Job scheduling efficiency: Reduced tail latency variance allowed the SLURM scheduler to pack more jobs onto the same set of nodes without performance interference, increasing overall cluster throughput by approximately 18%.

When the team later benchmarked the MQM8790-HS2F price against alternative switches from other vendors, they found that the total cost per Gb/s delivered was highly competitive—especially when factoring in the reduced management overhead and the switch's ability to support both HDR and HDR100 link speeds, protecting investments in mixed-speed environments.

From an operational perspective, the integrated UFM platform provided a single pane of glass for monitoring fabric health, traffic patterns, and link-level errors. This visibility enabled the team to proactively identify degrading cables and replace them during scheduled maintenance, avoiding unplanned downtime.

Summary and Outlook: A Blueprint for Next-Generation Low-Latency Fabrics

This case study demonstrates that the MQM8790-HS2F InfiniBand switch solution is more than a simple speed upgrade—it is a transformative component for organizations seeking to unlock the full performance potential of their AI and HPC infrastructure. By combining high port density, ultra-low latency, and in-network computing capabilities, the switch directly addresses the communication bottlenecks that have historically limited cluster scalability.

Looking ahead, the supercomputing center plans to expand its fabric to 2,400 nodes in the next fiscal year, leveraging the same MQM8790-HS2F architecture with additional spine switches. The team is also exploring the switch's support for SHARP v3.0's enhanced reduction algorithms, which promise to further accelerate emerging workloads such as graph neural networks and reinforcement learning.

For IT managers, network architects, and engineers evaluating interconnect options for their own cluster builds, the NVIDIA Mellanox MQM8790-HS2F offers a proven, field-validated path to achieving sub-microsecond latency, maximized GPU utilization, and simplified fabric management. Detailed MQM8790-HS2F specifications and reference designs are available from NVIDIA's technical documentation portal, and the switch is now widely available—search for MQM8790-HS2F for sale to locate a regional partner.