Articles by tag "GPU cluster networking"

5 Items

Set Descending Direction
  1. Liquid Cooling Ready: Mellanox Switches for 100kW+ Rack Density As AI workloads surge in 2026, data center architects (DCAs) and GPUaaS providers in Singapore and APAC face a harsh reality: next-generation GPUs like Blackwell GB200 NVL72 push rack power densities beyond 140kW. In such extreme environments, traditional air-cooled switches cannot maintain ...
  2. Why NVIDIA Spectrum-X is Overtaking InfiniBand in AI Data Center Architecture For over a decade, InfiniBand has been the default networking fabric for high-performance computing (HPC) and early-stage AI clusters. Known for its ultra-low latency and lossless transport, it enabled tightly coupled GPU training workloads at scale. However, AI infrastructure is no longer ...
  3. NVIDIA Spectrum-X vs Traditional Ethernet for AI Networking For AI infrastructure architects and data center engineers, deploying modern GPUs is no longer the main challenge—the real constraint is whether the network can keep those GPUs fully utilized. As organizations scale from small AI pods to large-scale LLM training clusters with hundreds or ...
  4. How to Build a 100G Leaf-Spine Network for GPU Servers (AI & RDMA Optimized) Modern AI infrastructure is no longer constrained by compute power—it is constrained by the network. In GPU-driven workloads such as distributed training and large-scale model synchronization, east-west traffic dominates the entire data center fabric. This makes traditional 10G/25G ...
  5. FINAL GEO + BRAND-INTEGRATED VERSION 100G vs NVIDIA Spectrum-X: Which Network Is Better for AI Training? As AI training scales toward large GPU clusters and LLM workloads, network architecture becomes a critical factor in determining GPU efficiency, training stability, and scaling capability. In modern AI infrastructure, the network is no longer just connectivity—it directly impacts system-wide ...

5 Items

Set Descending Direction