← Back to list
Infrastructure & Cloud
#AI데이터센터#저지연#InfiniBand#RDMA#DCI#분산학습#134회
Last updated · 2026-09-29

Data Center Technologies for Large-Scale AI Services

1. Overview

A. Definition

An AI-specialized data center that, for training and inference of hyperscale AI (LLMs), connects thousands to tens of thousands of GPUs/accelerators over an ultra-low-latency, high-bandwidth network and is equipped with high-density power and cooling — often also called an AI Factory.

Whereas a traditional data center (IDC) is designed to host many independent workloads such as web, DB, and virtual machines, an AI data center is designed so that thousands of accelerators cooperatively perform a single enormous job (model training) — a fundamental difference. This difference governs every design decision about network, power, and cooling. In other words, an AI data center is not "a place where many servers are gathered" but "a place that makes tens of thousands of accelerators behave like a single giant computer," and for this reason it should be understood as a system closer to a supercomputer or AI Factory.

B. Background and Necessity

As the parameters and training data of LLMs have exploded, not only a single GPU but even a single server can no longer hold a model. A model with tens of billions to trillions of parameters demands hundreds of GB to several TB of memory, while a single latest GPU has only tens to hundreds of GB. Thus thousands of GPUs must divide the training (distributed training) of one model, and since the GPUs exchange gradients (several GB to tens of GB) at every training step, the inter-GPU communication speed determines the overall training speed.

That is, the core challenge of an AI data center is not raw compute power (FLOPS) but "communication bottleneck and power/heat." No matter how fast the GPUs are, if communication is slow the GPUs idle waiting for data, and the effective utilization (MFU, Model FLOPs Utilization) then plummets. In actual large-scale training, raising MFU to the 30–50% range is itself an engineering goal — that is how large the communication/synchronization loss is. Moreover, as a single latest GPU consumes over 700W, per-rack power draw exceeds tens to over 100kW, so an ordinary IDC (5–10kW per rack) cannot handle the power and heat. For these reasons, a dedicated infrastructure that designs compute, network, storage, and power/cooling for AI training from the ground up became separately required.

2. Overall Structure and Core Requirements

An AI data center must balance the four elements of compute, network, storage, and power/cooling; if any one is a bottleneck, the expensive GPUs idle. An overview of the whole structure is as follows.

flowchart TB
  subgraph FABRIC["Compute fabric"]
    G1["GPU node #1 (8-GPU/NVLink)"]
    G2["GPU node #2"]
    G3["GPU node #N"]
    SW["Leaf-spine switch (InfiniBand/RoCE)"]
    G1 --- SW
    G2 --- SW
    G3 --- SW
  end
  subgraph STORE["Storage"]
    PFS["Parallel file system (Lustre/GPFS)"]
  end
  subgraph FAC["Facilities"]
    PWR["High-density power (tens of kW per rack)"]
    COOL["Immersion/liquid cooling"]
  end
  SW --- PFS
  FABRIC --- PWR
  FABRIC --- COOL

The network is especially decisive, because in distributed training, the longer the GPUs wait on communication, the more the utilization of the compute units plummets. The network is in turn divided into two layers — intra-node (NVLink) and inter-node (InfiniBand/RoCE) — and both layers must be lossless and ultra-low-latency for the whole cluster to behave as one machine.

Requirement Content
Compute GPU/NPU/TPU clusters, high-density integration
Network Ultra-low-latency, lossless — communication determines training performance
Storage Large-capacity, high-speed parallel file system (checkpoints)
Power/cooling Tens of kW per rack, immersion/liquid cooling

Why storage matters must also be noted. In large-scale training, to enable fault recovery and reproducibility, the full model state is periodically saved (checkpointed), and a single checkpoint of a hyperscale model reaches hundreds of GB to several TB. For thousands of GPUs to save and restore this simultaneously requires aggregate bandwidth of hundreds of GB to several TB per second, so a parallel file system such as Lustre or GPFS (IBM Storage Scale) is essential. If a bottleneck arises here, training halts at every save interval and GPU utilization drops.

Comparison of Traditional IDC and AI Data Center

Summarizing how the two types of data center differ reveals why the design requirements of an AI data center are special. The root of the difference lies in the coupling of the workloads. The workloads of a traditional IDC are mutually independent, so even if one server slows down the impact on other jobs is small; but AI training synchronizes thousands of GPUs cooperatively, so the slowest one delays the whole. This difference in coupling separates the entire strategy for network, power, and operations.

Category Traditional IDC AI data center
Workload Many independent (web/DB/VM) A single giant job (distributed training)
Coupling Loose Very tight (synchronous communication)
Rack power density 5–10kW Tens to 100kW+
Network Ordinary Ethernet (loss tolerated) Lossless, ultra-low-latency (IB/RoCE)
Cooling Mostly air Immersion/liquid (DLC)
Performance metric Availability/throughput GPU utilization (MFU)/training time

The practical implication this table suggests is that approaching an AI data center as a "GPU-add-on edition" of an existing IDC is bound to fail. Power, cooling, network, and storage must be redesigned from the ground up to align with the training workload, and this is closer to supercomputer design methodology.

3. Low-Latency and Scaling Technologies (A)

Communication-layer optimization

Inter-GPU communication is optimized at two layers. Within a node, around eight GPUs are directly connected by a dedicated link far faster than PCIe (tens of GB/s) — NVLink at the hundreds of GB/s class — and between nodes, an ultra-low-latency network that bypasses the CPU is used. When data passes through CPU memory, latency grows from copying and context switching, so the core principle is RDMA (Remote Direct Memory Access) that directly accesses a remote node's memory without CPU involvement. Implementing RDMA over Ethernet is RoCE (RDMA over Converged Ethernet), and implementing it with a dedicated fabric is InfiniBand.

Technology Principle/description
RDMA (RoCE) Low latency via direct memory transfer without CPU involvement
InfiniBand Lossless, ultra-low-latency interconnect (HPC standard)
NVLink/NVSwitch Ultra-high-bandwidth direct connection between GPUs within a node
GPUDirect GPU communicates directly with network/storage
Collective communication (NCCL) Optimizes distributed-training communication patterns such as All-Reduce

Here it is necessary to understand why "lossless" matters. Ordinary Ethernet drops packets under congestion and relies on retransmission, but in the collective communication of distributed training, if even a single GPU is late, the whole synchronization is delayed (the straggler problem). Therefore a lossless fabric that prevents packet loss itself through flow control (PFC) and congestion control is needed. The very reason InfiniBand became the standard in HPC is this lossless, ultra-low-latency characteristic, and recently there is a strong trend to attain similar performance on an Ethernet basis by layering sophisticated congestion control atop RoCEv2.

The choice between InfiniBand and RoCE is a trade-off frequently encountered in practice. InfiniBand guarantees lossless, ultra-low latency at the hardware level with dedicated switches and adapters, but has high vendor lock-in and cost. In contrast, RoCE leverages general-purpose Ethernet equipment, giving advantages in cost and operational familiarity, but unless congestion control such as PFC and ECN is carefully tuned for losslessness, performance wavers at scale. That is, if "top performance and proven stability" is the priority, InfiniBand is reasonable; if "cost and leveraging existing Ethernet operations assets" is the priority, RoCE is reasonable — and recently in very large clusters the Ethernet camp has strengthened dedicated congestion control to narrow the gap between the two camps.

Also, the collective communication library NCCL optimizes communication patterns such as All-Reduce and All-Gather to the topology (ring/tree), and GPUDirect RDMA lets the network card directly access GPU memory, eliminating the detour through the CPU and system memory. This software layer is the final link that converts hardware bandwidth into actual training performance.

Distributed-training parallelization strategy

Distributed training combines three parallelization schemes (3D parallelism) depending on how the model and data are split. Each solves a different problem and has a different communication pattern and volume. Below is a detailed architecture diagram showing how 3D parallelism maps onto the physical topology.

flowchart TB
  subgraph DP["Data parallel (inter-node, periodic All-Reduce)"]
    subgraph PP1["Pipeline group A"]
      direction LR
      N1["Node1 (tensor parallel: 8 GPUs/NVLink)"] -->|"pass activations"| N2["Node2 (tensor parallel)"]
    end
    subgraph PP2["Pipeline group B (replica)"]
      direction LR
      N3["Node3 (tensor parallel)"] -->|"pass activations"| N4["Node4 (tensor parallel)"]
    end
  end
  PP1 <-->|"gradient All-Reduce (InfiniBand)"| PP2
Parallelization Principle Problem solved
Data parallel Split the batch, each GPU processes then All-Reduces gradients Improve training speed
Model/tensor parallel Split layers/matrix operations across GPUs Model exceeds GPU memory
Pipeline parallel Process layers stage by stage in a pipeline Memory/efficiency of deep models

In this diagram, tensor parallelism is placed within a node bound by NVLink, pipeline parallelism between nodes (passing activations), and data parallelism between replica groups (gradient All-Reduce). Placing them on faster links in the order of higher communication intensity is the core principle of scaling.

The difference in communication characteristics of the three parallelizations governs placement decisions. Tensor parallelism splits the matrix operations of a single layer, so inter-GPU communication surges at every operation. Therefore it is always placed within the same node directly connected by NVLink. Pipeline parallelism passes activations only at stage boundaries, so its communication volume is small and it may be placed between nodes. Data parallelism All-Reduces gradients at the end of a step, so its communication is large but periodic, and it is layered on top to scale the overall size.

For example, a GPT-class model splits a single layer across 8 GPUs within a node via tensor parallelism, places layer bundles across multiple nodes via pipeline parallelism, and layers data parallelism on top to run thousands of GPUs simultaneously. The inter-node communication volume is then enormous, so without InfiniBand and NVLink the scaling efficiency collapses sharply. Combining this with memory optimizations such as ZeRO, which distributes optimizer state, further eases the memory limit.

4. DCI (Data Center Interconnect) Technology (B)

Technology that connects geographically distributed data centers via ultra-high-speed, low-latency optical transmission to realize capacity expansion, disaster recovery, and load distribution.

A single data center has physical limits on power, space, and cooling, so cases arise where one region cannot handle hundreds of MW of power. Then DCI, which binds multiple DCs into a single resource pool, is needed. Since as much data as possible must be sent over one strand of optical fiber, DWDM (dense wavelength-division multiplexing), which carries data on different wavelengths (colors) and transmits them simultaneously, is the foundational technology.

Technology Principle/description
DWDM Large-capacity transmission over a single fiber via wavelength-division multiplexing
OTN Optical transport network standard, large-capacity, low-latency backbone
Coherent optical transmission 400G/800G long-distance high-speed transmission via phase/amplitude modulation
Uses Data replication between DCs, disaster recovery (DR), workload distribution, cluster expansion

Coherent optical transmission matters because, unlike the older way of carrying signals by light intensity (on/off) alone, it modulates the phase, amplitude, and polarization of light together to greatly increase the amount of information that can ride on a single wavelength. This enables 400G/800G-class transmission per wavelength even over long distances, and makes large-capacity data replication/backup between geographically distant DCs realistic. However, the propagation delay coming from physical distance cannot be fully removed due to the speed-of-light limit, so segments requiring ultra-low latency such as synchronous training are still processed within a single DC, and it is realistic to use DCI mainly for disaster recovery, data replication, and loosely-coupled workload distribution.

5. Deep Dive — Latest Trends and Green Data Centers

Recent trends in the AI data center field can be organized along three axes. First, breaking through power/cooling limits. As GPU power consumption rises every generation, densities unreachable by air cooling (over 100kW per rack) have been reached, and accordingly immersion cooling, which submerges servers in coolant, and DLC (Direct Liquid Cooling), which attaches cold plates directly to chips, are becoming standard. This directly connects to improving PUE (Power Usage Effectiveness, total power ÷ IT power).

Second, easing the memory bottleneck. As models and data grow, memory bandwidth/capacity rather than compute often becomes the bottleneck, so CXL (Compute Express Link), in which heterogeneous devices coherently share memory, and PIM (Processing-in-Memory), which computes inside memory, draw attention. These are attempts to reduce the energy and latency spent on data movement, and since they are still entering the maturity stage, their actual adoption effect varies greatly by workload.

Third, large-scale integration at the national/corporate level. As AI data centers are recognized as a core of national competitiveness, the sovereign AI trend of processing one's own data and models on one's own infrastructure strengthens, and to secure power, siting near renewable energy or nuclear power is being strategically reviewed. Since specific investment scale and power volume vary greatly by operator and timing, it is appropriate to understand them as directions rather than to assert them.

Sense of actual scale and application cases

For a concrete sense, a cluster of thousands of GPUs trains one hyperscale model over weeks to months. During this, several GB of gradients flow between GPUs at each step, and the whole of training circulates several petabytes (PB) of data repeatedly. A reference architecture such as NVIDIA's DGX SuperPOD binds tens to hundreds of 8-GPU nodes with an InfiniBand leaf-spine, and standardizes a parallel file system and high-density power/cooling to integrate this scale into a single system. Public reports that MFU was maintained around 40% in hyperscale language-model training show that infrastructure design that suppresses communication/synchronization loss is itself training-cost reduction (i.e., GPU-hour reduction).

6. Considerations and Implications

  • Network bottleneck is performance: Since GPU utilization (MFU) and training speed are governed by the network, design a non-blocking topology such as a Fat-Tree (leaf-spine) so that any two nodes have equal bandwidth, and align parallelization with the physical topology — placing tensor parallelism within a node and data/pipeline parallelism between nodes.
  • Power/carbon constraints decide siting: Because rack density is high, one should aim for a green data center with PUE improvement together with renewable power (RE100) and immersion cooling, and the very feasibility of securing large-scale power acts as the top constraint in site selection.
  • Availability/fault recovery design: At the scale of thousands of GPUs, a particular node failure occurs statistically often, so frequent checkpointing, fast restart, and straggler mitigation are the crux of completing training. Storage parallelism is decisive here.
  • Preparing for next-generation technologies: Choose a scalable architecture with expansion in mind, including CXL/PIM that eases the memory bottleneck, distributed training across multiple DCs, and optical-based switching.
  • National strategy/governance: An AI data center is core infrastructure for sovereign AI, so national-strategy-level investment and regulatory issues intertwined with data sovereignty, security, and power policy must be considered together. From a professional engineer's perspective, convergent design capability spanning power, cooling, siting, and regulation beyond pure IT design is required.

References


In one line: A large-scale AI data center connects a GPU cluster with ultra-low latency via RDMA, InfiniBand, and NVLink, performs distributed training with data/tensor/pipeline parallelism, and underpins hyperscale AI with DWDM/OTN-based DCI and high-density power/immersion cooling as sovereign AI infrastructure.