AI lab, part 5: NVLink partitions, fabric health and telemetry
Table of Contents
NVLink partitions, fabric health and telemetry, and how a partition change reaches the scheduler. Part 4 broke individual NVLinks through the BMCs. This part steps back to the whole fabric. We’ll cover what an NVLink domain and its partitions are, how to query and change them through the lab’s partition controller, how a partition change ends up in the scheduler, and what the fabric looks like in Prometheus. The examples run in Slurm mode; the Kubernetes counterpart is noted where it differs. On a GB200 NVL72 rack each GPU has 18 NVLinks, one to each of the rack’s 18 NVSwitch chips (9 switch trays × 2 chips). The lab models a single switch tray, so each GPU uses 9 links on each of its 2 chips. Together, GPUs and switches form one NVLink domain: any GPU can reach any other at NVLink speed, even across trays, without touching the network. That’s why a 72-GPU rack can train like a single big machine. A domain is usually shared, so the switch trays can cut it into partitions, isolated groups of GPUs that can only talk to each other. NVIDIA’s NMX Controller (NMX-C) manages them. Every GPU reports its partition to software as a clique ID ( Schedulers care because a job whose ranks sit in different partitions can’t use NVLink between them. The scheduler therefore needs to know the partition layout and keep jobs inside one partition. In the lab, the domain is 8 GPUs, the switch tray has two NVSwitch chips with 72 ports each, and From the tray, NVML shows the fabric registration of every GPU: From the switch side, Out of the box, all 8 GPUs are in the default partition, ID 32766: Every GPU has all 18 links active, 9 to each of the two NVSwitch chips: Remember the two NVLinks disabled through the tray BMC in Part 4? With them down, the lab’s controller rates that GPU as degraded: The same appears in metrics as Give tray 2 its own partition. A GPU belongs to at most one partition, so creating the partition straight away fails: Take the tray out of the default partition first: Clique Nobody tells Slurm about this directly. Every minute, topograph on the controller collects Each block is named after the cluster UUID and clique it was built from. In Kubernetes mode the same pipeline relabels the nodes instead ( You might expect Ranks 0–3 are in clique 1 and ranks 4–7 in clique 7. Slurm’s Kubernetes behaves differently. The lab’s JobSets set Merge back by deleting the partition, which leaves its GPUs in no partition, and adding the tray to the default partition again. The default block returns within a minute: The controller reads per-GPU NVLink byte counters that the fake CUDA stack keeps for NCCL collectives and GPU-to-GPU copies, and exposes them alongside the link and partition state: With the 8-GPU DDP job from Part 2 running ( The Grafana dashboard Scheduler & NVLink fabric puts these next to the scheduler panels: unhealthy GPUs, switch ports down, GPUs per partition, and NVLink and NVSwitch throughput. Disable a switch port as in Part 4 while DDP runs, and the switch ports down panel and the GPU’s active-link count change within one scrape. Part 6: Observability puts all of this on dashboards: which layers of a GPU cluster to monitor, and how to read the graphs.The concepts
nvidia-smi --query-gpu=fabric.cliqueId).sched-nvswitch runs fakenmxc, an NMX-C-style controller. It speaks gRPC on port 9370 (with its own .proto, since NVIDIA’s is proprietary) and serves fabric metrics on port 9372.Observation
$ bin/ssh sched-worker1 nvidia-smi --query-gpu=index,fabric.clusterUuid,fabric.cliqueId --format=csv
index, fabric.cluster_uuid, fabric.clique_id
0, 7f3c2a10-5b4e-4d6a-9c1e-2b8f0e6d4a91, 1
…bin/nvlink asks the partition controller:$ bin/nvlink domain
domain nvl8 (7f3c2a10-5b4e-4d6a-9c1e-2b8f0e6d4a91)
state CONFIGURED
trays 2 compute, 1 switch
GPUs 8 x 18 NVLinks
partitions default 32766, at most 32765$ bin/nvlink partitions
ID NAME GPUS STATE HEALTH MEMBERS
32766 default 8 ACTIVE healthy sched-worker1:0-3 sched-worker2:0-3$ bin/nvlink topology
TRAY GPU ACTIVE NVSWITCH_0 NVSWITCH_1
sched-worker1 0 18/18 9 9
sched-worker1 1 18/18 9 9
sched-worker1 2 18/18 9 9
sched-worker1 3 18/18 9 9
sched-worker2 0 18/18 9 9
sched-worker2 1 18/18 9 9
sched-worker2 2 18/18 9 9
sched-worker2 3 18/18 9 9bin/nvlink is a thin client over the controller’s gRPC API: --json prints the raw responses, and bin/grpcurl -plaintext 10.107.111.34:9370 describe nmxlab.v1.NMXController lists every RPC.Health: a degraded GPU
$ bin/nvlink gpus sched-worker1
TRAY GPU UUID PARTITION CLIQUE NVLINKS HEALTH
sched-worker1 0 GPU-81501924-f170-f6d0-f2da-dc7595876881 32766 1 18 healthy
sched-worker1 1 GPU-8150a4ae-478e-5534-f66c-621a95f85f59 32766 1 18 healthy
sched-worker1 2 GPU-81502f39-9cad-de97-2723-3d7d71928458 32766 1 16 degraded
sched-worker1 3 GPU-8151bac3-f1cc-3cd0-015f-475439ee933a 32766 1 18 healthynvlink_gpu_active_links{host="sched-worker1",gpu="2"} 16 and nvlink_gpu_healthy{…} 0. min(nvlink_gpu_healthy) == 0 makes a natural first alert.degraded (NMX_GPU_HEALTH_DEGRADED_BANDWIDTH in the API) is the lab’s rating for a GPU with some links down. NVIDIA’s GB200 NVL Partition User’s Guide (§6.2) documents an access-link failure as marking the GPU NO_NVLINK, with the partition’s workload running into errors, and the real health enum also has a DEGRADED_BW value.Operations: split the domain
$ bin/nvlink create tray2 --id 7 sched-worker2
nvlink: CreatePartition: NMX_ST_GPU_IN_USE: GPU sched-worker2:0 is in partition 32766; remove it there first$ bin/nvlink remove default sched-worker2
GPUs removed: 32766 default
$ bin/nvlink partitions
ID NAME GPUS STATE HEALTH MEMBERS
32766 default 4 ACTIVE healthy sched-worker1:0-3
in no partition: sched-worker2:0-3
$ bin/ssh sched-worker2 nvidia-smi --query-gpu=index,fabric.cliqueId --format=csv
index, fabric.clique_id
0, 0
1, 0
2, 0
3, 00 means the GPU is in no partition. On real hardware that GPU now has no NVLink peers at all. Create the new partition:$ bin/nvlink create tray2 --id 7 sched-worker2
partition created: 7 tray2
$ bin/nvlink partitions
ID NAME GPUS STATE HEALTH MEMBERS
7 tray2 4 ACTIVE healthy sched-worker2:0-3
32766 default 4 ACTIVE healthy sched-worker1:0-3
$ bin/ssh sched-worker2 nvidia-smi --query-gpu=index,fabric.cliqueId --format=csv
index, fabric.clique_id
0, 7
1, 7
2, 7
3, 7How the scheduler finds out
ibnetdiscover and the NVML clique from every tray, and regenerates topology.conf when the result changes. I polled scontrol show topology every 10 seconds after creating the partition, and the change landed about 60 seconds later:12:38:21 partition 7 created
12:38:31 BlockName=block001 BlockIndex=0 Nodes=sched-worker[1-2]
…
12:39:23 BlockName=block001 BlockIndex=0 Nodes=sched-worker1
BlockName=block002 BlockIndex=1 Nodes=sched-worker2$ bin/ssh sched-control cat /etc/slurm/topology.conf
# block001=7f3c2a10-5b4e-4d6a-9c1e-2b8f0e6d4a91.1
BlockName=block001 Nodes=sched-worker1
# block002=7f3c2a10-5b4e-4d6a-9c1e-2b8f0e6d4a91.7
BlockName=block002 Nodes=sched-worker2accelerator.topograph.run/domain=<uuid>.<clique>, plus GPU Feature Discovery’s nvidia.com/gpu.clique). Part 3 shows those labels.The surprise: an 8-GPU job still runs
nvl8-hello (2 nodes × 4 GPUs) to wait now. It doesn’t:$ sbatch --wait nvl8-hello.sbatch && cat nvl8-hello-21.out
job 21 on sched-worker[1-2]: 8 tasks
0: rank 0 on sched-worker1 CUDA_VISIBLE_DEVICES=0: NVIDIA GB200, GPU-81501924-…, 1, scratch 198G
…
4: rank 4 on sched-worker2 CUDA_VISIBLE_DEVICES=0: NVIDIA GB200, GPU-81477906-…, 7, scratch 198G
…topology/block prefers to keep a job inside one block, but a job larger than any block is allowed to span several. Newer Slurm releases add --segment and BlockSizes to control that; the lab runs Ubuntu’s Slurm 23.11, which has neither. On real hardware, this job’s NCCL traffic between the halves would fall back to the network. The lab doesn’t model that: the fake NCCL ranks don’t communicate and ignore cliques, so the job’s timing is unchanged.kueue.x-k8s.io/podset-required-topology on the NVLink domain label, so Kueue keeps the same job queued rather than splitting it. The same partition change gives you two different scheduler behaviours, which makes the lab a good place to study them.$ bin/nvlink delete tray2
partition deleted; its GPUs are in no partition
$ bin/nvlink add default sched-worker2
GPUs added: 32766 default
$ bin/nvlink partitions
ID NAME GPUS STATE HEALTH MEMBERS
32766 default 8 ACTIVE healthy sched-worker1:0-3 sched-worker2:0-3Telemetry
$ curl -s http://10.107.111.34:9372/metrics | grep -E '^(nvlink_domain_info|nvlink_partition_gpus|nvswitch_ports_up)'
nvlink_domain_info{cluster_uuid="7f3c2a10-5b4e-4d6a-9c1e-2b8f0e6d4a91",domain="nvl8"} 1
nvlink_partition_gpus{default="true",name="default",partition_id="32766"} 8
nvswitch_ports_up{host="sched-nvswitch",switch="NVSwitch_0"} 72
nvswitch_ports_up{host="sched-nvswitch",switch="NVSwitch_1"} 72sbatch ddp-train.sbatch --steps 60000), Prometheus shows the all-reduce traffic:Query Reading sum by (host) (rate(nvlink_gpu_tx_bytes_total[1m]))~1.68 TB/s from each tray rate(nvswitch_tx_bytes_total[1m])~1.68 TB/s per NVSwitch chip; each GPU spreads its links over both sum(nvswitch_ports) - sum(nvswitch_ports_up)0 ports downmin(nvlink_gpu_healthy)1Ideas to build on this
Next in the series