AI lab, part 3: the same GPU cluster on Kubernetes, with Kueue and JobSet
Table of Contents
The same emulated GPU cluster on Kubernetes: GPU pods, Kueue topology-aware queueing, JobSet and monitoring. Part 2 ran the lab under Slurm. This part rebuilds the same emulated hardware (two trays of four fake GB200 GPUs, one NVLink domain) with Kubernetes. We’ll look at how GPUs and NVLink topology show up as Kubernetes objects, run the same jobs as pods, and see how Kueue queues them. To switch, set The two trays are worker nodes. The controller runs the control plane and add-ons, and is tainted so pods don’t land on it. Both trays carry the same NVLink clique, and two different components publish it. GFD adds about 25 more Queues: The device plugin is the lab’s own. NVIDIA’s device plugin and container toolkit expect a real driver installation, so the lab ships a small replacement that hands out GPUs by UUID through CDI: Kueue is where the topology awareness lives. Its A workload annotated with A single GPU pod. No Kueue and no JobSet, just a pod asking for two GPUs: The pod sees exactly two GPUs. They are physical GPUs 1 and 2 of I use The examples. The repository’s That’s eight pods, each holding one different GPU, all in clique 1. DDP works the same way: two pods with four GPUs each, and The 4.3 ms step time (5.3 ms on another run) is the simulated cost of the GEMMs and the all-reduce across 8 GPUs; the loss values are meaningless, since nothing is computed. Queueing. Start a longer DDP run (a copy of No pods are created while the workload is suspended, so nothing sits half-scheduled holding GPUs. When the DDP run finishes, Cordoning a tray. This is the Kubernetes half of the power-cycle runbook from Part 4: Kueue sees that only one tray’s worth of GPUs is schedulable inside the domain and holds the whole gang. After Cluster state reaches Prometheus through kube-state-metrics (NodePort 30808): nodes and their allocatable GPUs, pods’ GPU requests, Jobs and their status. Recording rules turn it into a small set of A Job that Kueue keeps suspended counts as The pods’ GPU work is visible on the trays as it would be for any process. This is the same 60,000-step run: The same worker appears as PID 14832 on the tray and PID 16 inside the pod: each viewer sees the process under the PID from its own namespace, as with the real driver. Part 6 walks through both dashboards and what to read from each graph. Part 4: BMC and Redfish goes below the scheduler, to the management controllers that power trays on and off and switch NVLinks.scheduler: k3s in inventory/group_vars/all.yml, then run make down && make up (Part 1). Commands prefixed with $ run on the login node as joe (bin/ssh login), whose kubeconfig and namespace are already set up. bin/kubectl on the host is cluster admin.Runtime information
$ kubectl get nodes -L nvidia.com/gpu.clique,accelerator.topograph.run/domain
NAME STATUS ROLES AGE VERSION GPU.CLIQUE DOMAIN
sched-control Ready control-plane 7m39s v1.36.5+k3s1
sched-worker1 Ready <none> 7m24s v1.36.5+k3s1 7f3c2a10-5b4e-4d6a-9c1e-2b8f0e6d4a91.1 7f3c2a10-5b4e-4d6a-9c1e-2b8f0e6d4a91.1
sched-worker2 Ready <none> 7m24s v1.36.5+k3s1 7f3c2a10-5b4e-4d6a-9c1e-2b8f0e6d4a91.1 7f3c2a10-5b4e-4d6a-9c1e-2b8f0e6d4a91.1nvidia.com/gpu.clique comes from NVIDIA’s GPU Feature Discovery (GFD), which reads it from NVML. accelerator.topograph.run/domain comes from topograph, which rebuilds it from the live NVLink and InfiniBand fabric every minute.$ kubectl get nodes -o custom-columns='NODE:.metadata.name,GPU:.status.allocatable.nvidia\.com/gpu,PRODUCT:.metadata.labels.nvidia\.com/gpu\.product,MEM:.metadata.labels.nvidia\.com/gpu\.memory'
NODE GPU PRODUCT MEM
sched-control <none> <none> <none>
sched-worker1 4 NVIDIA-GB200 189471
sched-worker2 4 NVIDIA-GB200 189471nvidia.com/* labels per tray, all derived from the fake NVML: gpu.family=blackwell, gpu.compute.major=10, cuda.driver-version.full=580.95.05, cuda.runtime-version.full=13.0 and so on. topograph adds the InfiniBand tiers (fabric.topograph.run/tier-0 and tier-1, the leaf and spine).$ kubectl get clusterqueues
NAME COHORT PENDING WORKLOADS
gpu 0
$ kubectl get localqueues
NAME CLUSTERQUEUE PENDING WORKLOADS ADMITTED WORKLOADS
default gpu 0 0
gpu gpu 0 0joe works in the joe namespace. They can create pods and JobSets there and read nodes and queues, but they can’t touch other namespaces or cluster-wide queue objects:$ kubectl auth can-i create jobsets.jobset.x-k8s.io
yes
$ kubectl auth can-i create pods -n kube-system
noHow it’s configured
Piece What it does k3s 1.36 control plane on sched-control, agents on the trays, all in unprivileged Incus containersfakedp + CDIdevice plugin advertising nvidia.com/gpu (4 per tray); each GPU is a CDI device in /etc/cdi/fakegpu.json that injects /dev/nvidia<n> and the fake driverNode Feature Discovery + GPU Feature Discovery the real NVIDIA/upstream components, labelling trays from NVML topograph labels NVLink domain and InfiniBand tiers every minute Kueue ClusterQueue gpu (8 CPUs, 18 GiB, 8 GPUs); in each user namespace a LocalQueue gpu and a default one for workloads without a queue label; topology-aware schedulingJobSet multi-pod jobs with stable DNS names for rank 0 $ bin/ssh sched-worker1 'jq -c "{kind, devices: [.devices[].name]}" /etc/cdi/fakegpu.json'
{"kind":"nvidia.com/gpu","devices":["0","GPU-81501924-f170-f6d0-f2da-dc7595876881","1","GPU-8150a4ae-…", …]}Topology object, nvl, orders the node labels from coarse to fine:$ bin/kubectl get topology nvl -o jsonpath='{.spec.levels}' | jq -c '[.[].nodeLabel]'
["fabric.topograph.run/tier-1","fabric.topograph.run/tier-0","accelerator.topograph.run/domain","kubernetes.io/hostname"]kueue.x-k8s.io/podset-required-topology: accelerator.topograph.run/domain is admitted only if all its pods fit inside one NVLink domain. The pods of a multi-pod job are admitted together, as a gang, or not at all.Simple jobs
apiVersion: v1
kind: Pod
metadata:
name: gpu-pod
spec:
restartPolicy: Never
containers:
- name: smi
image: mirror.gcr.io/library/buildpack-deps:noble
command: [nvidia-smi, -L]
resources:
limits: {nvidia.com/gpu: 2}$ kubectl apply -f gpu-pod.yaml
$ kubectl logs gpu-pod
GPU 0: NVIDIA GB200 (UUID: GPU-8150a4ae-478e-5534-f66c-621a95f85f59)
GPU 1: NVIDIA GB200 (UUID: GPU-81502f39-9cad-de97-2723-3d7d71928458)sched-worker1, renumbered 0 and 1, just as with the real driver. nvidia-smi and CUDA agree on what the pod has, because CDI only injects the allocated device nodes. The pod has no queue label, but Kueue still takes it: it goes to the namespace’s default LocalQueue and is admitted against the same ClusterQueue quota as everything else, so nobody can grab GPUs by skipping the queue:$ kubectl get workloads | grep gpu-pod
pod-gpu-pod-c7498 default gpu True True 3smirror.gcr.io because it serves the same image as Docker Hub (identical digest) without Docker Hub’s anonymous pull limits. The repository’s examples still say docker.io/library/buildpack-deps:noble, but the trays’ containerd is configured to fetch docker.io images through mirror.gcr.io first, so they avoid the limits as well.examples/kubernetes/ has four example jobs, written as JobSets queued in Kueue. submit is a small helper: it fills in your uid, gid and home, gives the run a unique name, and with --wait collects the logs into <name>.out.$ cd examples/kubernetes
$ ./submit --wait nvl8-hello.yaml
nvl8-hello-000472
$ cat nvl8-hello-000472.out
[pod/nvl8-hello-000472-rank-0-0-2bncb/hello] rank 0 on sched-worker1 NVIDIA_VISIBLE_DEVICES=GPU-8150a4ae-…: NVIDIA GB200, GPU-8150a4ae-478e-5534-f66c-621a95f85f59, 1
…
[pod/nvl8-hello-000472-rank-0-7-jdvdh/hello] rank 7 on sched-worker2 NVIDIA_VISIBLE_DEVICES=GPU-81477906-…: NVIDIA GB200, GPU-81477906-d580-de42-e868-ff62e9ef4317, 1torchrun doing rendezvous on pod 0 through the JobSet’s DNS name:$ ./submit --wait ddp-train.yaml
ddp-train-386106
$ grep -E "world=|step +50|rank 0/" ddp-train-386106.out
[pod/ddp-train-386106-node-0-0-t9m8f/torchrun] world=8 params=537M batch/rank=64 width=8192
[pod/ddp-train-386106-node-0-0-t9m8f/torchrun] step 50 4.3 ms 119008 samples/s 48 TFLOP/s/GPU
[pod/ddp-train-386106-node-0-0-t9m8f/torchrun] rank 0/8 on ddp-train-386106-node-0-0 cuda:0 (NVIDIA GB200) peak mem 7.1 GiBOperations
ddp-train.yaml with --steps 20000), then submit nvl8-hello behind it. Kueue suspends the second JobSet and says why:$ kubectl get workloads
NAME QUEUE RESERVED IN ADMITTED FINISHED AGE
jobset-ddp-long-408229-28e60 gpu gpu True 26s
jobset-nvl8-hello-912542-a34fa gpu False 21s
$ kubectl get workload jobset-nvl8-hello-912542-a34fa \
-o jsonpath='{.status.conditions[?(@.type=="QuotaReserved")].message}'
couldn't assign flavors to pod set rank: insufficient unused quota for nvidia.com/gpu in flavor gb200, 8 more needednvl8-hello is admitted and completes.$ bin/kubectl cordon sched-worker2
$ ./submit nvl8-hello.yaml
nvl8-hello-664894
$ kubectl get workload … -o jsonpath='{.status.conditions[?(@.type=="QuotaReserved")].message}'
couldn't assign flavors to pod set rank: topology "nvl" allows to fit only 4 out of 8 pod(s)bin/kubectl uncordon sched-worker2, the JobSet is admitted and completes within 10 seconds. You’ll see the same message again in Part 5, where splitting the NVLink domain has the same effect as cordoning.Monitoring
sched_* series (GPUs allocated, GPUs per user, jobs and nodes by state), which the Grafana dashboards are built on. With a 60,000-step DDP run holding all 8 GPUs and nvl8-hello queued behind it:Query Value sched_gpus_alloc8sched_user_gpus{namespace="joe", user="joe"} 8sched_jobs{state!="COMPLETED"}RUNNING 1, PENDING 1sched_nodesREADY 2sum by (instance) (nvidia_smi_utilization_gpu_ratio)3.48 per tray (4 GPUs × ~87 %)sum by (instance) (nvidia_smi_compute_apps)4 per trayPENDING, so the queue is visible next to the allocation. In Grafana’s Scheduler & NVLink fabric dashboard it shows up as a stacked band under the running job:
$ bin/ssh sched-worker1 nvidia-smi --query-gpu=index,utilization.gpu,memory.used,power.draw,temperature.gpu --format=csv,noheader
0, 86 %, 7850 MiB, 889.35 W, 69
…
$ bin/ssh sched-worker1 nvidia-smi --query-compute-apps=pid,process_name,used_memory --format=csv,noheader
14832, /shared/venv/bin/python, 7338 MiB
…
$ bin/kubectl -n joe exec ddp-long-205020-node-0-0-q8crt -- nvidia-smi --query-compute-apps=pid,process_name,used_memory --format=csv,noheader
16, /shared/venv/bin/python, 7338 MiB
…Next in the series