AI lab, part 4: BMCs and Redfish, the out-of-band side of a GPU rack
Table of Contents
Out-of-band management of GPU trays with Redfish BMCs: inventory, GPU sensors, NVLink faults and power cycles. Parts 2 and 3 used the cluster the way its users do, through a scheduler. This part covers the operator’s back door: the baseboard management controllers. We’ll walk the Redfish tree, break NVLinks on purpose, power-cycle a tray, and see what the scheduler makes of it. The examples run in Slurm mode, but the BMCs are identical with Kubernetes. A BMC is a small computer on every server board with its own network port. It stays up when the host is off or hung, and it can power the host on and off, report hardware inventory and health, and change firmware-level settings. In a GB200 rack every compute tray and every NVLink switch tray has one. Redfish is the DMTF’s REST/JSON API for BMCs, and it replaced IPMI. Everything is a resource under The lab has three BMCs: The credentials are OpenBMC’s defaults, The service root is the one resource you can read without logging in: The usual three branches are there. Systems is what runs (host, CPUs, GPUs). Chassis is the physical enclosure. Managers is the BMC itself. In the lab, the GPUs are processors of Don’t hard-code these paths. The lab uses a simplified layout. On real GB200 compute-tray BMCs the GPUs sit behind the HMC under Discover resources by following Each GPU has 18 NVLink ports, and each port has a pending-settings object ( The switch tray BMC describes the other end of the cable. Every switch port knows which tray, GPU and link it’s connected to: The tray BMCs also carry GPU sensors, which is how a BMC reports temperature and power when the host OS can’t. On a real tray the BMC reads them from the GPUs over its own sideband bus. In the lab the BMC reads the same GPU state NVML does and applies the same model, so the two agree. Captured during a training run: The temperature matches exactly. The power readings differ by about 20 W because the two commands ran a moment apart and a busy GPU’s power jitters from sample to sample (a minute earlier, one Stage two links of GPU 2 as disabled. The live port is unchanged until the reset: About 20 seconds later the tray is back up, and both the BMC and the GPU’s own view agree: Slurm didn’t notice anything. The tray restarted well within Switch-side changes apply immediately, with no reset. Port 19 on One emulation gap shows up here: the tray BMC still reports its end of that link as For link state from the switch side, trust the switch BMC or NVML in the lab, not the tray BMC. The tray is off, but Slurm still sees it as Within a minute topograph also drops the unreachable tray from Automation should use a session token rather than sending the password on every request. You I ran exactly this script against the lab. The node goes Part 5: NVLink moves from single links to the whole fabric: NVLink partitions, the controller that manages them, fabric telemetry, and how a partition change rewrites the scheduler’s topology.BMCs and Redfish in two paragraphs
/redfish/v1: you read with GET, change settings with PATCH, and trigger actions with POST. Vendors add their own fields under Oem. NVIDIA’s GB200 BMCs run a fork of OpenBMC’s bmcweb, and the lab’s fakebmc is modelled on it, Oem.Nvidia fields included.BMC Address Manages sched-worker1-bmc, sched-worker2-bmc.31, .32 a GPU tray: power, 4 GPUs, 18 NVLink ports each sched-nvswitch-bmc.33 the switch tray: 2 NVSwitch chips × 72 ports root / 0penBmc, and the certificates are self-signed. bin/redfish <tray> <path> [curl args] wraps curl -k with the credentials filled in.Observation
$ curl -sk https://10.107.111.31/redfish/v1 | jq -c '{RedfishVersion, Systems, Chassis, Managers}'
{"RedfishVersion":"1.17.0","Systems":{"@odata.id":"/redfish/v1/Systems"},"Chassis":{"@odata.id":"/redfish/v1/Chassis"},"Managers":{"@odata.id":"/redfish/v1/Managers"}}$ bin/redfish sched-worker1 /redfish/v1/Systems/System_0 | jq '{PowerState, Status, ProcessorSummary}'
{
"PowerState": "On",
"Status": { "Health": "OK", "State": "Enabled" },
"ProcessorSummary": { "Count": 4 }
}
$ bin/redfish sched-worker1 /redfish/v1/Chassis/Chassis_0 | jq -c '{ChassisType, Model, PowerState}'
{"ChassisType":"Sled","Model":"NVIDIA GB200 compute tray","PowerState":"On"}System_0. The BMC’s UUID matches the one nvidia-smi reports inside the tray, so inventory tooling can join out-of-band and in-band data on it:$ bin/redfish sched-worker1 /redfish/v1/Systems/System_0/Processors/GPU_0 | jq '{Model, UUID, SerialNumber, Oem}'
{
"Model": "NVIDIA GB200",
"UUID": "81501924-f170-f6d0-f2da-dc7595876881",
"SerialNumber": "16523240298025",
"Oem": {
"Nvidia": {
"@odata.type": "#NvidiaProcessor.v1_4_0.NvidiaGPU",
"FabricClique": { "CliqueId": 1, "ClusterUUID": "7f3c2a10-5b4e-4d6a-9c1e-2b8f0e6d4a91" },
"PCIeBusId": "00000000:18:00.0"
}
}
}
$ bin/ssh sched-worker1 nvidia-smi -L | head -1
GPU 0: NVIDIA GB200 (UUID: GPU-81501924-f170-f6d0-f2da-dc7595876881)/redfish/v1/Systems/HGX_Baseboard_0/Processors, with an HGX_GPU_<n> chassis per GPU and an HGX_BMC_0 manager. On the switch tray, the fabric is MGX_NVLinkFabric_0 (the lab’s is NVLinkFabric_0) and the switches have MGX_NVSwitch_<n> chassis. make test-bmc-conformance checks for the real layout and fails on every difference:FAILED tests/bmc/test_conformance.py::test_gpus_live_under_the_hgx_baseboard[sched-worker1]
FAILED tests/bmc/test_conformance.py::test_gpus_live_under_the_hgx_baseboard[sched-worker2]
FAILED tests/bmc/test_conformance.py::test_each_gpu_has_a_chassis_with_its_uuid[sched-worker1]
FAILED tests/bmc/test_conformance.py::test_each_gpu_has_a_chassis_with_its_uuid[sched-worker2]
FAILED tests/bmc/test_conformance.py::test_tray_has_host_bmc_and_hmc_managers[sched-worker1]
FAILED tests/bmc/test_conformance.py::test_tray_has_host_bmc_and_hmc_managers[sched-worker2]
FAILED tests/bmc/test_conformance.py::test_switch_tray_identifies_as_nvswitch
FAILED tests/bmc/test_conformance.py::test_switch_fabric_name - AssertionErro...Members links from the collections instead of building paths, and your tooling will work against both.@Redfish.Settings). That’s the standard Redfish pattern for changes that only take effect at the next reset:$ bin/redfish sched-worker1 /redfish/v1/Systems/System_0/Processors/GPU_2/Ports/NVLink_3 \
| jq -c '{LinkState, LinkStatus, CurrentSpeedGbps, Settings: ."@Redfish.Settings".SettingsObject}'
{"LinkState":"Enabled","LinkStatus":"LinkUp","CurrentSpeedGbps":200,"Settings":{"@odata.id":"/redfish/v1/Systems/System_0/Processors/GPU_2/Ports/NVLink_3/Settings"}}$ bin/redfish sched-nvswitch /redfish/v1/Fabrics/NVLinkFabric_0/Switches/NVSwitch_0/Ports/NVLink_19 \
| jq -c '{LinkStatus, RemoteEndpoint: .Oem.Nvidia.RemoteEndpoint}'
{"LinkStatus":"LinkUp","RemoteEndpoint":{"BMC":"sched-worker1-bmc","GPU":"GPU_2","GPUUUID":"81502f39-9cad-de97-2723-3d7d71928458","Host":"sched-worker1","Port":"NVLink_2"}}$ bin/redfish sched-worker1 /redfish/v1/Chassis/Chassis_0/Sensors | jq -r '.Members[]."@odata.id"'
/redfish/v1/Chassis/Chassis_0/Sensors/GPU_0_TEMP_0
/redfish/v1/Chassis/Chassis_0/Sensors/GPU_0_Power_0
…
/redfish/v1/Chassis/Chassis_0/Sensors/Total_GPU_Power_0
$ bin/redfish sched-worker1 /redfish/v1/Systems/System_0/Processors/GPU_0/EnvironmentMetrics \
| jq -c '{Temp: .TemperatureCelsius.Reading, Power: .PowerWatts.Reading}'
{"Temp":69,"Power":889.878}
$ bin/ssh sched-worker1 nvidia-smi -i 0 --query-gpu=temperature.gpu,power.draw --format=csv
temperature.gpu, power.draw [W]
69, 870.59 Wnvidia-smi call read the same tray’s four GPUs at 865–889 W). Total_GPU_Power_0 sums the tray’s four GPUs (3557.758 W at that moment). A Redfish exporter polls these sensors and the power state into Prometheus; Part 6 shows what that looks like during a power cycle.Operations
Disable NVLinks and reset the tray
$ for l in 3 4; do
bin/redfish sched-worker1 /redfish/v1/Systems/System_0/Processors/GPU_2/Ports/NVLink_$l/Settings \
-X PATCH -d '{"LinkState": "Disabled"}'
done
$ bin/redfish sched-worker1 …/GPU_2/Ports/NVLink_3 | jq -c '{LinkState, LinkStatus}'
{"LinkState":"Enabled","LinkStatus":"LinkUp"}
$ bin/redfish sched-worker1 /redfish/v1/Systems/System_0/Actions/ComputerSystem.Reset \
-X POST -d '{"ResetType": "ForceRestart"}'$ bin/redfish sched-worker1 …/GPU_2/Ports/NVLink_3 | jq -c '{LinkState, LinkStatus, Status}'
{"LinkState":"Disabled","LinkStatus":"LinkDown","Status":{"Health":"Warning","State":"Disabled"}}
$ bin/ssh sched-worker1 nvidia-smi topo -m
GPU0 GPU1 GPU2 GPU3 CPU Affinity NUMA Affinity GPU NUMA ID
GPU0 X NV18 NV16 NV18 0-3 0 N/A
GPU1 NV18 X NV16 NV18 0-3 0 N/A
GPU2 NV16 NV16 X NV16 0-3 0 N/A
GPU3 NV18 NV18 NV16 X 0-3 0 N/A
…SlurmdTimeout (300 s), so it stayed idle throughout. On real hardware you’d drain first, and the lab lets you find out what happens when somebody doesn’t. Set the links back to "Enabled" and reset again to restore NV18.Take a switch port down
NVSwitch_0 is GPU 2’s link 2:$ bin/redfish sched-nvswitch /redfish/v1/Fabrics/NVLinkFabric_0/Switches/NVSwitch_0/Ports/NVLink_19 \
-X PATCH -d '{"LinkState": "Disabled"}'
$ bin/ssh sched-worker1 'nvidia-smi topo -m | grep ^GPU2; nvidia-smi nvlink -s -i 2 | grep "Link 2:"'
GPU2 NV17 NV17 X NV17 0-3 0 N/A
Link 2: <inactive>LinkUp. On real hardware, a link that’s down is down at both ends, so the GPU’s port would report LinkDown too. The lab’s fakebmc doesn’t propagate switch-side changes to the tray BMC, and the conformance tier catches it:E AssertionError: assert 'LinkUp' == 'LinkDown'
FAILED tests/bmc/test_conformance.py::test_link_down_shows_at_both_ends - Ass...Power off without draining (what not to do)
$ bin/redfish sched-worker2 /redfish/v1/Systems/System_0/Actions/ComputerSystem.Reset \
-X POST -d '{"ResetType": "ForceOff"}'
$ bin/redfish sched-worker2 /redfish/v1/Systems/System_0 | jq -c '{PowerState}'
{"PowerState":"Off"}
$ sinfo -N -o "%N %T"
NODELIST STATE
sched-worker1 idle
sched-worker2 idleidle. Prometheus is quicker: up{job="gpu", instance="sched-worker2:9835"} drops to 0 within seconds. A two-tray job submitted now gets placed on the dead node and fails:$ sbatch -J blind -N2 --gpus-per-node=4 --wrap "srun hostname"
Submitted batch job 20
$ sacct -X -n -j 20 -o JobID,JobName,State
20 blind FAILEDtopology.conf (BlockName=block001 Nodes=sched-worker1). ResetType: On brings it back, and both the node and the block recover on their own.The proper power cycle, with a Redfish session
POST credentials once, get back X-Auth-Token, and DELETE the session when you’re done:BMC=https://10.107.111.32
TOKEN=$(curl -sk -D - -o /dev/null -X POST $BMC/redfish/v1/SessionService/Sessions \
-H 'Content-Type: application/json' -d '{"UserName":"root","Password":"0penBmc"}' \
| awk -F': ' 'tolower($1)=="x-auth-token"{print $2}' | tr -d '\r')
rf() { curl -sk -H "X-Auth-Token: $TOKEN" -H 'Content-Type: application/json' "$@"; }
bin/ssh sched-control 'scontrol update nodename=sched-worker2 state=drain reason="power cycle"'
rf -X POST $BMC/redfish/v1/Systems/System_0/Actions/ComputerSystem.Reset -d '{"ResetType":"GracefulShutdown"}'
until [ "$(rf $BMC/redfish/v1/Systems/System_0 | jq -r .PowerState)" = Off ]; do sleep 2; done
rf -X POST $BMC/redfish/v1/Systems/System_0/Actions/ComputerSystem.Reset -d '{"ResetType":"On"}'
until bin/ssh sched-worker2 systemctl is-active slurmd 2>/dev/null | grep -qx active; do sleep 3; done
bin/ssh sched-control 'scontrol update nodename=sched-worker2 state=resume'drained → idle* (resumed, waiting for slurmd to check in) → idle within about 20 seconds. Finish by deleting the session (rf -X DELETE $BMC/redfish/v1/SessionService/Sessions/<id>). The Kubernetes version swaps drain/resume for kubectl cordon/uncordon and is covered in Part 3.What to build against it
nvidia-smi and the scheduler’s node list.make test-bmc in the repo runs the lab’s own pytest suite, and make test-bmc-conformance lists where fakebmc differs from real GB200 BMCs.Next in the series