GPU Nodes
The primary hardware health view. Start here when investigating GPU issues.
Confirm node identity, GPU count, active alerts, and scheduler allocation.

Check CPU, memory, load, context switches, processes, swap, and PSI signals for host-level bottlenecks.
Job Analysis
Maps Slurm jobs to nodes and GPUs.
Slurm operations summaries help connect queue state, running jobs, reservations, and scheduler context.

Slurm status shows whether Slurm is running continuously or changing state during the investigation window.
slurm_job_state does not carry a per-node node label. For node-level detail, use slurm_job_cpus_allocated or slurm_job_gpus_allocated, which include node labels.Tenant Dashboard
Capacity and readiness overview.
Reservation views show whether the reserved cluster and node allocation match expectations.
tenant:slurm_nodes:bad is nonzero, open the GPU Nodes dashboard to identify which specific nodes are affected and why, such as ECC errors, thermal state, cordon, or drain.

Per-node reservation rows help identify which nodes are active and how GPU reservations changed over time.

Correlate training or checkpointing symptoms with Weka throughput, IOPS, pending I/O, and latency where Weka is deployed.