GPU
DCGM utilization, memory, thermal, ECC, NVLink, PCIe, and XID metrics.
Node
Aggregated CPU, memory, load, pressure, thermals, and EDAC metrics.
Container
CPU and memory metrics scoped to pods in your namespace.
Slurm
Job state, GPU allocation, CPU allocation, and assigned-capacity rules.
InfiniBand
Port counters, link rate, transmit wait, and fabric-level congestion metrics.
Kubernetes state
Node conditions, cordon state, pod metadata, and retained pod phase metrics.
Metric for dashboard or query matching, Unit for panel scale, and Description to decide whether the signal applies to your workload.
GPU (DCGM)
Labels:cluster, namespace, pod, node, gpu (index), modelName, Hostname.
SM profiling metrics (
DCGM_FI_PROF_SM_ACTIVE, DCGM_FI_PROF_SM_OCCUPANCY) and volatile ECC counters appear when they are enabled for your cluster. If a panel is empty after you confirm filters, ask Andromeda Support whether those signals are available for your environment.