Skip to main content
All metrics listed below are available through your Grafana dashboards. Queries are automatically scoped to your assigned nodes and namespaces.

GPU

DCGM utilization, memory, thermal, ECC, NVLink, PCIe, and XID metrics.

Node

Aggregated CPU, memory, load, pressure, thermals, and EDAC metrics.

Container

CPU and memory metrics scoped to pods in your namespace.

Slurm

Job state, GPU allocation, CPU allocation, and assigned-capacity rules.

InfiniBand

Port counters, link rate, transmit wait, and fabric-level congestion metrics.

Kubernetes state

Node conditions, cordon state, pod metadata, and retained pod phase metrics.
Use Metric for dashboard or query matching, Unit for panel scale, and Description to decide whether the signal applies to your workload.

GPU (DCGM)

Labels: cluster, namespace, pod, node, gpu (index), modelName, Hostname.
SM profiling metrics (DCGM_FI_PROF_SM_ACTIVE, DCGM_FI_PROF_SM_OCCUPANCY) and volatile ECC counters appear when they are enabled for your cluster. If a panel is empty after you confirm filters, ask Andromeda Support whether those signals are available for your environment.

Node (aggregated)

CPU metrics are summarized per host (and often per NUMA group) so charts stay readable, rather than listing every logical CPU separately.

Container

Scoped to pods in your namespace.

Slurm

Pre-computed recording rules

InfiniBand

Available where InfiniBand is deployed, not on RoCE clusters.

Weka (where deployed)

Kubernetes state

Beyond the default dashboard set

Standard dashboards emphasize GPU health, host pressure, Slurm placement, containers, InfiniBand, and Kubernetes state. If you need per-core CPU metrics, Kubernetes control-plane detail, or container network and disk statistics, contact Andromeda Support to discuss expanded coverage.