Skip to main content

Grafana access

Sign in at grafana.andromedacluster.xyz with Andromeda SSO credentials. The session lands in the Tenants org with the Viewer role. Three dashboards are available, pre-filtered to your assigned nodes and namespaces. You only see metrics and workloads for your organization’s reserved capacity.
  • GPU Nodes - GPU utilization, temperature, power, ECC, memory, node CPU/memory
  • Job Analysis - Slurm job state, GPU/CPU allocation, node mapping
  • Tenant Dashboard - capacity overview, node readiness, reservation status
Grafana reservation summary showing reserved nodes, reserved GPUs, and reserved clusters.

Confirm assigned capacity, ready nodes, and reservation status before drilling into a node or job.

Grafana node info panels showing hostname, cluster, OS type, GPU inventory, uptime, and active alerts.

Confirm host identity, GPU inventory, uptime, and active alerts.

For panel-level detail, see Dashboards. For custom dashboards, additional metrics, or a direct feed to an external monitoring stack, contact Andromeda Support.

Metric naming

Some node-level metrics use a tenant_ prefix to scope them to your assigned capacity: GPU metrics (DCGM_FI_*), container metrics, and Slurm metrics use their standard names and are already scoped by namespace and node assignment. Full list in Metrics Reference.
Grafana Explore view with PromQL queries and time-series output.

Use dashboard panels or supported queries to confirm metric names and labels before requesting additional access.

Permissions

Pre-built dashboards are read-only for all users. Your team can ask Andromeda Support to enable ad-hoc Explore queries or dashboard editing when you need those capabilities.