Skip to main content
Andromeda monitors clusters continuously and routes alerts to your team’s designated Slack channel. This page lists alert rules, meanings, routing behavior, and expected follow-up.
Grafana memory dashboard with memory usage, OOM kills, huge pages, EDAC memory errors, swap usage, and memory pressure panels.

Memory and pressure panels provide supporting context when host pressure or node alerts require impact review.

Alert tiers

Severity levels and examples.

GPU alerts

Thermal, ECC, NVLink, PCIe, row remap, and XID signals.

Node alerts

Readiness, reachability, cordon, drain, and pressure states.

Host alerts

CPU, memory, I/O, network, disk, and temperature pressure.

Telemetry health

Metrics and logs freshness checks.

Expected chains

Follow-on signals that should appear after hardware faults.

Alert tiers

Alert types

Slack behavior

Alerts fire once per state change. Slack does not post a separate “resolved” message for every alert; use the Alerts Overview dashboard in Grafana for current status. When more than 5 events of the same type fire within 60 seconds, they are aggregated into a single Slack message.

Cluster and capacity alerts

GPU alerts

Node alerts

Host alerts

Log-derived alerts

These alerts also evaluate host log signals in addition to metrics:

Telemetry health

Alert routing

When an alert fires, it is posted to your designated Slack channel. Related signals on the same node may be grouped so your channel stays readable.

Expected alert chains

Hardware faults typically produce a sequence of alerts: If a root cause fires but the expected consequence does not appear within the expected window, escalate to Andromeda Support.