Skip to main content

View firing alarms

The "is anything broken right now" page. Updates live; opens into Grafana for deeper inspection.

Where it is​

Two ways to land here:

  • Sidebar → Alarms → Firing tab
  • Direct URL: /alarms/firing

Both surfaces show the same data. The standalone URL is useful for bookmarking / pinning to a TV dashboard.

What you see​

A list of every alarm rule in the active org that is currently firing (state = Alerting):

  • Severity — coloured badge (red/orange/yellow/grey)
  • Name of the rule
  • Started at — when this firing instance began
  • Latest value — the metric value at the most recent evaluation, plus the threshold for context
  • Edge instance(s) — which of the org's edges the matching series came from

The list polls every 30 seconds and updates live via WebSocket when a rule's state flips.

Open a row​

Clicking a firing row opens the detail panel:

  • The PromQL the rule evaluates (with tenant scope already applied — what Grafana actually sees)
  • A small chart of the last 6 hours of values, with the threshold drawn as a dashed line
  • Actions:
    • Open in Grafana — deep-link to the rule in Grafana for silencing or for inspecting the panel/dashboard linked to it
    • Edit — opens the same rule edit drawer as the Rules tab

Silencing during an incident​

Three options, in order of recommendation:

  1. Fix the underlying issue. If the alarm is correct, the right answer is to fix what it's pointing at, not silence it.
  2. Open in Grafana → Silences → New silence. For a time-bound mute (e.g. 2-hour migration window). Survives sessions and deploys.
  3. Disable the rule in DT Edge Platform for a longer pause where you know the condition will keep firing. Sidebar → Alarms → Rules tab → row toggle.

DT Edge Platform has no built-in silence UI of its own — Grafana owns that mechanic.

What "edge instance(s)" means​

Alarms are tenant-scoped; the backend injects telemetry_id=~"<your-org's-edges>" into every PromQL query before sending to Grafana. The result is that only metric series from your org's edges can match. The "Edge instance(s)" column on a firing row tells you which subset of those edges actually contributed the matching data points right now — so for a kube_pod_container_status_restarts_total alarm, you can see that "edge-istanbul" is the one whose pods are restarting.

If an alarm fires without any edge listed, the PromQL probably returned a vector that doesn't carry the telemetry_id label (some aggregations strip it). The alarm is still tenant-scoped at the query level, but the UI can't attribute the firing to a specific edge. Common with summary metrics.

Why is X firing but I don't see it here​

Possible reasons:

  • Different org — switch the org selector. Each org sees only its own alarms.
  • Different cluster's Grafana — does your install have a per-edge Grafana set? (No — see Explanation: alarms and tenant scoping; there's one central Grafana per DT Edge Platform install.)
  • Stale browser — refresh. The list polls but a stuck WebSocket can show stale state.
  • Alert state isn't Alerting — Pending (within the for: window) doesn't show as firing. It will flip after the duration elapses.

Common errors​

  • Page is empty even though things are obviously broken — see "Why is X firing but I don't see it here" above
  • GRAFANA_NOT_CONFIGURED — admin needs to fill the central Grafana URL/creds in Admin → Settings
  • GRAFANA_LIST_FAILED — Grafana is unreachable; check status from admin settings → Test connection

See also​