3. Set up your first alarm
15 minutes. By the end you'll have a Grafana-backed alarm rule that fires when one of your apps misbehaves, plus you'll know where to see firing alerts.
What you need
- An organization selected
- At least one edge instance + at least one application running (Tutorial 2 covers both)
- The
alarms:createpermission in your org
Step 1 — Open the alarms page
Sidebar → Alarms.
You'll land on a list of rules already configured for this organization. New install? It'll be empty.
The page has two tabs:
- Rules — alarm definitions you (and others in the org) manage
- Firing — a slim view of which rules are currently triggering
alerts; same info as the standalone
/alarms/firingpage
Step 2 — Click "Add alarm"
Top-right of the Rules tab. The drawer opens with three sections.
a. Identity
- Name — short, descriptive. e.g.
myapp pod restart loop - Description — what it means and what to do about it. The on-call engineer at 3am will thank you.
- Severity — low / medium / high / critical. This shows up as a badge in the firing list.
b. Query
This is the PromQL the rule evaluates. Two things to know:
- Don't try to add tenant-scope labels yourself. The backend
injects
telemetry_id=~"<your-org's-edges>"into every VectorSelector before sending the query to Grafana. Just write the metric and the threshold logic. - Only metrics scraped by your edges' Prometheus stacks are available. kube-prometheus-stack (KPS) defaults are present on every standard edge — node/pod/container metrics, kube-state metrics, kubelet, cAdvisor. If your app exposes its own metrics, make sure it's being scraped.
Examples that work out of the box:
# Pod restarted more than 3 times in the last 15 minutes
increase(kube_pod_container_status_restarts_total[15m]) > 3
# Node CPU above 90% for 5 minutes
1 - avg(rate(node_cpu_seconds_total{mode="idle"}[5m])) > 0.9
# Any pod stuck in CrashLoopBackOff
kube_pod_container_status_waiting_reason{reason="CrashLoopBackOff"} > 0
The page shows a Test query button — click it to run the PromQL and see what value(s) come back without saving the rule. Saves you from saving an alarm that never matches anything.
c. Threshold + cadence
- Operator + value —
><>=<===!=and a numeric threshold - For — how long the condition must be true before the alarm fires. Use this to suppress flapping (e.g. "for 5m" means a brief blip won't page anyone).
You don't pick the evaluation interval — every rule in the org runs at the platform's default cadence (60s). Trying to get sub-minute resolution out of DT Edge Platform isn't supported.
Step 3 — Save and watch
Click Save. You'll see a toast confirmation; the rule appears in the Rules tab.
Behind the scenes, DT Edge Platform:
- Wraps your PromQL with a tenant-scope matcher
- Builds a Grafana rule with reduce(last) → threshold pipeline
- Stamps
dtedge_managed=true+dtedge_org_id=<your-org>labels so the next list call finds it - Saves it into Grafana's
dtedgefolder + group
Step 4 — Trigger it (optional but satisfying)
Cause the condition you're alarming on — restart a pod a few
times if you used the restart-loop example, or use
stress-ng if you wrote a CPU alarm.
Wait <for> plus one evaluation interval (default 60s + your
"for" window). Then:
- Sidebar → Alarms → Firing tab
- You should see your rule with a red 🔥 badge
If nothing fires, three things to check:
- Test query in the rule's edit drawer — what does PromQL say right now? If it's empty, the metric isn't being collected or your label filter is too tight.
- Time + threshold —
for: 5mmeans literally 5 minutes; you have to keep the condition true that long. - Tenant scoping — if your test run was on an edge that's not part of this org, the matcher filters it out. Check under Edge Instances that the edge belongs to the active org.
Step 5 — Mute / silence (when needed)
When you're working on a known issue and don't want pages, you have two options:
- Disable the rule — toggle in the row's actions menu. Stops evaluation entirely.
- Silence in Grafana — for a time-bound mute (e.g. "shut up for the next 2 hours during the migration"). Click the rule → "Open in Grafana" — silences are a Grafana feature, not a DT Edge Platform one.
What's next
- Tutorial 4 — Onboard your team — get more eyes on these alarms
- Reference: Alarms page — every field, every button
- Explanation: Alarms and tenant scoping — the why behind the PromQL injection