Skip to main content

Alarms and tenant scoping

Alarms are Grafana-backed alert rules. DT Edge Platform keeps orgs from seeing each other's data via PromQL injection — not by giving each org its own Grafana.

The single Grafana​

There's one Grafana instance per DT Edge Platform install. Every org's alarms live in the same Grafana, in the same folder (dtedge), in the same rule group. Older docs talked about per-edge Grafanas; that's gone.

So how do orgs not see each other's alarms?

Two layers of scoping​

Layer 1 — labels on the rule​

When DT Edge Platform creates an alarm rule, it stamps two labels:

  • dtedge_managed=true — distinguishes dtedge-owned rules from any rules someone wrote by hand directly in Grafana
  • dtedge_org_id=<this-org's-uuid> — the tenant identity

The list/get/update/delete handlers in DT Edge Platform filter by these labels. An org user calling GET /api/orgs/.../alarms only ever sees rules with their own dtedge_org_id. The Grafana behind might host a thousand rules across fifty orgs — the user sees only their own.

Layer 2 — PromQL injection​

A rule's PromQL doesn't carry an org identity by itself; it's just metric expressions. Without further scoping, a rule for org A could match data points emitted by org B's edges (because the central Prometheus has both).

So: DT Edge Platform rewrites every rule's PromQL before sending to Grafana. It walks the AST, finds every VectorSelector, and adds a label matcher:

user typed: rate(http_requests_total[5m]) > 0.9
backend sends: rate(http_requests_total{telemetry_id=~"<edge-uuid-1>|<edge-uuid-2>|..."}[5m]) > 0.9

The telemetry_id matcher lists every edge belonging to this org. The list is built fresh per request — adding an edge picks up matches on its data immediately; removing an edge stops matching the next eval cycle.

The user never types the matcher. The UI hides it on read (strip on read, inject on write — round-trip invisible). The server rejects any user PromQL containing a literal telemetry_id reference, as defence in depth.

Why this design​

Three options were on the table:

  1. One Grafana per org. Cleanest isolation, expensive operationally (many Grafanas to manage, many users with their own logins, hard to cross-correlate).
  2. One Grafana, DT Edge Platform mediates everything. What we have. PromQL injection enforces tenancy at the query AST level; labels enforce ownership at the rule list level.
  3. One Grafana, just trust users not to look at each other's rules. Not actually a multi-tenant story. Rejected.

Option 2 wins on operational simplicity (one Grafana to back up, upgrade, monitor) at the cost of having to be careful about the injection (if the AST walker misses a VectorSelector, that's a tenant-isolation bug). The injection lives in internal/services/grafana/promql_inject.go and has tests.

What "tenant" means here​

Tenant = dtedge_org_id. So:

  • Two members of the same org see the same alarms (yes)
  • A super-admin can switch org context and see different alarm sets (yes — same database, just filtered differently per request)
  • An admin user sees only their active org's alarms (yes — if you're in org A but you have org B selected, you see org B's alarms)
  • The "Public Edges" admin surface (when the license has it) doesn't have alarms of its own — alarms still need an org

Why no per-rule eval interval​

Every DT Edge Platform alarm runs at the platform's default eval cadence (60 seconds). The UI doesn't surface a "evaluate every X seconds" picker.

Reason: ops simplicity. With many orgs sharing one Grafana, a rule running every 5 seconds in one org affects everyone's load. The platform default is what we tune; rule authors don't each get to argue about it.

If you genuinely need sub-minute resolution for a specific condition, DT Edge Platform isn't the right tool — you'd run that rule in Prometheus's own alerting layer (with PrometheusRule manifest on the edge cluster) and have it page directly.

Alertmanager + silences​

Firing alerts go to Grafana's bundled Alertmanager, which is the notification routing layer. Silences (time-bound mutes) are an Alertmanager feature, not a DT Edge Platform one. DT Edge Platform doesn't have its own silence UI — you click "Open in Grafana" from the firing list and create a silence there.

Two reasons for this split:

  1. Silences are inherently tied to a notification window ("shut up for the next 2 hours during this migration"). DT Edge Platform's view is rule-centric, not silence-centric.
  2. Alertmanager already does silences well. Re-implementing them in DT Edge Platform would be re-skinning a perfectly good feature for no win.

What an alarm-receiving system would see​

If you wire Alertmanager to a paging system (PagerDuty / Opsgenie / Slack / custom webhook), the payload includes the rule's labels. So:

  • dtedge_managed=true — filter on this if you want only DT Edge Platform alarms (e.g. send those to one channel, manual alarms to another)
  • dtedge_org_id=<uuid> — route per-org to per-team channels
  • severity=... — route by severity to different escalation paths

The dtedge_org_id is opaque to the receiver — it's a UUID, not a friendly name. Build a tiny mapping in your Alertmanager config (or a downstream router) if you need readable org names in the page text.

See also​