Edge instances
An "edge instance" is a row in DT Edge Platform's database that represents one Kubernetes cluster you've registered. The row holds enough state to manage that cluster from DT Edge Platform — but not the cluster itself.
What's in the row
- Identity — name, description, optional categories,
telemetry_id(a UUID that ties this edge's metrics + logs to this row) - Ownership —
org_id(NULL for public edges) - kubeconfig — encrypted with the master AES-GCM key
- Status —
active/in_progress/failed - Health —
healthy/degraded/unreachable/unknown, plus the last probe timestamp + error message - Provisioner link —
provisioner_cluster_idif this edge was created via the upstream Provisioner
What's not in the row: per-edge Prometheus / OpenSearch /
Grafana URLs. Observability is centrally configured at the
system level; per-edge isolation happens at the metric label
level via the telemetry_id.
Two creation paths
Upload kubeconfig (always available)
The cluster already exists. You paste a kubeconfig; DT Edge Platform
encrypts + stores it; runs a probe to confirm reachability;
flips status to active.
This is the path you'll use 99% of the time, including for clusters you provisioned by hand, clusters from cloud k8s services (GKE / EKS / AKS), and edge clusters provisioned by some other tool.
Provision new cluster (when configured)
DT Edge Platform calls the upstream Provisioner service to spin up a
new k3s cluster from a list of VM IPs. Provisioner runs
ansible over SSH; when it's done, it webhooks DT Edge Platform with
cluster.ready and DT Edge Platform fetches the kubeconfig.
This path only surfaces when admin has configured provisioner_url
- admin key in system settings. Without that, the path doesn't appear in the UI.
What the kubeconfig encryption gives you
- At rest in the DB — the column is
kubeconfig_enc bytea, AES-GCM ciphertext. The plaintext kubeconfig isn't readable via SQL. - In memory — decrypted only inside the request handler that needs it (helm install, health probe, etc.), used immediately, not persisted to disk.
- In the UI — never. There's no "show me the kubeconfig" button; users see metadata about the edge but not the underlying credentials.
- In the audit log — sanitised. The middleware that records
every mutation strips
kubeconfigfields before storage.
If the master key (ENCRYPTION_KEY env) is rotated, every
encrypted column needs re-encryption — see backend operations
docs. Don't change it casually.
telemetry_id — the stable identifier
Every edge gets a telemetry_id UUID minted in BeforeCreate.
This UUID is:
- The label written by remote-write configs on each edge cluster's Prometheus, so DT Edge Platform's central Prometheus can scope queries
- The field added by fluent-bit on each edge cluster, so DT Edge Platform's central OpenSearch can scope log queries
- The matcher backend injects into PromQL alarm queries, so an alarm rule for org X scopes to org X's edges
Why a UUID and not the name? Because the name is editable.
If you rename staging-cluster to staging-cluster-deprecated,
you don't want to break every Grafana query and alarm rule
pinned to the old name. The UUID never changes.
Editing telemetry_id is allowed via the admin drawer with a
prominent warning — only do it during a controlled migration.
The "what breaks" footprint is wide.
Public vs org-scoped
org_id IS NULL makes an edge public — visible to every
org as an install target.
- Public edges are admin-published. Only super-admins can
create them, and only when the active license carries
MaxPublicEdges > 0. Useful for shared infrastructure (a dev/test cluster every team can deploy against). - Org-scoped edges (the common case) belong to one org. Only members of that org see them.
- Ownership of releases is tracked separately:
helm_release_snapshots.installed_by_org_idrecords which org installed each release. On a public edge, Org A's releases are visible to Org B (so they don't accidentally collide on names) but only Org A can mutate them.
Health probes
A periodic background job (default every 30 minutes) probes each edge's apiserver:
- Calls
/version+/api/v1/nodes - Records: response time, kube version, total node count, ready node count, error message if any
- Updates
health_status+last_checked_at+last_health_error
Probes don't need the user's permission — they run as DT Edge Platform.
Failures stamp the edge as unreachable, surfaced as a red
badge in the list.
A failed probe doesn't break anything else. The edge stays in the database; the next probe (or a manual recheck) can flip it back. Operations against an unreachable edge fail loudly when you try them, with a more useful error than "timeout".
When you remove an edge
- The DB row is soft-deleted (
deleted_atset) - The edge stops appearing in install dropdowns and lists
- Helm releases on the actual cluster are not touched.
DT Edge Platform just stops managing them. Use
helm uninstalldirectly on the cluster if you want to clean those up. - Alarm rules tied to this edge become orphaned (their
telemetry_idmatchers won't match anything live). Review alarms after removing edges. - The fleet capacity snapshot reflects the removal on the next aggregator pass (~15 min) — see How-to: handle fleet quota exceeded.
See also
- How-to: manage edge instances
- Reference: Edge Instances page
- Backend docs: Concept → Edge instances