Skip to main content

Edge instances

An "edge instance" is a row in DT Edge Platform's database that represents one Kubernetes cluster you've registered. The row holds enough state to manage that cluster from DT Edge Platform — but not the cluster itself.

What's in the row​

  • Identity — name, description, optional categories, telemetry_id (a UUID that ties this edge's metrics + logs to this row)
  • Ownership — org_id (NULL for public edges)
  • kubeconfig — encrypted with the master AES-GCM key
  • Status — active / in_progress / failed
  • Health — healthy / degraded / unreachable / unknown, plus the last probe timestamp + error message
  • Provisioner link — provisioner_cluster_id if this edge was created via the upstream Provisioner

What's not in the row: per-edge Prometheus / OpenSearch / Grafana URLs. Observability is centrally configured at the system level; per-edge isolation happens at the metric label level via the telemetry_id.

Two creation paths​

Upload kubeconfig (always available)​

The cluster already exists. You paste a kubeconfig; DT Edge Platform encrypts + stores it; runs a probe to confirm reachability; flips status to active.

This is the path you'll use 99% of the time, including for clusters you provisioned by hand, clusters from cloud k8s services (GKE / EKS / AKS), and edge clusters provisioned by some other tool.

Provision new cluster (when configured)​

DT Edge Platform calls the upstream Provisioner service to spin up a new k3s cluster from a list of VM IPs. Provisioner runs ansible over SSH; when it's done, it webhooks DT Edge Platform with cluster.ready and DT Edge Platform fetches the kubeconfig.

This path only surfaces when admin has configured provisioner_url

  • admin key in system settings. Without that, the path doesn't appear in the UI.

What the kubeconfig encryption gives you​

  • At rest in the DB — the column is kubeconfig_enc bytea, AES-GCM ciphertext. The plaintext kubeconfig isn't readable via SQL.
  • In memory — decrypted only inside the request handler that needs it (helm install, health probe, etc.), used immediately, not persisted to disk.
  • In the UI — never. There's no "show me the kubeconfig" button; users see metadata about the edge but not the underlying credentials.
  • In the audit log — sanitised. The middleware that records every mutation strips kubeconfig fields before storage.

If the master key (ENCRYPTION_KEY env) is rotated, every encrypted column needs re-encryption — see backend operations docs. Don't change it casually.

telemetry_id — the stable identifier​

Every edge gets a telemetry_id UUID minted in BeforeCreate. This UUID is:

  • The label written by remote-write configs on each edge cluster's Prometheus, so DT Edge Platform's central Prometheus can scope queries
  • The field added by fluent-bit on each edge cluster, so DT Edge Platform's central OpenSearch can scope log queries
  • The matcher backend injects into PromQL alarm queries, so an alarm rule for org X scopes to org X's edges

Why a UUID and not the name? Because the name is editable. If you rename staging-cluster to staging-cluster-deprecated, you don't want to break every Grafana query and alarm rule pinned to the old name. The UUID never changes.

Editing telemetry_id is allowed via the admin drawer with a prominent warning — only do it during a controlled migration. The "what breaks" footprint is wide.

Public vs org-scoped​

org_id IS NULL makes an edge public — visible to every org as an install target.

  • Public edges are admin-published. Only super-admins can create them, and only when the active license carries MaxPublicEdges > 0. Useful for shared infrastructure (a dev/test cluster every team can deploy against).
  • Org-scoped edges (the common case) belong to one org. Only members of that org see them.
  • Ownership of releases is tracked separately: helm_release_snapshots.installed_by_org_id records which org installed each release. On a public edge, Org A's releases are visible to Org B (so they don't accidentally collide on names) but only Org A can mutate them.

Health probes​

A periodic background job (default every 30 minutes) probes each edge's apiserver:

  • Calls /version + /api/v1/nodes
  • Records: response time, kube version, total node count, ready node count, error message if any
  • Updates health_status + last_checked_at + last_health_error

Probes don't need the user's permission — they run as DT Edge Platform. Failures stamp the edge as unreachable, surfaced as a red badge in the list.

A failed probe doesn't break anything else. The edge stays in the database; the next probe (or a manual recheck) can flip it back. Operations against an unreachable edge fail loudly when you try them, with a more useful error than "timeout".

When you remove an edge​

  • The DB row is soft-deleted (deleted_at set)
  • The edge stops appearing in install dropdowns and lists
  • Helm releases on the actual cluster are not touched. DT Edge Platform just stops managing them. Use helm uninstall directly on the cluster if you want to clean those up.
  • Alarm rules tied to this edge become orphaned (their telemetry_id matchers won't match anything live). Review alarms after removing edges.
  • The fleet capacity snapshot reflects the removal on the next aggregator pass (~15 min) — see How-to: handle fleet quota exceeded.

See also​