License states
DT Edge Platform ships license-gated. The API has four possible states based on what license is loaded; everything operational flows from these four.
The four states
| Status | What it means | API behaviour |
|---|---|---|
active | Signed, in date, fingerprint matches the cluster | Serves normally |
grace | Past exp but inside grace_days | Reads pass; mutations 503 with X-License-Status: grace |
missing | No license row in the DB | All mutations 503; only the bootstrap whitelist serves |
invalid | Verification failed (signature, fingerprint, schema, clock rollback, expired past grace) | Mutations 503; UI surfaces the specific reason |
Why so few
The state machine is intentionally narrow. There's no separate
"signature wrong" vs "fingerprint wrong" runtime state because
the operational answer is the same: refuse to serve until the
operator fixes it. The specific reason is exposed in
invalid_reason for the support flow.
What "the bootstrap whitelist" lets through
A fresh install has no license loaded. The operator needs to be able to upload one. So a small set of API paths bypass the gate:
/api/capabilities— so the UI can render the License Required landing page/api/auth/login+/api/auth/login/2fa— admin login flow/api/auth/registration-status— for the register page to decide what to show/api/admin/license/...— upload + fingerprint endpoints (still gated on super-admin)/healthz,/readyz— health probes (live outside/api, naturally excluded)
Every other path returns 503 until a license loads. This includes login for non-super-admin users — they can authenticate but can't reach any feature surface.
Grace, in practice
Grace is the window between exp and exp + grace_days (the
license's grace days field, typically 14 or 30).
In grace:
- Reads pass. Dashboards, lists, audit log — all work.
- Mutations don't. Install / upgrade / uninstall, alarm
create/edit, edge instance changes — all 503 with
X-License-Status: grace.
The yellow banner at the top of the UI says "License expires in N days; renew before grace ends." The intent is to make sure nobody is surprised — the lights are on, but you can't change anything.
Best practice: renew before grace starts. The audit log on
Admin → License has every customer's expires_at; any
calendar-or-monitoring tool can build "renewing in 60 days"
alerts on top of it.
Why the cluster fingerprint
Every license is bound to a specific cluster fingerprint:
fingerprint = sha256( kube_system_uid ‖ "\0" ‖ installation_id )
kube_system_uid— UID of thekube-systemnamespace; stable for the cluster's lifetimeinstallation_id— UUID stored in DT Edge Platform's DB on first boot
The license carries this fingerprint in
Binding.ClusterFingerprint. On every load and periodic recheck,
the verifier recomputes the local fingerprint and rejects the
license if they differ.
This means: a license issued for cluster A can't run cluster B. And if cluster A is wiped + re-created (kube-system UID changes) or DT Edge Platform's DB is reset (installation_id regenerates), the license stops verifying — you need a new one bound to the new fingerprint.
DT Edge Platform and Provisioner running on the same cluster have two
different fingerprints because they have different
installation_ids. Each product needs its own license bound to
its own fingerprint.
Schema versioning
const SchemaVersion = 1
The verifier rejects any license claiming a higher schema than this build supports. Deliberate asymmetry — a binary running an old DT Edge Platform image cannot interpret a license minted for behaviour the binary doesn't have. New verifiers always accept old licenses; old verifiers refuse new ones.
So: upgrade the DT Edge Platform image first, then load the schema-bumped license. Reverse order fails until you upgrade.
What this all isn't
- Online revocation. No CRL, no phone-home. The kill switch for a customer is to wait for expiry + not renew, OR to rotate the embedded public signing key (which retires every license signed by the old key in lockstep).
- Per-feature gates. Today the gate is signature + product +
expiry + fingerprint, then
Limitsfor capacity. There's noFeatures.X = true/falsemechanism. - Time integrity. We trust the host clock. A customer with
root who sets
dateback can keep an expired license alive. The monotonic anchor catches the obvious case — clock back further than monotonic uptime — but it's not a defence against a determined on-host adversary.
Multi-replica HA
Every API replica subscribes to a Redis pub/sub channel
(dtedge:license:reloaded). The instant any replica writes a
new license (admin upload), every replica drops its cached
snapshot and re-reads. There's never an "old license still
cached on replica 7" race.
Worker and scheduler do the same on boot, plus a periodic 60-second recheck to catch:
- Expiry crossing the wall during runtime (active → grace → invalid past grace)
- Clock rollback attempts
- DB-side license row edits applied via psql (e.g. operator deletes the row to force re-bootstrap)
See also
- Tenancy — the other axis the license drives
- How-to: upload a license — the admin flow
- How-to: handle fleet quota exceeded
- Backend docs: Concept → License