Skip to main content

Handle "fleet quota exceeded"

The license caps total fleet capacity (nodes / CPU cores / memory / GPUs). When an install or edge-create would push you over a cap, the API returns 403 FLEET_QUOTA_EXCEEDED and the UI shows a toast.

What you'll see​

  • A red toast: "Fleet quota exceeded — would add 32 cores, headroom 12" (numbers vary)
  • A banner on Admin → License when fleet usage is over 90%
  • Backend response includes which limit failed and by how much

Why this happens​

DT Edge Platform aggregates capacity across all managed edges every ~15 minutes into a snapshot. When you try to add an edge or install something that would grow the fleet, a preflight check overlays a fresh live read of the target edge over the snapshot and runs the comparison.

If the result would breach MaxTotalNodes, MaxTotalCPUCores, MaxTotalMemoryGi, or MaxTotalGPUs, the operation is refused before anything is changed.

Three paths to fix​

Path 1 — Free up capacity (fastest)​

Remove a managed edge or scale down nodes on an existing one.

  • Remove an edge instance → the next aggregator pass (≤15 min) drops it from the totals
  • Force-refresh the snapshot now if you don't want to wait: Admin → License → Refresh fleet usage (super-admin only) — triggers an immediate aggregation

If the over-limit dimension is GPUs, decommissioning a GPU node on an existing edge releases that allocation as soon as the snapshot refreshes.

Path 2 — Wait for the snapshot to catch up​

If you've already removed capacity but the snapshot still reports the old totals, wait up to 15 minutes for the next periodic aggregation, or use the manual refresh above.

Path 3 — Re-issue with bigger caps​

The customer's contract changed and they need more headroom. The operator who issues licenses re-mints with new --max-total-* values:

./dtedge-license-issue issue \
--customer "Acme Corp" \
--product dtedge \
--bind sha256:abc... \
--expires <unchanged> \
--tier <unchanged> \
--max-total-nodes 200 \
--max-total-cpu-cores 1600 \
--max-total-memory-gi 6400 \
--max-total-gpus 16 \
...

Customer uploads through Admin → License (see Upload a license). Within seconds the new caps apply; the install retries succeed.

Why an unreachable edge still counts​

If an edge is currently unreachable but still registered, it stays counted at its last known capacity. An edge going offline doesn't free up phantom budget — capacity is a contractual concept, the edge is still allocated to you.

If the edge is permanently gone (cluster destroyed, network permanently severed), delete the instance row to release the allocation.

What's measured​

For each managed edge:

  • Schedulable nodes — control-plane-only nodes are excluded
  • CPU cores — sum of node.status.allocatable.cpu (in milli-cores internally, displayed as cores)
  • Memory GiB — sum of node.status.allocatable.memory
  • GPUs — sum of node.status.allocatable["nvidia.com/gpu"]

What's not measured: pods, releases, disk, pending nodes, control-plane-only nodes.

Why my install fails when totals look fine​

The gate fires on delta, not absolute total. Even if total_cpu_cores < max, an install on edge X that would push total + delta > max is rejected. The error response includes delta_if_added so the toast can say which capacity dimension the operation would breach and by how much.

See also​