Cluster Onboarding
The Add cluster wizard (/clusters/add, reached from the Clusters page) takes a Kubernetes cluster from nothing to two separately-provable outcomes. It records the attempt, mints the enrollment token, prints the command you run yourself, and then measures what happened.
Nothing on this page is applied to your cluster by Console. Console has no inbound path to a customer's API server at all: the operator runs helm out of band on their own credentials, the collector dials out to NATS, and a live kubectl request is published on a NATS subject that the in-cluster agent answers. That is why there is no "test connection" button — the equivalent is a round-trip through the agent.
The two tracks, and why they are proved separately
| Track | Question it answers | Gates |
|---|---|---|
| Inventory | Does this cluster appear in Console at all — nodes, namespaces, workloads? | 4 |
| Live access | Can a person who should be able to run pc kube against this cluster actually do it? | 5 |
They are independent, and they are deliberately not collapsed into one progress bar. Whoever runs helm is usually not whoever owns RBAC — the first track is a release change inside the cluster, the second is a role change in Console plus a chart flag — and either outcome is useful without the other. A cluster with inventory and no pc kube is fully onboarded for monitoring, compliance, ArgoCD drift and the Clusters pages; that is the common case, not a half-finished install.
A client's own engineer can complete the inventory track alone. For them the access track is read-only and ends in a request naming the exact two asks, because granting access needs roles:write (which could widen access beyond this cluster) and a Helm upgrade. A button that appeared to grant access and then failed server-side would be worse than no button.
Permissions
There is no clusters:* permission and onboarding introduces none.
| Action | Permission |
|---|---|
| Record an attempt, mint its install token | hosts:write on the client and on the target environment |
| View attempts and recompute status | hosts:read on the attempt's client |
| Run the live round-trip probe | kubernetes:read on the cluster's client and environment — the same gate the kube proxy applies |
| Abandon an attempt | hosts:write on the attempt's client and environment |
Author a kube_spec | roles:write, in the role editor — the wizard never authors one |
| Rotate a collector's NATS credentials | agents:write (see Rotating Agent Credentials) |
1. Record the attempt
Pick the client, the environment and a display name. The slug is derived from the name and shown on screen: lowercased, hyphen-joined, 1–63 characters (^[a-z0-9](?:[a-z0-9-]{0,61}[a-z0-9])?$). It is half the attempt's identity — the first inventory sync is matched against (environment_id, cluster_slug) — and it cannot be changed afterwards.
Recording is idempotent per (environment, slug): a second attempt at the same cluster resumes the first rather than forking it, so a double-click, a reload, or a return a week later lands on the same row with its measured progress intact.
The attempt mints one install token:
- 24-hour expiry, the same as the Install Tokens page — the wizard is a different door onto the same credential.
max_uses25. Not single-use, because the collector runs as a StatefulSet and every replica enrolls separately — one use would admit one replica and strand the rest. Not unlimited either: this token is pasted into a terminal, a Slack thread or a CI log, and an unbounded one would let any holder enroll agents of any type into that environment for the whole 24 hours (install tokens are single-use by default and unlimited is an explicit opt-in, for exactly that reason). 25 covers an HA StatefulSet plus restarts — a replica's identity persists in its PVC, so re-enrolment is rare — and keeps the blast radius finite. The count is printed next to the plaintext, so you read what you are pasting. The token is also pinned to one environment.- The plaintext is shown once, in the response that minted it. Only a hash is stored, so a resumed attempt prints a
<install-token>placeholder rather than a blank you might paste.
A new token is minted again only to clear a dead end — no token recorded, or an expired token that nothing ever enrolled with. A live token is never replaced (you may be holding the only copy), and once a collector has enrolled the token's job is finished: agents.enrolled_by is the only attribution between a replica and an attempt, so overwriting it would make a working cluster's replicas invisible to the gates.
2. Install, out of band
The wizard prints the command for the method you pick. Both forms set the cluster slug explicitly rather than letting the agent derive it: the agent's own slugify has no 63-character bound, so on a long display name it would enroll under a slug the attempt is not recorded against, and the inventory gate would wait with nothing on screen explaining why.
- Helm (recommended) —
helm upgrade --install, withclusterAgent.clusterName,clusterAgent.clusterSlug,backendUrlandinstallToken.value. See Helm Chart. - Manifests — the single-replica, no-persistence quickstart from
infra/k8s/cluster-agent/, followed bykubectl set envfor the values the attempt is recorded under. The manifest has noPROXIMA_CLUSTER_SLUGkey at all, which is the other reason the wizard sets env rather than telling you to hand-editdeployment.yaml. See Cluster Agent → Deployment.
Every value is rendered inside a single-quoted shell argument. That is a boundary, not formatting: you run this command against your own cluster with cluster-admin credentials, and a resumed attempt renders a display name a teammate may have typed.
What the install grants inside your cluster: one read-only ClusterRole — get, list, watch — which does not name Secrets or ConfigMaps. That is a property of the manifest, not a setting. Live pc kube access is a separate thing and ships off; see below.
Node metrics are optional and are not a track: a cluster with no node agent is fully onboarded. Tick the box to add the DaemonSet, and tick "these nodes already run a host agent" if they do — that adds nodeAgent.hostDataDir. Both agents default to /var/lib/proxima-agent, and since v0.7.2 whichever starts second refuses to start rather than adopting the other's identity, so the flag pre-empts a startup failure rather than tidying a layout. See Node Agent on Kubernetes.
3. Verification: the five gate states
Each gate is evaluated server-side and carries the sentence that describes it. The UI renders those sentences verbatim and derives no verdict of its own — a UI-derived verdict eventually lies (a green "can be paged" badge in this product once asserted an outcome from two of about twenty-one gates).
| State | Meaning |
|---|---|
| Passed | Measured true. |
| Waiting | Not yet true, and not yet late. A first full inventory sync has a budget; silence inside it is healthy and is not a fault. Rendering this as failure would cry wolf on every healthy install, which is why it is toned as neutral, not as danger. |
| Blocked | Could not run because a prior gate failed. Nothing was attempted, and the sentence says so — a red that implies your cluster refused something it was never asked is worse than no red. |
| Warning | True, but not yet usable — or "this could not be read". A group allowlisted in one of three layers is a warning, never a pass. |
| Failed | Measured false, with the action that fixes it. |
Two rules hold across every gate:
- No gate reports a pass on configuration it has not observed. Where Console cannot see a layer, the gate says which layer and reports a warning.
unknownis not a sixth state. It is a warning whose sentence names what could not be read.
A track is done only when every one of its gates passed. A warning is exactly the state where something is true but not yet usable, so a track with one is not done.
The inventory track
| # | Gate | What it reads |
|---|---|---|
| 1 | Install token issued | An unexpired token for this attempt, or an expired one that something already enrolled with. A token that lapses unused is the commonest reason an install silently never appears. An expired token whose collector enrolled in time passes — enrolment does not lapse with the token that bought it. |
| 2 | Agent enrolled as a collector | An agent row whose enrolled_by is this attempt's token. Exact, not a guess at pod naming. |
| 3 | Heartbeat fresh | Last heartbeat inside the fleet's 5-minute staleness window. With more than one replica only the leader heartbeats, so a quiet standby is correct, not a fault. Silence is bounded: a collector that enrolled and then died is reported failed, never "waiting" forever. |
| 4 | First inventory sync landed | A clusters row for (environment_id, slug) published by this attempt's collector. |
Collectors create no host record
A collector publishes inventory straight to the Kubernetes worker; there is no hosts table entry. So a cluster never appears under Hosts, and the host wizard's way of proving success — diffing the host list before and after an install — is structurally blind to it.
Gates 2–4 therefore read the agent fleet, the clusters row and sync recency instead, and the copy on gate 2 says so out loud so you do not go looking in the wrong list.
The waiting budget
Gate 4's budget is twice PROXIMA_K8S_SYNC_INTERVAL — ten minutes at the 5-minute default. It is doubled because the first sync begins after the pod finishes starting, not at the instant it enrolls. The waiting sentence and the failure sentence quote the same number, so the gate cannot argue with itself on screen.
PROXIMA_K8S_SYNC_INTERVALThe cluster agent gets its value from Helm; the backend gets its own from its environment, and the backend never asks the agent what cadence it is on. Nothing reconciles them, and the gate copy quotes the backend's copy. Raise it for the agent alone and a healthy cluster reads as overdue. See the env-vars entry and Cluster Agent → Configuration.
If a cluster with this name already exists from another install, gate 4 says so rather than claiming this attempt's collector published it — and turns failed, not passed, once the budget has elapsed.
The live access track
| # | Gate | What it reads |
|---|---|---|
| 1 | Someone holds kubernetes:read for this client | A role carrying the permission, assigned to an active user. It names the role. This pages nobody and grants nothing; it only means the permission exists to be used. |
| 2 | A role's kube_spec matches this cluster | A kube_spec you hold that resolves a group set for this cluster's labels — resolution runs on your roles, while gate 1 above is client-wide, so every sentence here says whose access it measured. Resolution fails closed, so without it Console refuses the request here — nothing is sent to your cluster. Waiting until the first inventory sync, because a spec selects clusters by their labels and this cluster has none yet. |
| 3 | Something answers on kube.in.> | Whether a probe got a responder at all. Two causes of silence, named and not chosen between. |
| 4 | Requested groups accepted by the backend | Whether every group the spec asks for appears in PROXIMA_KUBE_ALLOWED_GROUPS — one allowlist of three. It reaches passed only once gate 5 has actually passed, and says how old that round-trip is; a recorded success beside a blocked round-trip is not a pass. Waiting before this cluster appears: no group has been requested, so nothing has been checked. |
| 5 | Live round-trip | What a real GET /version through the agent did, as the caller. Blocked while gate 1 or 2 is unresolved, or gate 4 has failed — a probe offered while nobody holds the permission, or while this backend will refuse the group the spec asks for, cannot succeed. A warning on gate 4 is deliberately not a blocker: that is its normal state before any round-trip, and the round-trip is the only thing that can clear it, so treating it as one would deadlock the two gates. While blocked, the copy states that nothing has been sent to your cluster. |
kubeAccess.enabled ships false
The chart's kubeAccess.enabled defaults to false, so a default install serves no kube proxy at all. This is the commonest reason access onboarding looks broken, and it is not an RBAC problem: nothing is listening. Setting it means --set kubeAccess.enabled=true with the Helm CLI, or the values of whatever manages the release — an Argo CD Application, for instance.
The chart's own S10 constraint. The read tier (proxima:kube-readonly) grants cluster-wide read, so on a multi-tenant cluster enabling kube access leaks data across Proxima clients. Enable it only on a cluster dedicated to a single client, until v2 namespace scoping lands. See Kubernetes Access → Security.
The three-layer impersonation allowlist
An impersonation group must be allowed in three places, and Console can observe exactly one of them:
| # | Layer | Where | Observable from Console? |
|---|---|---|---|
| 1 | In-cluster group bindings — what the group is actually granted on resources | ClusterRole + ClusterRoleBinding in your cluster (the chart installs them for proxima:kube-readonly / proxima:kube-exec when kubeAccess.enabled=true; a custom group is yours to bind) | No. Nothing reports them today. |
| 2 | The apiserver's resourceNames backstop — which groups the collector may ever assert | Helm kubeAccess.impersonateGroups, on the collector's impersonate ClusterRole | No. |
| 3 | What a kube_spec may request | Backend PROXIMA_KUBE_ALLOWED_GROUPS | Yes — this is gate 4. |
So gate 4 reports a warning by default: accepted by this backend, with the other two layers unread. It never claims them. (Custom impersonation groups covers adding a group to layers 2 and 3.)
What the round-trip proves — and what it does not
Gate 5 sends one GET /version through the agent, over the same NATS relay and with the same identity headers a real pc kube request uses, carrying Impersonate-Group for the group the kube_spec resolved. A success establishes, in one shot, what no amount of configuration inspection could:
- the agent is alive and serving the kube proxy;
- its NATS credential carries the
kube.in.>grant; - the caller's RBAC in Console permitted the call;
- the impersonation was permitted — so the apiserver's
resourceNamesbackstop (layer 2) accepts that group. A 403 there is something Console could never have predicted from configuration it cannot read.
It proves nothing about layer 1. /version is readable by any authenticated identity, so a success says the impersonation was allowed and says nothing about whether those groups are granted anything on real resources. The in-cluster bindings stay unproved until someone runs a real kubectl get, and gate 4 says so rather than banking a green it has not earned.
The first version of this design said a successful round-trip settled layers 1 and 2, and the gate copy quoting it inherited the error. It was corrected during the build. The honest position is the one above: the probe settles impersonation; the bindings remain unobserved.
A 403 that came back has two possible layers — the apiserver's impersonation backstop, and Kubernetes RBAC — and the gate names both without choosing. The agent's own enforcement is excluded on this path: /version is a non-resource URL that every kube_spec allows, which is pinned by a test against the agent's kuberbac package so a future agent that classifies it differently fails the test rather than quietly invalidating this sentence.
The kube.in.> grant caveat
When gate 3 reports that nothing answered, two causes look identical from the backend and the gate refuses to guess between them:
kubeAccess.enabledisfalse— the collector is serving no proxy.- The collector's NATS credential predates the
kube.in.>grant. An agent enrolled before that grant entered the enrollment template keeps its old JWT, its relay subscription is denied, and from this side that is indistinguishable from the first cause.
The pod's own log is what distinguishes them — a collector with the proxy enabled logs kube access proxy enabled once it is connected. The gate names both causes and the remedy for each, which costs it nothing: neither sentence claims which cause is in play.
For the second cause, an agent that is online can be renewed in place from Console: Rotate credentials on the Agents page, or POST /api/v1/fleet/agents/{agentID}/rotate-credentials (agents:write). It reaches a collector that has no kube access yet — which is the whole point here — because the cluster agent always runs its command dispatcher and registers the renew_credentials handler unconditionally; only the kube relay is gated on the flag. An offline agent cannot be — the command is a NATS request/reply and answers 503 — so it picks up the current template on its own next renewal, which happens when under 50% of its 7-day TTL remains. There is no way to shorten that wait for an agent Console cannot reach. See Rotating Agent Credentials, whose trigger table names this exact grant.
The probe is evidence, never a gate
The probe is the only part of onboarding that causes traffic inside a customer's cluster, so it is bounded on every side:
| Constraint | Value |
|---|---|
| Path | GET /version, fixed. Not a general proxy; the general proxy exists and is separately gated. |
| Authority | The caller's. No system context anywhere on this path — a probe that ran as the system would prove nothing about whether a person can reach the cluster. |
| Timeout | 5 seconds, no retry. One request tells the operator everything a second would. |
| Rate limit | One probe per 15 seconds per attempt (429, with Retry-After) — per attempt, not per caller, because the resource being protected is somebody else's cluster. |
| Single-flight | A probe already running is refused (409), never queued. |
| Outcomes | Four, kept distinct: ok, no responder, forbidden (something answered and did not succeed), timeout. |
And it decides nothing. A recorded outcome is read by the access-track gates and by nothing else: the kube proxy re-checks live RBAC on every real request and never consults a probe result, so a failed probe cannot stop a real pc kube request from being attempted, and a red access track takes nothing away from a cluster whose inventory track is done. Console's own pre-dial refusals — no permission, no resolved group set, a rate limit, a probe in flight — are never recorded as probe outcomes either: "forbidden" asserts that something in the cluster answered, and blaming a customer's cluster for our own refusal would be a lie in the one place the page exists to tell the truth.
A probe answers 404 until the cluster has appeared in Console, because a probe needs a cluster to address.
This is the same posture the product takes everywhere a probe exists: on the voice side, a path whose health probe is down is still dialed — whether a page is attempted depends on the box's ARI event loop, not on the probe, because the probe is evidence for a human and never a gate that could stop a page being placed (Voice trunk health). Nothing about cluster onboarding touches paging; the rule is quoted because it is the same rule, and a measurement that silently becomes a gate is the failure mode both pages are written against.
Every probe is audited at initiation
Each probe writes one kube_probe row against resource_type = kube_cluster before anything is dialled — the same posture the kube proxy takes for a proxied request. The row carries the cluster, client and environment, the onboarding attempt, the requested path, and the impersonated identity (the user the cluster's own audit log will show, not the free-text email of a service account).
It deliberately carries no outcome. At that moment there is none, and the division is the point: this row answers who reached into this cluster and when, while the attempt's status cache answers what came back. Moving the write after the probe to add an outcome would lose the record exactly when it matters most — the probe runs on the request's context, so a caller who abandons mid-probe cancels it, and an abandoned probe still leaves a record. See Audit Logging.
Resuming and abandoning
The Clusters page shows attempts that are started and not finished, and reopening one lands on its live state rather than at step one. Finished means the inventory track is proved — not both tracks. Access is optional by design (see the two tracks), and access_done_at is stamped only when all five access gates pass, which needs kubeAccess.enabled=true; waiting for it would keep every inventory-only cluster in the list forever under a banner saying it was never finished. The attempt row itself is never deleted and stays resumable. Status is always recomputed on read; the stored last_status is a cache for those list rows and is labelled as of its last_checked_at, never as the present. A persisted status would drift from reality the moment someone uninstalled the chart.
Abandoning an attempt stamps it abandoned and touches nothing else. It is a discarded intent, never a removed cluster: someone who abandons an attempt after a successful install loses the wizard entry and keeps their cluster, its inventory and its agent.
API
| Method | Path | Gate |
|---|---|---|
POST | /api/v1/clients/{id}/cluster-onboarding | hosts:write on the client + environment containment |
GET | /api/v1/clients/{id}/cluster-onboarding | hosts:read on the client |
GET | /api/v1/cluster-onboarding/{onboardingID}/status | hosts:read on the row's client |
POST | /api/v1/cluster-onboarding/{onboardingID}/probe | kubernetes:read on the cluster's client + environment, as the caller |
DELETE | /api/v1/cluster-onboarding/{onboardingID} | hosts:write on the row's client + environment |
Troubleshooting
| Symptom | Where to look |
|---|---|
| Nothing ever enrolls | Gate 1. An expired, never-used token is the commonest cause; the gate offers a new one. |
| The cluster is not under Hosts | It never will be — collectors create no host record. Look at the Clusters page and the Agents fleet. |
| "No inventory yet" for minutes | Normal inside twice the sync interval. Past it, the gate turns failed and points at the collector's log for Kubernetes API permission errors. |
| Enrolled, heartbeat fresh, no inventory, collector log clean | Check that the release really points at this cluster slug — a collector publishing under a different slug is a different cluster to Console. |
| Only one of several replicas heartbeats | Correct. Only the leader collects and heartbeats; see High availability. |
| Access gate 3 says nothing answered | kubeAccess.enabled, or a credential predating the kube.in.> grant. The pod log says which; the gate names the fix for both — set the flag, or rotate the agent's credentials. |
| Access gate 4 stays a warning | That is its normal state until a successful round-trip. It is not a blocker. |
| Round-trip returns 403 | The apiserver's impersonation backstop or Kubernetes RBAC. Nothing in Console can tell you which. |
See also
- Helm Chart — the recommended install, all values, GitOps
- Cluster Agent — what the collector collects, its ClusterRole, HA and storage
- Node Agent on Kubernetes — the optional per-node metrics DaemonSet
- Kubernetes Access (
pc kube) · Roles &kube_spec· Security model - Rotating Agent Credentials — the
kube.in.>grant, and when rotation can and cannot help - Kubernetes Clusters — what the pages show once inventory lands