Cluster Agent
The cluster agent is a standalone binary that runs inside Kubernetes clusters, queries the K8s API, and publishes inventory data to Proxima Console via NATS.
Architecture
- Workload: Kubernetes StatefulSet in the
proxima-systemnamespace — one replica by default, with optional leader-elected high availability - Data source: K8s API via
client-go(in-cluster service account) - Transport: NATS (same as the node agent, using JetStream)
- Collection: Full sync every 5 minutes (configurable via
PROXIMA_K8S_SYNC_INTERVAL— read by the backend too, and the two must match; see Configuration) - Enrollment: Self-enrolls with the backend via install token on first startup as
agent_type=collector - Identity: Each replica enrolls under its pod name (a stable StatefulSet identity). Enrollment binds a NATS identity once; the backend refuses to re-bind it (identity-takeover protection), so each pod keeps a PVC-backed
state.jsonto preserve its identity across restarts — see Storage - No host record: Collector agents do not create host records. Inventory is published directly to the Kubernetes worker without a corresponding
hoststable entry. - Heartbeat: Publishes periodic heartbeats for liveness detection (only the leader heartbeats)
Collected Resources
| Resource | K8s API | Data |
|---|---|---|
| Cluster info | Discovery API | Version, platform, cloud provider |
| Nodes | CoreV1 | Status, capacity, allocatable, addresses |
| Namespaces | CoreV1 | Status, labels, resource counts |
| Workloads | AppsV1 | Deployments, StatefulSets, DaemonSets |
| Argo Rollouts | Dynamic client (argoproj.io/v1alpha1) | Rollouts, as workload kind Rollout — only where the CRD is installed (see Argo Rollouts) |
| ArgoCD Applications | Dynamic client (argoproj.io/v1alpha1) | Identity, destination, sync/health, revisions, repo URL (credentials stripped), workload resource rows and resource totals — never Helm values, parameters or plugin env (see ArgoCD Applications) |
| Pods | CoreV1 | Phase, containers, restarts, resources, owner |
| Services | CoreV1 | Type, ports, selector, cluster IP, load balancer |
| Ingresses | NetworkingV1 | Rules, TLS, ingress class, load balancer |
| PVCs | CoreV1 | Phase, storage class, capacity, access modes |
| Events | CoreV1 | Type, reason, involved object, count, timestamps |
| Jobs/CronJobs | BatchV1 | Status, completions, schedule, containers |
| HPAs | AutoscalingV2 | Target, min/max/current replicas, metrics |
| Network Policies | NetworkingV1 | Pod selector, ingress/egress rules, policy types |
| Resource Quotas | CoreV1 | Hard limits, current usage per namespace |
| Endpoint Slices | DiscoveryV1 | Addresses, ports, service association |
Not collected: Secrets and ConfigMaps are deliberately excluded. The ClusterRole does not grant access to these resources.
Argo Rollouts
Argo Rollouts (argoproj.io/v1alpha1, kind Rollout) are collected as a fourth workload kind next to Deployments, StatefulSets and DaemonSets. Each collection cycle the agent first asks API discovery whether the cluster serves argoproj.io/v1alpha1 rollouts:
- CRD not installed (the usual case) — Rollouts are skipped silently. Nothing is logged and nothing else changes. A cluster that has Argo CD's
argoproj.io/v1alpha1group but not Argo Rollouts counts as not installed. - Access denied (403) — the ClusterRole is missing the rule, usually because the agent was installed from a chart older than this support. The agent logs one warning naming the fix and skips Rollouts. It warns again only if a later cycle succeeds and a cycle after that is denied again.
- Served and allowed — every Rollout is listed and mapped: desired replicas from
spec.replicas(Argo's default of 1 when unset), ready and available fromstatus.readyReplicas/status.availableReplicas, the strategy (canaryorblueGreen), the containers' names and images fromspec.template.spec.containers, conditions, labels and annotations. Annotations are scrubbed the same way as every other workload's.
None of these fails the collection: the other resources are collected either way.
A Rollout's pods are owned by a ReplicaSet named <rollout>-<hash> and carry Argo's rollouts-pod-template-hash label — not the pod-template-hash label a Deployment's pods carry. The agent uses that label to link each pod to its Rollout rather than to a Deployment of the same name that does not exist.
A Rollout that takes its pod template from a Deployment (spec.workloadRef) has no inline template, so it is shown with no containers.
The ClusterRole rule is part of the Helm chart. Upgrading only the cluster agent image on an older chart gives you the one 403 warning above and no Rollouts. Run helm upgrade with the new chart so the rule below is applied.
ArgoCD Applications
From v0.8.0 the agent also lists ArgoCD Application objects (argoproj.io/v1alpha1) in its own cluster, with its own ServiceAccount — Console needs no ArgoCD credentials — and reports their sync and health state for ArgoCD drift. The same three cases apply: CRD not installed is reported as crd_absent silently; a 403 (chart without the rule) is reported as forbidden and logged once; neither affects any other resource. Application values are never read: the agent decodes only identity, destination (server URL scrubbed), sync/health, revisions, repository URL/path (userinfo, query and fragment stripped, failing closed on unparseable URLs), the managed-resource rows of the workload kinds and totals over all rows — never Helm values/valuesObject, parameters, plugin env or kustomize patches. The ClusterRole rule is list on applications only.
Renamed ArgoCD in-cluster clusters
ArgoCD calls the cluster it runs in in-cluster (server https://kubernetes.default.svc), and
Console maps Applications targeting either to this cluster. ArgoCD lets you rename that
cluster, and Console never reads ArgoCD's cluster secrets, so it cannot tell that, say,
production is ArgoCD's own cluster: those Applications show as "unmapped destination
'production'" and their workloads' drift is unknown (unmapped_destination). Declare the name
in the chart values:
argocd:
inClusterNames:
- production
The chart sets PROXIMA_ARGOCD_IN_CLUSTER_NAMES=production on the cluster agent. Names are
1–253 characters of letters, digits, ., _, -, at most 20; an invalid one is logged once
and ignored (the agent keeps running) — ArgoCD allows any name, so one outside this charset cannot
be declared. A name that is another Console cluster's slug is ignored and flagged. The agent reports the names next to the Applications —
destinations are reported unchanged — and the backend treats them as in-cluster for this
cluster's own Applications only. Needs a cluster agent, chart and backend newer than v0.8.0.
See ArgoCD Drift → Renamed in-cluster clusters.
Deployment
The easiest way to deploy the cluster agent (and node agent) is via the Proxima Helm chart — it ships the collector as a StatefulSet with per-pod persistence and optional leader-elected HA. The manual kubectl apply steps below are a minimal single-replica, no-persistence quickstart for environments where Helm is not available.
Cluster Onboarding (/clusters/add) mints the install token, prints either form of the command with every value filled in — including the cluster slug, which the steps below leave for you to set — and then verifies enrollment, heartbeat and the first sync. The steps on this page are the reference and the fallback; the wizard is the path to prefer.
Prerequisites
- Kubernetes cluster with API access
- Proxima Console backend running and accessible
- NATS server accessible from the cluster
- Install token from Proxima Console
1. Create namespace and RBAC
kubectl apply -f infra/k8s/cluster-agent/namespace.yaml
kubectl apply -f infra/k8s/cluster-agent/serviceaccount.yaml
kubectl apply -f infra/k8s/cluster-agent/clusterrole.yaml
kubectl apply -f infra/k8s/cluster-agent/clusterrolebinding.yaml
2. Create install token secret
kubectl -n proxima-system create secret generic proxima-install-token \
--from-literal=token=YOUR_INSTALL_TOKEN
3. Deploy, then set the values this install is recorded under
Apply the manifest, then set the three values that identify the cluster. This is the form the onboarding wizard emits, and it is preferred over hand-editing the file for one concrete reason: the manifest has no PROXIMA_CLUSTER_SLUG key at all, so an edit that sets only the display name leaves the agent deriving the slug itself — see the note below.
kubectl apply -f infra/k8s/cluster-agent/deployment.yaml
kubectl -n proxima-system set env deployment/proxima-cluster-agent \
PROXIMA_CLUSTER_NAME='my-cluster' \
PROXIMA_CLUSTER_SLUG='my-cluster' \
PROXIMA_BACKEND_URL='https://api-console.prxm.uz'
Quote each value in single quotes. You are running this against your own cluster with cluster-admin credentials, and a display name is free text.
Fallback — hand-edit the manifest. Editing infra/k8s/cluster-agent/deployment.yaml before applying it works just as well, and is the better choice if the file is under version control or managed by GitOps. Set PROXIMA_CLUSTER_NAME and PROXIMA_BACKEND_URL, add a PROXIMA_CLUSTER_SLUG entry (the manifest ships without one), and update the image tag if needed.
The slug is worth setting explicitly
When PROXIMA_CLUSTER_SLUG is empty the agent derives a slug from the display name with its own slugify, which has no 63-character bound. A long or unusual display name can therefore enroll the cluster under a slug that differs from the one Console recorded, and Console matches inventory on (environment, slug) — so the cluster appears under an unexpected name, or an onboarding attempt waits for a sync that has already landed somewhere else. Setting the value removes the guess.
4. Verify
kubectl -n proxima-system logs deployment/proxima-cluster-agent
You should see enrollment and collection logs. The cluster appears in the Proxima Console API after the first full sync — allow up to two sync intervals, because the first sync begins once the pod has finished starting rather than at the instant it enrolls. That is the same budget the onboarding wizard treats as healthy before it calls a sync overdue.
Look for it on the Clusters page, not under Hosts: a collector creates no host record (see Architecture above), so the things worth checking are the agent fleet, the cluster row and how recently it synced.
Configuration
| Environment Variable | Default | Description |
|---|---|---|
PROXIMA_CLUSTER_NAME | (required) | Cluster display name |
PROXIMA_CLUSTER_SLUG | (derived from name) | URL-safe cluster identifier, and the key Console matches inventory on — (environment, slug). Set it explicitly. Left empty, the agent derives it from the display name with a slugify that has no 63-character bound, so a long name can enroll under a slug Console did not record; see step 3. The onboarding wizard and Helm's clusterAgent.clusterSlug both set it for you. |
PROXIMA_BACKEND_URL | (required) | Backend API for enrollment |
PROXIMA_INSTALL_TOKEN | (required) | Enrollment token |
PROXIMA_K8S_SYNC_INTERVAL | 5m | Full sync interval. Set the same value on the backend (PROXIMA_K8S_SYNC_INTERVAL there too) — the cluster-onboarding gates quote twice the backend's copy as the budget inside which "no inventory yet" is healthy, and nothing reconciles the two. Changing it here alone makes a healthy cluster read as overdue in the onboarding wizard. |
PROXIMA_K8S_POD_LIMIT | 5000 | Max pods per sync. When exceeded, non-Running pods are prioritized (sorted first), then by restart count descending. |
PROXIMA_K8S_EXCLUDE_NAMESPACES | (empty) | Comma-separated list of namespaces to exclude from collection. Resources in excluded namespaces are not collected or stored. |
PROXIMA_ARGOCD_IN_CLUSTER_NAMES | (empty) | Comma-separated ArgoCD cluster names that are this cluster, for an ArgoCD whose in-cluster cluster was renamed (chart value argocd.inClusterNames). See Renamed in-cluster clusters. |
PROXIMA_NATS_URL | (from enrollment) | Override NATS URL (takes priority over enrolled URL) |
PROXIMA_NATS_CA_FILE | (empty) | Path to NATS TLS CA certificate |
PROXIMA_AGENT_DATA_DIR | /var/lib/proxima-agent | Directory for enrollment state and CA certificate |
PROXIMA_LOG_LEVEL | info | Log verbosity (debug, info, warn, error) |
PROXIMA_METRICS_ADDR | :9090 | Prometheus metrics HTTP address |
Events are capped at 1000 per sync (sorted by LastTimestamp descending, most recent kept). This limit is hardcoded and not configurable via environment variable.
High Availability
Run multiple replicas (clusterAgent.replicas > 1 in Helm) for HA. Replicas elect a single leader through a coordination.k8s.io Lease; only the leader collects and publishes inventory and heartbeats, while standbys stay enrolled and connected, ready to take over within the lease window (~15s). This prevents duplicate inventory while giving much faster failover than waiting for a single pod to reschedule.
Each replica is a StatefulSet pod with its own PVC and a stable per-pod identity (its pod name), so no ReadWriteOnce volume is shared and every replica keeps a durable identity across restarts. Standbys appear as separate agents in the Console; exactly one — the leader — reports at a time.
Leader election is automatic and needs no configuration beyond replicas. With replicas: 1 (the default) the single pod is always the leader, so behavior is unchanged. Local/dev runs without the Kubernetes downward API (no POD_NAME) skip leader election and run as the sole collector.
Security Hardening
The deployment manifest includes security best practices:
securityContext:
runAsNonRoot: true
runAsUser: 65534 # nobody
fsGroup: 65534
seccompProfile:
type: RuntimeDefault # meets the restricted Pod Security Standard
containers:
- securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop: [ALL]
RBAC ClusterRole
The collector uses a read-only ClusterRole with minimal permissions:
rules:
- apiGroups: [""]
resources: [nodes, namespaces, pods, services, endpoints,
resourcequotas, events]
verbs: [get, list, watch]
- apiGroups: [apps]
resources: [deployments, statefulsets, daemonsets, replicasets]
verbs: [get, list, watch]
- apiGroups: [batch]
resources: [jobs, cronjobs]
verbs: [get, list, watch]
- apiGroups: [networking.k8s.io]
resources: [ingresses, networkpolicies]
verbs: [get, list, watch]
- apiGroups: [""]
resources: [persistentvolumeclaims]
verbs: [get, list, watch]
- apiGroups: [autoscaling]
resources: [horizontalpodautoscalers]
verbs: [get, list, watch]
- apiGroups: [discovery.k8s.io]
resources: [endpointslices]
verbs: [get, list, watch]
- apiGroups: [metrics.k8s.io]
resources: [nodes]
verbs: [get, list]
# Argo Rollouts: the collector only lists them, so "list" is the only verb.
- apiGroups: [argoproj.io]
resources: [rollouts]
verbs: [list]
# ArgoCD Applications (drift): list only; values are never read.
- apiGroups: [argoproj.io]
resources: [applications]
verbs: [list]
No access to secrets, configmaps, or write operations.
When deployed via Helm, the collector also gets a small namespaced Role granting get/create/update on its own coordination.k8s.io Lease — used only for leader election. Nothing cluster-wide, nothing else.
Graceful Shutdown
The collector handles SIGINT and SIGTERM signals for clean shutdown:
- The signal cancels the main context.
- The metrics HTTP server is shut down with a 5-second timeout.
- The scheduler waits for any in-progress sync to complete.
- The NATS connection is closed.
NATS max_payload
The entire K8s inventory is published as a single NATS message. For large clusters (5000+ pods) the payload can exceed the NATS default max_payload of 1MB.
The Proxima NATS server is configured with max_payload: 8MB to fit large inventories (see infra/nats/). If a payload still exceeds the server limit, the agent rejects it up front with a clear ErrPayloadTooLarge error (logged and counted) rather than silently dropping it — lower PROXIMA_K8S_POD_LIMIT or raise the server's max_payload if you hit this. Before giving up, the agent drops its optional parts one step at a time and retries after each: the per-workload spec facts (see Workload Spec Facts), then the ArgoCD Application workload rows, then the ArgoCD Applications themselves (see ArgoCD Drift) — which keeps the rest of the inventory flowing. Watch proxima_collector_inventory_payload_bytes for headroom and proxima_collector_inventory_too_large_total{outcome} for which step a sync needed.
Storage
The collector binds a NATS identity on its first enrollment, and the backend refuses to re-bind it (identity-takeover protection). A collector that loses its state.json therefore cannot re-enroll under the same name — it fails fast with 409 already enrolled. The collector needs durable state.
- Helm (recommended): the StatefulSet gives each replica its own PVC via a
volumeClaimTemplate(persistence.enabled: true, the default). State survives restarts, reschedules, node drains, and upgrades. Requires a (default) StorageClass. - Manual /
emptyDir: state is lost on pod restart, so the pod then crash-loops on re-enrollment. UseemptyDironly for short-lived testing; for anything real, back/var/lib/proxima-agentwith a PVC.
If you must reset a collector's identity (e.g. migrating off emptyDir), delete its agent record in Proxima Console first, then let the fresh pod enroll.
Monitoring
The collector exposes a Prometheus-compatible /metrics endpoint on port 9090:
proxima_collector_syncs_total— total sync countproxima_collector_sync_errors_total— error countproxima_collector_last_sync_duration_seconds— last sync durationproxima_collector_resources_collected{resource="..."}— per-type resource counts
The collector exposes two probe endpoints: /healthz backs the Kubernetes liveness probe (process alive, never gated on enrollment or NATS), and /readyz backs the readiness probe — it reports ready only once the collector is enrolled and connected to NATS.