Helm Chart
The proxima-agent Helm chart deploys both the node agent (DaemonSet) and the cluster agent (StatefulSet) in a single installation. This is the recommended method for Kubernetes deployments.
Cluster Onboarding (/clusters/add) mints the install token, renders the helm upgrade --install below with every value filled in — the cluster name and slug, the backend URL, the token, and nodeAgent.hostDataDir when those nodes already run a host agent — and then verifies enrollment, heartbeat and the first inventory sync. Use this page for the full value reference and for anything the wizard does not generate.
Prerequisites
- Kubernetes 1.26+
- Helm 3.x or 4.x
- Proxima Console backend running and accessible
- Install token from Proxima Console (Settings → Install Tokens)
Add the Helm Repository
The chart is published to the GitLab Package Registry.
# Add the Proxima Helm repository (requires authentication)
helm repo add proxima \
https://gitlab.prxm.uz/api/v4/projects/271/packages/helm/stable \
--username <token-name> \
--password <access-token>
helm repo update
Use a GitLab deploy token or personal access token with read_package_registry scope.
Quick Start
helm install proxima-agent proxima/proxima-agent \
--namespace proxima-system \
--create-namespace \
--set backendUrl=https://api-console.prxm.uz \
--set installToken.value=YOUR_INSTALL_TOKEN \
--set clusterAgent.clusterName=my-cluster
This deploys:
- A DaemonSet on every Linux node (node agent with host metrics + kubelet metrics)
- A StatefulSet (cluster agent for K8s API inventory) — one replica by default, each replica backed by its own PVC for a stable, restart-durable enrollment identity. Run more than one replica for high availability.
Configuration
Required Values
| Value | Description |
|---|---|
backendUrl | Proxima Console backend URL (e.g., https://api-console.prxm.uz) |
installToken.value | Install token for agent enrollment |
clusterAgent.clusterName | Display name for this cluster in the Console UI |
clusterAgent.clusterSlug | Not strictly required, but set it: the URL-safe identifier Console matches inventory on. Left empty the agent derives it from clusterName with no 63-character bound, so a long name can enroll under a slug Console did not record. |
Using an Existing Secret
Instead of passing the install token as a value, reference a pre-existing Kubernetes Secret:
# Create the secret first
kubectl -n proxima-system create secret generic proxima-install-token \
--from-literal=token=YOUR_INSTALL_TOKEN
# Install with existingSecret reference
helm install proxima-agent proxima/proxima-agent \
--namespace proxima-system \
--set backendUrl=https://api-console.prxm.uz \
--set installToken.existingSecret.name=proxima-install-token \
--set installToken.existingSecret.key=token \
--set clusterAgent.clusterName=my-cluster
Common Customizations
# custom-values.yaml
backendUrl: "https://api-console.prxm.uz"
installToken:
existingSecret:
name: proxima-install-token
key: token
# Node Agent
nodeAgent:
logLevel: "info"
metricsInterval: "60s"
hostMetrics: true
# Schedule on all nodes including control-plane
tolerations:
- operator: Exists
resources:
requests:
cpu: "100m"
memory: "128Mi"
limits:
cpu: "400m"
memory: "512Mi"
# Cluster Agent
clusterAgent:
clusterName: "production-rke2"
# Set the backend's PROXIMA_K8S_SYNC_INTERVAL to the same value: it reads the key too,
# nothing reconciles the two, and the onboarding gates quote the backend's copy.
syncInterval: "5m"
heartbeatInterval: "30s"
podLimit: 5000
excludeNamespaces:
- kube-system
helm install proxima-agent proxima/proxima-agent \
--namespace proxima-system \
-f custom-values.yaml
Running Alongside a Host Agent
If you install the standalone Proxima host agent on K8s nodes (for full inventory, terminal access, and runbooks), the two agents can coexist — but only if you give them separate data directories. This needs a configuration change; it is not the default.
A data directory holds state.json, which is the agent's identity — its agent_id and its private nkey seed. An agent only enrolls when no valid state file is present, so an agent that finds one adopts it wholesale, ignoring its configured agentType.
Setting nodeAgent.agentType: k8s-node-monitor does not prevent this. That setting applies at enrollment, and this is precisely the case where enrollment never happens.
Two processes on one identity fail in ways that are hard to trace back:
- Reported versions flap. Both heartbeat under the same
agent_id, so each overwrites the other's version. The console shows whichever wrote last, and it changes on refresh. - Fleet rollouts fail against healthy hosts. One
self_updatecommand reaches both processes. The pod tries to replace/usr/local/bin/proxima-agenton its read-only image layer and reportsbackup current binary: rename …: read-only file system— recorded against a host that may have updated successfully. Enough of these trip the rollout's failure threshold and stall the wave for every remaining agent.
Since v0.7.2 an agent refuses to start when the persisted agent_type differs from the type it is configured as, so this misconfiguration surfaces as a clear startup error instead of silent identity sharing. It does not delete the state file it found — on a shared directory that file belongs to the other agent.
The host agent installs at the default /var/lib/proxima-agent and is usually there first, so move the DaemonSet's host path instead:
helm upgrade --install proxima-agent proxima/proxima-agent \
--namespace proxima-system \
--set nodeAgent.hostDataDir=/var/lib/proxima-node-agent \
...
hostDataDir changes only the host side of the mount; inside the container the path stays /var/lib/proxima-agent. The default is unchanged, so clusters where the DaemonSet is the only agent on the node keep their existing identity.
hostDataDir on a live cluster is not always safeThe pod abandons the identity in the old directory and enrolls fresh. Enrollment is unique on (environment, agent_type, name), so it succeeds only where nothing already holds that combination:
- Node also runs a host agent — the pod had adopted the host identity, so enrolling as
k8s-node-monitoris a free name. This is the case the setting fixes, and it works. - Node runs only the DaemonSet — it already owned a
k8s-node-monitorrecord under that node's name. Re-enrollment collides and fails with HTTP 409 indefinitely. The agent retries instead of exiting, so the pod staysReadywhile never enrolling: it looks healthy and reports nothing.
Most nodes now recover on their own. Enrollment re-binds a machine onto its own prior record when that record has been quiet for more than five minutes and the machine still presents an identity signal the record already carries. A DaemonSet pod that mounts the host filesystem reads the host's /etc/machine-id and SMBIOS/DMI uuid, and neither changes when hostDataDir does — so the abandoned record is recognised as this node's own and re-bound, keeping its id, its links and its history. No operator action, and no 409.
Three things have to hold, and where any of them does not, the node still needs the manual path below:
- The pod can read the host's identity. Without a host-root mount its only identity was the UUID in the data directory — the thing that just moved — so it presents nothing the old record shares.
- The old record has identity evidence. Records acquire it at enrollment (agents released after 2026-09-20) or at their first credential renewal. A record that has done neither since has no signals to match against.
- The old record is quiet. If something is still heartbeating as it, that is a second live process rather than this node coming back, and re-binding is refused by design.
To recover a node that is not re-bound, free the old record's name with POST /api/v1/fleet/agents/{agentID}/release-identity (requires agents:write on that agent's client and environment). It clears the agent's NATS identity so the next enrollment re-binds the same record — keeping its id, its links and its history. It refuses a revoked agent, and one seen in the last 5 minutes, so a live agent's name can never be freed underneath it.
rotate-credentials does not help here: it needs the agent online holding its identity, so it cannot recover one whose state is gone.
Before changing this on a live cluster, list which nodes already hold a k8s-node-monitor record under their own name, and plan to clear those first.
To verify no node is sharing an identity:
# Should print nothing. Any output is a node running both agents on one directory.
kubectl get ds -n proxima-system proxima-agent-node-agent \
-o jsonpath='{range .spec.template.spec.volumes[?(@.name=="data")]}{.hostPath.path}{"\n"}{end}' \
| grep -x '/var/lib/proxima-agent'
See Dual-Agent Setup for details.
Disabling Components
# Node agent only (no cluster inventory)
helm install proxima-agent proxima/proxima-agent \
--set nodeAgent.enabled=true \
--set clusterAgent.enabled=false \
...
# Cluster agent only (no per-node metrics)
helm install proxima-agent proxima/proxima-agent \
--set nodeAgent.enabled=false \
--set clusterAgent.enabled=true \
...
High Availability (leader election)
The cluster agent can run as multiple replicas for high availability. Replicas elect a single leader through a Kubernetes coordination.k8s.io Lease; only the leader collects and publishes inventory and heartbeats, while the others stay enrolled, connected, and ready to take over. This avoids duplicate inventory while giving fast failover (seconds, versus waiting for a single pod to reschedule).
clusterAgent:
replicas: 2 # >1 enables leader-elected HA
persistence:
enabled: true # required for HA — each replica gets its own PVC
- Each replica is a StatefulSet pod with its own PVC (
volumeClaimTemplate) and a stable, restart-durable enrollment identity (its pod name). No ReadWriteOnce volume is ever shared between pods. - Standby replicas appear as separate agents in the Console; only the leader publishes, so exactly one is actively reporting at any time. Standbys log
standing by; another collector holds leadership. - Leader election needs a namespaced Role for its Lease — the chart creates it automatically.
replicas: 1 (the default) runs without contention (the single pod is always the leader) and behaves exactly like a single collector.
All Values
See the full values.yaml with documentation comments:
helm show values proxima/proxima-agent
Key sections:
| Section | Description |
|---|---|
backendUrl | Backend API URL |
installToken | Token config (plaintext or existing Secret) |
namespace | Namespace creation and naming |
commonLabels / commonAnnotations | Applied to all resources |
nodeAgent.* | DaemonSet image, resources, tolerations, security context |
clusterAgent.* | StatefulSet image, replicas, sync interval, pod limit, persistence, security |
GitOps (Argo CD)
Deploy the chart declaratively with an Argo CD Application that points at the published Helm chart. Keep the install token in a Secret managed by your secrets operator (e.g. External Secrets / Vault) and reference it with installToken.existingSecret — never commit the token to Git.
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: proxima-agent
namespace: argocd
spec:
project: default
source:
repoURL: https://gitlab.prxm.uz/api/v4/projects/271/packages/helm/stable
chart: proxima-agent
targetRevision: 0.6.15 # pin to a published chart version
helm:
releaseName: proxima-agent
values: |
backendUrl: https://api-console.prxm.uz
clusterAgent:
clusterName: production-rke2
replicas: 2 # leader-elected HA
installToken:
existingSecret:
name: proxima-install-token
key: token
destination:
server: https://kubernetes.default.svc
namespace: proxima-system
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=true
helm templateThe chart's required-value checks (backendUrl, install token, clusterName) run during template rendering — so a misconfigured Application fails the Argo CD sync with a clear error instead of deploying a broken collector.
The GitLab Helm registry requires authentication. Register it as an Argo CD repository (Settings → Repositories) with a deploy/access token that has read_package_registry scope. Published chart versions track the release tag (the release pipeline pins the chart version/appVersion to the tag), so targetRevision can be any released agent version.
Upgrading
helm repo update
helm upgrade proxima-agent proxima/proxima-agent \
--namespace proxima-system \
-f custom-values.yaml
Uninstalling
helm uninstall proxima-agent --namespace proxima-system
Uninstalling removes the DaemonSet and the cluster-agent StatefulSet but does not remove agent records from the Proxima Console database. Agents will appear as "offline" after their stale timeout.
Security
Both components follow security best practices:
- Node agent:
readOnlyRootFilesystem, capabilities dropped, no privilege escalation. Runs as root (required for/procand/sysaccess). - Cluster agent:
runAsNonRoot(UID 65532), read-only filesystem, all capabilities dropped,seccompProfile: RuntimeDefault(meets the restricted Pod Security Standard). Its ClusterRole is read-only — no access to Secrets or ConfigMaps; it includeslistonargoproj.iorolloutsso Argo Rollouts can be collected — plus a small namespaced Role for its leader-election Lease only.
Troubleshooting
# Check pod status
kubectl -n proxima-system get pods
# Node agent logs
kubectl -n proxima-system logs daemonset/proxima-agent-node-agent
# Cluster agent logs (leader runs collection; standbys log "standing by")
kubectl -n proxima-system logs statefulset/proxima-agent-cluster-agent
# Verify enrollment
kubectl -n proxima-system logs daemonset/proxima-agent-node-agent | grep "enrolled"
# Check NATS connectivity
kubectl -n proxima-system logs daemonset/proxima-agent-node-agent | grep "nats"
Common Issues
| Issue | Cause | Fix |
|---|---|---|
Pod stuck in ImagePullBackOff | Private registry, no pull secret | Add imagePullSecrets |
connection refused to backend | Wrong backendUrl or network policy | Verify URL and cluster egress rules |
| Kubelet metrics not flowing | hostMetrics: false or kubelet RBAC | Set hostMetrics: true, check ClusterRole |
| Cluster not appearing in Console | Collector not enrolling | Check logs for enrollment errors, verify install token |
| Argo Rollouts missing from the Workloads tab; cluster agent logged "forbidden to read argo rollouts" | Chart older than Argo Rollout support: the ClusterRole has no argoproj.io rollouts rule | helm upgrade to the current chart (upgrading the image alone does not add the rule). See Cluster Agent → Argo Rollouts |