Agent Troubleshooting
This page covers common issues, edge cases, and recovery procedures for the Proxima Agent.
Quick Diagnostics
# Agent status and recent logs
systemctl status proxima-agent
journalctl -u proxima-agent --since "10 min ago" -f
# Check state.json
cat /var/lib/proxima-agent/state.json | jq .
# Check agent version
proxima-agent --version
# Run preflight checks
proxima-agent preflight --backend-url https://api-console.prxm.uz
Enrollment Issues
Agent crash-loops on first start
Symptom: Agent starts, fails enrollment, systemd restarts it repeatedly.
Diagnosis:
journalctl -u proxima-agent | grep "enrollment failed"
Common causes:
| Log message | Cause | Fix |
|---|---|---|
http 401: invalid install token | Token is wrong, expired, deleted, or already used | Generate a new token in the Console UI (or, for an agent that is already enrolled, use Fleet → Agents → Re-issue enrolment) |
http request: dial tcp: connection refused | Backend is unreachable | Check PROXIMA_BACKEND_URL in systemd unit, verify backend is running |
enrollment canceled: context deadline exceeded | Backend responding too slowly | Check backend health, network connectivity |
Note: The agent retries a transient enrollment failure (network error, 5xx) with exponential backoff (5s doubling, up to 5 minutes); check logs for retry_in. A refusal no retry can fix — 401 (token unknown, expired or used up) or 409 (name conflict, revoked agent, or a re-issued token that belongs to a different agent) — is logged at ERROR as enrollment refused; retrying hourly until an operator acts, with an action field saying what to do, and retried hourly. A host agent stays active; the cluster collector also fails /healthz, so its pod crash-loops visibly.
Re-enrollment is rejected with 409 Conflict
Symptom: Enrollment fails with http 409: an agent with this name is already enrolled in this environment — typically after state.json was lost or wiped.
Why: Agent identity is keyed by agent_id and protected against takeover (AUTH H-1). An ordinary install token will not rebind the NATS key of an already-enrolled (environment_id, agent_type, name) — that returns 409 rather than silently reusing or overwriting the existing agent record.
Fix: Don't wipe state.json to rotate credentials — the agent renews its own NATS JWT automatically (NATS-first, HTTPS fallback), including a JWT that expired up to 90 days ago. To bring back an agent that lost its state, or to re-key one (e.g., a compromised seed), use Fleet → Agents → row menu → Re-issue enrolment for that agent, put the token in PROXIMA_INSTALL_TOKEN, and restart. The re-issued token re-binds the same agent record (same id and history) and revokes its previous credential. See Recovering an agent whose credential expired.
systemctl daemon-reload && systemctl restart proxima-agent
409 bound_agent_mismatch: the configured token was re-issued for a different agent (another name, type or environment). It is not spent and it cannot be used here; re-issue enrolment for this agent instead.
If a stale host record was left behind under a previous identity, deactivate it via the API:
curl -X POST https://api-console.prxm.uz/api/v1/hosts/<old-host-id>/deactivate \
-H "Authorization: Bearer <token>"
Expired install token
Symptom: http 401: install token expired in logs.
Fix: Generate a new token (TTL up to 30 days). Tokens are single-use by default — pass "max_uses": 0 for an unlimited/fleet token (or "max_uses": N for a bounded count):
# Via API
curl -X POST https://api-console.prxm.uz/api/v1/clients/<slug>/install-tokens \
-H "Authorization: Bearer <token>" \
-H "Content-Type: application/json" \
-d '{"description":"fleet install","expires_in":2592000,"max_uses":0}'
Update the agent's environment:
# Edit the systemd environment file
vim /opt/proxima/.env # or wherever PROXIMA_INSTALL_TOKEN is set
systemctl restart proxima-agent
State & Credential Issues
Corrupt state.json
Symptom: Agent logs enrolling with the install token on startup, with why set to missing, corrupt: … or invalid: ….
What happens: The agent validates state.json on load. If the file is missing, unreadable as JSON, or any required field is missing or malformed (bad UUID, missing JWT, wrong nkey seed prefix), it enrolls with PROXIMA_INSTALL_TOKEN. The bad file is left in place and replaced atomically by the enrollment. Without an install token the agent cannot start (no usable state.json … and no install token).
Validated fields:
agent_id— must be a valid UUIDagent_type— must be a known type (host,collector,k8s-node-monitor, orprobe)nkey_seed— must start withSU(NATS user seed prefix)user_jwt— must be non-emptyclient_slug,environment_slug,nats_url— must be non-empty
If enrollment is then refused: the agent's name is usually still enrolled, so an ordinary token gets 409. Use Re-issue enrolment for this agent as described under Re-enrollment is rejected with 409 Conflict. Deleting state.json is never needed.
JWT expired — agent can't communicate
Symptom: Agent is running but not sending data. Logs show JWT renewal failed via all channels.
How it should work: The agent checks JWT expiry hourly and renews when less than 50% TTL remains (default: renews after 3.5 days of a 7-day JWT). Renewal uses NATS first, then HTTPS fallback.
If the JWT has expired: the agent keeps trying to renew it — at boot before connecting to NATS, then hourly, and after a transient failure again 30s later, doubling up to the hour. The backend renews a JWT that expired up to 90 days ago (PROXIMA_AGENT_RENEW_EXPIRED_GRACE) over HTTPS, so restoring backend reachability is usually enough; no restart is needed.
If renewal is refused (credential refused; this agent cannot recover on its own, with a reason and an action):
expired_too_long(expired more than 90 days ago) oragent_not_found: the agent enrolls withPROXIMA_INSTALL_TOKENif one is set. Otherwise, or if that token is refused, use Fleet → Agents → Re-issue enrolment for this agent, put the token inPROXIMA_INSTALL_TOKEN, and restart. Do not deletestate.jsonand use a fresh ordinary token: the agent's name is still enrolled, so that ends in409.revoked: the agent was revoked, which is terminal; re-issue is refused for it.jwt_invalidornkey_mismatch: the stored credential is not the one the backend has on record; re-issue enrolment for this agent.
See Recovering an agent whose credential expired for the full playbook, including where each agent kind reads its token.
If renewal fails but JWT is not yet expired:
- Check backend health (
curl https://api-console.prxm.uz/healthz) - Check NATS connectivity (
nats server pingfrom the host) - Check TLS certificate expiry on the NATS server
- The agent will retry hourly — no action needed unless the JWT is about to expire
NATS TLS certificate expired
Symptom: Agent logs tls: failed to verify certificate: x509: certificate has expired.
Fix: Reissue the NATS server certificate:
# On the backend/infrastructure host
source /opt/proxima/infra/vault/.vault-keys
export VAULT_ADDR=http://127.0.0.1:8200 VAULT_TOKEN
./scripts/nats-cert-issue.sh
cp infra/nats/certs/* /opt/proxima/infra/nats/certs/
docker restart proxima-nats
systemctl restart proxima-backend
The agent will automatically reconnect once NATS is back with a valid certificate.
Agent has ca.pem but NATS CA was rotated
Symptom: Agent fails to reconnect after NATS CA rotation.
Fix: The agent's ca.pem was written during enrollment. If the CA was rotated, give the agent the new CA and restart — either replace /var/lib/proxima-agent/ca.pem with it, or point PROXIMA_NATS_CA_FILE at it (a configured CA takes precedence over the stored ca.pem):
systemctl stop proxima-agent
install -m 0644 /path/to/new-ca.pem /var/lib/proxima-agent/ca.pem
systemctl start proxima-agent
Do not delete state.json for this: the agent's name is still enrolled, so an ordinary install token would end in 409.
NATS Connectivity Issues
Agent connected but not sending data
Symptom: Host shows as "online" (heartbeat works) but no metrics, inventory, or processes.
Diagnosis:
# Check scheduler is running
journalctl -u proxima-agent | grep "scheduler"
# Check for outbox full (data loss under NATS pressure)
journalctl -u proxima-agent | grep "outbox full"
If outbox is full: The agent's publish queue (256 messages) is saturated. This indicates the NATS connection is slow or blocked. Check:
- NATS server health and JetStream status
- Network latency between agent and NATS
- JetStream storage usage (may need cleanup)
Permission violation on NATS publish
Symptom: nats async error: permissions violation in agent logs.
Cause: The agent's JWT doesn't have permission for the subject it's trying to publish to. This happens if:
- The host was moved to a different client/environment but the JWT still has the old scopes
- The JWT was manually tampered with
Fix: Get the agent a JWT minted from its current record: use Fleet → Agents → Re-issue enrolment for this agent, put the token in PROXIMA_INSTALL_TOKEN, and restart. The re-issued token re-binds the same agent record and revokes the old credential (see Recovering an agent whose credential expired). Deleting state.json and enrolling with an ordinary token ends in 409, because the agent's name is still enrolled.
Update Issues
Self-update command failed
Symptom: Backend sent self_update command but agent didn't update.
Diagnosis:
journalctl -u proxima-agent | grep "self-update\|self_update"
Common causes:
- Download failed (network issue, R2 unreachable)
- Checksum mismatch (corrupted download)
- Permission denied (binary not writable)
- Disk full
malformed checksum file — fixed in v0.6.1+Older agents failed self-update with malformed checksum file because the
verifier required the two-field sha256sum form while the release pipeline
publishes a bare hash. This was fixed in v0.6.1+: the agent now accepts
the published bare-hash checksum. An agent still hitting this error is on an
older build — reinstall it once via the install script
to pick up the fix, after which self-updates (including console
fleet rollouts) work normally.
Manual update:
proxima-agent self-update # update to latest
proxima-agent self-update v0.5.1 # pin specific version
systemctl restart proxima-agent
For the recommended console-driven flow, see Updating Agents (Fleet Management).
Agent running old version after update
Symptom: proxima-agent --version shows old version.
Cause: The self_update handler swaps the binary and calls systemctl restart, but if the restart failed, the old process is still running.
Fix:
systemctl restart proxima-agent
proxima-agent --version
Data & Configuration Issues
Config not applied after push from backend
Symptom: Backend pushed a config update but agent behavior didn't change.
Diagnosis:
# Check config handler logs
journalctl -u proxima-agent | grep "config_update\|config applied\|config handler"
# Check config cache
cat /var/lib/proxima-agent/config.cache.json | jq .
The agent publishes a config apply ACK to proxima.system.config.apply.<agent_id>. Check the backend for the ACK status.
Agent collecting stale collector data
Symptom: Metrics reference collectors that no longer exist (e.g., a removed PostgreSQL instance).
Cause: Collector configs are pushed from the backend. If the backend still has the old config, the agent will keep trying.
Fix: Update the collector configuration in the Console UI. The change is pushed to the agent in real-time.
File Layout Reference
/usr/local/bin/proxima-agent # Binary
/etc/systemd/system/proxima-agent.service # Systemd unit
/var/lib/proxima-agent/
state.json # All enrollment state (agent_id, JWT, nkey seed, slugs) — 0600
ca.pem # NATS CA certificate (only with private CA) — 0644
config.cache.json # Cached config from backend
file_hashes.json # Change detection baseline
journald-cursor # Log collection cursor
logs/ # Log collection state
Environment Variables
| Variable | Required | Description |
|---|---|---|
PROXIMA_BACKEND_URL | Yes | Backend API URL for enrollment and renewal |
PROXIMA_INSTALL_TOKEN | First boot | Install token — needed when there is no usable state.json. Also used, without a restart, when renewal is refused as expired_too_long (JWT expired more than 90 days ago) or agent_not_found. A token from Re-issue enrolment re-binds the same agent. |
PROXIMA_LOG_LEVEL | No | debug, info (default), warn, error |
PROXIMA_AGENT_DATA_DIR | No | Default: /var/lib/proxima-agent |
PROXIMA_NATS_CA_FILE | No | Override CA path (normally from enrollment) |
PROXIMA_NATS_URL | No | Override NATS URL (normally from enrollment) |
Recovery Decision Tree
Agent not sending data
├── systemctl is-active proxima-agent → inactive
│ └── Check: journalctl -u proxima-agent -e
│ ├── "no usable state.json … and no install token" → Set PROXIMA_INSTALL_TOKEN, restart
│ ├── "enrollment failed, retrying" → Backend down or slow, check retry_in
│ ├── "enrollment refused; retrying hourly" → Read `action`; usually Re-issue enrolment
│ └── Other error → Check specific error message above
├── active, but no data in Console
│ ├── Check: journalctl | grep "outbox full" → NATS backpressure
│ ├── Check: journalctl | grep "credential refused" → Read `reason` + `action` (see JWT expired)
│ ├── Check: journalctl | grep "permission" → JWT scope mismatch, Re-issue enrolment
│ └── Check: journalctl | grep "tls" → Certificate expired, reissue
└── active, data flowing, but wrong host
└── Duplicate host_id → Deactivate old host via API