Skip to main content

Agent Troubleshooting

This page covers common issues, edge cases, and recovery procedures for the Proxima Agent.

Quick Diagnostics​

# Agent status and recent logs
systemctl status proxima-agent
journalctl -u proxima-agent --since "10 min ago" -f

# Check state.json
cat /var/lib/proxima-agent/state.json | jq .

# Check agent version
proxima-agent --version

# Run preflight checks
proxima-agent preflight --backend-url https://api-console.prxm.uz

Enrollment Issues​

Agent crash-loops on first start​

Symptom: Agent starts, fails enrollment, systemd restarts it repeatedly.

Diagnosis:

journalctl -u proxima-agent | grep "enrollment failed"

Common causes:

Log messageCauseFix
http 401: invalid install tokenToken is wrong, expired, deleted, or already usedGenerate a new token in the Console UI (or, for an agent that is already enrolled, use Fleet → Agents → Re-issue enrolment)
http request: dial tcp: connection refusedBackend is unreachableCheck PROXIMA_BACKEND_URL in systemd unit, verify backend is running
enrollment canceled: context deadline exceededBackend responding too slowlyCheck backend health, network connectivity

Note: The agent retries a transient enrollment failure (network error, 5xx) with exponential backoff (5s doubling, up to 5 minutes); check logs for retry_in. A refusal no retry can fix — 401 (token unknown, expired or used up) or 409 (name conflict, revoked agent, or a re-issued token that belongs to a different agent) — is logged at ERROR as enrollment refused; retrying hourly until an operator acts, with an action field saying what to do, and retried hourly. A host agent stays active; the cluster collector also fails /healthz, so its pod crash-loops visibly.

Re-enrollment is rejected with 409 Conflict​

Symptom: Enrollment fails with http 409: an agent with this name is already enrolled in this environment — typically after state.json was lost or wiped.

Why: Agent identity is keyed by agent_id and protected against takeover (AUTH H-1). An ordinary install token will not rebind the NATS key of an already-enrolled (environment_id, agent_type, name) — that returns 409 rather than silently reusing or overwriting the existing agent record.

Fix: Don't wipe state.json to rotate credentials — the agent renews its own NATS JWT automatically (NATS-first, HTTPS fallback), including a JWT that expired up to 90 days ago. To bring back an agent that lost its state, or to re-key one (e.g., a compromised seed), use Fleet → Agents → row menu → Re-issue enrolment for that agent, put the token in PROXIMA_INSTALL_TOKEN, and restart. The re-issued token re-binds the same agent record (same id and history) and revokes its previous credential. See Recovering an agent whose credential expired.

systemctl daemon-reload && systemctl restart proxima-agent

409 bound_agent_mismatch: the configured token was re-issued for a different agent (another name, type or environment). It is not spent and it cannot be used here; re-issue enrolment for this agent instead.

If a stale host record was left behind under a previous identity, deactivate it via the API:

curl -X POST https://api-console.prxm.uz/api/v1/hosts/<old-host-id>/deactivate \
-H "Authorization: Bearer <token>"

Expired install token​

Symptom: http 401: install token expired in logs.

Fix: Generate a new token (TTL up to 30 days). Tokens are single-use by default — pass "max_uses": 0 for an unlimited/fleet token (or "max_uses": N for a bounded count):

# Via API
curl -X POST https://api-console.prxm.uz/api/v1/clients/<slug>/install-tokens \
-H "Authorization: Bearer <token>" \
-H "Content-Type: application/json" \
-d '{"description":"fleet install","expires_in":2592000,"max_uses":0}'

Update the agent's environment:

# Edit the systemd environment file
vim /opt/proxima/.env # or wherever PROXIMA_INSTALL_TOKEN is set
systemctl restart proxima-agent

State & Credential Issues​

Corrupt state.json​

Symptom: Agent logs enrolling with the install token on startup, with why set to missing, corrupt: … or invalid: ….

What happens: The agent validates state.json on load. If the file is missing, unreadable as JSON, or any required field is missing or malformed (bad UUID, missing JWT, wrong nkey seed prefix), it enrolls with PROXIMA_INSTALL_TOKEN. The bad file is left in place and replaced atomically by the enrollment. Without an install token the agent cannot start (no usable state.json … and no install token).

Validated fields:

  • agent_id — must be a valid UUID
  • agent_type — must be a known type (host, collector, k8s-node-monitor, or probe)
  • nkey_seed — must start with SU (NATS user seed prefix)
  • user_jwt — must be non-empty
  • client_slug, environment_slug, nats_url — must be non-empty

If enrollment is then refused: the agent's name is usually still enrolled, so an ordinary token gets 409. Use Re-issue enrolment for this agent as described under Re-enrollment is rejected with 409 Conflict. Deleting state.json is never needed.

JWT expired — agent can't communicate​

Symptom: Agent is running but not sending data. Logs show JWT renewal failed via all channels.

How it should work: The agent checks JWT expiry hourly and renews when less than 50% TTL remains (default: renews after 3.5 days of a 7-day JWT). Renewal uses NATS first, then HTTPS fallback.

If the JWT has expired: the agent keeps trying to renew it — at boot before connecting to NATS, then hourly, and after a transient failure again 30s later, doubling up to the hour. The backend renews a JWT that expired up to 90 days ago (PROXIMA_AGENT_RENEW_EXPIRED_GRACE) over HTTPS, so restoring backend reachability is usually enough; no restart is needed.

If renewal is refused (credential refused; this agent cannot recover on its own, with a reason and an action):

  • expired_too_long (expired more than 90 days ago) or agent_not_found: the agent enrolls with PROXIMA_INSTALL_TOKEN if one is set. Otherwise, or if that token is refused, use Fleet → Agents → Re-issue enrolment for this agent, put the token in PROXIMA_INSTALL_TOKEN, and restart. Do not delete state.json and use a fresh ordinary token: the agent's name is still enrolled, so that ends in 409.
  • revoked: the agent was revoked, which is terminal; re-issue is refused for it.
  • jwt_invalid or nkey_mismatch: the stored credential is not the one the backend has on record; re-issue enrolment for this agent.

See Recovering an agent whose credential expired for the full playbook, including where each agent kind reads its token.

If renewal fails but JWT is not yet expired:

  • Check backend health (curl https://api-console.prxm.uz/healthz)
  • Check NATS connectivity (nats server ping from the host)
  • Check TLS certificate expiry on the NATS server
  • The agent will retry hourly — no action needed unless the JWT is about to expire

NATS TLS certificate expired​

Symptom: Agent logs tls: failed to verify certificate: x509: certificate has expired.

Fix: Reissue the NATS server certificate:

# On the backend/infrastructure host
source /opt/proxima/infra/vault/.vault-keys
export VAULT_ADDR=http://127.0.0.1:8200 VAULT_TOKEN
./scripts/nats-cert-issue.sh
cp infra/nats/certs/* /opt/proxima/infra/nats/certs/
docker restart proxima-nats
systemctl restart proxima-backend

The agent will automatically reconnect once NATS is back with a valid certificate.

Agent has ca.pem but NATS CA was rotated​

Symptom: Agent fails to reconnect after NATS CA rotation.

Fix: The agent's ca.pem was written during enrollment. If the CA was rotated, give the agent the new CA and restart — either replace /var/lib/proxima-agent/ca.pem with it, or point PROXIMA_NATS_CA_FILE at it (a configured CA takes precedence over the stored ca.pem):

systemctl stop proxima-agent
install -m 0644 /path/to/new-ca.pem /var/lib/proxima-agent/ca.pem
systemctl start proxima-agent

Do not delete state.json for this: the agent's name is still enrolled, so an ordinary install token would end in 409.

NATS Connectivity Issues​

Agent connected but not sending data​

Symptom: Host shows as "online" (heartbeat works) but no metrics, inventory, or processes.

Diagnosis:

# Check scheduler is running
journalctl -u proxima-agent | grep "scheduler"

# Check for outbox full (data loss under NATS pressure)
journalctl -u proxima-agent | grep "outbox full"

If outbox is full: The agent's publish queue (256 messages) is saturated. This indicates the NATS connection is slow or blocked. Check:

  • NATS server health and JetStream status
  • Network latency between agent and NATS
  • JetStream storage usage (may need cleanup)

Permission violation on NATS publish​

Symptom: nats async error: permissions violation in agent logs.

Cause: The agent's JWT doesn't have permission for the subject it's trying to publish to. This happens if:

  • The host was moved to a different client/environment but the JWT still has the old scopes
  • The JWT was manually tampered with

Fix: Get the agent a JWT minted from its current record: use Fleet → Agents → Re-issue enrolment for this agent, put the token in PROXIMA_INSTALL_TOKEN, and restart. The re-issued token re-binds the same agent record and revokes the old credential (see Recovering an agent whose credential expired). Deleting state.json and enrolling with an ordinary token ends in 409, because the agent's name is still enrolled.

Update Issues​

Self-update command failed​

Symptom: Backend sent self_update command but agent didn't update.

Diagnosis:

journalctl -u proxima-agent | grep "self-update\|self_update"

Common causes:

  • Download failed (network issue, R2 unreachable)
  • Checksum mismatch (corrupted download)
  • Permission denied (binary not writable)
  • Disk full
malformed checksum file — fixed in v0.6.1+

Older agents failed self-update with malformed checksum file because the verifier required the two-field sha256sum form while the release pipeline publishes a bare hash. This was fixed in v0.6.1+: the agent now accepts the published bare-hash checksum. An agent still hitting this error is on an older build — reinstall it once via the install script to pick up the fix, after which self-updates (including console fleet rollouts) work normally.

Manual update:

proxima-agent self-update          # update to latest
proxima-agent self-update v0.5.1 # pin specific version
systemctl restart proxima-agent

For the recommended console-driven flow, see Updating Agents (Fleet Management).

Agent running old version after update​

Symptom: proxima-agent --version shows old version.

Cause: The self_update handler swaps the binary and calls systemctl restart, but if the restart failed, the old process is still running.

Fix:

systemctl restart proxima-agent
proxima-agent --version

Data & Configuration Issues​

Config not applied after push from backend​

Symptom: Backend pushed a config update but agent behavior didn't change.

Diagnosis:

# Check config handler logs
journalctl -u proxima-agent | grep "config_update\|config applied\|config handler"

# Check config cache
cat /var/lib/proxima-agent/config.cache.json | jq .

The agent publishes a config apply ACK to proxima.system.config.apply.<agent_id>. Check the backend for the ACK status.

Agent collecting stale collector data​

Symptom: Metrics reference collectors that no longer exist (e.g., a removed PostgreSQL instance).

Cause: Collector configs are pushed from the backend. If the backend still has the old config, the agent will keep trying.

Fix: Update the collector configuration in the Console UI. The change is pushed to the agent in real-time.

File Layout Reference​

/usr/local/bin/proxima-agent              # Binary
/etc/systemd/system/proxima-agent.service # Systemd unit
/var/lib/proxima-agent/
state.json # All enrollment state (agent_id, JWT, nkey seed, slugs) — 0600
ca.pem # NATS CA certificate (only with private CA) — 0644
config.cache.json # Cached config from backend
file_hashes.json # Change detection baseline
journald-cursor # Log collection cursor
logs/ # Log collection state

Environment Variables​

VariableRequiredDescription
PROXIMA_BACKEND_URLYesBackend API URL for enrollment and renewal
PROXIMA_INSTALL_TOKENFirst bootInstall token — needed when there is no usable state.json. Also used, without a restart, when renewal is refused as expired_too_long (JWT expired more than 90 days ago) or agent_not_found. A token from Re-issue enrolment re-binds the same agent.
PROXIMA_LOG_LEVELNodebug, info (default), warn, error
PROXIMA_AGENT_DATA_DIRNoDefault: /var/lib/proxima-agent
PROXIMA_NATS_CA_FILENoOverride CA path (normally from enrollment)
PROXIMA_NATS_URLNoOverride NATS URL (normally from enrollment)

Recovery Decision Tree​

Agent not sending data
├── systemctl is-active proxima-agent → inactive
│ └── Check: journalctl -u proxima-agent -e
│ ├── "no usable state.json … and no install token" → Set PROXIMA_INSTALL_TOKEN, restart
│ ├── "enrollment failed, retrying" → Backend down or slow, check retry_in
│ ├── "enrollment refused; retrying hourly" → Read `action`; usually Re-issue enrolment
│ └── Other error → Check specific error message above
├── active, but no data in Console
│ ├── Check: journalctl | grep "outbox full" → NATS backpressure
│ ├── Check: journalctl | grep "credential refused" → Read `reason` + `action` (see JWT expired)
│ ├── Check: journalctl | grep "permission" → JWT scope mismatch, Re-issue enrolment
│ └── Check: journalctl | grep "tls" → Certificate expired, reissue
└── active, data flowing, but wrong host
└── Duplicate host_id → Deactivate old host via API