NATS Security
Proxima Console uses a production-grade JWT/operator model for NATS authentication, issuing per-agent credentials with scoped permissions.
Why Per-Agent JWT?
Each agent receives a unique JWT with permissions scoped to its own subjects. This ensures that if one agent is compromised, the attacker cannot publish or subscribe to any other agent's NATS subjects.
Key security properties:
| Property | Per-Agent JWT |
|---|---|
| Authentication | Unique nkey + JWT per agent |
| Authorization | Scoped to agent-specific subjects only |
| Revocation | Revoke a single agent (instant disconnect) |
| Blast radius | Single agent compromised |
| Key storage | Nkey seed in agent state.json (0600); operator/account seeds in the deployment's key store (Vault by default, files on disk when PROXIMA_KEYSTORE=disk) |
| Transport | Mandatory mutual TLS 1.2+ |
Architecture
NATS Account Model
JWT mode uses the NATS operator/account/user hierarchy:
| Level | Name | Purpose |
|---|---|---|
| Operator | proxima-operator | Root of trust. Signs account JWTs. Nkey seed stored in the key store — Vault KV v2 by default, <keys>/nats/seeds/operator.seed under PROXIMA_KEYSTORE=disk. |
| Account | proxima | Application traffic. All agent and backend users belong here. |
| Account | SYS | NATS system account. Used for admin operations (JWT revocation push). |
| User | backend | Backend service user with full proxima.> permissions. |
| User | sys-admin | System admin user for $SYS.> operations. |
| User | agent-{uuid} | Per-agent user with scoped permissions (one per enrolled agent). |
Per-Agent User Permissions
Each agent JWT is scoped to subjects containing its own client, environment, and agent identifiers:
| Direction | Subjects | Purpose |
|---|---|---|
| Publish | proxima.{client}.{env}.{agent_id}.> | Telemetry (inventory, metrics, heartbeat, processes, changes), plus tunnel/SSH outbound bytes |
| Publish | proxima.system.agent.renew | JWT renewal requests |
| Publish | proxima.system.agent.renew.nonce | Fetch a server-issued single-use renewal challenge nonce (AUTH F2 — replay freshness) |
| Publish | proxima.system.agent.config.{agent_id} | Config fetch request-reply at startup (per-agent subject — the broker enforces identity; the worker derives the requesting agent from the subject's last token, so an agent can only fetch its own config) |
| Publish | proxima.system.config.apply.{agent_id} | Config apply ACK (success/failure per config type) |
| Subscribe | proxima.system.commands.{agent_id} | Commands from backend (config updates, terminal, runbooks) |
| Subscribe | proxima.system.file.{agent_id}.> | File transfer data (upload chunks, download acks) |
| Subscribe | proxima.{client}.{env}.{agent_id}.tunnel.in.> | Inbound tunnel data from backend (port forwarding) |
| Subscribe | proxima.{client}.{env}.{agent_id}.ssh.in.> | Inbound SSH bytes from the backend relay (OpenSSH ProxySSH) |
| Subscribe | _INBOX.{agent_id}.> | Per-agent reply inbox for request-reply. The agent sets a matching custom inbox prefix (nats.CustomInboxPrefix) on its connection, so a compromised agent cannot subscribe to (or eavesdrop on) another agent's reply inbox — the shared _INBOX.> wildcard is not granted. |
| Response | Max 1 message, 30s expiry | Limits reply scope |
All other subjects are denied by default. An agent cannot publish to or subscribe to another agent's subjects.
Probe agents hold no tenant grant at all
An agent enrolled as agent_type: probe (a synthetic-monitoring pop — see Uptime Monitors) gets a different JWT: the whole proxima.{client}.{env}.{agent_id}.> publish grant and the tunnel/SSH/kube subscriptions are withheld, and it receives only two extra subjects of its own.
| Direction | Subjects | Purpose |
|---|---|---|
| Publish | proxima.probe.results.{agent_id} | Check results |
| Publish | proxima.system.probe.assign.{agent_id} | Assignment request-reply ("what should I be checking?") |
The shared identity subjects (commands, file transfer, credential renewal, reply inbox) are granted as above — without them a pop could not be updated and would fall off the fleet at its first TTL.
The reason is that a public pop serves every tenant: the tenant grant is exactly what a compromised pop would use to publish forged inventory or metrics into a client's namespace, and it has no legitimate use for it. It also means the reporting location can be derived from the subject's agent id and never from the payload — the broker enforces the identity the backend then reads off the subject, so a compromised pop can lie about what it saw but never about which pop it is.
The trade-off is stated plainly in the uptime docs: a prober ships no metrics and no logs either, since both of those subjects are tenant-scoped too.
Infrastructure Bootstrap
Bootstrap is a one-time process that creates the operator, accounts, and backend credentials.
Prerequisites
- HashiCorp Vault running and unsealed — unless the deployment sets
PROXIMA_KEYSTORE=disk, in which case Steps 1 and 3 are replaced (see Disk keystore below) and there is no Vault at all - NATS server not yet started in JWT mode
Step 1: Initialize Vault
./scripts/vault-init.sh
This script configures Vault with:
| Component | Details |
|---|---|
| PKI engine | Root CA (EC P-256), nats-server role for server certificates (30-day max TTL) |
| KV v2 engine | At secret/, stores nkey seeds under proxima/nats/ prefix |
| AppRole auth | proxima-backend role with 1h token TTL, 4h max TTL |
| Policy | proxima policy granting KV read/write access |
The script outputs VAULT_ROLE_ID and VAULT_SECRET_ID for the backend.
Step 2: Generate NATS JWTs
make nats-jwt-init
Runs backend/cmd/nats-jwt-init/main.go, which:
- Generates (or loads from the key store) nkey pairs for operator,
proximaaccount, andSYSaccount - Creates self-signed operator JWT
- Creates account JWTs signed by the operator (with JetStream enabled)
- Creates backend user credentials with full
proxima.>permissions - Creates sys-admin user credentials with
$SYS.>permissions - Generates
resolver_preload.confwith account public key → JWT mappings for NATS bootstrap - Templates the
system_accountvalue innats-server-jwt.confwith the current SYS public key
Output files:
| File | Purpose | Used by |
|---|---|---|
infra/nats/jwt/operator.jwt | Operator identity | NATS server |
infra/nats/jwt/proxima.jwt | Proxima account identity | NATS server (resolver preload) |
infra/nats/jwt/sys.jwt | SYS account identity | NATS server (resolver preload) |
infra/nats/jwt/resolver_preload.conf | Account JWT preload config | NATS server config include |
infra/nats/creds/backend.creds | Backend user JWT + nkey seed | Backend service |
infra/nats/creds/sys-admin.creds | Sys-admin JWT + nkey seed | Backend (revocation operations) |
Vault KV paths:
| Path | Contents |
|---|---|
secret/data/proxima/nats/operator | Operator nkey seed |
secret/data/proxima/nats/account-proxima | Proxima account nkey seed |
secret/data/proxima/nats/account-SYS | SYS account nkey seed |
Disk keystore
The bootstrap tool is not Vault-only. --keystore=disk persists the same three
seeds as files on the keys volume and constructs no Vault client at all, so a
single-VM deployment can stand up JWT mode without running Vault:
PROXIMA_KEYS_DIR=/var/lib/proxima/keys \
sh -c 'cd backend && go run ./cmd/nats-jwt-init --keystore=disk'
Give PROXIMA_KEYS_DIR an absolute path. This tool runs from backend/ and
the server runs from the deployment root, so the relative default ./keys names
a different directory for each: the seeds would land in
backend/keys/nats/seeds/ while the server reads ./keys/nats/seeds/ and finds
no trust chain. The tool logs its resolved seed directory absolute so the
disagreement shows at bootstrap; --seed-dir overrides it for scripted use.
The --keystore flag defaults to $PROXIMA_KEYSTORE, so a disk deployment's own
environment cannot bootstrap seeds into Vault while its server looks for them on
disk. Seed locations resolve through the same function the server uses
(PROXIMA_KEYS_DIR + PROXIMA_NATS_SEED_DIR, 0700).
Path (under PROXIMA_KEYS_DIR) | Contents |
|---|---|
nats/seeds/operator.seed | Operator nkey seed |
nats/seeds/account-proxima.seed | Proxima account nkey seed |
nats/seeds/account-SYS.seed | SYS account nkey seed |
nats/seeds/revocations-<account>.json | Per-account JWT revocation list |
nats/seeds/revocations-<account>.iat | iat of the last account JWT built from that list (back it up and restore it together with the .json; a missing file reads as 0) |
Everything else the tool emits — the JWTs, both creds files,
resolver_preload.conf and the templated system_account in
nats-server-jwt.conf — is identical in both modes, so the output tables
above apply unchanged. Step 3 does change: Vault PKI is not available to issue
the NATS server certificate, so that material is supplied by the operator.
Minting is guarded on the whole seed set, not just the operator seed. With all three present the tool reuses them; with all three absent it mints; with some present and some missing it refuses, naming both lists. That last case — a partial restore, an interrupted copy — used to mint a new operator and overwrite the surviving account seeds, which silently invalidates every enrolled agent's credential. Restore the missing seeds from backup; emptying the seed directory is the deliberate way to re-root and re-enroll the fleet.
Rotation on this path is the same command. The backend and sys-admin creds
carry a bounded 90-day expiry, and nats-user-mint requires Vault, so re-running
nats-jwt-init --keystore=disk is how a Vault-free deployment refreshes them: it
reuses the seeds, so the operator and both accounts are unchanged and no agent
re-enrolls, while the two service users get fresh nkeys and a new expiry. Full
procedure in Rotating the service
credentials.
Step 3: Issue NATS Server TLS Certificate
./scripts/nats-cert-issue.sh [common_name] [ttl]
# Defaults: CN="nats.proxima.local", TTL="720h" (30 days)
Output files in infra/nats/certs/:
| File | Permissions | Purpose |
|---|---|---|
server.pem | 644 | Server TLS certificate |
server-key.pem | 600 | Server private key |
ca.pem | 644 | CA certificate (agents use this for verification) |
Step 4: Start NATS
The NATS server loads the operator JWT, preloads account JWTs, and enables TLS with mutual authentication using the nats-server-jwt.conf configuration.
Step 5: Start the Backend
Configure the backend with:
PROXIMA_NATS_CREDS_FILE=infra/nats/creds/backend.creds
PROXIMA_NATS_SYS_CREDS_FILE=infra/nats/creds/sys-admin.creds
PROXIMA_NATS_CA_FILE=infra/nats/certs/ca.pem
VAULT_ADDR=http://localhost:8200
VAULT_ROLE_ID=<from vault-init output>
VAULT_SECRET_ID=<from vault-init output>
Deployment: Local Dev / Bootstrap
In local development the bootstrap writes files under the repo's infra/nats/ directory and Docker Compose mounts them into the NATS container:
| Host path | Container path |
|---|---|
infra/nats/nats-server-jwt.conf | /etc/nats/nats-server-jwt.conf |
infra/nats/jwt/ | /etc/nats/jwt/ |
infra/nats/creds/ | /etc/nats/creds/ |
infra/nats/certs/ | /etc/nats/certs/ |
Deployment: Production (Kubernetes / ArgoCD)
The Console application tier runs on the proxima-production Kubernetes cluster (GitOps via ArgoCD),
but NATS itself does not: the broker runs under systemd on the dedicated nats01 VM. That
split decides how each side is configured.
- Backend side (Kubernetes). The backend's NATS credentials live in Vault and are injected
into the backend pods by the External-Secrets Operator (ESO). Nothing sensitive lives in
manifests, there is no
/opt/proximabox and no Docker Compose. Rotation flows through Vault plus an ArgoCD re-sync (or pod restart) of the backend. - Broker side (nats01). The operator/account JWTs, the resolver config and the server TLS
material are files on nats01 under
/etc/nats/. ESO does not reach this host. Rotation means writing the new material to/etc/nats/and reloading the unit —systemctl reload nats(the server reopens onSIGHUP) orsystemctl restart natswhen the change requires it.
Regenerating JWTs
If you need to regenerate JWTs (e.g., after key rotation or accidental deletion), follow this procedure. The nkey seeds in the key store are preserved, so account public keys remain the same and existing agent credentials continue to work.
Regenerating JWTs requires restarting NATS. All connected agents will briefly disconnect and automatically reconnect. Plan for a short maintenance window.
Local dev / bootstrap:
# 1. Ensure Vault is unsealed and accessible
export VAULT_ADDR=http://127.0.0.1:8200 VAULT_TOKEN=<token>
# 2. Regenerate JWTs, creds, resolver_preload.conf, and nats-server-jwt.conf
make nats-jwt-init
# 3. Restart NATS (Docker Compose) to pick up new JWTs
docker restart proxima-nats
# 4. Restart the backend to use new credentials, then verify it reconnected
Production: regeneration writes the refreshed JWTs and creds back to Vault. The two sides then diverge:
- Backend — ESO syncs the new credentials into the backend pods; an ArgoCD re-sync or pod restart picks them up. No file copies.
- nats01 — the refreshed operator/account JWTs and
resolver_preload.confmust be placed under/etc/nats/on the VM and the unit reloaded (systemctl reload nats, orrestartif the change needs it). This host is outside the cluster, so neither ESO nor ArgoCD touches it.
What make nats-jwt-init does on re-run:
- Loads existing nkey seeds from the key store — Vault, or the keys volume with
--keystore=disk(does not generate new keys) - Creates the account JWTs (
proxima.jwt,sys.jwt) from the stored revocation list at its storediat— the same JWT the backend last pushed — and writes nothing to the key store. The server applies a preload only when itsiatis not older than the JWT it holds, so a re-run can never un-revoke a credential - Regenerates
resolver_preload.confwith those account JWTs - Updates
system_accountinnats-server-jwt.conf(handles both placeholder and existing values) - Regenerates backend and sys-admin credential files with new user nkeys
What happens to connected agents:
- Agents disconnect briefly when NATS restarts
- Agents reconnect automatically (NATS client has built-in reconnection)
- Existing agent JWTs remain valid (same account keys); revoked ones stay revoked
- Agent JWT renewal continues to work (backend has new account signing keys from the key store)
Agent Enrollment Flow
When an agent starts for the first time, it performs a single-call enrollment with the backend over HTTPS:
Step-by-Step
-
Enroll: Agent sends
POST /api/v1/agents/enrollwith the install token and optionallyagent_type("host","collector","k8s-node-monitor", or"probe", default"host"). The backend validates the token (SHA-256 hash lookup, not expired), generates an nkey pair server-side, resolves client and environment from the token, upserts an agent record (not a host), creates a scoped JWT namedagent-{uuid}, and returns everything the agent needs. -
Persist: Agent writes
state.jsoncontaining the JWT, nkey seed, NATS URL,agent_id,agent_type, and client/environment slugs. If a CA PEM is included in the response (private CA deployments), it writesca.pemalongside. -
Connect: Agent connects to NATS using
nats.UserJWT()callbacks that re-read the current JWT + nkey seed fromstate.jsonon every (re)connect. This is deliberate: a JWT renewed in state is then picked up on the next connect, so the renewer can apply a fresh credential to the live connection (see JWT Renewal) instead of the connection being pinned to the JWT captured at process start. -
Host record creation (
agent_type=hostonly): For a host agent that supplies ahostname, the backend pre-links thehostsrow at enrollment (best-effortINSERT … ON CONFLICT,agent_status='online') so per-host config resolves on the agent's very first config fetch; theInventoryWorkerlater upserts the same row (preserving theagent_idviaCOALESCE). When no hostname is supplied (or for collectors, which own no host row), the host record is instead created on the first inventory message. -
Subsequent starts:
state.jsonwins only while it is usable. If it is present and valid, the agent skips enrollment, starts its renewal loop before connecting to NATS (so a JWT that expired while the agent was down renews over HTTPS first), and connects; credential refresh is handled by the renewal loop, not by re-enrolling. Ifstate.jsonis missing, corrupt or invalid, or a renewal is refused withexpired_too_longoragent_not_found, the agent enrolls with its configured install token instead (see Recovering an agent whose credential expired). A second enrollment of an already-enrolled identity (sameenvironment_id,agent_type,name) that maps to a different agent is rejected with409 Conflict— identity-takeover protection (AUTH H-1): an unbound install token cannot rebind an existing agent's nkey. Re-keying an existing agent is done with a bound token from Fleet → Agents → Re-issue enrolment, which re-binds that same agent row.
Install Token Properties
| Property | Value |
|---|---|
| Format | 32 cryptographically random bytes, hex-encoded |
| Storage | SHA-256 hash in database (plaintext shown once at creation) |
| Usage | Single-use by default (max_uses=1). Reuse is an explicit opt-in at creation: a positive max_uses bounds it to N enrollments, and max_uses=0 opts into unlimited use (e.g. a fleet/reusable token). The single-use default ensures one leaked token cannot enroll a whole fleet. |
| Default expiry | 24 hours |
| Maximum expiry | 30 days |
| Enrollment tracking | enrolled_by column on agents table |
| API | POST /api/v1/clients/{slug}/install-tokens (create), GET (list), DELETE (revoke) |
JWT Renewal
Agent JWTs have a 7-day default TTL. A background renewal loop ensures uninterrupted connectivity:
Renewal Details
| Parameter | Value |
|---|---|
| Check interval | Every 1 hour |
| Renewal threshold | When remaining TTL < 50% of original (e.g., < 3.5 days for a 7-day JWT) |
| Challenge nonce | Before renewing, the agent fetches a server-issued single-use nonce (NATS proxima.system.agent.renew.nonce, or HTTPS POST /api/v1/agents/credentials/renew/nonce) — gives each renewal replay freshness (AUTH F2). Older backends without the endpoint fall back to a legacy nonce-less path. |
| Primary channel | NATS request-reply on proxima.system.agent.renew (30s timeout) |
| Fallback channel | HTTPS POST /api/v1/agents/credentials/renew |
| Proof of identity | Agent signs the server-issued nonce with its nkey seed from state; the backend verifies the signature and cross-checks the nkey against the enrolled agent |
| Identity lookup | Backend extracts agent_id from JWT name (agent-{uuid}) and validates against the agents table |
| Legacy JWT rejection | JWTs with the host-{uuid} name prefix never renew. Within the grace window the refusal is jwt_invalid; past it — every one still in the field, since they were minted before 2026-04-13 with a 7-day TTL — it is expired_too_long, which makes an agent with an install token enroll with it. Both answer HTTP 400, and only the JWT's own exp decides between them |
| On success | state.json atomically updated first; only once it is on disk does the new JWT reach the agent's memory (if the save fails — a full disk — the agent keeps its current credentials and retries the renewal as a transient failure, so it never presents one JWT while believing it holds another). Then the agent tells its NATS connection the credentials changed (transport.CredentialsRenewed): a connected connection reconnects at once and presents the new JWT (the nats.UserJWT callbacks read from state on each connect); one that is still reconnecting has its auth-error backoff reset, so its next attempt uses the new JWT within about one reconnect wait. Without this the connection would stay pinned to the old JWT until it expired and the server dropped it. |
| Agent JWT TTL | 7 days (natsjwt.CreateUserJWT) |
| Status gate | active and inactive agents renew (inactive is only "silent for 300s" — exactly when renewal is needed); revoked is refused. A successful renewal never changes status; the next heartbeat reactivates a stale agent |
| Expired JWT | Renews if it expired at most PROXIMA_AGENT_RENEW_EXPIRED_GRACE ago (default 2160h, 90 days), and only with a server-issued nonce signature — the legacy signature over the JWT is refused for an expired JWT. Over NATS an expired JWT cannot connect, so this is the HTTPS path in practice. Past the grace window: re-issue enrolment |
| Revocation list | A JWT whose subject is on the account revocation list with a revocation time at or after the JWT's iat is refused (revoked), on both transports |
| Mint under the row lock | After the rules pass, the new JWT is minted while the agent row is held FOR SHARE, after re-checking that the row is not revoked and still holds the request's nkey. A re-bind or revoke that commits first is seen (nkey_mismatch, revoked); one that arrives during the mint waits for it, so its revocation is stamped after the minted JWT's iat and covers it |
| Refusal reasons | Every refusal carries reason: jwt_invalid, expired_too_long, revoked, nkey_mismatch, nonce_invalid, agent_not_found (plus bad_request, internal_error). NATS: reason beside error in the reply. HTTPS: error.details.reason; status codes unchanged: expired_too_long and nkey_mismatch are 401, except that a legacy host- JWT past the grace window answers expired_too_long with 400; jwt_invalid is 401 for a JWT that fails validation and 400 when its name carries no usable agent id (including a legacy host- JWT within the grace window); nonce_invalid is 401 for a failed or stale nonce signature and 400 for an undecodable signature or a missing nonce under PROXIMA_REQUIRE_RENEWAL_NONCE; bad_request is 400; revoked is 403; agent_not_found is 404; internal_error is 500, or 503 when the renewal challenge store is not available. Counted on proxima_agent_renewal_refused_total{reason}, except internal_error, which is logged at ERROR instead — a database or Vault failure during renewal is internal_error, never agent_not_found |
| On failure | Transient failures (network, 5xx, internal_error, nonce_invalid, bad_request, no reason) retry at the next hourly check. If the JWT has already expired when a transient failure happens, the agent retries sooner instead: 30s after that failure, doubling after each further one, capped at the hourly check. expired_too_long / agent_not_found with an install token configured → the agent enrolls with the token. revoked, jwt_invalid, nkey_mismatch, or those two without a token → terminal: one ERROR per hourly attempt naming the operator action; the collector's /healthz returns 503 |
| Broker refusals | The NATS connection no longer closes for good after repeated auth errors. While the broker keeps refusing the credential, reconnect attempts back off exponentially up to 5 minutes, so a credential fixed out of band (not through a renewal by this agent) can take up to 5 minutes to take effect |
Agent Lifecycle Management
Administrators manage hosts and agents through three actions. Only Revoke acts on the agent identity (the agents table); Deactivate and Activate act on the host record and its asset.
| Action | API Endpoint | DB Effect | JWT Effect | NATS Effect |
|---|---|---|---|---|
| Deactivate | POST /api/v1/hosts/{id}/deactivate | hosts.is_active = false and asset inactive; the agent row is not touched | None: renewal continues | None |
| Activate | POST /api/v1/hosts/{id}/activate | hosts.is_active = true and asset active; the agent row is not touched | None | None |
| Revoke | POST /api/v1/hosts/{id}/revoke (host and its agent), or POST /api/v1/fleet/agents/{id}/revoke (any agent, including one with no host, such as a prober) | status = 'revoked' on the agent (terminal) | Revocation entry added to account JWT; renewal refused (revoked) | Immediate disconnect |
Revoke is the single audited way to stop an agent. Deactivating a host never blocks its agent's credential renewal.
Revocation
When a host is revoked, the backend:
- Adds a revocation entry (the host's nkey public key + timestamp) to the
proximaaccount JWT - Signs the updated account JWT with the account nkey
- Pushes the updated JWT to the NATS server via
$SYS.REQ.CLAIMS.UPDATE(using sys-admin credentials) - The NATS server immediately rejects the agent's connection and any reconnection attempts
This provides instant, server-enforced disconnection without waiting for the JWT to expire.
The revocation list is one record per account (Vault KV v2, or revocations-<account>.json on the disk keystore), and every write is a merge into it. On Vault each write is a check-and-set against the version it read: a write that loses a race with another revoke (or a re-bind revoking the nkey it replaced) re-reads, re-applies its entry and retries, up to 5 attempts, and then fails — the revoke reports an error rather than dropping the entry. The disk keystore holds one lock across the read, the merge and the write. Before this, two overlapping revokes could each overwrite the other's entry, leaving an agent shown as revoked whose JWT NATS still accepted. The same record holds the iat of the last account JWT built from it, and every write takes a strictly greater one. The NATS server applies a direct claims-update push without comparing iat, and two crossing pushes can leave it enforcing the older list, so every push — from every backend replica and from nats-user-revoke — holds a per-account Postgres advisory lock across reading the committed list, pushing it, and the server's reply; a push that times out keeps the lock and is re-sent on the same connection (the server handles one connection in order, so the reply proves the delayed push was applied first), up to 3 attempts or 15 s, after which it is an error. The caller giving up does not shorten that hold: an HTTP client disconnecting during a revoke, or a re-bind's 15 s budget running out, ends only the wait for the lock, never a hold whose push is in flight (otherwise the next holder could push, be told it succeeded, and then be overtaken by the delayed bytes). A re-bind whose request has ended by then rolls back instead of committing, so its token stays unspent. If the push connection reconnected between attempts, the re-push stops with an error, because a reply on a new socket proves nothing about bytes held on the old one. A refusal is an error at once, and the push connection never replays a failed push after a reconnect. A server stalled for longer than that cap can still apply a delayed push late; a periodic check of the enforced account state is a planned follow-up. nats-user-revoke takes the same lock and therefore needs the Console database; it has no override for a database outage, because it cannot tell whether a backend replica can still push. During a rolling upgrade a replica on an older release can still push unlocked and drop the recorded iat; ordering holds once every replica runs this release.
Recovering an agent whose credential expired
With this agent release and backend, most expired credentials recover without an operator:
| Situation | What happens |
|---|---|
JWT expired, agent on this release, HTTPS reachable, expired for at most 90 days (PROXIMA_AGENT_RENEW_EXPIRED_GRACE) | Renews on its own: immediately at boot (the renewer runs before NATS connects), or at the next hourly check while running. If that attempt fails transiently while the JWT is already expired, the next one comes 30s later, doubling up to the hour. |
| JWT expired, agent older than this release | Upgrade the agent first: put the new binary or image in place and restart it. Self-update cannot reach it, because updates arrive over NATS. On the new version, a JWT expired at most 90 days with state.json intact renews on its own at boot — no Re-issue. (An old process that is still running, with its hourly renewer live, may also renew once the backend is deployed, but its NATS connection may have closed for good.) Re-issue is needed only past the grace window, after state loss, or for an agent that cannot be upgraded. |
| Expired for more than 90 days, state lost, or an old dead collector | Operator: Fleet → Agents → Re-issue enrolment → put the token where the agent reads it (below) → restart the agent or pod (on the new image, for collectors). It re-binds the same agent. |
Collector refused with 409 (its name is already enrolled and its token is unbound) | Terminal: the collector logs one ERROR per hourly attempt and fails /healthz, so the pod crash-loops visibly. Re-issue enrolment for that agent: a bound token re-binds the existing row, so the name collision does not apply. |
| k8s-node-monitor expired for more than 90 days, or state lost | Known limitation. The DaemonSet reads one shared install-token Secret on every node, so a per-node re-issued token cannot be delivered through it (see below). Recovering node monitors in bulk needs part B of this work — freeing an agent's identity, built on the existing release-identity API — which is parked. |
| Agent revoked | Terminal (revoked). Revocation is not undone by re-issue (it answers 409 agent_revoked); bring the machine back as a new agent, or through release-identity. |
Re-issue enrolment
Fleet → Agents → row menu → Re-issue enrolment (API: POST /api/v1/fleet/agents/{agentID}/reissue-enrollment, answered 201 with the token).
- It mints a single-use, 24-hour install token bound to that agent. The secret is shown once (
Cache-Control: no-store) and the issue is audited asagent.enrollment_reissued, without the secret. - Only the newest re-issued token for an agent works: issuing one supersedes every earlier unused re-issued token for the same agent.
- Issuing changes nothing on the agent. An agent seen in the last 5 minutes is refused with
409 agent_liveunless the request confirms it (confirm_live: true), because whoever enrolls with the token takes over the identity and the running process's credential is revoked at that moment. A revoked agent is refused with409 agent_revoked. - Who may re-issue: an unrestricted super admin, or a user holding
agents:writeon the agent's client and environment — the same gate as Revoke. The Fleet routes also require a client-wideagents:read, so an operator whose grants are only environment-scoped is refused with403today (a known gap, tracked as a follow-up); ask a client-wide administrator.
Where to put the token:
| Agent | Location |
|---|---|
| Collector (Helm chart) | The install-token Secret: <release fullname>-install-token, or your installToken.existingSecret, under its key (installToken.existingSecret.key, default token). If the chart is managed by GitOps (Argo CD), change it at the GitOps source; a hand-edited Secret is reverted. Once the collector has enrolled, put the original token back: the re-issued one is spent, and a spent token is refused (401, terminal) the next time this collector needs one |
| k8s-node-monitor (Helm chart DaemonSet) | Not through the shared Secret. Every node's pod reads the same Secret, and a re-issued token belongs to exactly one agent: any other node that presents it is refused (409 bound_agent_mismatch, the token stays unspent), and while it replaces the reusable fleet token no new node can enrol. Keep the reusable fleet token in the Secret. See the known limitation above |
| Host agent or prober (Linux) | Environment=PROXIMA_INSTALL_TOKEN=… in the systemd unit the installer wrote, then systemctl daemon-reload and restart |
| Host agent (Windows) | The service's PROXIMA_INSTALL_TOKEN environment variable, then restart the service |
What the agent does with it. A bound token is one agent's identity: it is accepted only from an agent of the same type, in the same environment, with the same name (its hostname, or the collector's pod name) as the agent it was issued for. Anything else is refused with 409 bound_agent_mismatch, and the token is not spent and the previous credential stays valid; the refused agent logs, hourly, that the token was re-issued for a different agent. Enrolling with a bound token re-binds the same agent id (same name and history), with a new nkey; the previous nkey is revoked before the re-bind commits. The agent replaces its state and reconnects in place. An unbound install token that yields a new agent id makes the agent write the new state and exit with code 75; its supervisor (systemd or the kubelet) restarts it under the new id.
The agent uses a configured install token only when stored state cannot recover (missing, corrupt or invalid state.json, or a renewal refused expired_too_long / agent_not_found), never on a network error, a 5xx or a refusal without a reason.
Do not use Re-issue enrolment until every backend replica runs this version (and so after its migration). A replica from before it does not know bound tokens: it treats a re-issued token as an ordinary install token, so the agent is refused with 409 (terminal for an hour) or, on a quiet row whose machine id matches, re-keyed without revoking its previous nkey.
Before the first deploy of this backend, audit the account revocation list against the nkeys of active agents: renewal now checks that list on both transports, so an active agent with a stale revocation entry will be refused as revoked.
File Layout Reference
Agent Files
| File | Path | Permissions | Purpose | Created |
|---|---|---|---|---|
state.json | {data_dir}/state.json | 0600 | All agent state: JWT, nkey seed, agent_id, agent_type, client/environment slugs, NATS URL, version | After enrollment |
ca.pem | {data_dir}/ca.pem | 0644 | NATS server CA certificate (only with private CA) | After enrollment (if private CA) |
Default data_dir: /var/lib/proxima-agent
Infrastructure Files
| File | Path | Purpose |
|---|---|---|
operator.jwt | infra/nats/jwt/operator.jwt | NATS operator identity |
proxima.jwt | infra/nats/jwt/proxima.jwt | Proxima account JWT |
sys.jwt | infra/nats/jwt/sys.jwt | SYS account JWT |
resolver_preload.conf | infra/nats/jwt/resolver_preload.conf | Account JWT preload for NATS config |
backend.creds | infra/nats/creds/backend.creds | Backend user credentials |
sys-admin.creds | infra/nats/creds/sys-admin.creds | Sys-admin user credentials |
server.pem | infra/nats/certs/server.pem | NATS server TLS certificate |
server-key.pem | infra/nats/certs/server-key.pem | NATS server TLS private key |
ca.pem | infra/nats/certs/ca.pem | TLS CA certificate |
Security Properties
- Server-side key generation: The nkey pair is generated on the backend. The seed is sent once over TLS during enrollment, which is equivalent security to sending the JWT over the same connection.
- Single-use install tokens by default: A token is consumed by one enrollment unless created with a higher
max_usesbound (ormax_uses=0for an unlimited fleet/reusable token). Authorization is the install token itself — possession of a valid, unconsumed token grants enrollment — so the single-use default limits the blast radius of a leaked token. - Default-deny permissions: Each agent JWT allows only its own agent-scoped subjects. Publishing to another agent's subjects is rejected by the NATS server.
- Identity-takeover protection (AUTH H-1): The first enrollment binds the agent's nkey. A second enrollment of an already-enrolled
(environment_id, agent_type, name)mapping to a different agent is rejected with409 Conflict— a fresh, unbound install token cannot rebind (hijack) an existing agent's NATS identity. Pod restarts and re-deployments are safe because the agent reuses its persistedstate.jsonwhile it is usable (credential refresh is handled by renewal). The one deliberate way to re-key an existing agent is a bound token from Re-issue enrolment (single use, 24 hours, gated like Revoke and audited), which re-binds the same agent row — only for an enrolling agent with that row's type, environment and name — and revokes its previous nkey before it commits. - Mandatory TLS: TLS 1.2+ with strong cipher suites is required for all NATS connections. The agent verifies the server certificate against the CA the backend hands it at enrollment — issued by Vault PKI by default, or read from
PROXIMA_CA_CERT_FILEon the keys volume in disk mode. - Key-store seed storage: Operator and account nkey seeds are stored in Vault KV v2 by default. Only the bootstrap tool and backend access them. With
PROXIMA_KEYSTORE=diskthey are files under<keys>/nats/seeds/in a0700directory instead — that directory holds the operator's private key, so it belongs on an encrypted volume and must never leave the host in an unencrypted backup. - Rate limiting: Credential endpoints are rate-limited per IP (configurable via
PROXIMA_CREDENTIAL_RATE_LIMIT, default 10/min) to prevent brute-force attacks on install tokens. - Immediate revocation: Compromised agents can be instantly disconnected by pushing a revocation entry to the NATS server, without waiting for JWT expiry.
Configuration Reference
Backend Environment Variables
| Variable | Default | Description |
|---|---|---|
PROXIMA_NATS_CREDS_FILE | — | Path to backend NATS credentials file |
PROXIMA_NATS_SYS_CREDS_FILE | — | Path to sys-admin NATS credentials file (for JWT revocation) |
PROXIMA_NATS_CA_FILE | — | Path to NATS TLS CA certificate |
PROXIMA_AGENT_RENEW_EXPIRED_GRACE | 2160h | How long after its JWT expired an agent may still renew it (nonce proof only). Past it: expired_too_long and Re-issue enrolment |
VAULT_ADDR | — | Vault server address |
VAULT_TOKEN | — | Vault root token (from vault-init.sh output or infra/vault/.vault-keys) |
VAULT_ROLE_ID | — | Vault AppRole role ID (production — use instead of VAULT_TOKEN) |
VAULT_SECRET_ID | — | Vault AppRole secret ID (production — use instead of VAULT_TOKEN) |
Agent Environment Variables
| Variable | Default | Description |
|---|---|---|
PROXIMA_BACKEND_URL | — | Backend HTTPS URL for enrollment and renewal. |
PROXIMA_INSTALL_TOKEN | — | Install token for enrollment. Also used after enrollment when stored state cannot recover (see Recovering an agent whose credential expired); a token from Re-issue enrolment re-binds the same agent. |
PROXIMA_NATS_CA_FILE | — | Path to NATS TLS CA certificate |
PROXIMA_AGENT_DATA_DIR | /var/lib/proxima-agent | Directory for persistent agent data (state.json) |
PROXIMA_CLIENT_SLUG and PROXIMA_ENVIRONMENT_SLUG are not required. The client and environment are determined during enrollment from the install token's associated client.