Skip to main content

Disk Keystore (Vault-free Backend)

PROXIMA_KEYSTORE=disk runs the backend with its secret material in files on a local volume instead of in HashiCorp Vault. It exists for a single-VM deployment, where running, unsealing and backing up Vault costs more than it buys.

What this is, and what it is not

This page documents a backend capability: the server, and the NATS bootstrap tool, run without Vault. It does not document an installer. There is no single-VM package, no Compose bundle and no console-ctl — an operator creates the keys volume by hand today. Production Console remains the Vault-backed Kubernetes deployment described in Production (Kubernetes).

Cluster and Vault are the defaults everywhere. A deployment that sets nothing new is byte-for-byte the deployment it was.

What the disk path supplies​

With PROXIMA_KEYSTORE=disk the boot switch selects, in place of the Vault equivalents:

ConcernVault pathDisk path
NATS operator/account nkey seedsVault KV v2files under <keys>/nats/seeds/
TLS CA an agent verifies the NATS server againstVault PKIPROXIMA_CA_CERT_FILE + PROXIMA_CA_KEY_FILE
SSH CA (terminal certificates)sshca.VaultSigner (Vault SSH secrets engine)sshca.PersistentLocalSigner — ed25519 keys sealed under the KEK in <keys>/ssh/
Stored-secret encryption (SecretEncryptor)Vault TransitAES-256-GCM under the KEK
Credential resolution (Provider)Vault Transitthe same KEK, wrapping the encryptor

Both credential seams work Vault-free. SecretEncryptor covers alert-source tokens, webhook signing secrets, pull-source credentials and the sealed SSH CA keys; Provider covers the /api/v1/clients/{id}/credentials surface and the NATS config-fetch path the agent resolves its collector credentials through.

The failure this shape guards against is a partial one, which is why it is worth stating. With no provider registered, the credentials API is not uniformly broken:

  • Create fails with 400 whose body reads unknown credential provider: local.
  • Update fails the same way, but only when the request changes the secret; a metadata-only edit succeeds.
  • Test returns 200 carrying {"status":"error","message":"unknown provider: local"} — the lookup is the first step of credential.Tester.
  • List, read and delete succeed. Delete does reach for the provider, but only to clean up a ProviderRef, and neither provider populates one — both return a SecretRef carrying EncryptedData alone — so the branch is never entered and delete returns 204.

So two surfaces name the cause outright, create and test, and three look untroubled. A credential list that renders normally is not evidence the provider is wired.

Meanwhile the agent's config fetch resolves through the same registry, and one unresolvable credential costs it its whole configuration, of every type — not just that credential. Boot therefore registers the provider and the encryptor once, in the switch, and passes both into the worker construction, so there is no second gate to drift out of step with the API's.

What the disk path does not supply​

AI Chat is unavailable on the disk keystore

Chat encrypts its per-client data-encryption key (getOrCreateDEK(ctx, clientID)) through Vault Transit specifically, not through the SecretEncryptor seam, and has no local equivalent. AdminDeps.VaultClient is documented as "nil = chat disabled", and the encryption service, model router, LLM client, message builder, tool service and both chat handlers are constructed inside if adminDeps.VaultClient != nil && …. With PROXIMA_KEYSTORE=disk there is no Vault client, so none of them is built and the chat routes are never mounted. (The four chat stores are built above that guard, unconditionally; nothing reaches them.)

The blast radius is wider than chat. That same if adminDeps.VaultClient != nil && … block also builds llmClient and the Telegram reactive ops-bot's engage handler, and both are hoisted out of it precisely because something else consumes them. So on the disk path, two more surfaces go with chat:

  • POST /api/v1/search/translate answers 503 translate_unavailable. The natural-language→PQL translator is built only when llmClient != nil, and the handler returns that status when its translator is nil. Structured search itself is unaffected — only the NL front door.
  • POST /api/v1/telegram/ops/engage is never mounted. The route is registered inside if cfg.telegramEngage != nil, and that handler is constructed in the Vault-gated block from the same llmClient. The reactive ops-bot therefore cannot be engaged; the rest of the Telegram integration is not built from that block and is unaffected.

How you would find out: one line at startup — WARN AI chat disabled ... reasons=["Vault not configured"]. That is the whole list: ChatCfg is always non-nil and PROXIMA_LITELLM_URL falls back to a non-empty default, so the other possible reason cannot fire. And that warning is the only signal — no runtime error, no /readyz degradation, no UI banner.

And it is only absent while the client's ai_chat feature stays off, which is its default (LaunchDefaults ships ai_chat: false). The sidebar gates the link on chat:read and that flag alone, with no knowledge of whether the backend mounted the routes — so turning the flag on for a client of a Vault-free install gives them a visible AI Chat link that leads to a page calling routes that do not exist. Leave it off.

This is a known gap, not a misconfiguration you can fix with settings. Giving chat a local DEK path is real implementation work and belongs with the single-VM package (Phase B). Until then, do not sell or plan around AI Chat, natural-language search or the reactive ops-bot on a Vault-free install.

Two other features had the same shape and were fixed in this phase rather than documented — the credentials API (which resolves through a Provider, now implemented locally) and the MetricsQL query proxy (which hand-built cluster-only URLs). Chat is the one that remains, because it needs a new key path rather than a switch.

The keys volume​

Everything lives under PROXIMA_KEYS_DIR (default ./keys). Each file name is joined under that directory unless it is absolute, so a single artifact can be relocated without moving the volume. See Keystore in the environment-variable reference for the full table.

keys/                       # 0700
├── kek.bin # 0600 — 32 raw bytes; group/other-readable FAILS THE BOOT
├── ca/
│ ├── ca.pem # TLS CA certificate
│ └── ca-key.pem # SEC1 `EC PRIVATE KEY` block — PKCS#8 and ed25519 are rejected
├── nats/seeds/ # 0700
│ ├── operator.seed
│ ├── account-proxima.seed
│ └── account-SYS.seed
└── ssh/ # created by the server on first boot
├── user_ca.sealed
└── host_ca.sealed

The server reads this volume; with one exception it does not create it. The exception is ssh/, which the server generates on first boot and then requires on every boot after.

Creating the keys volume​

There is no bootstrap tool for this yet — console-ctl bootstrap is Phase B. Until it ships, the KEK and the CA pair are created by hand. The commands below are the ones the end-to-end fixture runs (scripts/e2e/setup_suite.bash:545-563), so they are exercised on every disk-keystore test run rather than written from memory:

KEYS=/var/lib/proxima/keys           # absolute — see the warning below
SVC=proxima # the user the backend process runs as

mkdir -p "$KEYS/ca" "$KEYS/nats/seeds" "$KEYS/ssh"
chmod 700 "$KEYS" "$KEYS/ca" "$KEYS/nats" "$KEYS/nats/seeds" "$KEYS/ssh"

# 32-byte key-encryption key. Must not be group- or world-readable.
head -c 32 /dev/urandom > "$KEYS/kek.bin"
chmod 600 "$KEYS/kek.bin"

# TLS CA. -noout is REQUIRED (see below).
openssl ecparam -name prime256v1 -genkey -noout -out "$KEYS/ca/ca-key.pem"
openssl req -x509 -new -key "$KEYS/ca/ca-key.pem" -days 365 \
-subj "/CN=proxima-ca" -out "$KEYS/ca/ca.pem"
chmod 600 "$KEYS/ca/ca-key.pem"
-noout is load-bearing

Without it, openssl ecparam -genkey writes an EC PARAMETERS block before the EC PRIVATE KEY block. ca.NewStaticProviderFromPEM decodes only the first PEM block in the file, so the boot fails with parse CA key — a message that reads like a corrupt key rather than an extra header.

These are the fixture's commands with two deliberate changes, both cosmetic: the subject CN (it uses proxima-e2e-ca) and the lifetime — 30 days there because it is a test rig, 365 here. Pick a lifetime you can live with.

Replacing this CA later is not free: the PEM is handed to an agent only in the enrollment response — credential renewal does not carry it — so an already-enrolled agent keeps the CA it was given until it is re-enrolled or the file is replaced on the host out of band.

ssh/ is left empty here. The server generates and seals its own SSH CA keys into it on first boot, so it needs no content — only the right owner, which is the last step.

The NATS chain is generated by the bootstrap tool, which follows the same switch:

PROXIMA_KEYS_DIR=/var/lib/proxima/keys \
sh -c 'cd backend && go run ./cmd/nats-jwt-init --keystore=disk'

PROXIMA_KEYS_DIR must be absolute here. This tool runs from backend/ and the server runs from the deployment root, so the relative default ./keys names a different directory for each: the bootstrap would write the operator seed to backend/keys/nats/seeds/ and the server would look in ./keys/nats/seeds/ and find no trust chain. The tool logs its resolved seed directory absolute so the disagreement is visible at bootstrap rather than much later.

The --keystore flag defaults to $PROXIMA_KEYSTORE, and the seed directory is resolved by the same function the server resolves it with, so a disk deployment cannot bootstrap seeds into Vault while its server looks for them on disk. Everything else the tool emits — JWTs, creds files, resolver_preload.conf, the templated system_account — is identical in both modes. See NATS Security.

The NATS server's own TLS material is not produced by any of this: Vault PKI issued it before, and on the disk path an operator supplies it.

Ownership — the last step​

The fixture creates and consumes this volume as one user, so it needs no ownership step. A real install does: the directories are created by root and the backend runs as a service user, and at that point the 0700 directories and the 0600 KEK are unreachable to the process.

Do this last, after nats-jwt-init above has written the seeds. Running it before that leaves the seeds root-owned inside a service-user volume, which reproduces the problem one step after fixing it:

chown -R "$SVC:$SVC" "$KEYS"

Everything under $KEYS must be readable by the service user. Two subdirectories must additionally be writable by it at runtime:

  • ssh/ — the server generates and seals its own SSH CA keys there on first boot, and refuses to start if it cannot.

  • nats/seeds/ — revoking an agent writes the account's revocation list back into the seed directory as revocations-proxima.json. Exactly two operator actions reach it:

    • POST /api/v1/hosts/{hostID}/revoke — the host Revoke action.
    • DELETE /api/v1/probe-locations/{id} — deleting a probe location, which revokes its prober's credentials on the way out.

    Deleting a host does not. DELETE /api/v1/hosts/{hostID} removes the row and marks the asset decommissioned; it never touches the keystore. So testing this hazard by deleting a host returns 204 and proves nothing — use Revoke.

    A read-only seed directory boots fine and then fails the first revocation, which is the operation you least want to discover is broken.

Where a permissions problem actually surfaces: not where you would guess. The boot touches this volume in four places, in this order — ca.NewStaticProvider (cmd/server/main.go:322), the NATS operator and account seed reads via natsjwt.NewManager (:399), credential.LoadKEK (:485), and finally the sealed SSH CA keys (:754).

So the first message is disk CA provider init failed: read CA cert: permission denied. Fix only that and the next boot stops on the seeds — and not quickly: NewManager retries five times with 1/2/4/8-second backoff, so you get four WARN NATS JWT manager init failed, retrying lines over about fifteen seconds before the fatal NATS JWT manager init failed after 5 attempts: … permission denied. The KEK error is reachable only once both of those are readable, and the SSH CA only after that. Chasing them one at a time costs a boot cycle each and grepping the logs for LoadKEK finds nothing. Fix ownership for the whole volume at once.

Do not widen the modes instead of fixing ownership. LoadKEK refuses a KEK that group or other can read, and fails the boot rather than warning — on a single VM that file is all that separates an unprivileged local account from every stored secret.

Operational consequences​

Four of these are new behavior, and each is deliberate. Read them before you run this mode, not during the incident.

Rotating the service credentials​

This is a dated fuse, and nothing reminds you

nats-jwt-init mints the backend and sys-admin creds with a bounded 90-day expiry (defaultServiceCredsTTL). On day 90 the broker stops accepting them and the backend's messaging goes down — agent config pushes, JetStream consumption, JWT revocation. There is no warning, no /readyz degradation and no metric on the remaining lifetime.

The rotation command on the Vault path is make nats-user-mint, and it does not work here: that tool hard-requires VAULT_ADDR/VAULT_TOKEN and has no --keystore flag, so a disk install gets VAULT_ADDR is required. Giving it one is Phase B.

The remedy on the disk path is to re-run the bootstrap, which is safe because minting is guarded on the seed set:

PROXIMA_KEYS_DIR=/var/lib/proxima/keys   sh -c 'cd backend && go run ./cmd/nats-jwt-init --keystore=disk'
# then install the two new creds files where the server reads them, and restart it:
# PROXIMA_NATS_CREDS_FILE -> infra/nats/creds/backend.creds
# PROXIMA_NATS_SYS_CREDS_FILE -> infra/nats/creds/sys-admin.creds

It does not re-root the trust chain. With a complete seed set present the tool reuses the seeds and re-derives everything from them: the operator subject, the system account and the proxima account that signs the backend's user JWT are byte-identical across runs. What changes is the two service users — each gets a fresh user nkey and an expiry of now + TTL. Every enrolled agent's credential is signed by the same unchanged account, so nothing re-enrolls, and the broker's trust anchor is the same operator it already has.

Two things to know before you run it:

  • The seed set must be complete. If some seeds are present and others are missing — a partial restore, an interrupted copy — the tool now refuses, naming which it found and which it did not, rather than minting a new operator over the survivors. Restore the missing seeds from backup. Emptying the seed directory is the deliberate way to re-root, and it re-enrolls the whole fleet.
  • Paths are relative to the working directory. The tool writes its JWTs and creds to ../infra/nats/{jwt,creds} and templates system_account into ../infra/nats/nats-server-jwt.conf. Run it from a checkout's backend/ as above and copy the two creds files out, or run it from a scratch tree seeded with a copy of that config — which is what the e2e fixture does, so the repository's tracked config is never edited. Only --seed-dir / PROXIMA_KEYS_DIR points at the keys volume.

Bring the cadence forward with PROXIMA_NATS_SERVICE_CREDS_TTL. It is parsed as a Go duration, so use hours: 2160h is the 90-day default and 720h is 30 days. 90d is not valid — the loader logs invalid duration env var, using fallback and silently keeps 90 days.

<keys>/ssh/ is boot-blocking, not just backup-critical​

The SSH CA keys are generated on first boot and sealed under the KEK. From then on, a sealed key file that exists but cannot be read, unsealed or parsed fails the boot. Nothing on that path rewrites the file, so a CA that is merely unreadable today — wrong KEK, wrong file mode, wrong volume mounted — stays recoverable once the cause is fixed.

Refusing to start is the point. The alternative, which is what this replaced, silently minted a fresh in-memory CA per process: the API would keep answering 200 to certificate mints while every agent, which pins the user CA in its authorized_keys line, rejected every one of them. One unreadable file would have presented as a fleet-wide permission denied (publickey) with no obvious cause.

So: back up <keys>/ssh/ with the rest of the volume, and treat a boot failure naming the SSH CA as a key-custody problem, never as something to clear by deleting the file.

Losing the keys volume is unrecoverable​

There is no escrow and no derivation. If the volume is lost:

  • The KEK cannot be regenerated. Every value sealed under it — stored credentials, alert-source tokens, webhook signing secrets, pull-source credentials, the SSH CA — is permanently unreadable. Those secrets must be re-entered from their sources.
  • Every agent must re-enroll. The NATS operator seed is the root of the trust chain; without it no account JWT can be re-signed and no agent credential renewed.

Back the volume up encrypted, off the host, and restore-test it. It holds the NATS operator's private key: leaking it is equivalent to leaking the operator, which is why the seed directory is 0700 and the KEK refuses to load from a file that group or other can read.

There is no Vault → disk migration​

Ciphertexts written by Vault Transit begin vault:v1:. The local encryptor writes local:v1: and cannot read the other.

A deployment that switches PROXIMA_KEYSTORE from vault to disk therefore cannot read any secret it already stored — credential rows, webhook signing secrets, alert-source tokens, pull-source credentials. Re-entering them is the only path.

Where the prefix shows up matters, because it is not everywhere. The SecretEncryptor surfaces pass the error through — a pull source's last_error reads decrypt credentials: credential: ciphertext is not local:v1-encrypted, and the SSH CA names it in the boot failure. The credential surfaces do not. Tester.Test substitutes failed to decrypt credential data, and the agent's credential resolver substitutes decryption failed; on both, the prefix survives only inside an OpenTelemetry span.

What an operator sees is worse than the substitute, because it takes work to see anything at all. The edit dialog's connection test is automatic but gated: it clears the form's secret fields on open and stays idle — rendering no result row whatsoever — until every required field is non-empty. For postgresql (dsn), redis (addr) and nginx (url) that means a blank panel on open; the operator must retype the secret and wait out an 800 ms debounce before failed to decrypt credential data appears. Only docker and generic, which have no required fields, test on open. And that message never names the prefix — it is in Tempo, not in the response.

This is a product decision, not an oversight: a silent fallback that tried Vault on a vault:v1: value would reintroduce the dependency the mode exists to remove. Choose the keystore when the deployment is created.

Two other variables belong to the same single-VM story and are documented in full in the environment-variable reference:

  • PROXIMA_VM_SINGLE_NODE — target a single-node VictoriaMetrics binary instead of a cluster. A precondition, not a guarantee: nothing enforces one client, and where more than one exists metric reads are not tenant-isolated. The backend logs a warning at boot if it sees more than one client in this mode, which catches the obvious misconfiguration but cannot prevent a client created afterwards.

    Five per-client read sites widen, and two of them return metric values, not just names: the client metric-name listing, the client series listing, the client-level aggregate chart (GET /api/v1/clients/{id}/metrics/{metricName}) and its derivative variant, and the per-client recorded series count in the usage recorder. The aggregate chart is the sharpest: with a client-wide grant the selector carries neither client_id nor environment_id, so the average a client-scoped user sees is taken over every client's hosts.

    The MetricsQL proxy (POST /api/v1/metrics/query) is not in that set. It rejects a request carrying neither host_id nor environment_id, and always appends the matching extra_filters[], so it stays label-isolated with no tenant path segment to lose.

  • PROXIMA_MIGRATE_ON_START — let the server apply its own embedded migrations, since a single-VM install runs no separate migrate job. Off by default. A failed migration leaves the ledger dirty and the server crash-looping — normally with the schema untouched, since each migration runs in one transaction; the recovery is in docs/ROLLBACK.md.