Comprehensive reference for all environment variables used across Proxima Console components.
Loading order: environment variables > hardcoded defaults. All components (backend, agent, MCP server) are configured via environment variables only — there is no YAML config file.
How config is supplied in production (Kubernetes)
In the proxima-production cluster (namespace console-system), the backend's environment
comes from two places, both managed via GitOps in the argocd-infra repo:
Non-secret defaults and endpoints are set as plain env on the backend Rollout pod spec
(e.g. PROXIMA_SERVER_BIND=0.0.0.0, the VictoriaMetrics/VictoriaLogs URLs below).
Secrets are injected by the External-Secrets Operator (ESO), which syncs them from the
in-cluster Vault (vault.prxm.uz) into a Kubernetes Secret the pod loads via envFrom.
No secret values appear in any manifest.
The variable names and meanings documented below are unchanged — only the source differs (pod
env + ESO-from-Vault in production, an .env file or shell in local dev). The
production endpoints they point at are called out in the relevant sections (
VictoriaMetrics, VictoriaLogs, OpenTelemetry,
Core). See Backend Deployment for the full topology.
NATS server URL. Production:10.10.4.15:4222 — the backend dials nats01 privately. Agents get nats.prxm.uz:4222 instead, via PROXIMA_NATS_PUBLIC_URL.
PROXIMA_SERVER_PORT
8080
No
HTTP server listening port
PROXIMA_SERVER_BIND
127.0.0.1
No
Address the HTTP server binds to. Production: set to 0.0.0.0 so kubelet probes and the in-cluster Service can reach :8080 (the istio gateway terminates TLS, not a same-host nginx).
PROXIMA_LOG_LEVEL
info
No
Log level: debug, info, warn, error
PROXIMA_CORS_ORIGINS
http://localhost:5173
No
Comma-separated allowed CORS origins
PROXIMA_GRPC_PORT
0 (disabled)
No
gRPC server port for CLI terminal access. Defaults to 0, which disables the gRPC server — set a non-zero port to enable it. Production: TLS is terminated at the istio Gateway and gRPC traffic on api-console.prxm.uz:443 (HTTP/2) is routed to this port.
PROXIMA_GRPC_BIND
127.0.0.1
No
Address the gRPC server binds to. Production:0.0.0.0 so the Service can reach the gRPC port.
PROXIMA_GRPC_TLS_CERT
—
No
Path to TLS certificate for gRPC server (empty = plaintext; in production the istio Gateway terminates TLS).
PROXIMA_GRPC_TLS_KEY
—
No
Path to TLS private key for gRPC server
PROXIMA_DOCS_URL
https://docs-console.prxm.uz
No
Base URL of the docs site. Used when the backend renders links into docs (help surfaces, notification bodies).
The audit log is partitioned by month. Months older than the retention period are archived to
the R2 bucket above (audit-archive/audit_log/YYYY/MM/, gzip NDJSON plus a manifest that
records how to verify it), read back and checksummed, and only then removed from PostgreSQL.
Without R2 nothing is ever removed.
Variable
Default
Required
Description
PROXIMA_AUDIT_RETENTION_ENABLED
false
No
Allow months past the retention period to be archived and removed. Off by default so no audit data leaves the database by surprise; partitions are maintained and overdue months reported either way.
PROXIMA_AUDIT_RETENTION_MONTHS
24
No
Retention period in months. A month is removed only once its last entry is this old. Minimum 12; the backend refuses to start below it.
Apply the migrations embedded in the server binary at startup, before anything touches a table. Leave this off on Kubernetes, where migrations are applied by an ArgoCD PreSync hook — two mechanisms racing to migrate one database during a blue-green rollout is worse than either alone. A single-VM install, which runs no separate migrate job, is the case this exists for. Only up-migrations are reachable this way; rolling one back stays a deliberate operator action with the migrate CLI (see docs/ROLLBACK.md). Failure mode to know before you turn this on: a failed migration leaves golang-migrate's schema_migrations ledger marked dirty, every subsequent start refuses to migrate, and the server crash-loops. The schema itself is normally untouched — each migration runs inside one transaction — so the usual remedy is to fix the SQL and force back to the previous version. Clearing it is a manual migrate force <version>, written up under "Recovering a dirty migration ledger" in docs/ROLLBACK.md.
Required. JWT signing secret. Must be at least 32 characters — the check is unconditional, so a shorter value makes the backend refuse to start. Generate with openssl rand -base64 48.
PROXIMA_JWT_ACCESS_TTL
900
Access token TTL in seconds (15 minutes). Must be 60–3600; out of range is a startup failure, not a clamp — access tokens are not individually revocable.
PROXIMA_JWT_REFRESH_TTL
604800
Refresh token TTL in seconds (7 days). Must be 300–2592000; out of range is a startup failure.
Absolute lifetime of a session measured from creation (30 days). Refresh is refused once a session is older than this, regardless of rotation. 0 disables the cap.
PROXIMA_SESSION_IDLE_TIMEOUT_SECONDS
1209600
Idle timeout — the maximum allowed gap between refreshes (14 days). Refresh is refused if the session has not been used within this window. 0 disables the idle timeout.
AI Chat requires LiteLLM gateway and Vault Transit for message encryption. Chat is automatically disabled if either is unavailable.
Variable
Default
Required
Description
PROXIMA_LITELLM_URL
http://litellm:4000
For AI chat
LiteLLM gateway URL
PROXIMA_LITELLM_MASTER_KEY
—
For AI chat
LiteLLM admin API key
PROXIMA_DEEPSEEK_API_KEY
—
For AI chat
Platform DeepSeek API key (standard tier)
PROXIMA_KIMI_API_KEY
—
For AI chat
Platform Kimi/Moonshot API key (complex tier)
PROXIMA_GOOGLE_API_KEY
—
For AI chat
Platform Google AI API key (expert tier)
PROXIMA_CHAT_RATE_LIMIT
30
No
Chat messages per minute per user
PROXIMA_CHAT_MAX_TOOL_CALLS
10
No
Max MCP tool calls per message (safety limit)
PROXIMA_CHAT_MAX_CONVERSATION_MESSAGES
100
No
Soft limit on messages per conversation
PROXIMA_CHAT_MAX_MESSAGE_LENGTH
10000
No
Maximum chat message length in characters
PROXIMA_CHAT_CONTEXT_WINDOW
50
No
Max messages sent to LLM as context (sliding window)
PROXIMA_CHAT_VAULT_KEK_NAME
proxima-chat-kek
No
Vault Transit key name for DEK encryption
PROXIMA_MCP_BASE_URL
http://localhost:8081
No
MCP server URL for tool execution
PROXIMA_LITELLM_TIMEOUT
15m
No
HTTP timeout for calls to the LiteLLM gateway. Must exceed the longest triage or chat completion.
PROXIMA_LITELLM_DB_URL
—
No
Optional read-only DSN to LiteLLM's own PostgreSQL (LiteLLM_SpendLogs). When set, the triage cost reconciler reads real spend/tokens/cache directly, because the gateway /spend/logs API is proxy-admin only. Empty falls back to the gateway API (works only where the key has admin).
PROXIMA_EMBED_MODEL
proxima-embed
No
Embedding model alias used for L1 prior-resolution search.
PROXIMA_CHAT_MODEL_STANDARD
—
No
Overrides the standard-tier chat model. Empty uses the chat package default. Set to a central LiteLLM alias.
PROXIMA_CHAT_MODEL_COMPLEX
—
No
Overrides the complex-tier chat model.
PROXIMA_CHAT_MODEL_EXPERT
—
No
Overrides the expert-tier chat model. Also the fallback L1 triage model when a client's selected triage alias is not configured in the gateway.
PROXIMA_CHAT_MODEL_TRANSLATE
claude-haiku-4-5-20251001
No
Fast/cheap model for structured translation tasks (e.g. natural language → PQL).
Where this deployment's secret material lives: the NATS operator and account nkey seeds, the TLS CA that agents verify the NATS server certificate against, the SSH CA, and the key-encryption key that seals stored credentials.
Variable
Default
Description
PROXIMA_KEYSTORE
vault
vault or disk. vault is the production posture and the default, so existing deployments are unaffected. disk keeps the same material in files on a local volume, for a single-VM deployment where running Vault is not warranted. An unrecognized value fails the boot rather than falling back. Note that VAULT_ADDR has a non-empty default, so unsetting it does not disable Vault — this variable is the only switch.
PROXIMA_KEYS_DIR
./keys
Root of the keys volume. Read only when PROXIMA_KEYSTORE=disk. Set it to an absolute path. The default is relative, and nats-jwt-init runs from backend/ while the server runs from the deployment root — so with the default the bootstrap writes seeds to backend/keys/nats/seeds/ and the server looks in ./keys/nats/seeds/. The tool logs its resolved directory absolute so the disagreement is visible at bootstrap rather than later, as a missing trust chain.
PROXIMA_CA_CERT_FILE
ca/ca.pem
TLS CA certificate. The backend hands this PEM to an agent at enrollment so the agent can verify the NATS server's certificate (nats.RootCAs); the broker sets verify: false and no client certificate is issued to anyone. Joined under PROXIMA_KEYS_DIR unless absolute.
PROXIMA_CA_KEY_FILE
ca/ca-key.pem
TLS CA private key. Required to construct the CA provider — the boot fails without it — though the backend calls no signing operation on it today. Must be a SEC1 EC key: a lone EC PRIVATE KEY PEM block. ca.NewStaticProviderFromPEM decodes only the first block, so openssl ecparam -genkeywithout -noout emits an EC PARAMETERS block first and the boot fails with parse CA key — see the worked example in Disk Keystore. A PKCS#8 (PRIVATE KEY) or ed25519 key is rejected too. Joined under PROXIMA_KEYS_DIR unless absolute.
PROXIMA_NATS_SEED_DIR
nats/seeds
Directory holding operator.seed, account-proxima.seed, account-SYS.seed and the per-account revocation lists. Created 0700. Joined under PROXIMA_KEYS_DIR unless absolute.
PROXIMA_KEK_FILE
kek.bin
32-byte key-encryption key for AES-256-GCM credential encryption and the sealed SSH CA keys. The file must not be group- or world-readable; loose permissions fail the boot. Joined under PROXIMA_KEYS_DIR unless absolute.
Disk mode does not create this material
The server reads the keys volume; nothing in it writes the CA pair, the KEK or the NATS seeds. An operator creates them (and runs nats-jwt-init --keystore=disk for the NATS chain) before the server starts — see Disk Keystore. The exceptions are the sealed SSH CA keys under <keys>/ssh/, which the server generates on first boot and then requires on every boot after.
There is no Vault → disk migration
Ciphertexts written by Vault Transit begin vault:v1:; the local encryptor writes local:v1: and cannot read the other. A deployment that switches PROXIMA_KEYSTORE from vault to diskcannot read any secret it already stored — credential rows, webhook signing secrets, alert-source tokens and pull-source credentials. Where the prefix actually shows up matters, because it is not everywhere. The SecretEncryptor surfaces pass the error through: a pull source's last_error reads decrypt credentials: credential: ciphertext is not local:v1-encrypted, and the SSH CA names it in the boot failure. The credential surfaces do not.Tester.Test replaces it with failed to decrypt credential data, and the agent's credential resolver with decryption failed — on both, the prefix survives only inside an OpenTelemetry span. What an operator sees is worse than the substitute, because it takes work to see anything at all. The edit dialog's connection test is automatic but gated: it clears the form's secret fields on open and stays idle — rendering no result row whatsoever — until every required field is non-empty. For postgresql (dsn), redis (addr) and nginx (url) that means a blank panel on open; the operator has to retype the secret and wait out an 800 ms debounce before failed to decrypt credential data appears. Only docker and generic, which have no required fields, test on open. And that message never names the prefix — it is in Tempo, not in the response. Re-entering the affected secrets is the only path. This is a deliberate product decision, not an oversight.
These variables configure NATS JWT/operator mode authentication and credential encryption. Seed storage and credential encryption go through the keystore selected by PROXIMA_KEYSTORE above: Vault KV v2 plus Transit on the Vault path, files on the keys volume on the disk path. The VAULT_* variables below are read only on the Vault path. See NATS Security and Credentials for details.
Variable
Default
Description
PROXIMA_NATS_CREDS_FILE
—
Path to backend NATS credentials file
PROXIMA_NATS_SYS_CREDS_FILE
—
Path to sys-admin NATS credentials file (for JWT revocation)
PROXIMA_NATS_CA_FILE
—
Path to NATS TLS CA certificate
VAULT_ADDR
http://127.0.0.1:8200
Vault server address (for nkey seed storage and credential encryption via Transit)
VAULT_TOKEN
—
Vault root token (from vault-init.sh output or infra/vault/.vault-keys)
VAULT_ROLE_ID
—
Vault AppRole role ID (production — use instead of VAULT_TOKEN)
VAULT_SECRET_ID
—
Vault AppRole secret ID (production — use instead of VAULT_TOKEN)
PROXIMA_NATS_PUBLIC_URL
value of PROXIMA_NATS_URL
The NATS URL handed to agents during enrollment. Set it when agents reach NATS on a different address than the backend does (e.g. backend uses the in-cluster Service, agents use nats-client-console.prxm.uz).
PROXIMA_NATS_VAULT_MOUNT
secret
Vault KV v2 mount holding per-host NATS nkey seeds.
PROXIMA_NATS_VAULT_PREFIX
proxima/nats
Path prefix under the KV mount for NATS secrets.
PROXIMA_NATS_PKI_MOUNT
pki
Vault PKI mount used to issue NATS server TLS certificates.
PROXIMA_NATS_PKI_ROLE
nats-server
Vault PKI role used when issuing those certificates.
PROXIMA_REQUIRE_RENEWAL_NONCE
false
Rejects legacy nonce-less agent credential renewals. Ships false so a live fleet keeps working during rollout; set true once every agent sends a server-issued nonce, which makes replay-fresh renewal mandatory.
PROXIMA_AGENT_RENEW_EXPIRED_GRACE
2160h
How long after its JWT expired an agent may still renew it (with a fresh server-issued nonce signature). Past it the agent is refused with expired_too_long and an operator re-issues enrolment from Fleet → Agents.
PROXIMA_VAULT_CREDS_KEK_NAME
proxima-creds
Vault Transit key name used as the KEK for stored credential encryption.
Google OAuth 2.0 client ID. When set (along with secret), the "Sign in with Google" button appears on the login page.
PROXIMA_GOOGLE_CLIENT_SECRET
—
Google OAuth 2.0 client secret.
PROXIMA_GOOGLE_ALLOWED_DOMAIN
proximaops.io
Google Workspace domain (hd) that SSO logins are restricted to; users outside it are rejected. Also gates auto-provisioning: when unset or empty, the backend will not auto-create or email-link accounts from a verified Google identity — SSO then succeeds only for accounts already explicitly linked by provider ID. Set it to enable safe open provisioning.
PROXIMA_BACKEND_URL
http://localhost:8080
Backend public URL, used to build OIDC redirect URIs.
PROXIMA_GOOGLE_SERVICE_ACCOUNT_KEY and PROXIMA_GOOGLE_ADMIN_EMAIL were documented ahead of a
feature that was designed but never built. No Go code reads either name — there is no Directory
API client and no group-sync worker. Setting them has no effect; group membership does not flow into
RBAC. Assign roles and clients explicitly instead. Tracked as a planned item in SPEC.md.
Enable /swagger and /metrics debug endpoints. Set to false in production to hide these endpoints.
PROXIMA_METRICS_TOKEN
—
When set, /metrics requires Authorization: Bearer <token>. Leave empty to serve Prometheus metrics unauthenticated (rely on network policy instead).
PROXIMA_TRUSTED_PROXY_COUNT
1
Number of trusted reverse-proxy hops in front of the backend. Controls how many entries are trimmed off X-Forwarded-For when deriving the client IP for rate limiting and audit logs. Set it to the real hop count — too high lets a client spoof its IP, too low rate-limits the proxy. Status pages use it for their IP allowlist: see Status Pages before raising it.
PROXIMA_MAX_REQUEST_BODY_BYTES
4194304 (4 MiB)
Maximum accepted request body size. Larger requests are rejected before handler dispatch.
PROXIMA_KUBE_ALLOWED_GROUPS
proxima:kube-readonly,proxima:kube-exec
Comma-separated allowlist of Kubernetes impersonation groups a role's kube_spec may request. Defaults to the two groups the Helm chart binds.
Max API requests per minute per IP on all /api/v1/ routes.
PROXIMA_AUTH_RATE_LIMIT
5
Max public auth requests per minute per IP (login, refresh, forgot-password, reset-password, accept-invite). Protected auth routes (sessions, me, mfa) use the general API limit.
PROXIMA_CREDENTIAL_RATE_LIMIT
10
Max agent credential requests per minute per IP (/agents/enroll, /agents/credentials/renew, and /agents/credentials/renew/nonce).
PROXIMA_WEBHOOK_RATE_LIMIT
60
Max inbound webhook requests per minute per IP (alert receivers, change-source webhooks, Jira/JSM callbacks).
Email for the initial super-admin user seeded at startup. Skipped if empty.
PROXIMA_ADMIN_PASSWORD
—
Password for the initial super-admin user. Must meet the password policy (12+ characters).
First Run
Set both variables on the first startup to create a super-admin account. The seed is idempotent — if a user with that email already exists, it is skipped. You can unset them after the first run.
Atlassian instance URL (e.g. https://yourorg.atlassian.net). When empty, the feedback widget logs submissions to structured logs instead of creating Jira issues.
PROXIMA_JIRA_EMAIL
—
Service account email for Jira API authentication (Basic Auth).
PROXIMA_JIRA_API_TOKEN
—
Jira API token (paired with email for Basic Auth).
PROXIMA_JIRA_PROJECT_KEY
—
Target Jira project key for feedback issues (e.g. PDD).
The worker ingests the proxima-wiki per-project change logs. They feed the staff-only change notes in Service Desk (the "Change note" card on a ticket, the request-list marker and the "Change noted, still open" filter) and the change_notes section of the alert context pack.
Variable
Default
Description
PROXIMA_WIKI_GITLAB_URL
—
GitLab base URL. Unset disables the wiki worklog worker.
PROXIMA_WIKI_GITLAB_TOKEN
—
Token with read_repository on the wiki project only. Unset disables the worker.
How long an upgrade target may stay in preflight, applying, verifying or rolling back. Past it a preflight is blocked (no_preflight_response) and every later state becomes unknown for a person to check. See Service Upgrades.
PROXIMA_UPGRADE_WORKER_INTERVAL
5s
How often the upgrade worker ticks over its claimed upgrades.
Public client status pages (see Status Pages). Turning them on is
an infrastructure rollout, not just a variable: follow docs/runbooks/status-pages-rollout.md.
Variable
Default
Description
PROXIMA_STATUS_PAGE_DOMAIN
—
Serve status pages at <slug>.<domain>, e.g. prxm.uz. Lower-cased, with any leading or trailing dot trimmed. A value that is not a domain name stops the backend at startup. Unset = the public handler is off; the Console screens, Preview and the Console login audience still work. Needs the origin lockdown, wildcard DNS and gateway route in the rollout runbook first
PROXIMA_STATUS_PAGE_RESERVED
—
Comma-separated labels no page may use, added to the built-in list in backend/internal/statuspage/slug.go (infrastructure hosts, common service names, brand names, and anything ending in -console or -rke2). Each entry must be in the slug format, or the backend stops at startup, since a typo would silently reserve nothing. A reserved label is never served as a status page: its requests go to the API router as before
PROXIMA_TRUSTED_PROXY_COUNT decides who a status page visitor is
A page's IP allowlist and the 120-requests-a-minute limit both use the client IP derived with
PROXIMA_TRUSTED_PROXY_COUNT. With Cloudflare in front of an ingress that appends to
X-Forwarded-For, the count is 2; at 1, every visitor shares one Cloudflare edge IP's budget. But raise it only after the origin accepts traffic from
Cloudflare alone: with a count of 2, a request sent straight to the gateway chooses its own
X-Forwarded-For entry and so its own IP, which defeats the allowlist. The rollout runbook makes
the origin lockdown its first step.
These variables configure the connection to VictoriaMetrics Cluster for time-series metrics storage.
Variable
Default
Description
VM_INSERT_URL
http://localhost:8480
vminsert endpoint for metrics ingestion. Production: the in-cluster VM cluster fronted by vmauth-proxy.monitoring:8427 (full path: vmauth-proxy.monitoring.svc.cluster.local:8427).
VM_SELECT_URL
http://localhost:8481
vmselect endpoint for metrics queries. Production: also vmauth-proxy.monitoring:8427.
VM_WRITE_TIMEOUT
10
Write timeout in seconds
VM_QUERY_TIMEOUT
30
Query timeout in seconds
VM_PROJECT_ID
0 (local dev)
Production: 42 — Console's isolated VM projectID. Combined with each client's per-tenant accountID, Console writes/reads tenant <accountID>:42, separate from the shared cluster's :0 tenants.
VM_AUTH_USERNAME
—
Production: vmuser-console — the VMUser that vmauth-proxy authenticates Console as.
VM_AUTH_PASSWORD
—
Password for the vmuser-console VMUser. Injected via ESO from Vault; never in manifests.
PROXIMA_VM_SINGLE_NODE
false
Target a single-node VictoriaMetrics binary instead of a cluster. Single-node VM has no tenancy, so URLs carry no /insert/<tenant>/ or /select/<tenant>/ segment and VM_PROJECT_ID is inert; point both VM_INSERT_URL and VM_SELECT_URL at the same :8428 address. Precondition, not a guarantee: enable only on an install with exactly one client. Nothing enforces it. With more than one client, metric reads are not tenant-isolated — every client's series share one namespace. Five per-client read sites widen: the metric-name and series listings, the recorded series count, and — returning metric values, not just names — the client-level aggregate chart (GET /api/v1/clients/{id}/metrics/{metricName}) and its derivative variant, whose average is then taken over every client's hosts. The MetricsQL proxy (POST /api/v1/metrics/query) is not affected: it requires a host_id or environment_id and always sends the matching extra_filters[]. The backend logs a warning at boot if it sees more than one client in this mode. Intended for the self-hosted single-VM install.
Local dev vs production
In local dev the backend talks to the docker-compose VM cluster directly on :8480/:8481
with no auth and no project isolation. In production all metric reads and writes go through
vmauth-proxy:8427 authenticated as vmuser-console, scoped to projectID 42. See
VictoriaMetrics Architecture.
How often the backend reads metrics for each enabled Hetzner Cloud pull source's load balancers and its servers without a Console agent, and writes them to the client's VictoriaMetrics account (the topology details panel reads them). It is also how late a cloud object's "now" can be. Each object costs one API call per interval: 12 an hour at 5m, against Hetzner's 3,600 requests an hour per Hetzner project (not per token), a budget shared with everything else that calls the API for that project. Below 1m, or unparseable, is refused and 5m stays in force. Off when no credential encryption (Vault Transit or the disk keystore) is configured.
One backend replica polls: the holder of a PostgreSQL session advisory lock, held across
intervals. The other replicas poll nothing, so the cost per object does not grow with the
replica count. The others ask for the lock every 30 seconds, so after a rolling deploy a new
replica polls within about 30 seconds.
Hetzner's limit is 3,600 requests an hour per Hetzner project, not per token. Every token
of the project draws on the same budget, and so does everything else that calls the API for
that project: the hcloud cloud controller manager (which reconciles the load balancers of the
Kubernetes clusters behind them), the CSI driver, Terraform, the structural poll (which lists
the project's servers, load balancers, networks, firewalls and other resources, a few calls
each, more for a large project, on the source's own poll interval) and any other Console source
on the same project. Without the lock, two replicas would each read every object: with 100
objects in one project at 5m, 2 × 100 × 12 = 2,400 metrics calls an hour, two thirds of the
project's limit. With it, the same fleet costs 1,200, a third. 300 objects at 5m
(3,600 ÷ 12) would take the whole budget and leave the cloud controller manager none; a
shorter interval lowers that in proportion, to 60 objects at 1m.
So the poller also guards the budget. Each Hetzner answer carries RateLimit-Limit and
RateLimit-Remaining; once fewer than a quarter of the project's requests are left, the
poller stops that source for the cycle, logs a WARN and counts it in
proxima_hetzner_metrics_rate_limited_total. It resumes at the next interval, from where each
object's last write ended (up to an hour back), so a cycle cut short loses no points.
If the polling replica's node dies without closing its database connection, PostgreSQL
keeps the lock until TCP keepalive notices, about 2 hours at the Linux defaults. No replica
polls in that time, and the replica that takes over backfills at most the last hour of each
object, so the rest of that gap stays empty.
The offline country file (DB-IP "IP to Country Lite", CC-BY 4.0) the topology map uses as the last source of a machine's location: the country of its first public address. The backend image ships it at this path, unpacked at image build from the repository's backend/data/dbip-country-lite.mmdb.gz and refreshed monthly. The path names the unpacked file: to see countries when running the backend outside the image, unpack the .gz and point this at the result. Missing or unreadable: one warning at startup, and no machine is placed by its address; the map still works. The file's build date is logged at startup, with a warning when it is over two months old. No address leaves Console.
Connection settings for VictoriaLogs (log ingestion and querying).
Variable
Default
Description
PROXIMA_VICTORIALOGS_URL
http://localhost:9428
VictoriaLogs base URL for log ingestion and queries. Production: VictoriaLogs runs under systemd on vmstorage02 (:9428); set PROXIMA_VL_PROJECT_ID=42 alongside it.
PROXIMA_VL_PROJECT_ID
0
VictoriaLogs projectID for tenant isolation. 0 keeps the legacy bare-accountID path.
OTLP exporter endpoint (gRPC). Production:Tempo runs in-cluster and receives OTLP on :4317; its blocks are stored in Cloudflare R2.
PROXIMA_OTEL_SERVICE_NAME
proxima-backend
Service name for trace identification
PROXIMA_OTEL_TRACES_SAMPLER
always_on
Trace sampler strategy
PROXIMA_OTEL_TRACES_RATIO
1.0
Trace sampling ratio (0.0 to 1.0). Only used when sampler is traceidratio.
Trace-ingest reuse
PROXIMA_OTEL_EXPORTER_ENDPOINT is also the endpoint the trace-ingest worker forwards agent spans to (agent spans arrive over NATS; the backend re-exports them to the same Tempo it sends its own spans to). The worker no-ops cleanly if this is unset. See Distributed Tracing.
The native escalation engine replaced Grafana OnCall. It arms a per-alert-group timer, walks
the matched escalation policy's steps, and hands each step to the per-user notification chain.
See Escalation and Paging Control.
Variable
Default
Description
PROXIMA_ESCALATION_POLL_INTERVAL
2s
How often the escalation timer worker ticks its planner over due (armed) alert-group timers. A non-positive value falls back to the default.
PROXIMA_ESCALATION_AUDITOR_INTERVAL
60s
How often the escalation auditor sweeps for steps that should already have fired.
PROXIMA_ESCALATION_AUDITOR_OVERDUE_THRESHOLD
5m
A step overdue by this much enters the healing zone (the auditor tries to recover it).
PROXIMA_ESCALATION_AUDITOR_DIRTY_THRESHOLD
15m
A step still overdue by this much is unrecoverable and is marked dirty. Must be greater thanPROXIMA_ESCALATION_AUDITOR_OVERDUE_THRESHOLD.
PROXIMA_ESCALATION_AUDITOR_HEARTBEAT_URL
—
External dead-man's-switch URL pinged on each clean sweep. Empty disables the ping (the gauge is still published). See Escalation Dead Man's Switch.
PROXIMA_NOTIFICATION_CHAIN_INTERVAL
2s
How often the notification-chain timer worker ticks its planner over due per-user chains. Mirrors the escalation poll interval. See Notification Chains.
PROXIMA_NOTIFICATION_TRANSIENT_MAX_ATTEMPTS
3
Maximum delivery attempts per channel before a transient (retryable) send error is treated as terminal.
PROXIMA_SCHEDULE_AUDITOR_INTERVAL
1h
How often the on-call schedule auditor sweeps for coverage gaps and empty shifts.
PROXIMA_SCHEDULE_AUDITOR_WINDOW
336h (14d)
Forward look-ahead over which schedule gaps are detected.
PROXIMA_SCHEDULE_AUDITOR_GAP_TOLERANCE
60s
A coverage hole at or below this length is a back-to-back handoff, not a gap.
PROXIMA_DISPATCH_AUDITOR_INTERVAL
1m
How often the dispatch auditor sweeps for deliveries that never reached a terminal state.
PROXIMA_DISPATCH_AUDITOR_CUTOFF
5m
A sending claim older than this with no terminal state is an orphan. The default safely exceeds the fast Send return of every channel.
The readiness drill is a human dead-man's switch: it voice-calls the current on-call at a
deterministic-random instant (at most once per on-call per 24h) to confirm reachability, and
pages the covering team through their critical escalation policy when the call goes unconfirmed
past the grace window. The worker is neither constructed nor started unless
PROXIMA_ONCALL_READINESS_DRILL=onand the ops-bot token is set. See
Readiness Drill.
Variable
Default
Description
PROXIMA_ONCALL_READINESS_DRILL
off
Master switch (on / off).
PROXIMA_ONCALL_READINESS_INTERVAL
1m
How often the drill worker ticks its firing/miss/escalate sweep.
PROXIMA_ONCALL_READINESS_GRACE
10m
Nudge-to-escalate grace window: after a missed call and the readiness nudge, the on-call has this long to confirm before the covering team is paged.
PROXIMA_ONCALL_READINESS_WINDOW
— (any hour)
Local clock band the drill's target instant is restricted to, so a readiness call never rings at 03:00. Format HH:MM-HH:MM, e.g. 09:00-21:00.
PROXIMA_ONCALL_READINESS_COVERAGE_ROLE
critical
Inert since 2026-10-05 — the coverage-gap arm is unwired, so nothing reads this. Formerly: which team escalation-policy role the coverage-gap page was armed through. See the readiness drill.
PROXIMA_ONCALL_REROUTE
off
Soft-absence reroute. When on, a real incident whose resolved on-call user is readiness-degraded also pages the next escalation tier — strictly additive, the primary is always paged first. When off, the fire path never consults the degraded check.
The shadow-poll drift worker compares Console's on-call resolution against a still-running
Grafana OnCall during migration. It is constructed only when bothBASE_URL and
API_TOKEN are set.
Variable
Default
Description
PROXIMA_ONCALL_BASE_URL
—
Grafana OnCall API root (e.g. https://oncall.example.com). Empty disables the shadow worker.
PROXIMA_ONCALL_API_TOKEN
—
Grafana OnCall API token, sent as the Authorization header. Empty disables the shadow worker.
PROXIMA_ONCALL_SHADOW_POLL_INTERVAL
60s
How often the shadow worker polls OnCall and compares. The default matches the flip gate's clean-poll-count math.
Strict alert-ingestion contract. When on, a webhook delivery whose parsed payload carries no alert content at all is refused with a 400 instead of accepted as an empty alert, and a delivery that carries content but no grouping key gets a per-content synthesized key. Observation mode is always wired: both conditions are counted on every delivery regardless of this flag, so the blast radius is measurable per source before you enable it — read the Would-Reject Ratio by Source panel first. See docs/standards/alert-ingestion.md.
Recovers firing alert groups whose current episode never reached a paging decision (a lost hand-off from the alert worker to correlation, or an arm error). See Paging safety nets.
Variable
Default
Description
PROXIMA_CORRELATION_WATCHDOG_INTERVAL
60s
How often one replica sweeps (advisory lock) for undecided firing episodes. A non-positive value is refused at load and the default stays in force.
PROXIMA_CORRELATION_WATCHDOG_GRACE
2m
How long an episode may stay undecided before its correlation is republished; also the minimum spacing between two retries of the same episode, across all replicas. Raising it delays every recovered page by the same amount. A non-positive value is refused at load.
PROXIMA_CORRELATION_WATCHDOG_MAX_RETRIES
5
Republishes per firing episode before giving up. Must be >= 1; a lower value is refused at load and the default stays in force. A failed publish is refunded, so only republishes actually sent count. When spent, the episode will not page: an ERROR is logged and proxima_alert_correlation_unrecovered_total is incremented once — alert on it.
PROXIMA_VOICE_PROVIDER selects between the self-hosted Asterisk path and Twilio;
empty auto-selects based on which credentials are configured. Voice is disabled entirely when
neither provider is configured. See Voice Paging and
Voice Trunk Health.
Variable
Default
Description
PROXIMA_VOICE_PROVIDER
— (auto)
asterisk, twilio, or empty to auto-select.
PROXIMA_VOICE_FROM_DEFAULT
—
Fallback caller ID in E.164 format.
PROXIMA_VOICE_FROM_998
—
Caller ID used for Uzbekistan (+998) destinations, in E.164. On the Sarkor trunk this is the trunk DID.
PROXIMA_VOICE_PUBLIC_BASE_URL
—
Publicly reachable base URL for provider callbacks (Twilio TwiML fetch, status callbacks).
PROXIMA_VOICE_VERIFIED_GATE
warn
warn logs and proceeds when a destination number is unverified; enforce blocks the call.
PROXIMA_VOICE_TTS_RENDER_URL
—
/render endpoint of the proxima-tts service (a cluster-internal Service, not a process on the Asterisk VM). Empty disables rendered speech: every voice page plays the pre-recorded prompt.
PROXIMA_VOICE_TTS_MAX_TEXT_CHARS
125
How many characters of the alert summary are spoken. Refused above 2000 (the render service rejects longer text outright).
PROXIMA_VOICE_TTS_FALLBACK_URI
sound:proxima-alert
ARI media value played when TTS rendering fails, is not configured, or the prefetched sound is not confirmed. Must exist in Asterisk's sounds directory.
PROXIMA_VOICE_FETCHER_TOKEN
—
Bearer token for the two sound-prefetch endpoints the fetcher on the Asterisk VM calls (GET /api/v1/voice/pending-sounds, POST /api/v1/voice/pending-sounds/{name}/ready). The only authenticator on those routes, and not a user session. Empty → the routes are not registered, nothing is ever confirmed, and every voice page speaks the generic prompt. Set-but-unusable is refused at startup: fewer than 16 printable non-space characters, or a value that could not survive an HTTP header (surrounding whitespace, or a trailing newline — what vault kv get -field=… > file produces). Must match the value on the voice VM exactly. See Voice Paging.
PROXIMA_VOICE_HEALTH_INTERVAL
30s
How often the voice trunk health probe runs.
PROXIMA_VOICE_LEADER_ACQUIRE_INTERVAL
5s
How often a non-leader replica retries acquiring the voice leader lock. Only the leader holds the ARI WebSocket, so exactly one replica drives calls.
PROXIMA_ASTERISK_ARI_URL
—
ARI base URL (wss/https), e.g. https://voice.prxm.uz:8089. Empty disables the Asterisk voice path.
PROXIMA_ASTERISK_ARI_USER
—
ARI username.
PROXIMA_ASTERISK_ARI_PASSWORD
—
ARI password. Never logged.
PROXIMA_ASTERISK_ARI_APP
proxima-voice
Stasis application name the backend subscribes to.
PROXIMA_ASTERISK_TRUNK
sarkor-trunk
PJSIP trunk (endpoint) name used for outbound calls.
PROXIMA_ASTERISK_SIP_DOMAIN
—
Provider host for explicit-URI dialing (e.g. bell.uz). Empty falls back to the AOR-lookup form.
PROXIMA_ASTERISK_WS_PING_INTERVAL
20s
ARI WebSocket keepalive ping interval.
PROXIMA_ASTERISK_WS_READ_TIMEOUT
60s
ARI WebSocket read timeout. An unset or non-positive value defaults to 60s.
PROXIMA_ASTERISK_WS_WRITE_TIMEOUT
10s
ARI WebSocket write timeout.
PROXIMA_TWILIO_ACCOUNT_SID
—
Twilio account SID for outbound calls.
PROXIMA_TWILIO_AUTH_TOKEN
—
Twilio auth token. Also the HMAC key used to verify the X-Twilio-Signature header on inbound callbacks. Never logged.
Bounds on the L1 incident agent's investigation loop. The defaults fit a deep grounded
investigation (roughly 17 LLM round-trips) plus the verify pass, while staying under the
JetStream AckWait. See Alerting.
Variable
Default
Description
PROXIMA_TRIAGE_DEADLINE
300s
Per-triage wall-clock bound. Too low and the run force-concludes to an unverified stub instead of a real RCA.
PROXIMA_TRIAGE_TOKEN_BUDGET
120000
Cumulative token bound per triage.
PROXIMA_TRIAGE_MAX_LOOPS
20
Maximum investigation tool-loop iterations.
PROXIMA_TRIAGE_OPERATIONAL_NAK_BACKOFF
5m
Fixed NAK delay for an operational LLM failure (credit, quota, overloaded, rate limit) — a longer floor than the ramping transient backoff, because such errors do not self-heal in seconds.
PROXIMA_TRIAGE_VERIFY_DEADLINE
60s
Per-verify wall-clock bound. The verify pass gets its own deadline so a long investigation cannot starve it into a fail-open "unverified".
PROXIMA_TRIAGE_VERIFY_TOKEN_BUDGET
8000
Token bound for the verify pass.
PROXIMA_TRIAGE_VERIFY_MAX_RETRIES
1
Extra verify attempts when a verify pass times out. Each retry gets a fresh per-attempt deadline bounded by the parent triage context. 0 disables retries.
PROXIMA_TRIAGE_TRUST_MIN_EVIDENCE
2
Cited-evidence count that floors a failed verify at unconfirmed instead of unverified. 0 disables the floor.
PROXIMA_TRIAGE_SHORTCIRCUIT_MAX_LOOPS
2
Tool-loop cap for the cheap confirm run taken when a strong prior resolution matches.
PROXIMA_TRIAGE_SHORTCIRCUIT_TOKEN_BUDGET
15000
Token bound for that confirm run.
PROXIMA_TRIAGE_SHORTCIRCUIT_DEADLINE
60s
Wall-clock bound for that confirm run.
PROXIMA_TRIAGE_TEMPERATURE
0.1
LLM temperature for both the investigation and verify passes.
PROXIMA_TRIAGE_CONFIDENCE_THRESHOLD
0.5
Minimum calibrated confidence for a verified trust level. Below it the RCA cannot reach verified even when the verifier supports the claims.
PROXIMA_TRIAGE_MEMORY_SHORTCIRCUIT_DISTANCE
0 (disabled)
Cosine-distance ceiling at or below which a held, non-recurred prior resolution may short-circuit the full investigation with a cheap confirm run. 0 disables the short-circuit entirely — with the default, the three SHORTCIRCUIT budgets below are never reached. Production uses 0.15.
Comma-separated change_events.source values produced by cloud-pull workers. These rows have a null host_id by construction, so host-scoped correlation never reaches them — this list is what surfaces them client-scoped.
PROXIMA_CORRELATION_CHANGE_WINDOW
24h
How far before an alert's fire time the correlation worker looks for host changes and Kubernetes events. When the window is empty the worker falls back to the host's last-N changes labeled with their true offset.
PROXIMA_CORRELATION_DEPLOY_STALE_THRESHOLD
72h
Age above which the client's newest deploy/git change is flagged stale — observability only: a WARN and proxima_correlation_deploy_feed_stale_total fire so a non-delivering deploy webhook is visible. 0 disables the check.
PROXIMA_TOPOLOGY_EDGE_TTL
168h (7d)
Stale-prune window for the L1 dependency graph: network_edges and external_edges rows whose last_seen is older than this are dropped, so stopped dependencies age out.
PROXIMA_TOPOLOGY_WEAK_EDGE_TTL
168h (7d)
Cutoff for weak edges (see below). Equal to PROXIMA_TOPOLOGY_EDGE_TTL by default, so the weak prune deletes nothing extra: the topology graph offers display windows up to 7d, and windowing cannot resurrect a pruned row, so retention must outlive the longest offered window or the window control shows less than it claims. Transient noise is handled at draw time (chatter sidelining, edge verification) instead. Lower it only if you accept that lightly-observed edges vanish from the wider windows.
PROXIMA_TOPOLOGY_WEAK_OBS_MAX
2
An edge seen at most this many collection cycles counts as weak/transient.
PROXIMA_TOPOLOGY_EXTERNAL_EDGE_CAP
500
Flood guard on external_edges (observed connections to destinations that are not a Console host): the most rows one source host may hold, one row per destination address, port, process and source namespace/workload — so a host can reach it with fewer distinct addresses. At the cap a new row is skipped (counted as proxima_topology_edges_dropped_total{reason="external_cap"}), including a known address seen on a new port, process or source namespace/workload, while existing rows keep refreshing. External rows age out on PROXIMA_TOPOLOGY_EDGE_TTL. Zero or negative falls back to the default.
PROXIMA_TRIAGE_CHANGE_CAUSE_MIN_SCORE
0.35
Candidate changes scoring below this are dropped rather than surfaced as possible causes.
ivfflat.probes applied per prior-resolution similarity query. The PostgreSQL default of 1 usually probes an empty cell over the small resolution-stub corpus and silently drops the true nearest neighbour.
PROXIMA_RECURRENCE_CHRONIC_WINDOW
168h (7d)
Lookback window over which recurrences are counted.
PROXIMA_RECURRENCE_CHRONIC_THRESHOLD
3
Occurrences within that window that mark an alert family as chronic.
PROXIMA_RECURRENCE_FIX_REGRESSION_WINDOW
48h
How long after a confirmed fix a fresh occurrence counts as the fix having regressed.
PROXIMA_MEMORY_RETRIEVAL_MAX_DISTANCE
0.25
Cosine-distance ceiling above which a retrieved prior resolution is discarded as too dissimilar — the junk filter.
PROXIMA_TRIAGE_MEMORY_STRONG_DISTANCE
0.12
Tighter ceiling at or below which a retrieved prior resolution (also not-recurred and not-failed) is marked strong and elevated as a leading candidate cause.
Stale-incident reaper window: an open incident with no firing member activity in this window is auto-resolved, closing the active-grouping blackhole. A chronic alert that keeps re-firing keeps its members fresh and is never reaped. The same TTL drives the stale alert-group reaper: a firing group not updated within it is resolved. Alertmanager's repeat_interval must be shorter than this, or a still-firing alert is reaped between repeats.
When PROXIMA_RETENTION_ENABLED is false the retention worker no-ops every tick and deletes
nothing for any client. Per-client overrides (client_retention_settings) inherit any field
left unset by resolving against these global defaults.
Variable
Default
Description
PROXIMA_RETENTION_ENABLED
false
Master kill-switch for change-detection data retention.
PROXIMA_RETENTION_TIMELINE_DAYS
365
change_events horizon in days; events older than this are pruned. Floor 7.
PROXIMA_RETENTION_CONTENT_DAYS
90
file_versions content horizon in days; versions older than this and beyond the last-K are pruned. Floor 1.
PROXIMA_RETENTION_CONTENT_KEEP_LAST
20
Per-file count of newest versions always kept regardless of age. Floor 1.
PROXIMA_RETENTION_INTERVAL
6h
Worker sweep cadence.
PROXIMA_RETENTION_BATCH_SIZE
5000
Bounds each batched DELETE so prune statements stay short.
The in-app activity feed behind the bell in the header — not the paging chain. Nothing here
decides whether anyone is woken up, and shortening the window cannot suppress a page. See
Notification Feed.
Zero disables the sweep, it does not mean "keep nothing"
A value of 0 (or any negative number) turns the retention sweep off entirely and
notification_events then grows without bound. This knob is independent ofPROXIMA_RETENTION_ENABLED above — that kill-switch governs change-detection data only, and it
ships false, so coupling the two would mean the feed never expired on most installs.
Variable
Default
Description
PROXIMA_NOTIFICATION_RETENTION_DAYS
30
Days of in-app notification feed events kept before the sweep deletes them. Zero or negative disables the sweep entirely. The sweep runs every 6h (not configurable) and reports proxima_retention_rows_pruned_total{table="notification_events"} plus proxima_notification_retention_last_success_timestamp_seconds — pair them, since the counter records nothing for a sweep that deleted zero rows.
The Service Desk bot / Mini App and the Ops bot are distinct deployments with distinct
tokens. The backend's Telegram auth bridge is registered only whenPROXIMA_TELEGRAM_BOT_TOKEN is set — with it empty, the /api/v1/auth/telegram* routes do not
exist. See Telegram Overview and Ops Bot.
Variable
Default
Description
PROXIMA_TELEGRAM_BOT_TOKEN
—
Service Desk bot token from @BotFather. The backend uses it to validate Mini App initData via HMAC-SHA256. Empty leaves the Telegram auth routes unregistered.
PROXIMA_TELEGRAM_BOT_USERNAME
—
Service Desk bot username, used when the backend renders link-invitation and welcome-email deep links. Not the paging bot — see PROXIMA_OPS_TELEGRAM_BOT_USERNAME.
PROXIMA_TELEGRAM_BOT_SA_ID
—
Service account ID the Service Desk bot authenticates as when calling the Console API.
PROXIMA_OPS_TELEGRAM_BOT_TOKEN
—
Ops bot token. The bot that delivers per-user pages, and the master gate for the native on-call stack. Also gates the readiness drill — the drill worker does not start without it.
PROXIMA_OPS_TELEGRAM_BOT_USERNAME
—
Ops bot username. Builds the paging-verification deep link (t.me/<bot>?start=vfy_…). It must name the same bot as the token above — a Telegram bot may only DM someone who started a chat with it, so a username pointing at the Service Desk bot would let a responder verify and still be undeliverable. Empty makes POST /api/v1/telegram/verify/start answer 503 rather than emit a broken link.
PROXIMA_OPS_BOT_SA_ID
—
Service account ID the Ops bot authenticates as. Pins both the alert-action callback and paging-verification redemption.
Base URL of the internal invoice service (POST /api/v1/invoices/prefill, client-facing billable hours).
PROXIMA_INVOICE_JWT_SECRET
—
Invoice service HS256 shared secret used to mint the per-call service token. Empty disables the client — callers fail soft and fall back. Never logged.
One-time install token for enrollment. Not needed after initial enrollment.
PROXIMA_AGENT_TYPE
host
No
Agent type: host (full monitoring), k8s-node-monitor (kubelet metrics only), collector (K8s API inventory), or probe (an uptime-monitor pop). The Helm chart sets k8s-node-monitor for DaemonSet agents.
PROXIMA_PROBE_ALLOW_PRIVATE_TARGETS
false
No
Probe agents only. Lets this prober dial private address space: RFC1918, unique-local fc00::/7, carrier-grade NAT 100.64.0.0/10, and the special-purpose IPv4 ranges 240.0.0.0/4 (apart from 255.255.255.255, which stays refused), 198.18.0.0/15 and 0.0.0.0/8 (apart from 0.0.0.0, which stays refused). Only the exact values true and 1 enable it; any other value, TRUE included, leaves it off. Set it only on a prober you run inside the network it monitors, never on a shared public pop — monitor targets are tenant-authored. Loopback, link-local (including 169.254.169.254) and the other always-refused classes stay refused either way. Read once at startup, so it takes a restart; the prober logs probe dial policy installed with allow_private_targets. See What a prober will refuse to dial.
PROXIMA_PROBE_DEFAULT_RESOLVER
—
No
Probe agents only. The DNS server a dns check asks when its monitor names no resolver, written like a monitor's resolver: a host, host:port, or a bracketed IPv6 address such as [2606:4700:4700::1111]:53; a bare host gets port 53. Empty uses this host's system resolver. A prober that serves more than one tenant must set it to a public resolver. It is dialed through the same address guard as a monitor's resolver. Probe mode refuses to start on a value it cannot use: one it cannot parse, a port that is not a number from 1 to 65535, or an IP address its dial policy refuses — a private address unless PROXIMA_PROBE_ALLOW_PRIVATE_TARGETS is set, or loopback on any prober. A hostname is judged by the guard when it is dialed. Read once at startup, so it takes a restart; when it is set, the prober logs probe default resolver installed with the host:port it will dial in default_resolver, and it warns when neither this nor PROXIMA_PROBE_ALLOW_PRIVATE_TARGETS is set. See What a prober will refuse to dial.
PROXIMA_AGENT_UPGRADES_ENABLED
off
No
Opts this host in to service upgrades: the agent installs and rolls back apt packages as root when the backend asks. Only the exact values true and 1 enable it; any other value, TRUE included, leaves it off. Honoured only on host agents: a k8s-node-monitor agent refuses with upgrades_disabled, and probe and collector agents never handle upgrade commands (their hosts are blocked with no_preflight_response). Read once at startup, so it takes a restart; the agent logs service upgrades with enabled.
PROXIMA_NATS_CA_FILE
—
No
Path to NATS TLS CA certificate (received during enrollment)
PROXIMA_LOG_LEVEL
info
No
Log level: debug, info, warn, error
Enrollment
PROXIMA_BACKEND_URL and PROXIMA_INSTALL_TOKEN are required for the first run. The agent enrolls over HTTPS and receives NATS credentials, client/environment slugs, and the NATS URL from the backend. On subsequent runs, the agent loads credentials from disk and PROXIMA_INSTALL_TOKEN is no longer needed. See NATS Security.
The agent no longer reads a YAML configuration file. It is configured purely via environment variables (plus built-in defaults); the former PROXIMA_AGENT_CONFIG / /etc/proxima-agent/config.yaml has been removed. Operational settings beyond startup are delivered as remote config from the backend over NATS.
The agent exports its spans to Tempo over NATS (no OTLP ingress needed). Tracing is off by default and the preferred control is the per-host tracing remote config (Agent config → Distributed tracing in the UI), which hot-reloads without a restart. The env vars below are an override for local/OTLP debugging. See Distributed Tracing.
Variable
Default
Description
PROXIMA_AGENT_OTEL_ENABLED
false
Override that forces NATS span export on (also flips the tracing block on), overriding the per-host UI config. Set to true or 1.
PROXIMA_AGENT_OTEL_SAMPLE_RATIO
0.1
Head-sampling ratio for agent-root traces, applied via ParentBased(TraceIDRatioBased(ratio)). Float in [0,1]; invalid values are ignored with a warning. Also tunable from the UI.
Static labels as a JSON object. Example: {"team":"platform","env":"prod"}
Labels from this variable are merged with any labels delivered via the agent's remote labels config (Agent config in the UI). On key conflict, the environment variable wins.
Agent-side log collection bootstrap settings. These environment variables set the initial log-collection config; richer settings (watched files, journald units) are delivered via the agent's remote log_collection config from the backend.
Node name used as the agent enrollment name. Set via spec.nodeName in the DaemonSet for stable, idempotent re-enrollment.
PROXIMA_CLUSTER_NAME
(required)
Cluster display name (cluster agent only).
PROXIMA_CLUSTER_SLUG
(derived from name)
URL-safe cluster identifier used as the stable agent name for idempotent enrollment. Auto-derived from PROXIMA_CLUSTER_NAME if not set.
PROXIMA_K8S_SYNC_INTERVAL
5m
Full K8s inventory sync interval. Read by BOTH the cluster agent and the backend, and the two must match. Nothing reconciles them: the agent gets its value from Helm, the backend from its own environment, and the backend never asks the agent what cadence it is on. The cluster-onboarding gates quote twice this number as the budget inside which "no inventory yet" is healthy, so raising it for the agent alone makes a healthy cluster report "No inventory after 15m, past the 10m a first full sync should take". The backend's resolved value is shown on the read-only settings page.
PROXIMA_K8S_HEARTBEAT_INTERVAL
60s
Interval between liveness heartbeats published by the cluster agent.
PROXIMA_K8S_POD_LIMIT
5000
Max pods per sync. When exceeded, non-Running pods are prioritized (cluster agent only).
PROXIMA_K8S_EXCLUDE_NAMESPACES
(empty)
Comma-separated namespaces to exclude from collection (cluster agent only).