Skip to main content

Change Detection

Change detection answers one question during an incident: what changed just before this broke?

Every change Proxima Console knows about — a config file edited on a host, a package upgraded, a GitLab pipeline that deployed, a Hetzner firewall rule someone touched in the cloud console — lands in one table, change_events, with a common shape. That single timeline is what the L1 incident agent ranks against an alert's fire time when it looks for a root cause.

The pipeline​

Four producers, one timeline:

  1. The agent's file scanner — file integrity monitoring, covered below.
  2. The backend's inventory differ — when an agent inventory snapshot arrives, the backend diffs it against the last-known state and bulk-inserts the resulting events inside the snapshot transaction. This is what produces the whole Inventory group: packages, services, ports, users, firewall rules, kernel modules, cron, certificates, container images, interfaces. It is gated on change detection being enabled.
  3. Webhook sources and 4. pull sources — both on External Change Sources.

File integrity monitoring​

The agent's filewatch scanner walks the configured paths on an interval, hashes every file, and compares each hash against a local hash database. A file whose content hash moved always produces a change event.

A content-identical file whose metadata moved is deliberately noisier to report than to suppress, so it emits only when the change is privilege- or exposure-relevant: the mode gains setuid, setgid, or any world rwx bit; the owner or group crosses the root boundary; or the file matches a sensitive pattern and any of mode/owner/group changed. A cosmetic re-chmod, a permission tightening (0644 → 0600), and a uid/gid that re-resolves between numeric and name form are all suppressed — that filter is what stops a file_metadata_changed flood.

SettingDefaultMeaning
enabledfalseChange detection ships off; enable it per host or environment.
paths["/etc"]Directories walked on each scan.
interval60sHow often the scan runs.
exclude_patternssee belowExtra globs to skip, on top of the built-in exclusions.
sensitive_filessee belowExtra globs whose content is never transmitted, on top of the built-in list.
max_file_size1 MiBFiles larger than this are skipped entirely — not read, not hashed, not tracked. See the warning below.
max_message_size800 KiBUpper bound on a single change message published to NATS.
redact_secretstrueValue-level secret scrubbing on transmitted content.

Configuration arrives from two different places, and the split matters:

  • change_detection turns the scanner on and sets interval (backend-validated, minimum 30s).
  • watch_files pushes paths, exclude_patterns, sensitive_files, and max_file_size.
  • max_message_size and redact_secrets are agent-local only. They cannot be pushed from the Console — a PUT that includes redact_secrets: false under watch_files is a silent no-op.

See Config Sync and Agent Configuration.

max_file_size is a monitoring blind spot, not a content limit

A file over the limit returns before it is hashed, so it never enters the scanned path set at all — it is not integrity-monitored in any form. Worse, a file that was being tracked and then grows past the limit disappears from the path set, and the deletion sweep emits a file_deleted event for a file that is still on disk. Raise the limit for large configs you actually need watched.

What the agent stores​

TableHolds
watched_filesOne row per unique file per host: current hash, size, mode, owner, group, symlink target, plus the is_sensitive, is_binary, and deleted flags.
file_versionsContent history for a tracked file — the full content, the unified diff from the previous version, and the metadata at that version. content is NULL for sensitive and binary files.
change_eventsThe source-agnostic timeline row (source, event_type, summary, details, severity).

Secrets never leave the host in cleartext​

Two independent defenses, and they are not interchangeable:

  1. Path-based skip. A file matching a sensitive pattern is tracked by hash and metadata only — its content is never transmitted, so it never reaches the scrubber. Built-in patterns cover shadow, gshadow, *.key, *.pem, *.p12, *.pfx, ssl/private/*, sudoers.d/*, .env / .env.* / *.env, SSH private keys (id_rsa, id_ed25519, …), .netrc, .pgpass, htpasswd, and secrets/* / secret/*.

  2. Value-level redaction. For content that is transmitted, the scrubber replaces secret values in assignments, matching the key token case-insensitively. This is on by default; redact_secrets: false is an escape hatch that lets any secret not caught by a path pattern travel in cleartext.

    It recognizes key = value and key: value, quoted JSON keys ("password": "…"), PEM private-key blocks, connection strings with inline credentials (scheme://user:PASSWORD@host), JWTs, and a Shannon-entropy backstop for long random tokens carrying no key at all.

    What it does not catch, stated plainly because the path-based skip above is the defense you should rely on for a file that holds secrets:

    • A secret with no key and no recognizable shape — a short bare token below the entropy threshold.
    • An assignment written mid-sentence. The match is anchored at the start of the value, so api_key = … as a whole line is redacted while reminder: rotate api_key = … next week is not.
    • A value separated from its key by anything other than = or : — notably a flag-attached secret such as mysqldump -pSECRET, which is why file content, not command lines, is what this scrubber is wired to.

    This is a precision-biased design: it exists alongside diffs an engineer reads, and a scrubber that redacted aggressively would destroy the change history the feature is for. The two defenses are not interchangeable, and this is the reason.

caution

Adding a path to sensitive_files stops content collection for that file going forward. It does not retroactively purge file_versions rows captured before the pattern was added — prune those through retention, or delete the versions explicitly.

Built-in exclusions (never scanned at all) cover the churn that would otherwise flood the timeline: ld.so.cache, *.db, *.swp, *.swo, *~, *.bak, the dpkg/rpm backup suffixes, .pwd.lock, adjtime, and mtab.

The event catalog​

Every event type is registered in one place (backend/internal/domain/changecatalog) with a group, an owning source, and a weight class that determines how strongly L1 treats it as a candidate cause.

Weight classScoreTypical events
Deploy1.0Push, merge request, pull request, tag, release, pipeline/workflow completed or failed, ArgoCD sync
Config0.9Configuration changes; Kubernetes NetworkPolicy created / changed / deleted; Kubernetes workload changed (image, resources, env names, env/volume refs, strategy)
Cloud0.7Cloud server / firewall / network / DNS / WAF / load-balancer changes
Package0.6Package installed, removed, upgraded
Neutral0.5Services, ports, users, firewall rules, kernel modules, cron, interfaces; Kubernetes workload scaled by hand (fetched separately for triage, under its own limit; scaling by an HPA is not recorded)
Routine0.2File created / modified / deleted / metadata changed, certificate add / remove / rotate

Groups the UI filters by: Files, Git, CI & Deploy, Inventory, Cloud, Kubernetes.

The Kubernetes group is the k8s-config source: changes the backend finds by diffing each cluster-agent inventory sync against the stored one. Kubernetes emits no event when someone edits an object's spec, so these cannot come from the event stream. See Kubernetes inventory → Workload change records.

Because the catalog is the single registry, a new pull provider cannot be silently excluded from L1 causality — the cloud-source list used by triage is derived from it.

GET /api/v1/changes/catalog returns the live catalog, which is what the UI filter menus render from.

From change to cited cause​

When L1 triages an alert it ranks recent changes on the alerting host (and its environment) against the fire time. A change that survives ranking and is cited in the RCA is recorded in change_citations — the alert↔change edge, with the verdict the verification pass reached (unverified, supported, refuted, or unconfirmed).

That edge is queryable in reverse: GET /api/v1/changes/:id/citations answers "which incidents did this change cause?".

The ranking window and caps are tunable — see PROXIMA_TRIAGE_CHANGE_CAUSE_WINDOW, PROXIMA_TRIAGE_CHANGE_CAUSE_SKEW, PROXIMA_TRIAGE_CHANGE_CAUSE_CAP, and PROXIMA_CORRELATION_CHANGE_WINDOW in Environment Variables.

Retention​

Change data is the fastest-growing table set in the product, so retention is explicit and opt-in.

VariableDefaultPrunes
PROXIMA_RETENTION_ENABLEDfalseMaster switch. While false the worker no-ops every tick and deletes nothing, for any client.
PROXIMA_RETENTION_TIMELINE_DAYS365change_events older than this. Floor 7.
PROXIMA_RETENTION_CONTENT_DAYS90file_versions older than this and beyond the last-K. Floor 1.
PROXIMA_RETENTION_CONTENT_KEEP_LAST20Newest versions per file always kept regardless of age. Floor 1.
PROXIMA_RETENTION_INTERVAL6hSweep cadence.
PROXIMA_RETENTION_BATCH_SIZE5000Rows per batched DELETE, so prune statements stay short.

Per-client overrides live in client_retention_settings and are super-admin only. Each column is nullable: a NULL means inherit the global default, so partial overrides work and a global bump instantly applies to every inheriting client. A global enabled=false always wins — a client cannot opt itself into pruning while the master switch is off.

GET  /api/v1/clients/:id/retention-settings
PUT /api/v1/clients/:id/retention-settings

API​

MethodPathDescription
GET/api/v1/changesGlobal change timeline (filter by host_id, client_id, environment_id)
GET/api/v1/changes/:idChange detail
GET/api/v1/changes/catalogThe event catalog (sources, event types, groups)
GET/api/v1/changes/:id/citationsIncidents that cited this change
GET/api/v1/hosts/:hostID/changesPer-host timeline
GET/api/v1/hosts/:hostID/filesTracked files on a host
GET/api/v1/hosts/:hostID/files/:fileID/versionsVersion history for a tracked file

The full set, with permissions, is in the API Endpoints reference.

See also​