Change Detection
Change detection answers one question during an incident: what changed just before this broke?
Every change Proxima Console knows about — a config file edited on a host, a package upgraded, a
GitLab pipeline that deployed, a Hetzner firewall rule someone touched in the cloud console —
lands in one table, change_events, with a common shape. That single timeline is what the
L1 incident agent ranks against an alert's fire time when it
looks for a root cause.
The pipeline
Four producers, one timeline:
- The agent's file scanner — file integrity monitoring, covered below.
- The backend's inventory differ — when an agent inventory snapshot arrives, the backend diffs it against the last-known state and bulk-inserts the resulting events inside the snapshot transaction. This is what produces the whole Inventory group: packages, services, ports, users, firewall rules, kernel modules, cron, certificates, container images, interfaces. It is gated on change detection being enabled.
- Webhook sources and 4. pull sources — both on External Change Sources.
File integrity monitoring
The agent's filewatch scanner walks the configured paths on an interval, hashes every file, and
compares each hash against a local hash database. A file whose content hash moved always
produces a change event.
A content-identical file whose metadata moved is deliberately noisier to report than to
suppress, so it emits only when the change is privilege- or exposure-relevant: the mode gains
setuid, setgid, or any world rwx bit; the owner or group crosses the root boundary; or the file
matches a sensitive pattern and any of mode/owner/group changed. A cosmetic re-chmod, a permission
tightening (0644 → 0600), and a uid/gid that re-resolves between numeric and name form are all
suppressed — that filter is what stops a file_metadata_changed flood.
| Setting | Default | Meaning |
|---|---|---|
enabled | false | Change detection ships off; enable it per host or environment. |
paths | ["/etc"] | Directories walked on each scan. |
interval | 60s | How often the scan runs. |
exclude_patterns | see below | Extra globs to skip, on top of the built-in exclusions. |
sensitive_files | see below | Extra globs whose content is never transmitted, on top of the built-in list. |
max_file_size | 1 MiB | Files larger than this are skipped entirely — not read, not hashed, not tracked. See the warning below. |
max_message_size | 800 KiB | Upper bound on a single change message published to NATS. |
redact_secrets | true | Value-level secret scrubbing on transmitted content. |
Configuration arrives from two different places, and the split matters:
change_detectionturns the scanner on and setsinterval(backend-validated, minimum 30s).watch_filespushespaths,exclude_patterns,sensitive_files, andmax_file_size.max_message_sizeandredact_secretsare agent-local only. They cannot be pushed from the Console — aPUTthat includesredact_secrets: falseunderwatch_filesis a silent no-op.
See Config Sync and Agent Configuration.
max_file_size is a monitoring blind spot, not a content limitA file over the limit returns before it is hashed, so it never enters the scanned path set at all —
it is not integrity-monitored in any form. Worse, a file that was being tracked and then grows
past the limit disappears from the path set, and the deletion sweep emits a file_deleted event for
a file that is still on disk. Raise the limit for large configs you actually need watched.
What the agent stores
| Table | Holds |
|---|---|
watched_files | One row per unique file per host: current hash, size, mode, owner, group, symlink target, plus the is_sensitive, is_binary, and deleted flags. |
file_versions | Content history for a tracked file — the full content, the unified diff from the previous version, and the metadata at that version. content is NULL for sensitive and binary files. |
change_events | The source-agnostic timeline row (source, event_type, summary, details, severity). |
Secrets never leave the host in cleartext
Two independent defenses, and they are not interchangeable:
-
Path-based skip. A file matching a sensitive pattern is tracked by hash and metadata only — its content is never transmitted, so it never reaches the scrubber. Built-in patterns cover
shadow,gshadow,*.key,*.pem,*.p12,*.pfx,ssl/private/*,sudoers.d/*,.env/.env.*/*.env, SSH private keys (id_rsa,id_ed25519, …),.netrc,.pgpass,htpasswd, andsecrets/*/secret/*. -
Value-level redaction. For content that is transmitted, the scrubber replaces secret values in assignments, matching the key token case-insensitively. This is on by default;
redact_secrets: falseis an escape hatch that lets any secret not caught by a path pattern travel in cleartext.It recognizes
key = valueandkey: value, quoted JSON keys ("password": "…"), PEM private-key blocks, connection strings with inline credentials (scheme://user:PASSWORD@host), JWTs, and a Shannon-entropy backstop for long random tokens carrying no key at all.What it does not catch, stated plainly because the path-based skip above is the defense you should rely on for a file that holds secrets:
- A secret with no key and no recognizable shape — a short bare token below the entropy threshold.
- An assignment written mid-sentence. The match is anchored at the start of the value, so
api_key = …as a whole line is redacted whilereminder: rotate api_key = … next weekis not. - A value separated from its key by anything other than
=or:— notably a flag-attached secret such asmysqldump -pSECRET, which is why file content, not command lines, is what this scrubber is wired to.
This is a precision-biased design: it exists alongside diffs an engineer reads, and a scrubber that redacted aggressively would destroy the change history the feature is for. The two defenses are not interchangeable, and this is the reason.
Adding a path to sensitive_files stops content collection for that file going forward. It does
not retroactively purge file_versions rows captured before the pattern was added — prune those
through retention, or delete the versions explicitly.
Built-in exclusions (never scanned at all) cover the churn that would otherwise flood the
timeline: ld.so.cache, *.db, *.swp, *.swo, *~, *.bak, the dpkg/rpm backup
suffixes, .pwd.lock, adjtime, and mtab.
The event catalog
Every event type is registered in one place (backend/internal/domain/changecatalog) with a
group, an owning source, and a weight class that determines how strongly L1 treats it as a
candidate cause.
| Weight class | Score | Typical events |
|---|---|---|
| Deploy | 1.0 | Push, merge request, pull request, tag, release, pipeline/workflow completed or failed, ArgoCD sync |
| Config | 0.9 | Configuration changes; Kubernetes NetworkPolicy created / changed / deleted; Kubernetes workload changed (image, resources, env names, env/volume refs, strategy) |
| Cloud | 0.7 | Cloud server / firewall / network / DNS / WAF / load-balancer changes |
| Package | 0.6 | Package installed, removed, upgraded |
| Neutral | 0.5 | Services, ports, users, firewall rules, kernel modules, cron, interfaces; Kubernetes workload scaled by hand (fetched separately for triage, under its own limit; scaling by an HPA is not recorded) |
| Routine | 0.2 | File created / modified / deleted / metadata changed, certificate add / remove / rotate |
Groups the UI filters by: Files, Git, CI & Deploy, Inventory, Cloud, Kubernetes.
The Kubernetes group is the k8s-config source: changes the backend finds by diffing each
cluster-agent inventory sync against the stored one. Kubernetes emits no event when someone edits
an object's spec, so these cannot come from the event stream. See
Kubernetes inventory → Workload change records.
Because the catalog is the single registry, a new pull provider cannot be silently excluded from L1 causality — the cloud-source list used by triage is derived from it.
GET /api/v1/changes/catalog returns the live catalog, which is what the UI filter menus render
from.
From change to cited cause
When L1 triages an alert it ranks recent changes on the alerting host (and its environment)
against the fire time. A change that survives ranking and is cited in the RCA is recorded in
change_citations — the alert↔change edge, with the verdict the verification pass reached
(unverified, supported, refuted, or unconfirmed).
That edge is queryable in reverse: GET /api/v1/changes/:id/citations answers "which incidents
did this change cause?".
The ranking window and caps are tunable — see PROXIMA_TRIAGE_CHANGE_CAUSE_WINDOW,
PROXIMA_TRIAGE_CHANGE_CAUSE_SKEW, PROXIMA_TRIAGE_CHANGE_CAUSE_CAP, and
PROXIMA_CORRELATION_CHANGE_WINDOW in Environment Variables.
Retention
Change data is the fastest-growing table set in the product, so retention is explicit and opt-in.
| Variable | Default | Prunes |
|---|---|---|
PROXIMA_RETENTION_ENABLED | false | Master switch. While false the worker no-ops every tick and deletes nothing, for any client. |
PROXIMA_RETENTION_TIMELINE_DAYS | 365 | change_events older than this. Floor 7. |
PROXIMA_RETENTION_CONTENT_DAYS | 90 | file_versions older than this and beyond the last-K. Floor 1. |
PROXIMA_RETENTION_CONTENT_KEEP_LAST | 20 | Newest versions per file always kept regardless of age. Floor 1. |
PROXIMA_RETENTION_INTERVAL | 6h | Sweep cadence. |
PROXIMA_RETENTION_BATCH_SIZE | 5000 | Rows per batched DELETE, so prune statements stay short. |
Per-client overrides live in client_retention_settings and are super-admin only. Each column is
nullable: a NULL means inherit the global default, so partial overrides work and a global bump
instantly applies to every inheriting client. A global enabled=false always wins — a client
cannot opt itself into pruning while the master switch is off.
GET /api/v1/clients/:id/retention-settings
PUT /api/v1/clients/:id/retention-settings
API
| Method | Path | Description |
|---|---|---|
| GET | /api/v1/changes | Global change timeline (filter by host_id, client_id, environment_id) |
| GET | /api/v1/changes/:id | Change detail |
| GET | /api/v1/changes/catalog | The event catalog (sources, event types, groups) |
| GET | /api/v1/changes/:id/citations | Incidents that cited this change |
| GET | /api/v1/hosts/:hostID/changes | Per-host timeline |
| GET | /api/v1/hosts/:hostID/files | Tracked files on a host |
| GET | /api/v1/hosts/:hostID/files/:fileID/versions | Version history for a tracked file |
The full set, with permissions, is in the API Endpoints reference.
See also
- External Change Sources — webhooks and cloud pull sources
- L1 Incident Agent — how a change becomes a cited root cause
- Agent Configuration — the
watch_filesconfig type - Webhooks — delivering external change events