Skip to main content

External Change Sources

Not every change happens on a host the agent can see. A deploy pipeline, a merge request, a firewall rule edited in the Hetzner console, a DNS record flipped in Cloudflare — all of them can break production, and none of them touch a file on a monitored server.

Console ingests those two ways: webhook sources (the provider pushes) and pull sources (Console polls and diffs). Both land in the same change_events timeline as agent-detected changes, so L1 triage ranks them side by side.

Webhook sourcesPull sources
DirectionProvider → ConsoleConsole → provider
AuthPer-source token embedded in the receiver URLPer-source encrypted API credential
Providersgitlab, github, argocdhetzner_cloud, cloudflare
LatencyImmediatePoll interval (default 600s)
DetectsThe events the provider chose to sendAny structural drift, including changes made outside CI

Webhook sources​

A webhook source is a per-client receiver with its own token. Creating one returns the token once — it is embedded in the receiver URL you paste into the provider.

POST   /api/v1/webhook-sources          # create (returns token once)
GET /api/v1/webhook-sources # list
GET /api/v1/webhook-sources/:id # get
PUT /api/v1/webhook-sources/:id # update
DELETE /api/v1/webhook-sources/:id # delete
POST /api/v1/webhook-sources/:id/rotate # rotate token (invalidates the old URL)
GET /api/v1/webhook-sources/:id/deliveries # recent deliveries, for debugging

The receiver itself is public and token-authenticated — no JWT, no API key. It resolves the source by the URL token, verifies the provider signature, rejects replays, and then writes the change event to PostgreSQL synchronously, inside the HTTP request. There is no queue and no worker on this path, so a slow or unavailable database surfaces to the provider as a 500 rather than being absorbed — check delivery history first when events go missing.

Delivery history is the first place to look when events stop arriving: GET /api/v1/webhook-sources/:id/deliveries shows what actually reached Console.

Setup details per provider are on Webhooks.

What arrives​

SourceEvents
gitlabPush, merge request, tag, pipeline completed, pipeline failed
githubPull request, release, workflow completed, workflow failed
deployment (ArgoCD)Deployed, deploy failed, health degraded

The catalog also registers a k8s-event source. It is not a webhook you configure — Kubernetes events reach Console through the cluster agent's event sync, and the correlation worker reads them alongside host changes in the pre-fire window.

Every one of the webhook events above is weight class Deploy (1.0) — the highest causal weight in the catalog — because "something deployed" is the single most common answer to what changed?

Deploy-feed staleness is monitored

If a client's newest deploy or git change is older than PROXIMA_CORRELATION_DEPLOY_STALE_THRESHOLD (default 72h), the correlation worker logs a WARN and increments proxima_correlation_deploy_feed_stale_total. That is almost always a webhook that stopped delivering, not a team that stopped deploying. Alert on it.

Pull sources​

A pull source is a per-client, per-provider poller. Console holds an encrypted API token, fetches the provider's current state on an interval, and diffs it against the last-known snapshot. Any structural difference becomes a change event.

GET    /api/v1/pull-sources        # list (admin)
POST /api/v1/pull-sources # create (admin)
GET /api/v1/pull-sources/:id # get (admin)
PUT /api/v1/pull-sources/:id # update (admin)
DELETE /api/v1/pull-sources/:id # delete (admin)
FieldMeaning
source_typehetzner_cloud or cloudflare
credentials_encryptedThe provider API token, encrypted at rest (see Credentials)
configProvider options, e.g. {"resource_types": ["server", "firewall"]}
poll_interval_secondsDefault 600
enabledPause polling without deleting the credential
last_polled_at / last_status / last_errorPoll health, surfaced in the admin UI

The last-known state lives in external_resources, one row per (source, resource type, resource key), holding normalized structural attributes only — volatile fields are dropped at normalization time so a metrics tick or a changed timestamp does not manufacture a phantom change every poll.

A row keeps its id across polls for as long as the provider keeps the object, so anything keyed on it (the metrics below) keeps its history.

Metrics (Hetzner Cloud)​

A hetzner_cloud source's token is also used to read metrics for the source's load balancers and for its servers that run no Console agent, every PROXIMA_HETZNER_METRICS_INTERVAL (default 5 minutes). They are written to the client's VictoriaMetrics account and shown in the topology map's details panel and on a load balancer's entry line (Topology: the Traffic tab). That is about 12 API calls per object per hour. Hetzner's limit, 3,600 requests per hour, is per Hetzner project, not per token: it is shared with every other client of that project (the cloud controller manager, the CSI driver, Terraform, and any other source on the same project), so 100 objects take a third of it. The metrics poller stops a source's cycle once less than a quarter of the project's budget is left, and resumes at the next interval. A disabled source is not read. A metrics failure (a rate limit, a revoked token) delays that source's metrics until the next successful cycle, which back-fills the gap up to an hour back, and never touches the structural poll above, its last_status or its change events.

Why polling, not webhooks, for cloud​

Cloud providers either do not offer control-plane webhooks or offer them only on higher tiers, and — more importantly — a poller catches changes made outside any pipeline: the manual firewall edit at 2am, the DNS record someone flipped by hand. Those are exactly the changes nobody writes a deploy note for, and exactly the ones that cause incidents.

Detected changes​

ProviderResources tracked
hetzner_cloudServers, firewalls, networks, floating IPs, load balancers, SSH keys, volumes, certificates
cloudflareDNS records, zone settings, custom rules, managed WAF, page rules, redirect rules

Cloud changes are stored with a NULL host_id — they are client-scoped, not host-scoped, because a firewall rule belongs to the account and not to any one server. Host-scoped correlation therefore never reaches them; L1 pulls them in through a separate client-scoped lookup driven by PROXIMA_TRIAGE_CLOUD_SOURCES, whose default is derived from the change catalog.

The catalog is the guard

PROXIMA_TRIAGE_CLOUD_SOURCES defaults to every source flagged as cloud in the catalog, and a test asserts that set covers every supported pull provider. That is deliberate: when Cloudflare was added as a provider without being added to the cloud-source list, its changes were silently invisible to L1 triage. The guard exists so that cannot recur.

Adding a provider​

Both paths run through the same registry, backend/internal/domain/changecatalog:

  1. Register the source (ID, label, kind, and the IsCloud / IsDeploy flags).
  2. Register its event types with a group and a weight class.
  3. For a pull provider, add the type to the pull_sources source_type check constraint and implement normalization — structural fields only.

Getting step 1 right is what keeps the provider visible to L1 correlation, the UI filters, and GET /api/v1/changes/catalog without touching any of them individually.

See also​