Skip to main content

Alerting & Correlation

Proxima Console ingests alerts from Alertmanager, enriches them with infrastructure context (metrics, changes, related alerts, compliance), and presents a correlated view for faster incident response.

Overview​

The alerting system sits alongside your existing on-call tooling. Rather than replacing Alertmanager or Grafana OnCall, it acts as a parallel receiver that adds the infrastructure context operators need when they get paged at 3am:

  • Metrics snapshot around the alert firing time (baseline vs. current values)
  • Recent changes from both agent file monitoring and external webhooks (GitLab, GitHub, ArgoCD)
  • Related alerts on the same host or environment
  • Compliance state of the affected host
  • Host context (hostname, OS, architecture)
VMAlert --> Alertmanager --+--> Grafana OnCall  (keeps handling on-call)
|
+--> Proxima Console (adds infrastructure context)

Architecture​

The pipeline is fully asynchronous. The webhook endpoint validates the token, parses the payload, publishes to NATS, and returns 202 Accepted immediately. Processing happens in two worker stages:

  1. AlertWorker (fast path): resolves labels to Proxima entities, deduplicates alerts by fingerprint, finds or creates alert groups, handles resolve/reopen lifecycle
  2. CorrelationWorker (enrichment): queries VictoriaMetrics for metrics, PostgreSQL for changes/related alerts/compliance, and writes the correlation context back to the alert group

Alert Source Setup​

An alert source is a per-client integration endpoint that receives webhooks from Alertmanager.

Creating an Alert Source​

curl -X POST https://api-console.prxm.uz/api/v1/clients/{client_id}/alert-sources \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-d '{
"name": "Production Alertmanager",
"source_type": "alertmanager",
"enabled": true
}'

Response (201):

{
"data": {
"id": "a1b2c3d4-...",
"client_id": "...",
"source_type": "alertmanager",
"name": "Production Alertmanager",
"enabled": true,
"token": "pxm_as_7f3a8b2c1d...",
"created_at": "2026-03-15T10:00:00Z",
"updated_at": "2026-03-15T10:00:00Z"
}
}
caution

The token is a credential — store it securely and never put it in a URL. It is returned in full by this create response only, but it is not lost if you close the dialog: the Show webhook token action on the Alert Sources page (/oncall/alert-sources) reveals it again at any time via GET /api/v1/alert-sources/{sourceID}/token (alertsources:read, audit-logged). Only a source with no stored ciphertext — created before Console kept one, or on a deployment with no Vault Transit encryptor — has to be recreated.

Alertmanager Configuration​

Add a webhook_configs entry to your Alertmanager configuration pointing at the Proxima webhook receiver:

# alertmanager.yml
receivers:
- name: "proxima"
webhook_configs:
- url: "https://api-console.prxm.uz/api/v1/alerts/webhook/alertmanager/pxm_as_7f3a8b2c1d..."
send_resolved: true

route:
receiver: "default"
routes:
- receiver: "proxima"
continue: true # Important: continue to other receivers (OnCall, etc.)
matchers: [] # Match all alerts (or restrict with matchers)

Key points:

  • Set send_resolved: true so Proxima can track alert resolution
  • Use continue: true if you want alerts to also reach other receivers (e.g., Grafana OnCall)
  • The token in the URL authenticates the request (no additional headers needed)
  • Keep the route's repeat_interval (Alertmanager's default is 4h) shorter than PROXIMA_INCIDENT_STALE_TTL (default 6h) — see below
repeat_interval must be shorter than the stale TTL

While an alert keeps firing, Alertmanager re-sends it every repeat_interval. Those repeats are how Console knows the outage is still live: each one reaches the alert worker and refreshes the group. A firing group that nothing has refreshed for PROXIMA_INCIDENT_STALE_TTL is resolved by the stale-group reaper — escalation stops while the problem is still there, and the next repeat opens it again and pages again.

So if you raise repeat_interval (for example to 12h to quiet another receiver), raise PROXIMA_INCIDENT_STALE_TTL above it as well, or give the Proxima receiver its own route with a shorter interval. Violations show up per source type on proxima_alert_group_reaped_by_source_total{source_type}: a source type whose groups are routinely reaped is one whose repeats are not arriving while it is still firing.

Repeats only count if they get through the webhook replay guard. On the alertmanager (and generic) door its window is 60 seconds, so only an HA pair's duplicate send or an immediate retry is skipped; every repeat_interval resend reaches the worker. Before this was fixed the Alertmanager door used a 24-hour window, swallowed every byte-identical repeat, and auto-resolved outages that were still firing after 6 hours.

Each repeat that gets through is processed like any other delivery, so it writes an Updated and a Correlated entry on the group's timeline. With a short repeat_interval (1–5 minutes) a long outage's timeline fills with these; that is expected.

The webhook path is /alerts/webhook/, not /webhooks/

The receiver lives at /api/v1/alerts/webhook/alertmanager/{token}. Older revisions of this page published /api/v1/webhooks/alertmanager/{token}, which 404s — Alertmanager reports a delivery failure your alerts never recover from. If you copied a URL from an older doc, fix it.

Grafana 8.x legacy (panel) alerting posts to /api/v1/alerts/webhook/grafana/{token} on a source whose source_type is grafana_legacy.

proxima_probe — Console's own first-party source​

Uptime monitors publish onto the same ingest path a webhook does, so each client that has one gets an alert_sources row with source_type = 'proxima_probe'. It behaves differently from the two webhook types in five ways worth knowing before you look at one:

  • It is created on demand, not authored. The probe result worker creates the row on its first emission for that client. There is no webhook parser for it, so it is not creatable, not editable and not deletable through the API — an update or delete answers 409 conflict naming what the row is, because it sits in the operator's list beside their own integrations and the UI offers the same controls on it. Deleting it would stop every monitor in that tenant from paging until the next emission recreated it; re-pointing its source_type would break emission for that client permanently, because the reserved per-client token would be held by a row the create path can no longer match.
  • It has no webhook and no token. It stores a reserved sentinel in token_hash instead — shaped so nothing an attacker can send could hash to it — purely as the per-client uniqueness key.
  • It will always read "Never delivered." Per-source ingest health is written by RecordIngest, whose only caller is the webhook handler, and probe emission never touches it. See the caveat under Per-source ingest health.
  • The row can appear before anything has ever paged. Probe resolves are deliberately ungated, so a client's first monitor recovery creates the row even with paging_enabled off everywhere.
  • A monitor group pages through the same row. A group is a named service that pages once for a whole set of monitors, and it is not a new source type — its alert arrives on this client's proxima_probe source under a distinct alertname, ProximaMonitorGroupDown, and a grouping key of proxima_probe_group:<group_id> rather than proxima_probe:<monitor_id>. The two are deliberately distinct so filtering or silencing one does not catch the other. A group page carries the group's environment_id and, by construction, no host_id — so a host-scoped route can never match one.

Full behaviour in Uptime Monitors, and the model in Monitor groups.

Changing a source's type​

PUT /api/v1/alert-sources/{sourceID} accepts source_type, so a source created with the wrong type can be repaired without production SQL. A proxima_probe source is exempt — it refuses every write. For the two webhook types, two things to know:

  • The existing webhook URL keeps working. The handler resolves the source by token hash, so the /alertmanager/ vs /grafana/ path segment is cosmetic — both segments reach the same source, and the type stored on the source is what selects the parser. You do not have to reconfigure the sender.
  • It changes how every future payload is interpreted, so the flip is recorded in the application log with the actor, and the old and new type. Nothing already stored is rewritten, and another PUT reverses it.

Ingestion Contract​

Everything above gets a webhook delivered. This section is about what Proxima does with what is inside it — and, just as importantly, what it refuses to guess.

Why this section exists

A severity-mapping break ran in production for 16 days undetected. Deliveries never stopped, so every liveness signal stayed green while the alerts were silently wrong: unlabelled alerts were laundered into plausible, pageable warnings, and every delivery that arrived without a grouping key collapsed into a single alert group. Nothing counted what a parse actually extracted, so nothing could notice. The contract below, and the per-source health and ratio alerts at the end of it, are what make that class of break visible.

Severity mapping​

Proxima stores its own priority tiers, P1–P5. An incoming severity label is matched case-insensitively (leading/trailing whitespace is trimmed) against this table:

Proxima tierMeaningAccepted labels
P1Page immediatelyfatal, emergency, emerg, critical, crit, sev0, sev1, p1, page, panic, alert
P2Higherror, err, high, major, sev2, p2
P3Warningwarning, warn, minor, sev3, p3
P4Informationalinfo, informational, information, notice, low, sev4, p4
P5Least urgentdebug, trace, none, sev5, p5
unknownWe could not tellanything absent from the table, and a missing label

unknown is a real tier — we never fabricate a severity​

An alert whose severity label is missing or unrecognized is stored as unknown. It is never rounded up, never rounded down, and never given a plausible-looking default.

That is the whole point. A fabricated P3 looks exactly like a real warning, so it survives every review and every dashboard; a fabricated tier on a genuine P1 silently downgrades a page. unknown is honest, visible in the UI, countable, and ranks least urgent internally, so a value we could not identify can never masquerade as urgent.

unknown gets no routing special-case. An unknown-severity alert matches:

  • an escalation route whose severity is explicitly set to unknown, or
  • a catch-all (default) route,

and nothing else. There is no "treat unknown as P1 to be safe" — urgency inflation on a value we cannot identify would page people for debug messages, which is how a pager gets ignored.

unknown is selectable in the route editor at On-Call → Escalation, alongside P1–P5, and in the severity filter on the Alert Groups page, so you can list exactly the alerts whose severity Proxima could not identify.

This changed paging for one route configuration — check yours

Before this release, an alert whose severity label was missing or unrecognized was given a fabricated P3. P3 is a real, pageable tier, so that alert matched an explicit P3 escalation route and paged. It is now unknown, which matches only an explicit unknown route or a catch-all.

If your escalation routes are all explicit severities with no catch-all, those alerts are now captured, not paged: they are stored and visible on the Alert Groups page, but nobody is notified. Nothing is lost — but nobody is woken either, and there is no flag to turn this off.

The fix is one route. Add a catch-all (default) route pointing at whichever escalation policy should own "we don't know what this is", or add an explicit unknown route to handle it separately. If you already have a catch-all, you still get paged — but through the catch-all's policy rather than the explicit P3 route's, so check that the catch-all points at a team you want woken for an alert nobody could classify.

Sweep for the clients that need one with the catch-all audit below. Afterwards, the EscalationUnroutedUnknownSeverity alert rule fires on proxima_escalation_unrouted_total{severity="unknown"} whenever it happens — but the counter carries no client id, so the client comes from the accompanying escalation: alert matched no route warning in the logs.

Two operational traps​

1. Route severity matching is case-sensitive. The matcher compares the stored route severity to the alert's severity with an exact string comparison. Ingestion has already mapped the wire label, so a route must store a canonical tier — P1, not critical, and not p1. A route stored as critical or p1 matches nothing, forever, and nothing says so: no error, no metric, no UI signal.

The API rejects a non-canonical severity on every route write with a 400 naming the whole allowed set. That validation is write-time only — see the pre-cut-over audit for rows written before it existed.

2. Alertmanager fabricates severities too. Alertmanager's own PagerDuty receiver defaults severity to error when the alert carries no severity label. If Proxima and PagerDuty disagree about an alert's urgency, this is usually why: PagerDuty is showing you Alertmanager's invented error, Proxima is showing you the honest unknown. Fix it at the source by labelling the rule.

What returns a 400​

Eleven machine-readable reason codes. Seven are unconditional; four are gated by the strict ingestion contract (below). The last six apply only to the generic door.

ReasonAlways a 400?What it meansFix
unreadable_bodyyesThe body could not be read, or exceeded the size limitCheck the sender and any proxy in front of it
empty_bodyyesZero-length bodyThe sender posted nothing
malformed_payloadyesThe parser rejected the payloadWrong source_type, or the sender's format changed
no_contentonly in strict modeParsed, but no alert carried an alertname, labels or annotationsThe sender's payload shape drifted
no_identity_legacyonly in strict modeA grafana_legacy payload with neither ruleId nor ruleNameGrafana is sending an unidentifiable rule
missing_dedup_keyyesA generic payload with no dedup_key, on a source with no dedup_key_from derivationSend a dedup_key, or configure the derivation on the source
dedup_key_unresolvedyesThe source derives its key from fields this payload does not carry as non-empty top-level valuesFix the sender, or the source's dedup_key_from list (matched exactly and case-sensitively)
dedup_key_too_longyesA generic dedup key over 512 bytes (bytes, not characters — a CJK key is 3 bytes per character)Shorten the key, or derive it from fewer fields
invalid_statusyesA generic payload whose status is neither firing nor resolvedFix the sender
missing_summaryonly in strict modeA generic payload with no summary textFix the sender — this is the line an on-call engineer reads at 3am
unknown_severityonly in strict modeA generic payload whose severity is absent or maps to no Proxima tierLabel the alert, or ask for the alias to be added

We answer with a named reason rather than a silent 202 so the sender learns its payload shape broke. A 202 teaches a drifting sender that it is successfully paging you.

A label value is never one of these. On the generic door every labels value is rendered to text — numbers, booleans and null through the same renderer the dedup key uses, objects and arrays as their compact JSON — so {"labels":{"port":8080,"code":503}} is ingested, not refused. A whole page is never dropped over the least important field in the payload. Only a labels field that is not a JSON object at all is a parse error.

Authentication failures are a 401, not one of these, and they are counted separately (proxima_alert_webhook_auth_refused_total{reason}). Two of them — the source is disabled, or a legacy source's token was presented at the generic door — are also stamped on that source's health, so the Alert Sources page shows a refusal with its reason instead of "No deliveries yet". Everything else about the 401 is deliberately uninformative: five auth outcomes answer one identical body, so the response cannot be used to probe which token is real.

A group-level resolve is exempt. A grafana_legacy recovery (state=ok or paused) legitimately carries zero alerts — its content is the grouping key. Those are never refused; refusing them would leave every group from that source firing forever.

What gets a synthesized identity​

A missing grouping key is repaired, never refused.

An empty grouping key is not "no key" downstream: sha256("") is a perfectly valid hash, so every keyless delivery from a source hashed to the same value and collapsed into one alert group. That was the 16-day incident.

In strict mode Proxima derives a key from the group's identity — its common labels plus the sorted, de-duplicated set of alertnames — and stores it with a synthesized: prefix, so you can always tell an invented key from a source-provided one. Keying on identity (not on the whole alert batch) means repeats of the same group still land in the existing group, while two different groups can never merge.

PROXIMA_ALERT_INGEST_STRICT​

A single backend environment variable, default off. Only the literal string on enables it. It is per-deployment, not per-source.

Counting is not behavior. The strict contract's verdicts are evaluated and counted in both modes, and acted on only when the flag is on. With the flag off nothing is rejected, nothing is synthesized, and every byte published is identical to the pre-hardening path — only the counters, the WARN logs and the trace attributes differ.

That is what makes the flag safe to flip: you can read exactly what strict mode would refuse, per source, before touching it.

Flag off (default)Flag on
no_content delivery202, ingested, counted as would-refuse400 no_content
no_identity_legacy delivery202, ingested, counted as would-refuse400 no_identity_legacy
generic delivery with no summary202, ingested, counted as would-refuse400 missing_summary
generic delivery with an unmappable severity202, ingested as unknown, counted as would-refuse400 unknown_severity
Missing grouping keypublished with the empty key (collapses), countedkey synthesized, then published
Unconditional rejections400 — unchanged400 — unchanged

Rollout order: watch the ratios for a day with the flag off → confirm the would-reject ratio per source is what you expect → flip the flag → re-check the same ratios.

Per-source ingest health​

The Alert Sources page (/oncall/alert-sources; /admin/alert-sources redirects there) shows, per source: when it last delivered, when it last produced an accepted alert, its accepted / rejected / unknown-severity counts, and the reason and time of its last actual refusal.

StateWhat it means
ProducingRecent deliveries are producing accepted alerts
RejectingStill accepting, but a delivery in the last 24h was actually refused (the reason is shown)
Not producingDeliveries are arriving but none has produced an accepted alert for over two days
No deliveries yetNothing has ever been delivered. After 24h from creation this becomes a warning — usually "the sender was never pointed at this webhook"
Read this before trusting a green row

"Not producing" catches a total ingest collapse — two days of deliveries that produced no accepted alert at all. That is real and worth catching.

It does NOT catch the failure class behind the 16-day incident. Mapping drift means alerts that are accepted but wrong — a severity we could not map, a fabricated identity. Those keep last_accepted_at fresh, so the row stays green throughout. An operator who believes a green row rules that out is worse off than one who knows the limit.

That class shows up as the unknown-severity count on this page, and is detected by the ratio alerts below — not by the health state.

Two further caveats, both deliberate:

  • These counters are best-effort. They are written off the ingest hot path in a detached goroutine, so a health write can be lost without affecting the alert. Treat them as a health signal, never as an accounting ledger. The authoritative numbers are the Prometheus counters.
  • "Quiet" is not a state. A source that has not delivered for a while is normal — Alertmanager only talks when something fires — so silence after a first delivery is reported as timestamps, not as a fault.
A proxima_probe row shows "No deliveries yet" forever, and that is not a fault

These columns are written by RecordIngest, whose only caller is the webhook handler. Uptime monitors publish straight onto the ingest subject and never pass through it, so last_delivery_at, last_accepted_at and every counter on Console's own probe source stay null and zero however many pages it has carried — and after 24 hours the row turns the warning colour that normally means "the sender was never pointed at this webhook".

Judge probe health by proxima_probe_paged_total and the emission counters in Uptime Monitors, never by this row's status.

Ingest metrics and the ratio alerts​

Six per-source counters, all labelled source_id and source_type (full reference in the internal docs/standards/metrics.md):

MetricMeaning
proxima_alert_ingest_received_totalDeliveries past the token/enabled check — the denominator, less replayed
proxima_alert_ingest_accepted_totalDeliveries actually published for processing
proxima_alert_ingest_replayed_totalRepeat deliveries the replay guard suppressed (answered 200, nothing published)
proxima_alert_ingest_rejected_totalDeliveries judged unacceptable — plus reason and refused
proxima_alert_ingest_unknown_severity_totalSeverities that mapped to no known tier — plus kind
proxima_alert_ingest_synthesized_identity_totalDeliveries whose empty grouping key was (or would be) replaced
proxima_alert_webhook_auth_refused_totalDeliveries refused at the authentication stage — plus reason. These never reach received, so without this a sender failing auth 100% of the time moved no counter at all

Alert on the ratio, not the count. An absolute count alerts on volume; a ratio alerts on correctness.

# would-reject ratio — read per source BEFORE flipping PROXIMA_ALERT_INGEST_STRICT on it
sum by (source_id) (rate(proxima_alert_ingest_rejected_total{refused="false"}[1h])
or rate(proxima_alert_ingest_received_total[1h]) * 0)
/ sum by (source_id) (rate(proxima_alert_ingest_received_total[1h])
- (rate(proxima_alert_ingest_replayed_total[1h])
or rate(proxima_alert_ingest_received_total[1h]) * 0))

# internal-failure gap — publish failure, marshal failure, or a JetStream-less backend
sum by (source_id) (rate(proxima_alert_ingest_received_total[5m]))
- sum by (source_id) (rate(proxima_alert_ingest_replayed_total[5m])
or rate(proxima_alert_ingest_received_total[5m]) * 0)
- sum by (source_id) (rate(proxima_alert_ingest_accepted_total[5m])
or rate(proxima_alert_ingest_received_total[5m]) * 0)
- sum by (source_id) (rate(proxima_alert_ingest_rejected_total[5m])
or rate(proxima_alert_ingest_received_total[5m]) * 0)
Do not simplify these to a bare subtraction

The or … * 0 clauses and the per-term sum by (source_id) are load-bearing. rejected_total carries extra labels (reason, refused), and Prometheus arithmetic matches on the full label set — so a bare rate(received) - rate(rejected) matches nothing and the whole expression returns empty. Separately, a missing series drops its element rather than counting as zero, so a source that has never had an accepted delivery — exactly the outage the gap exists to catch — would vanish from the result. An expression that silently returns nothing looks identical to a healthy one.

That second point is why the would-reject ratio's numerator carries the fallback too: a source with no would-rejects has no rejected{refused="false"} series, and without it "safe to flip the flag" and "the expression is broken" would render identically — on the one number the flip is gated on. It must read 0, not nothing.

Three things about these that are easy to get wrong:

  • Both expressions subtract replayed. A replay-skip is routine, not a loss: an Alertmanager HA pair's duplicate send and a sender's immediate retry both land there, and on the deprecated grafana_legacy door (24h window) so does every byte-identical repeat. Leave them in the denominator and the would-reject ratio understates what strict mode would refuse on exactly the sources that repeat most. The window is 60 seconds on the generic and alertmanager doors, not 24 hours. generic has no time-varying field, so a cron-style sender's two distinct firings are byte-identical and a day-long window would swallow the second outage and page nobody; Alertmanager's repeat_interval resends are byte-identical and are the group's proof of life (see the repeat_interval rule). Past a client's immediate retry, the (source_id, grouping_key) partial-unique index already makes a repeated trigger join the open group instead of duplicating it.
  • refused is not "was the flag on". refused="true" means the delivery really was refused with a 400; refused="false" means it was judged refusable but ingested with a 202. An unconditional rejection emits refused="true" with the flag off. Filter reason=no_content|no_identity_legacy to isolate the flag-gated verdicts.
  • Never sum the two unknown_severity kinds. kind="unrecognized" is a label we do not speak (fix: add an alias). kind="absent" is a sender not labelling at all (fix: fix the sender). On a source that never labels severity, absent tracks alert volume — a permanently high constant that reads as chronic breakage and gets tuned out.

unknown_severity and synthesized_identity are emitted only for deliveries that were actually published, so both are a true share of accepted.

Where to watch this. The Alert Ingestion RED row on the Proxima Backend Observability Grafana dashboard renders all of the above, and the provisioned proxima-alert-ingest alert group fires on the unknown-severity signature, the internal-failure gap, a sustained refusal rate, and unknown-severity alerts that matched no escalation route (EscalationUnroutedUnknownSeverity, over proxima_escalation_unrouted_total{severity="unknown"} — the Unrouted Alerts by Severity panel on the same row is its dashboard half).

Pre-cut-over audit: stored route severities​

Route severity is validated at write time only. There is no database constraint and no backfill, so any route stored before that validation existed — critical, p1, P9 — is still there, still dead, and still silent.

Run this before cutting a client over, and treat a non-empty result as a blocker:

SELECT id, client_id, severity FROM escalation_routes
WHERE severity IS NOT NULL AND severity NOT IN ('P1','P2','P3','P4','P5','unknown');

Why no metric will tell you instead. You would expect proxima_escalation_unrouted_total to catch a route that matches nothing. It does not. That counter only increments when no route matched — and when a catch-all route exists, the matcher short-circuits on the catch-all before it ever compares severity, so it always returns a route. The alert therefore arms the wrong escalation policy, the unrouted counter never moves, and the one metric an operator would think to check stays at zero.

Every row this query returns is a route that has never matched anything and never will. Fix them with a PUT through the API — which validates — not with direct SQL.

Pre-cut-over audit: clients with no catch-all route​

The other half of the same sweep, and the one that follows from the paging change described under "unknown is a real tier — we never fabricate a severity" above. A client whose escalation routes are all explicit severities pages nobody for an alert whose severity Proxima could not map.

Run this before cutting a client over:

SELECT client_id, count(*) FILTER (WHERE is_default) AS catch_all_routes
FROM escalation_routes
GROUP BY client_id
HAVING count(*) FILTER (WHERE is_default) = 0;

Every client it returns will have its unknown-severity alerts captured but not paged. Give each one a catch-all route — or an explicit unknown route — before repointing its integrations.

Unlike the severity audit above, this one is watched after the fact: proxima_escalation_unrouted_total{severity="unknown"} moves as soon as it happens and the EscalationUnroutedUnknownSeverity rule fires on it. Run the query anyway — the counter tells you afterwards, at the cost of the alerts in between.

Label Mapping​

Label mappings define how Alertmanager labels resolve to Proxima infrastructure entities. Each client configures their own mappings based on their labeling conventions.

How It Works​

  1. When an alert arrives, the AlertWorker loads the client's label mappings (cached, refreshed every 60s)
  2. For each mapping, it extracts the source_label value from the alert's labels
  3. Applies the configured transform (exact match, strip port, or regex)
  4. Looks up the transformed value in the corresponding Proxima entity table

Creating Label Mappings​

curl -X POST https://api-console.prxm.uz/api/v1/clients/{client_id}/alert-label-mappings \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-d '{
"source_label": "instance",
"target_field": "host_address",
"transform": "strip_port",
"priority": 10
}'

Target Fields​

Target FieldMatches AgainstExample
client_slugclients.slugcluster=acme-prod maps to client acme-prod
environment_slugenvironments.slugnamespace=staging maps to environment staging
host_addresshosts.ip_address or hosts.hostnameinstance=10.0.1.5:9100 maps to host 10.0.1.5
team_nameteams.nameteam=platform maps to team platform

Transforms​

TransformBehaviorExample
exactUse label value as-iscluster=production matches production
strip_portRemove :port suffixinstance=10.0.1.5:9100 becomes 10.0.1.5
regexRegex extractionExtract hostname from FQDN patterns

Common Mapping Configurations​

Kubernetes cluster setup:

[
{"source_label": "cluster", "target_field": "client_slug", "transform": "exact", "priority": 10},
{"source_label": "namespace", "target_field": "environment_slug", "transform": "exact", "priority": 20},
{"source_label": "instance", "target_field": "host_address", "transform": "strip_port", "priority": 30},
{"source_label": "team", "target_field": "team_name", "transform": "exact", "priority": 40}
]

Bare-metal/VM setup:

[
{"source_label": "datacenter", "target_field": "environment_slug", "transform": "exact", "priority": 10},
{"source_label": "instance", "target_field": "host_address", "transform": "strip_port", "priority": 20}
]
Unresolved Alerts

If label resolution fails (no matching host, environment, or client), the alert is stored with host_id = NULL. It still appears in the global alerts list and client-scoped views.

Correlation​

When the AlertWorker finishes processing an alert, it publishes a message to proxima.alerts.correlate. The CorrelationWorker gathers infrastructure context and writes it to the correlation JSONB column on the alert group.

What Gets Correlated​

ContextSourceWindow
Metrics snapshotVictoriaMetrics±30 minutes around first_fired_at
Recent changesPostgreSQL (change_events)PROXIMA_CORRELATION_CHANGE_WINDOW before the alert — 24 h by default. When that window is empty, falls back to the host's last-N changes, each labeled with its true offset.
Related alertsPostgreSQL (alert_groups)Firing alerts on same host/environment (limit 10)
Compliance scorePostgreSQL (control_evaluations)Average across all frameworks at correlation time
Host snapshotPostgreSQL (hosts)Hostname, OS, architecture

How Correlation Works​

When the AlertWorker finishes processing an alert group, it publishes the alert_group_id to proxima.alerts.correlate. The CorrelationWorker picks it up and:

  1. Loads the alert group from PostgreSQL
  2. If host_id is set (label resolution succeeded):
    • Resolves the VM tenant via tenantCache (host → client → vm_account_id)
    • Queries VictoriaMetrics for up to 3 metrics chosen from the alert's own name and labels in a ±30-minute window around first_fired_at with 60-second step (see below)
    • Computes trends: separates data points into baseline (before alert) and at-alert (closest point to fire time). Trend = spike if at-alert > 1.5× baseline, drop if < 0.5× baseline, stable otherwise
    • Queries change events from the correlation window (24 h by default, limit 10) — includes agent-detected file changes, backend inventory diffs, and external webhook events (GitLab, GitHub, ArgoCD)
    • Queries compliance score — averages across all frameworks for the host
    • Captures host snapshot — hostname, OS, architecture
  3. Queries related alerts — other firing alert groups on the same host or environment (limit 10, excludes self), with relative time offsets
  4. Stores the correlation JSON in alert_groups.correlation and logs a "correlated" action to the audit trail

If any individual enrichment step fails (e.g., VM query timeout, change store error), it logs a warning and continues with the other steps. Partial correlation data is better than none.

Correlation JSON Structure​

The correlation column on alert_groups contains this JSON structure:

{
"metrics_snapshot": {
"window": {
"from": "2026-03-15T10:16:00Z",
"to": "2026-03-15T11:16:00Z"
},
"series": {
"cpu_usage_ratio": {"baseline": 0.341, "at_alert": 0.872, "trend": "spike"},
"memory_used_ratio": {"baseline": 0.623, "at_alert": 0.910, "trend": "spike"},
"disk_used_ratio": {"baseline": 0.779, "at_alert": 0.785, "trend": "stable"}
}
},
"recent_changes": [
{
"id": "d4e5f6a7-...",
"source": "gitlab",
"event_type": "deployment",
"summary": "Merged MR !482: reduce memory limits",
"file_path": "/etc/kubernetes/manifests/api-server.yaml",
"severity": "high",
"created_at": "2026-03-15T10:28:00Z"
}
],
"related_alerts": [
{
"id": "a1b2c3d4-...",
"grouping_key": "DiskSpaceFull",
"status": "firing",
"severity": "P1",
"first_fired_at": "2026-03-15T10:34:00Z",
"relative_offset": "-12m"
}
],
"compliance_score": 0.72,
"host_snapshot": {
"id": "e3f4a5b6-...",
"hostname": "prod-k8s-01",
"os": "Ubuntu 24.04",
"arch": "x86_64"
}
}

Field reference:

FieldTypeDescription
metrics_snapshot.windowobjectTime range queried (from/to in RFC 3339)
metrics_snapshot.seriesmapMetric name → trend object. The (up to 3) metrics are selected per alert — see the keyword table below. Values are 0.0–1.0 ratios.
metrics_snapshot.series.*.baselinefloatAverage value of data points before first_fired_at
metrics_snapshot.series.*.at_alertfloatValue of the data point closest to first_fired_at
metrics_snapshot.series.*.trendstring"spike" (above 1.5× baseline), "drop" (below 0.5× baseline), "stable", or "unknown" (the series has samples but none before first_fired_at, so there is no baseline; baseline is then a placeholder 0)
metrics_snapshot.series.*.has_databoolfalse when the metric had no sample in the window at all; the other fields are then placeholders and the UI shows "no data"

Metric selection is alert-aware. A disk, iowait, memory or database alert is investigated with the right series rather than always cpu/mem/disk. The worker scans the distinct alert names, the grouping key, and the common-labels blob for the first matching keyword:

KeywordMetrics queried (most relevant first)
iowaitcpu_iowait_ratio, disk_used_ratio, cpu_usage_ratio
diskdisk_used_ratio, disk_inodes_used_ratio, cpu_iowait_ratio
inodedisk_inodes_used_ratio, disk_used_ratio, memory_used_ratio
swapmemory_swap_used_ratio, memory_used_ratio, cpu_usage_ratio
oom, memory, memmemory_used_ratio, memory_swap_used_ratio, cpu_usage_ratio
cpu, loadcpu_usage_ratio, cpu_iowait_ratio, memory_used_ratio
latencycpu_usage_ratio, cpu_iowait_ratio, disk_used_ratio
postgrespg_cache_hit_ratio, memory_used_ratio, cpu_usage_ratio
databasepg_cache_hit_ratio, memory_used_ratio, disk_used_ratio

The first matching keyword wins, and the result is bounded to ~3 metrics. cpu_usage_ratio, memory_used_ratio, disk_used_ratio is only the no-match fallback, not the standard set. | recent_changes | array | Up to 10 most recent change events in the correlation window (24 h by default) before the alert | | recent_changes[].source | string | "agent", "gitlab", "github", "argocd" | | recent_changes[].file_path | string | File path (if applicable, omitted otherwise) | | related_alerts | array | Up to 10 other firing alert groups on the same host/environment | | related_alerts[].relative_offset | string | Time offset from this alert's first_fired_at (e.g., "-12m" = fired 12 min before, "+2m" = fired 2 min after) | | compliance_score | float | Average compliance score across all frameworks (0.0–1.0). null if no evaluations exist. | | host_snapshot | object | Host info at correlation time. null if host_id was not resolved. |

Metric values are ratios

All metric values in metrics_snapshot.series are 0.0–1.0 ratios (not percentages). The frontend multiplies by 100 for display. This matches the OpenMetrics naming convention used throughout the agent.

Null vs empty
  • metrics_snapshot is null when: the host has no VM data, the tenant cache is unavailable, or all metric queries failed.
  • recent_changes and related_alerts are always arrays (empty [] when none found, never null).
  • compliance_score is null when no compliance evaluations exist for the host.
  • host_snapshot is null when host_id was not resolved by label mapping.

Alert Lifecycle​

Alert groups follow a four-state lifecycle:

StatusDescription
firingOne or more alerts are actively firing
acknowledgedAn operator has acknowledged the alert group
resolvedClosed: every alert in the group resolved, a person resolved it, the stale-group reaper closed it, or the silence sweeper resolved a silence that had ended over 24h earlier
silencedPaging suppressed until a set time, when the silence sweeper ends it

State Transitions​

  • Acknowledge: Sets the status to acknowledged — from firing or silenced — with acknowledged_by and acknowledged_at. Requires alerts:write permission. When the group's escalation is armed, every acknowledgement starts the "still on it?" loop: the acknowledger is asked to confirm, and if nobody confirms within the confirm window the group is auto-unacknowledged — it returns to firing and escalates again (ack_auto_unacked on the timeline). Nothing else returns an acknowledged group to firing: a further firing alert leaves it acknowledged.
  • Resolve: Manually resolves the alert group from any open status, with its firing alerts, and halts its escalation. Does not reach back to Alertmanager. A person's Resolve is an assertion and is written as given — only a source's resolve is corroborated first.
  • A source's resolve is corroborated before it is written. When a delivery carrying no firing alert would close an open group, Console first looks for independent evidence that the entity is still alive — the host's check-in time, and the uptime monitors bound to that host. Evidence of trouble refuses the resolve: nothing is written, the group stays firing with its escalation untouched, and a resolve_not_corroborated row records why. Anything else writes the resolve and records whether a signal confirmed it (resolve_verified) or nothing could speak (resolve_unverified, the common case). The check is bounded at 5s and fails open, so it can never hold a resolve back on a failure of its own. See Resolution Corroboration.
  • Stale-group reaper: a firing group that nothing has updated for the reaper's TTL (PROXIMA_INCIDENT_STALE_TTL, default 6h) is resolved by the system, with its firing alerts. It never touches an acknowledged or silenced group. A source that keeps re-sending a still-firing alert keeps its group fresh — which is why Alertmanager's repeat_interval must be shorter than the TTL. Reaps are counted per source type on proxima_alert_group_reaped_by_source_total.
  • Silence: Sets silenced_until to a future timestamp and suppresses paging until then. The silence ends on its own at that time: within 30 seconds the silence sweeper returns the group to the status it had when it was silenced — acknowledged if it was acknowledged, otherwise firing — and paging resumes where it stopped (an armed escalation picks up on the timer's next tick; an acknowledged group's "still on it?" loop resumes). The status to return to is recorded as prior_status on the silenced timeline entry; silencing an already-silenced group carries the earlier one forward, and a silence recorded before this rule existed returns to firing. The timeline gets a system Silence ended entry saying which status the group went back to.
  • A silence that ended more than 24 hours before it was noticed resolves the group instead. This covers silences that never ended under earlier releases, and a sweeper that was down for a day: resuming them would page everyone at once for incidents that are days or months old. The group is resolved by the system with the same writes a manual resolve makes (its firing alerts resolved, its escalation halted), and the timeline says why. Nothing is lost: if the problem is still firing, its next alert opens a fresh group, which can page again. proxima_alert_silences_ended_total{outcome="expired"|"resolved_stale"} counts both.
  • Unsilence: Ends the silence early by the same rule — back to acknowledged if it was acknowledged, otherwise firing. On a group that is not silenced it clears any leftover silenced_until and changes no status.
  • Reopen: New firing alerts do not join a resolved group — the next firing for its scope opens a new group. A resolved group reopens in place only in a race: a delivery that had already found the group open writes its firing alert just after the resolve commits, and the recount, judging the status under its row lock, returns the group to firing with a new firing epoch, resolved_at cleared and its escalation state reset.

A resolution never creates an alert group​

Only a delivery carrying at least one alert with status firing can create an alert group. A delivery with no firing alert — an Alertmanager batch of resolves, or a group-level recovery that carries no alerts at all — can only close a group:

DeliveryOpen group for (source, grouping key)?Result
Any alert firing—Group found or created, exactly as before (this is the paging path)
Resolves onlyYesThe group resolves — unchanged behavior
Resolves onlyNoDropped. Nothing to resolve. Logged at INFO, counted, and ACKed — not an error and never retried

Dropping is the correct outcome, not a lost alert: the source is resolving something Proxima has already closed. It matches the contract every pager uses (PagerDuty's Events API accepts a resolve for an unknown dedup key and creates nothing).

Watch the drop rate with:

# Redundant resolves, by source type. A steady low rate is normal for a flapping source.
increase(proxima_alert_resolve_no_open_group_total[15m])
source_type="proxima_probe" is expected to be non-zero right now

Uptime monitors gate the firing half of an emission on paging_enabled and never gate a resolve, deliberately — a held resolve leaves a group open that the monitor's next outage merges into, which would page nobody at all. So while paging ships off, every monitor recovery publishes a resolve that finds no open group, gets dropped, and ticks this counter in its own series. Roughly one tick per recovery is the healthy shadow-window shape. It only becomes a question once a monitor has paging_enabled on and its resolves still find no group.

Why this matters

Before this rule, a resolve for an already-closed group fell into the find-or-create path, matched nothing (its predicate excludes resolved groups), and inserted a new firing group with an invented P5 severity and zero alerts — which the recount then resolved in the same second. The result was a stream of empty "phantom" alert groups that had never fired.

A phantom could not page anyone directly — escalation only arms a group whose status is still firing, and the recount resolves it first. But it was not inert: a phantom still reached incident assignment, where it could open an incident it was then unable to close (it had no incident at ingest time, so the member-resolve hook never ran for it). Under ACTIVE grouping, a later real alert joining that scope is suppressed as a follower. So phantoms could not raise a page, but they could hold an incident scope open and suppress one.

Paging safety nets​

Two mechanisms make sure a firing alert that should page does page, even when a step along the way fails.

The correlation watchdog​

An accepted alert reaches escalation only through correlation, and the hand-off from the alert worker to correlation is a NATS publish made after the alert is saved. If that publish failed, or arming escalation failed inside correlation, the alert used to sit firing and page nobody until the stale-group reaper closed it.

Correlation now records, per firing episode, that it reached a paging decision — armed, not routed, paging off for the client, team in shadow, held by a maintenance window, or a duplicate whose paging belongs to its incident's leader. A decision not to page counts: those groups are decided too. The correlation watchdog looks for firing groups whose current episode has no decision and re-sends their correlation. Re-sending is safe: correlation can run any number of times, and escalation is armed only once per episode.

SettingDefaultWhat it does
PROXIMA_CORRELATION_WATCHDOG_INTERVAL60sHow often one replica sweeps.
PROXIMA_CORRELATION_WATCHDOG_GRACE2mHow long an episode may stay undecided before it is re-sent, and the minimum gap between two re-sends of the same episode.
PROXIMA_CORRELATION_WATCHDOG_MAX_RETRIES5Re-sends per episode before giving up. Must be at least 1.

Worst case, a lost hand-off pages about GRACE + INTERVAL (≈3 minutes) late instead of never. GRACE is measured from the group's last update, and every escalation step updates the group. So when a re-arm fails on a group that is already paging, the retry waits until the running chain pauses for GRACE; until then the group keeps paging on its current chain. At most 200 groups are re-sent per sweep; the rest follow on the next sweep. A re-send that fails (NATS down) is not counted against the episode and ends that sweep, so an outage cannot use up the budget. A new firing episode always starts with a fresh budget.

Watch two counters:

  • proxima_alert_correlation_recovered_total — re-sends. A few after a NATS blip is the watchdog working. A steady rate means hand-offs are being lost; each one is logged at WARN with alert_group_id, client_id and firing_epoch.
  • proxima_alert_correlation_unrecovered_total — episodes the watchdog gave up on after MAX_RETRIES re-sends. An episode that was never armed will not page; one that was armed and only failed a later re-arm keeps paging on the chain it already had. Alert on any increase, and read the ERROR log beside it (watchdog retries exhausted) for the group.
After the upgrade that adds the watchdog

Groups already firing at upgrade time that had armed escalation are marked decided by the migration. Firing groups that were never armed — paging off, no matching route, shadow team — start undecided, so the watchdog re-sends their correlation, and correlation then marks them decided — normally on the first re-send. Expect a one-off bump of proxima_alert_correlation_recovered_total after the deploy.

This can page. The re-sent correlation uses the client's configuration as it is now. A group that did not page when it fired but would now — on-call turned on for the client since, a route that now matches, a policy fixed since — arms and pages a few minutes after the upgrade, possibly several at once. Before upgrading, count them with the query in the alert ingestion standard (docs/standards/alert-ingestion.md, "Liveness and the paging decision", Rollout), and resolve stale ones or warn the on-call teams.

A more severe alert re-arms escalation​

Alertmanager groups alerts by its group_by labels, so a warning and a critical for the same rule can land in one Console alert group. The group's severity rises to the most severe firing alert, but its escalation used to stay on the route it was first armed on — a critical joining a group armed on a P3 Telegram-only route never got the P1 voice chain.

Now, when correlation finds an armed group that has become strictly more urgent than the severity it was armed at and now matches a different route:

  • firing group — escalation is re-armed on the new route, once per raise. The old chain is fenced: its pending steps can no longer page. The new chain starts at its first step now. The timeline shows Escalation re-armed: severity raised P3 → P1 (the two tiers vary), and proxima_escalation_rearmed_total{to_severity} counts it.
  • acknowledged group — nothing is re-paged; someone already owns it. The timeline shows Severity raised P3 → P1 while acknowledged — not re-paged, once per episode and severity, and proxima_escalation_severity_raised_while_acked_total counts it.

A drop in severity never changes a running escalation, and a raise that still matches the same route changes nothing — it is the same chain. Groups armed before this release keep the old behaviour until they resolve: their escalation does not record the severity it was armed at, so there is nothing to compare a raise against, and they are never re-armed. The alert card in Telegram is not re-posted when the new route belongs to another team; the pages follow the new route's escalation policy.

REST API Reference​

Alert Source Management​

Requires alertsources:read, alertsources:write, or alertsources:delete permissions.

MethodPathDescription
POST/api/v1/clients/{id}/alert-sourcesCreate alert source (returns token once)
GET/api/v1/clients/{id}/alert-sourcesList alert sources
GET/api/v1/clients/{id}/alert-sources/{id}Get alert source
PUT/api/v1/clients/{id}/alert-sources/{id}Update alert source — name, source_type, config, enabled; every field optional, omitted fields unchanged
DELETE/api/v1/clients/{id}/alert-sources/{id}Delete alert source

The same handlers are also mounted unscoped at /api/v1/alert-sources and /api/v1/alert-sources/{sourceID}. Updating is a PUT (not PATCH); see Changing a source's type for what happens when source_type changes.

GET /api/v1/alert-sources takes an optional client_id:

client_idWhat is listedScope
PresentThat one client, exactly as beforeThe handler re-checks Can(alertsources:read, client_id) itself — membership is not enough, and the middleware falls back to a union check on the path-client mount
AbsentEvery client the caller may read sources onClientIDsWithPermission(alertsources:read) — not AllowedClientIDs: holding some other permission on a client is not permission to read that client's ingest-token health, refusal reasons and counters

A super-admin's unscoped set means every client; a caller who holds the permission on no client gets [], never everything. On the nested mount /api/v1/clients/{id}/alert-sources the path {id} supplies the client when the query param is absent, so a URL that names a client in its own path still lists exactly that client rather than quietly widening.

Alert Label Mappings​

Requires alertsources:read, alertsources:write, or alertsources:delete permissions.

MethodPathDescription
GET/api/v1/clients/{id}/alert-label-mappingsList label mappings
POST/api/v1/clients/{id}/alert-label-mappingsCreate label mapping
PUT/api/v1/clients/{id}/alert-label-mappings/{id}Update label mapping
DELETE/api/v1/clients/{id}/alert-label-mappings/{id}Delete label mapping

Webhook Receiver​

Public endpoints. The two legacy doors authenticate via a token in the URL path; the generic door is header-only — see below.

MethodPathDescription
POST/api/v1/alerts/webhook/alertmanager/{token}Receive Alertmanager webhook (returns 202)
POST/api/v1/alerts/webhook/grafana/{token}Receive a Grafana 8.x legacy (panel) alerting webhook (returns 202)
POST/api/v1/alerts/webhook/genericReceive a payload from any sender that can POST JSON (returns 202). Authenticated by Authorization: Bearer <token>, never a URL token — full contract, worked examples and response codes in Generic Alert Webhook

The two legacy paths resolve the source by token hash, so the path segment is cosmetic — the source's stored source_type selects the parser. A 400 carries one of the reason codes; on these two doors an unknown or disabled token is always 404 (never 401, so the endpoint does not confirm which tokens exist). The generic door answers 401 for every authentication failure instead, with an identical body regardless of reason — see Generic Alert Webhook for why. A repeat delivery inside the replay guard's window (60s on alertmanager and generic, 24h on grafana_legacy) answers 200 {"status":"skipped"} without republishing.

Alert Browsing​

Requires alerts:read permission.

MethodPathDescription
GET/api/v1/alertsList alert groups (filterable)
GET/api/v1/alerts/{id}Alert group detail with correlation, alerts, and log
GET/api/v1/alerts/countsAlert counts by status (scoped to user's clients) — returns all four buckets: firing, acknowledged, resolved, silenced

Query Parameters for /alerts​

ParameterDefaultDescription
page1Page number
per_page20Items per page (max 100)
client_id—Filter by client
environment_id—Filter by environment
host_id—Filter by host
status—Filter by status: firing, acknowledged, resolved, silenced. Repeatable — pass ?status=firing&status=acknowledged to match any of several statuses (server-side status = ANY(...)). A single value still works for back-compat.
severity—Filter by severity: P1, P2, P3, P4, P5, or unknown. The UI dropdown offers P1–P5 only; pass unknown via the API to find alerts whose wire severity could not be mapped.
search—Free-text search (case-insensitive) over the alert name (common_labels.alertname) and the grouping key. Backed server-side.
from—Start time (RFC3339) — matches first_fired_at >= from
to—End time (RFC3339) — matches first_fired_at <= to

Enriched List Fields​

Every list item in GET /api/v1/alerts carries three read-time enrichment fields (computed by JOIN / COALESCE / correlated subquery — no schema change, no migration). They power the Client / Team / Users columns on the Alert Groups page:

FieldTypeMeaning
client_namestringThe owning client's display name (LEFT JOIN clients).
team_namestringThe resolved page-owner team, in COALESCE precedence: the team-owned escalation policy → the frozen escalation_snapshot.team_id (survives a policy delete) → the client's covering team. Empty string on an unrouted alert with no covering team — an accepted data gap, not a bug.
involved_usersarray of {id, name}The DISTINCT union of the users who touched the alert group's current firing episode: its acknowledger ∪ the users paged this episode (notification chains for this firing_epoch) ∪ the human audit actors (audit-log entries with actor_type=user). Always an array ([] when none), never null. Deleted or service-account ids simply drop out (no users row to match).

Bulk Actions​

Three bulk-mutation endpoints back the Alert Groups page's selection toolbar. Each takes a list of alert-group ids and applies the action to every one, enforcing alerts:write per id (route-level alerts:write plus a per-id tenant check). Ids the caller can't write, or that don't exist, are skipped, not failed — the batch returns 200 with a partial-success split.

MethodPathDescription
POST/api/v1/alerts/bulk/acknowledgeAcknowledge every id in ids
POST/api/v1/alerts/bulk/resolveResolve every id in ids
POST/api/v1/alerts/bulk/silenceSilence every id in ids until until (required, RFC3339)

Request body:

{
"ids": ["a1b2c3d4-...", "e5f6a7b8-..."],
"until": "2026-03-16T08:00:00Z"
}

until is required only for bulk/silence. A single request carries at most 500 ids.

Response (200):

{
"data": {
"succeeded": ["a1b2c3d4-..."],
"failed": [
{ "id": "e5f6a7b8-...", "reason": "forbidden" }
]
}
}

A non-empty failed array is expected, not an error — it is how per-id RBAC surfaces. reason is one of not_found, forbidden, resolved, changed, or an internal error string. resolved means acknowledge or silence refused a group that is already resolved and changed nothing. changed means the group's status kept changing while the action was being applied, so nothing was written; retrying that id may succeed. The frontend renders forbidden, not_found and errors as "M skipped (no access)", resolved separately as "K already resolved", and changed as "K changed meanwhile — try again".

Alert Actions​

Requires alerts:write permission. All actions are local to Proxima only — they do not reach back to Alertmanager or OnCall.

Acknowledge, Silence and Unsilence refuse a resolved group with 409 Conflict — "alert group is resolved" — and change nothing; Resolve stays idempotent. Each of the three is also a compare-and-set on the status it read: when another change lands between the read and the write — a person, the "still on it?" loop's auto-unacknowledge, the silence sweeper — nothing is written, the group is read again, and the action is decided again from what the group now is (a Silence records the status the group really had, so it ends by restoring that). After three such attempts the action answers 409 Conflict — "alert group changed, try again" — having changed nothing. The guard is in the database write, so a resolve that lands while the action is in flight still wins. Reviving a resolved group would make it the target of the next firing, which would then join it instead of opening a new group and page nobody.

Every path runs through the same service, and each tells its caller which refusal it was:

PathGroup already resolvedGroup kept changing
Web single action409 "alert group is resolved" — the alert detail page's Acknowledge and Silence toasts and the incident page's Acknowledge toast repeat it and refresh the page409 "alert group changed, try again" — repeated the same way
Web bulk actionthe id under failed with reason: resolvedthe id under failed with reason: changed
Telegram buttons (Ack, silences, Unsilence)Already resolvedChanged meanwhile — try again
Voice keypad, press 1 (Asterisk)"This alert has already been resolved.""Sorry, that could not be done. Please use the console."
Voice keypad, press 1 (Twilio)"This alert has already resolved. Goodbye.""The alert changed while you pressed. Please try again or open Console."
Triage chat's acknowledge toolan answer: already resolved, not acknowledgedan answer: changed while being acknowledged, not acknowledged

Neither keypad refusal stamps ack_source, and a late press of either is counted refused.

MethodPathDescription
POST/api/v1/alerts/{id}/acknowledgeAcknowledge alert group
POST/api/v1/alerts/{id}/resolveManually resolve alert group
POST/api/v1/alerts/{id}/silenceSilence until specified timestamp
POST/api/v1/alerts/{id}/unsilenceRemove silence

Silence Request Body​

{
"until": "2026-03-16T08:00:00Z"
}

Deduplication​

Alerts are deduplicated using Alertmanager's fingerprint field:

  • A partial unique index on (alert_group_id, fingerprint) WHERE status = 'firing' enforces at most one active alert per fingerprint per group
  • If a firing alert with the same fingerprint already exists, updated_at is refreshed (no duplicate created)
  • Resolved alerts free the fingerprint for future firing alerts
  • Alert groups are capped at 1000 alerts each; additional alerts are rejected with a log warning

Permissions​

PermissionScopeDescription
alerts:readBrowseView alert groups, alert detail, correlation context
alerts:writeActionsAcknowledge, resolve, silence, unsilence alert groups
alertsources:readConfigView alert sources and label mappings
alertsources:writeConfigCreate and update alert sources and label mappings
alertsources:deleteConfigDelete alert sources and label mappings

All endpoints are scoped to the authenticated user's accessible clients. Super admins can see all alerts across all clients.

Frontend​

Alert Groups Page​

Navigate to On-Call → Alert Groups to view all alert groups. The page is designed for Grafana OnCall parity — a filter row, clickable stat cards, a bulk-action toolbar, and a rich table — so on-call responders can triage a wall of alerts the way they already do in OnCall.

Filter row​

A single filter row drives the list; all filter state lives in the URL, so a filtered view is shareable and survives a reload:

  • Client — narrow to one client (multi-tenant). Defaults to "All clients".
  • Status chips — a multi-select row of togglable Firing / Acknowledged / Resolved / Silenced chips. Click to add or remove a status; the list shows the union of the selected statuses. Deselecting the last chip lands on "all statuses" (not a rebound to Firing). The chips stay in sync with the stat cards.
  • Search — free-text, case-insensitive, matched server-side against the alert name and grouping key (debounced).
  • Date-range — a preset picker (24h / 7d / 30d / custom). "Custom" reveals From/To datetime-local inputs. Matched against first_fired_at.
  • Severity — P1–P5.

Stat cards​

Four clickable cards — Firing / Acknowledged / Resolved / Silenced — show the live count per status (from GET /alerts/counts) and double as the primary status filter: clicking a card toggles that status through the same path as the chips, so a card and its chip are always in the same state. The active card is shown filled. This replaces the old plain-text count summary.

Bulk actions​

Writers (alerts:write) get a selection checkbox column plus a header select-all. Selecting one or more rows raises a floating bulk toolbar with Resolve / Acknowledge / Silence over the whole selection (Silence opens an until-picker). The toolbar calls the bulk endpoints above.

Partial success is expected. Because each id is checked against alerts:write on its own tenant, a batch can partly succeed — the toast reads "N done, M skipped (no access)". That is per-id RBAC working, not an error.

Rich table​

The table is a wide, expandable grid: checkbox · expand · ID · Severity · Status · Alert · Integration · Client · Team · Host · Users · Count · Age. It scrolls horizontally on narrow screens, and the Integration and Host columns collapse on smaller breakpoints. Notable columns:

  • Integration — the alert source type (e.g. alertmanager).
  • Client — the owning client (client_name).
  • Team — the resolved page-owner team (team_name). This can be blank (—) on an unrouted alert with no covering team — an accepted data gap, not a bug.
  • Users — an avatar cluster of the involved users (involved_users): the acknowledger, everyone paged this firing episode, and any human who acted in the audit log — deduplicated. Avatars show initials with a deterministic per-user color; beyond three, the rest collapse into a +N chip, and the whole cluster is labelled with every name for screen readers.

Other page behavior:

  • Hybrid expandable table — click (or keyboard-activate) any row to expand inline and see the correlation preview, summary, and action buttons. The Investigate lifecycle lives in the expanded row and on the detail page.
  • P1–P5 severity badges — color-coded priority levels (P1=Critical red, P2=High orange, P3=Medium amber, P4=Low blue, P5=Info gray).
  • Sortable columns — Severity, Status, Age (server-side sort).
  • 30-second polling — alerts and counts refresh automatically.
  • Sidebar badge — red count of firing alerts on the On-Call navigation entry.
  • In maintenance · not paged — shown on the row and the detail page while a maintenance window holds the alert's page; links to the window. The timeline row Not paged: in maintenance names the window; if the alert is still firing when the window ends it pages at once and the timeline records Maintenance ended: paging resumed.

Alert Detail Page​

Click "View Full Context →" on any alert to open the detail page (/alerts/:id). Two-column layout:

Left column — Correlation Data:

  • Metrics around the alert — the host's own series charted across the alert window (see below)
  • Metric cards with the value at alert time, the baseline, and a SPIKE / DROP / STABLE / UNKNOWN badge. An unknown card says "no baseline — no samples before the alert fired" rather than showing a baseline of 0%
  • Recent changes with source badges (GitLab, Agent, ArgoCD) and time offsets
  • Related alerts in Before/After card layout
  • Host information bar
  • Individual alerts in group table

Right sidebar — Activity Timeline:

  • Unified chronological feed mixing all event types (newest first):
    • Alert status changes (fired, acknowledged, resolved, silenced)
    • Correlation events (config changes detected, related alerts)
    • Operator notes
  • Notes composer at bottom for adding investigation notes (Markdown supported)
  • Notes are preserved as post-mortem evidence

Metrics around the alert​

Above the metric cards, the page charts the metrics the correlation worker picked for this alert, using the host's real series rather than the three numbers in the snapshot. The charts use the incident style from the charts standard (direction C):

  • Series: drawn in slate ink.
  • Firing window: the part of each line between first_fired_at and resolved_at (or now) is copper, over a faint copper band.
  • Baseline: a dotted line at the snapshot's pre-fire baseline, labelled "baseline 31%".
  • Summary banner: one line naming the largest spike or drop, e.g. "CPU Usage rose to 92% from a baseline of 31% — firing since 14:05".

Copper marks the firing window, not a threshold. Alerts arrive by webhook and Console stores no alert rule or threshold, so the chart never draws a threshold line it would have to invent.

AspectBehaviour
Who sees itThe alert must be resolved to a host, and the viewer needs metrics:read for the alert's project and environment. Otherwise the section is absent and only the metric cards show.
WindowStarts at the snapshot's window.from, or 30 minutes before first_fired_at if there is no snapshot. While the alert is open (firing, acknowledged or silenced), it ends at now and slides forward once a minute. Once resolved, it ends 10 minutes after resolved_at (or at the snapshot's end, if later) and no longer moves. Capped at 30 days.
Which metricsUp to three: anomalous metrics first, largest change first, then the rest in snapshot order.
Disk and inode metricsOne line per mount, up to the five fullest (ranked by their latest value in the last 15 minutes of the window). The rest are counted as "+N more mounts", and a small legend names each line.
Baseline on multi-mount metricsNot shown. The snapshot's baseline and at-alert value describe only one mount (the first series the metrics store returned), so a chart with several mounts gets no baseline line, no "at alert · baseline" subtitle, and no banner. The banner names the strongest change among single-series metrics instead, if there is one.
MarkersOther alerts on the host appear on the marker rail in copper; deploys and agent-offline events keep their usual style. This alert itself is left off, because the highlighted window already shows it.
No dataA metric with no samples in the window shows "No samples for this host in the alert window". If one mount fails to load, the others still show; only when every request for a metric fails does its card say "Failed to load metric data". If no metric has any sample, or the metrics request fails, the section is hidden and the metric cards remain. The page never shows an empty or NaN chart.

The data comes from the existing host metric endpoints (GET /hosts/:id/metrics/:name/series and GET /hosts/:id/metrics/:name). No new API is involved.

Investigation Lifecycle​

An on-demand L1 investigation can be started and tracked directly from the alert detail page — you no longer have to run it from the alerts-list expanded row. The Investigation section renders the full lifecycle in one slot:

  • Not investigated → an Investigate button (shown only to viewers with alerts:write on the alert's client; read-only viewers see the state without the button).
  • Investigating… → clicking Investigate removes the button and shows an "Investigating…" status line until the verdict lands. This state is durable: it survives a page reload and is visible to other viewers looking at the same alert, because it is backed by a real server-side signal (triage_in_progress on the alert detail) derived from the existing triage claim — not a local, per-browser flag. A crashed or abandoned run self-expires after the claim lease (30 minutes), so a stuck run never leaves a false "Investigating…".
  • Verdict → once the L1 analysis concludes, the section flips to the trust verdict block automatically (the page polls every 30 seconds).
  • Already analyzed → if this firing episode was already analyzed, the section says so and offers a Re-run anyway button (with a confirmation, since it spends an LLM call).

The detail page and the alerts-list expanded row share the same triage logic, so the two views stay consistent — they differ only in what happens on a fresh trigger (the list row navigates to the linked incident; the detail page stays put and lets polling surface the verdict).

Admin Pages​

  • Alert Sources (/oncall/alert-sources; /admin/alert-sources redirects there) — manage Alertmanager and Grafana-legacy webhook integrations across projects. Token-based authentication with one-time display. Each row also carries a per-source ingest health state (Producing / Rejecting / Not producing / No deliveries yet) with the last delivery and accepted timestamps, the accepted / rejected / unknown-severity counts, and the last refusal reason — see Per-source ingest health, including what that health state does and does not catch. See The Alert Sources page spans projects for the filter and the create flow.
  • Alert Label Mappings (/admin/alert-label-mappings) — configure how Alertmanager labels map to Proxima entities (client, environment, host, team).

The Alert Sources page spans projects​

The project selector is a filter, not a gate. The page lands on All projects and lists every source on every project the caller may read sources on, so the first thing it renders is data rather than an instruction to pick something. There is no state in which the page shows nothing but a prompt.

  • A Project column, first and always rendered. It names each row's project in both states — filtered and unfiltered — so the column never appears and disappears under you, and a single-project view still confirms which project you are looking at.
  • The search matches project too. It has always matched the project name; on a list that spans projects that finally means something. The placeholder says so: name, type or project.
  • Switching the filter cannot serve stale rows. The filter travels in the query key, so the previous project's rows are never served from cache for another project.
  • Create inherits the filter. With a project filtered, the create dialog does not ask the question a second time: it shows that project as a read-only row with a Change control that re-enables the picker — not a disabled dropdown, which reads as broken. Creating for a different project stays possible. From All projects, the dialog asks which project, and submit is blocked until one is chosen; that is the one case where the question is asked, because it is the one case where nothing answered it.

What is listed is bounded by permission, not by the selector: see GET /api/v1/alert-sources for the scoping rules.

Host Detail — Alerts Tab​

The host detail page includes an Alerts tab showing all alert groups for that specific host.

Permissions​

PermissionAccess
alerts:readView alerts, alert detail, host alerts tab
alerts:writeAcknowledge, resolve, silence alerts, add notes
alertsources:readView alert source and mapping admin pages
alertsources:writeCreate/edit alert sources and mappings
alertsources:deleteDelete alert sources and mappings

Incident Response​

The full incident-response layer that originally sat behind Grafana OnCall now ships in Proxima Console as the L1 incident agent:

  • Escalation policies — multi-step escalation with notify-users, round-robin queues, on-call-schedule resolution, and channel notifications (escalation_store, escalation_timer_worker)
  • On-call schedules and rotations — per-client schedules that resolve who is on call at notification time (oncall_store)
  • Telegram notifications — escalation steps fan notifications out to bound Telegram chats (notification worker)
  • Automated triage — the L1 agent triages incoming alert groups against an incident memory loop (triage worker)

Console is mid-cutover from Grafana OnCall to this native stack; per-client cutover state is tracked in the backend.

What's Deferred​

FeatureTimeline
Internal alert sources (healthcheck failures, host offline)Post-launch
Inhibition and silence rules (label matchers)Post-launch