Alerting & Correlation
Proxima Console ingests alerts from Alertmanager, enriches them with infrastructure context (metrics, changes, related alerts, compliance), and presents a correlated view for faster incident response.
Overview
The alerting system sits alongside your existing on-call tooling. Rather than replacing Alertmanager or Grafana OnCall, it acts as a parallel receiver that adds the infrastructure context operators need when they get paged at 3am:
- Metrics snapshot around the alert firing time (baseline vs. current values)
- Recent changes from both agent file monitoring and external webhooks (GitLab, GitHub, ArgoCD)
- Related alerts on the same host or environment
- Compliance state of the affected host
- Host context (hostname, OS, architecture)
VMAlert --> Alertmanager --+--> Grafana OnCall (keeps handling on-call)
|
+--> Proxima Console (adds infrastructure context)
Architecture
The pipeline is fully asynchronous. The webhook endpoint validates the token, parses the payload, publishes to NATS, and returns 202 Accepted immediately. Processing happens in two worker stages:
- AlertWorker (fast path): resolves labels to Proxima entities, deduplicates alerts by fingerprint, finds or creates alert groups, handles resolve/reopen lifecycle
- CorrelationWorker (enrichment): queries VictoriaMetrics for metrics, PostgreSQL for changes/related alerts/compliance, and writes the correlation context back to the alert group
Alert Source Setup
An alert source is a per-client integration endpoint that receives webhooks from Alertmanager.
Creating an Alert Source
curl -X POST https://api-console.prxm.uz/api/v1/clients/{client_id}/alert-sources \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-d '{
"name": "Production Alertmanager",
"source_type": "alertmanager",
"enabled": true
}'
Response (201):
{
"data": {
"id": "a1b2c3d4-...",
"client_id": "...",
"source_type": "alertmanager",
"name": "Production Alertmanager",
"enabled": true,
"token": "pxm_as_7f3a8b2c1d...",
"created_at": "2026-03-15T10:00:00Z",
"updated_at": "2026-03-15T10:00:00Z"
}
}
The token is a credential — store it securely and never put it in a URL. It is returned in
full by this create response only, but it is not lost if you close the dialog: the
Show webhook token action on the Alert Sources page (/oncall/alert-sources) reveals it
again at any time via GET /api/v1/alert-sources/{sourceID}/token (alertsources:read,
audit-logged). Only a source with no stored ciphertext — created before Console kept one, or on
a deployment with no Vault Transit encryptor — has to be recreated.
Alertmanager Configuration
Add a webhook_configs entry to your Alertmanager configuration pointing at the Proxima webhook receiver:
# alertmanager.yml
receivers:
- name: "proxima"
webhook_configs:
- url: "https://api-console.prxm.uz/api/v1/alerts/webhook/alertmanager/pxm_as_7f3a8b2c1d..."
send_resolved: true
route:
receiver: "default"
routes:
- receiver: "proxima"
continue: true # Important: continue to other receivers (OnCall, etc.)
matchers: [] # Match all alerts (or restrict with matchers)
Key points:
- Set
send_resolved: trueso Proxima can track alert resolution - Use
continue: trueif you want alerts to also reach other receivers (e.g., Grafana OnCall) - The token in the URL authenticates the request (no additional headers needed)
- Keep the route's
repeat_interval(Alertmanager's default is4h) shorter thanPROXIMA_INCIDENT_STALE_TTL(default6h) — see below
repeat_interval must be shorter than the stale TTLWhile an alert keeps firing, Alertmanager re-sends it every repeat_interval. Those repeats are
how Console knows the outage is still live: each one reaches the alert worker and refreshes the
group. A firing group that nothing has refreshed for PROXIMA_INCIDENT_STALE_TTL is resolved
by the stale-group reaper — escalation stops while the problem is still
there, and the next repeat opens it again and pages again.
So if you raise repeat_interval (for example to 12h to quiet another receiver), raise
PROXIMA_INCIDENT_STALE_TTL above it as well, or give the Proxima receiver its own route with a
shorter interval. Violations show up per source type on
proxima_alert_group_reaped_by_source_total{source_type}: a source type whose groups are
routinely reaped is one whose repeats are not arriving while it is still firing.
Repeats only count if they get through the webhook replay guard. On the alertmanager (and
generic) door its window is 60 seconds, so only an HA pair's duplicate send or an immediate
retry is skipped; every repeat_interval resend reaches the worker. Before this was fixed the
Alertmanager door used a 24-hour window, swallowed every byte-identical repeat, and auto-resolved
outages that were still firing after 6 hours.
Each repeat that gets through is processed like any other delivery, so it writes an Updated
and a Correlated entry on the group's timeline. With a short repeat_interval (1–5 minutes)
a long outage's timeline fills with these; that is expected.
/alerts/webhook/, not /webhooks/The receiver lives at /api/v1/alerts/webhook/alertmanager/{token}. Older revisions of this
page published /api/v1/webhooks/alertmanager/{token}, which 404s — Alertmanager reports a
delivery failure your alerts never recover from. If you copied a URL from an older doc, fix it.
Grafana 8.x legacy (panel) alerting posts to
/api/v1/alerts/webhook/grafana/{token} on a source whose source_type is grafana_legacy.
proxima_probe — Console's own first-party source
Uptime monitors publish onto the same ingest path a webhook does, so each client that has one
gets an alert_sources row with source_type = 'proxima_probe'. It behaves differently from
the two webhook types in five ways worth knowing before you look at one:
- It is created on demand, not authored. The probe result worker creates the row on its
first emission for that client. There is no webhook parser for it, so it is not creatable,
not editable and not deletable through the API — an update or delete answers
409 conflictnaming what the row is, because it sits in the operator's list beside their own integrations and the UI offers the same controls on it. Deleting it would stop every monitor in that tenant from paging until the next emission recreated it; re-pointing itssource_typewould break emission for that client permanently, because the reserved per-client token would be held by a row the create path can no longer match. - It has no webhook and no token. It stores a reserved sentinel in
token_hashinstead — shaped so nothing an attacker can send could hash to it — purely as the per-client uniqueness key. - It will always read "Never delivered." Per-source ingest health is written by
RecordIngest, whose only caller is the webhook handler, and probe emission never touches it. See the caveat under Per-source ingest health. - The row can appear before anything has ever paged. Probe resolves are deliberately
ungated, so a client's first monitor recovery creates the row even with
paging_enabledoff everywhere. - A monitor group pages through the same row. A group is a named service that pages once
for a whole set of monitors, and it is not a new source type — its alert arrives on this
client's
proxima_probesource under a distinctalertname,ProximaMonitorGroupDown, and a grouping key ofproxima_probe_group:<group_id>rather thanproxima_probe:<monitor_id>. The two are deliberately distinct so filtering or silencing one does not catch the other. A group page carries the group'senvironment_idand, by construction, nohost_id— so a host-scoped route can never match one.
Full behaviour in Uptime Monitors, and the model in Monitor groups.
Changing a source's type
PUT /api/v1/alert-sources/{sourceID} accepts source_type, so a source created with the
wrong type can be repaired without production SQL. A proxima_probe source is exempt — it
refuses every write. For the two webhook types, two things to know:
- The existing webhook URL keeps working. The handler resolves the source by token hash,
so the
/alertmanager/vs/grafana/path segment is cosmetic — both segments reach the same source, and the type stored on the source is what selects the parser. You do not have to reconfigure the sender. - It changes how every future payload is interpreted, so the flip is recorded in the
application log with the actor, and the old and new type. Nothing already stored is rewritten,
and another
PUTreverses it.
Ingestion Contract
Everything above gets a webhook delivered. This section is about what Proxima does with what is inside it — and, just as importantly, what it refuses to guess.
A severity-mapping break ran in production for 16 days undetected. Deliveries never stopped, so every liveness signal stayed green while the alerts were silently wrong: unlabelled alerts were laundered into plausible, pageable warnings, and every delivery that arrived without a grouping key collapsed into a single alert group. Nothing counted what a parse actually extracted, so nothing could notice. The contract below, and the per-source health and ratio alerts at the end of it, are what make that class of break visible.
Severity mapping
Proxima stores its own priority tiers, P1–P5. An incoming severity label is matched
case-insensitively (leading/trailing whitespace is trimmed) against this table:
| Proxima tier | Meaning | Accepted labels |
|---|---|---|
P1 | Page immediately | fatal, emergency, emerg, critical, crit, sev0, sev1, p1, page, panic, alert |
P2 | High | error, err, high, major, sev2, p2 |
P3 | Warning | warning, warn, minor, sev3, p3 |
P4 | Informational | info, informational, information, notice, low, sev4, p4 |
P5 | Least urgent | debug, trace, none, sev5, p5 |
unknown | We could not tell | anything absent from the table, and a missing label |
unknown is a real tier — we never fabricate a severity
An alert whose severity label is missing or unrecognized is stored as unknown. It is
never rounded up, never rounded down, and never given a plausible-looking default.
That is the whole point. A fabricated P3 looks exactly like a real warning, so it survives
every review and every dashboard; a fabricated tier on a genuine P1 silently downgrades a
page. unknown is honest, visible in the UI, countable, and ranks least urgent internally,
so a value we could not identify can never masquerade as urgent.
unknown gets no routing special-case. An unknown-severity alert matches:
- an escalation route whose severity is explicitly set to
unknown, or - a catch-all (default) route,
and nothing else. There is no "treat unknown as P1 to be safe" — urgency inflation on a value we
cannot identify would page people for debug messages, which is how a pager gets ignored.
unknown is selectable in the route editor at On-Call → Escalation, alongside P1–P5, and
in the severity filter on the Alert Groups page, so you can list exactly the alerts whose
severity Proxima could not identify.
Before this release, an alert whose severity label was missing or unrecognized was given a
fabricated P3. P3 is a real, pageable tier, so that alert matched an explicit P3
escalation route and paged. It is now unknown, which matches only an explicit unknown
route or a catch-all.
If your escalation routes are all explicit severities with no catch-all, those alerts are now captured, not paged: they are stored and visible on the Alert Groups page, but nobody is notified. Nothing is lost — but nobody is woken either, and there is no flag to turn this off.
The fix is one route. Add a catch-all (default) route pointing at whichever escalation
policy should own "we don't know what this is", or add an explicit unknown route to handle it
separately. If you already have a catch-all, you still get paged — but through the catch-all's
policy rather than the explicit P3 route's, so check that the catch-all points at a team you
want woken for an alert nobody could classify.
Sweep for the clients that need one with the
catch-all audit below. Afterwards, the
EscalationUnroutedUnknownSeverity alert rule fires on
proxima_escalation_unrouted_total{severity="unknown"} whenever it happens — but the counter
carries no client id, so the client comes from the accompanying escalation: alert matched no route warning in the logs.
Two operational traps
1. Route severity matching is case-sensitive. The matcher compares the stored route severity
to the alert's severity with an exact string comparison. Ingestion has already mapped the wire
label, so a route must store a canonical tier — P1, not critical, and not p1. A route
stored as critical or p1 matches nothing, forever, and nothing says so: no error, no
metric, no UI signal.
The API rejects a non-canonical severity on every route write with a 400 naming the whole
allowed set. That validation is write-time only — see
the pre-cut-over audit for rows written before it
existed.
2. Alertmanager fabricates severities too. Alertmanager's own PagerDuty receiver defaults
severity to error when the alert carries no severity label. If Proxima and PagerDuty
disagree about an alert's urgency, this is usually why: PagerDuty is showing you Alertmanager's
invented error, Proxima is showing you the honest unknown. Fix it at the source by labelling
the rule.
What returns a 400
Eleven machine-readable reason codes. Seven are unconditional; four are gated by the strict
ingestion contract (below). The last six apply only to the generic door.
| Reason | Always a 400? | What it means | Fix |
|---|---|---|---|
unreadable_body | yes | The body could not be read, or exceeded the size limit | Check the sender and any proxy in front of it |
empty_body | yes | Zero-length body | The sender posted nothing |
malformed_payload | yes | The parser rejected the payload | Wrong source_type, or the sender's format changed |
no_content | only in strict mode | Parsed, but no alert carried an alertname, labels or annotations | The sender's payload shape drifted |
no_identity_legacy | only in strict mode | A grafana_legacy payload with neither ruleId nor ruleName | Grafana is sending an unidentifiable rule |
missing_dedup_key | yes | A generic payload with no dedup_key, on a source with no dedup_key_from derivation | Send a dedup_key, or configure the derivation on the source |
dedup_key_unresolved | yes | The source derives its key from fields this payload does not carry as non-empty top-level values | Fix the sender, or the source's dedup_key_from list (matched exactly and case-sensitively) |
dedup_key_too_long | yes | A generic dedup key over 512 bytes (bytes, not characters — a CJK key is 3 bytes per character) | Shorten the key, or derive it from fewer fields |
invalid_status | yes | A generic payload whose status is neither firing nor resolved | Fix the sender |
missing_summary | only in strict mode | A generic payload with no summary text | Fix the sender — this is the line an on-call engineer reads at 3am |
unknown_severity | only in strict mode | A generic payload whose severity is absent or maps to no Proxima tier | Label the alert, or ask for the alias to be added |
We answer with a named reason rather than a silent 202 so the sender learns its payload
shape broke. A 202 teaches a drifting sender that it is successfully paging you.
A label value is never one of these. On the generic door every labels value is rendered
to text — numbers, booleans and null through the same renderer the dedup key uses, objects and
arrays as their compact JSON — so {"labels":{"port":8080,"code":503}} is ingested, not
refused. A whole page is never dropped over the least important field in the payload. Only a
labels field that is not a JSON object at all is a parse error.
Authentication failures are a 401, not one of these, and they are counted separately
(proxima_alert_webhook_auth_refused_total{reason}). Two of them — the source is disabled,
or a legacy source's token was presented at the generic door — are also stamped on that
source's health, so the Alert Sources page shows a refusal with its reason instead of
"No deliveries yet". Everything else about the 401 is deliberately uninformative: five auth
outcomes answer one identical body, so the response cannot be used to probe which token is real.
A group-level resolve is exempt. A grafana_legacy recovery (state=ok or paused)
legitimately carries zero alerts — its content is the grouping key. Those are never
refused; refusing them would leave every group from that source firing forever.
What gets a synthesized identity
A missing grouping key is repaired, never refused.
An empty grouping key is not "no key" downstream: sha256("") is a perfectly valid hash, so
every keyless delivery from a source hashed to the same value and collapsed into one
alert group. That was the 16-day incident.
In strict mode Proxima derives a key from the group's identity — its common labels plus the
sorted, de-duplicated set of alertnames — and stores it with a synthesized: prefix, so you can
always tell an invented key from a source-provided one. Keying on identity (not on the whole
alert batch) means repeats of the same group still land in the existing group, while two
different groups can never merge.
PROXIMA_ALERT_INGEST_STRICT
A single backend environment variable, default off. Only the literal string on enables
it. It is per-deployment, not per-source.
Counting is not behavior. The strict contract's verdicts are evaluated and counted in both
modes, and acted on only when the flag is on. With the flag off nothing is rejected,
nothing is synthesized, and every byte published is identical to the pre-hardening path — only
the counters, the WARN logs and the trace attributes differ.
That is what makes the flag safe to flip: you can read exactly what strict mode would refuse, per source, before touching it.
| Flag off (default) | Flag on | |
|---|---|---|
no_content delivery | 202, ingested, counted as would-refuse | 400 no_content |
no_identity_legacy delivery | 202, ingested, counted as would-refuse | 400 no_identity_legacy |
generic delivery with no summary | 202, ingested, counted as would-refuse | 400 missing_summary |
generic delivery with an unmappable severity | 202, ingested as unknown, counted as would-refuse | 400 unknown_severity |
| Missing grouping key | published with the empty key (collapses), counted | key synthesized, then published |
| Unconditional rejections | 400 — unchanged | 400 — unchanged |
Rollout order: watch the ratios for a day with the flag off → confirm the would-reject ratio per source is what you expect → flip the flag → re-check the same ratios.
Per-source ingest health
The Alert Sources page (/oncall/alert-sources; /admin/alert-sources redirects there)
shows, per source: when it last delivered, when it last produced an accepted alert, its
accepted / rejected / unknown-severity counts, and the reason and time of its last actual
refusal.
| State | What it means |
|---|---|
| Producing | Recent deliveries are producing accepted alerts |
| Rejecting | Still accepting, but a delivery in the last 24h was actually refused (the reason is shown) |
| Not producing | Deliveries are arriving but none has produced an accepted alert for over two days |
| No deliveries yet | Nothing has ever been delivered. After 24h from creation this becomes a warning — usually "the sender was never pointed at this webhook" |
"Not producing" catches a total ingest collapse — two days of deliveries that produced no accepted alert at all. That is real and worth catching.
It does NOT catch the failure class behind the 16-day incident. Mapping drift means alerts
that are accepted but wrong — a severity we could not map, a fabricated identity. Those keep
last_accepted_at fresh, so the row stays green throughout. An operator who believes a green
row rules that out is worse off than one who knows the limit.
That class shows up as the unknown-severity count on this page, and is detected by the ratio alerts below — not by the health state.
Two further caveats, both deliberate:
- These counters are best-effort. They are written off the ingest hot path in a detached goroutine, so a health write can be lost without affecting the alert. Treat them as a health signal, never as an accounting ledger. The authoritative numbers are the Prometheus counters.
- "Quiet" is not a state. A source that has not delivered for a while is normal — Alertmanager only talks when something fires — so silence after a first delivery is reported as timestamps, not as a fault.
proxima_probe row shows "No deliveries yet" forever, and that is not a faultThese columns are written by RecordIngest, whose only caller is the webhook handler.
Uptime monitors publish straight onto the ingest subject and never pass through it, so
last_delivery_at, last_accepted_at and every counter on Console's own probe source stay
null and zero however many pages it has carried — and after 24 hours the row turns the
warning colour that normally means "the sender was never pointed at this webhook".
Judge probe health by proxima_probe_paged_total and the emission counters in
Uptime Monitors, never by this row's status.
Ingest metrics and the ratio alerts
Six per-source counters, all labelled source_id and source_type
(full reference in the internal docs/standards/metrics.md):
| Metric | Meaning |
|---|---|
proxima_alert_ingest_received_total | Deliveries past the token/enabled check — the denominator, less replayed |
proxima_alert_ingest_accepted_total | Deliveries actually published for processing |
proxima_alert_ingest_replayed_total | Repeat deliveries the replay guard suppressed (answered 200, nothing published) |
proxima_alert_ingest_rejected_total | Deliveries judged unacceptable — plus reason and refused |
proxima_alert_ingest_unknown_severity_total | Severities that mapped to no known tier — plus kind |
proxima_alert_ingest_synthesized_identity_total | Deliveries whose empty grouping key was (or would be) replaced |
proxima_alert_webhook_auth_refused_total | Deliveries refused at the authentication stage — plus reason. These never reach received, so without this a sender failing auth 100% of the time moved no counter at all |
Alert on the ratio, not the count. An absolute count alerts on volume; a ratio alerts on correctness.
# would-reject ratio — read per source BEFORE flipping PROXIMA_ALERT_INGEST_STRICT on it
sum by (source_id) (rate(proxima_alert_ingest_rejected_total{refused="false"}[1h])
or rate(proxima_alert_ingest_received_total[1h]) * 0)
/ sum by (source_id) (rate(proxima_alert_ingest_received_total[1h])
- (rate(proxima_alert_ingest_replayed_total[1h])
or rate(proxima_alert_ingest_received_total[1h]) * 0))
# internal-failure gap — publish failure, marshal failure, or a JetStream-less backend
sum by (source_id) (rate(proxima_alert_ingest_received_total[5m]))
- sum by (source_id) (rate(proxima_alert_ingest_replayed_total[5m])
or rate(proxima_alert_ingest_received_total[5m]) * 0)
- sum by (source_id) (rate(proxima_alert_ingest_accepted_total[5m])
or rate(proxima_alert_ingest_received_total[5m]) * 0)
- sum by (source_id) (rate(proxima_alert_ingest_rejected_total[5m])
or rate(proxima_alert_ingest_received_total[5m]) * 0)
The or … * 0 clauses and the per-term sum by (source_id) are load-bearing. rejected_total
carries extra labels (reason, refused), and Prometheus arithmetic matches on the full
label set — so a bare rate(received) - rate(rejected) matches nothing and the whole expression
returns empty. Separately, a missing series drops its element rather than counting as
zero, so a source that has never had an accepted delivery — exactly the outage the gap exists to
catch — would vanish from the result. An expression that silently returns nothing looks
identical to a healthy one.
That second point is why the would-reject ratio's numerator carries the fallback too: a
source with no would-rejects has no rejected{refused="false"} series, and without it "safe to
flip the flag" and "the expression is broken" would render identically — on the one number the
flip is gated on. It must read 0, not nothing.
Three things about these that are easy to get wrong:
- Both expressions subtract
replayed. A replay-skip is routine, not a loss: an Alertmanager HA pair's duplicate send and a sender's immediate retry both land there, and on the deprecatedgrafana_legacydoor (24h window) so does every byte-identical repeat. Leave them in the denominator and the would-reject ratio understates what strict mode would refuse on exactly the sources that repeat most. The window is 60 seconds on thegenericandalertmanagerdoors, not 24 hours.generichas no time-varying field, so a cron-style sender's two distinct firings are byte-identical and a day-long window would swallow the second outage and page nobody; Alertmanager'srepeat_intervalresends are byte-identical and are the group's proof of life (see therepeat_intervalrule). Past a client's immediate retry, the(source_id, grouping_key)partial-unique index already makes a repeated trigger join the open group instead of duplicating it. refusedis not "was the flag on".refused="true"means the delivery really was refused with a400;refused="false"means it was judged refusable but ingested with a202. An unconditional rejection emitsrefused="true"with the flag off. Filterreason=no_content|no_identity_legacyto isolate the flag-gated verdicts.- Never sum the two
unknown_severitykinds.kind="unrecognized"is a label we do not speak (fix: add an alias).kind="absent"is a sender not labelling at all (fix: fix the sender). On a source that never labels severity,absenttracks alert volume — a permanently high constant that reads as chronic breakage and gets tuned out.
unknown_severity and synthesized_identity are emitted only for deliveries that were actually
published, so both are a true share of accepted.
Where to watch this. The Alert Ingestion RED row on the Proxima Backend Observability
Grafana dashboard renders all of the above, and the provisioned proxima-alert-ingest alert
group fires on the unknown-severity signature, the internal-failure gap, a sustained refusal
rate, and unknown-severity alerts that matched no escalation route
(EscalationUnroutedUnknownSeverity, over proxima_escalation_unrouted_total{severity="unknown"}
— the Unrouted Alerts by Severity panel on the same row is its dashboard half).
Pre-cut-over audit: stored route severities
Route severity is validated at write time only. There is no database constraint and no
backfill, so any route stored before that validation existed — critical, p1, P9 — is still
there, still dead, and still silent.
Run this before cutting a client over, and treat a non-empty result as a blocker:
SELECT id, client_id, severity FROM escalation_routes
WHERE severity IS NOT NULL AND severity NOT IN ('P1','P2','P3','P4','P5','unknown');
Why no metric will tell you instead. You would expect proxima_escalation_unrouted_total to
catch a route that matches nothing. It does not. That counter only increments when no route
matched — and when a catch-all route exists, the matcher short-circuits on the catch-all before
it ever compares severity, so it always returns a route. The alert therefore arms the wrong
escalation policy, the unrouted counter never moves, and the one metric an operator would think
to check stays at zero.
Every row this query returns is a route that has never matched anything and never will. Fix them
with a PUT through the API — which validates — not with direct SQL.
Pre-cut-over audit: clients with no catch-all route
The other half of the same sweep, and the one that follows from the paging change described
under "unknown is a real tier — we never fabricate a severity" above. A client whose
escalation routes are all explicit severities pages nobody for an alert whose severity Proxima
could not map.
Run this before cutting a client over:
SELECT client_id, count(*) FILTER (WHERE is_default) AS catch_all_routes
FROM escalation_routes
GROUP BY client_id
HAVING count(*) FILTER (WHERE is_default) = 0;
Every client it returns will have its unknown-severity alerts captured but not paged. Give each
one a catch-all route — or an explicit unknown route — before repointing its integrations.
Unlike the severity audit above, this one is watched after the fact:
proxima_escalation_unrouted_total{severity="unknown"} moves as soon as it happens and the
EscalationUnroutedUnknownSeverity rule fires on it. Run the query anyway — the counter tells
you afterwards, at the cost of the alerts in between.
Label Mapping
Label mappings define how Alertmanager labels resolve to Proxima infrastructure entities. Each client configures their own mappings based on their labeling conventions.
How It Works
- When an alert arrives, the AlertWorker loads the client's label mappings (cached, refreshed every 60s)
- For each mapping, it extracts the
source_labelvalue from the alert's labels - Applies the configured
transform(exact match, strip port, or regex) - Looks up the transformed value in the corresponding Proxima entity table
Creating Label Mappings
curl -X POST https://api-console.prxm.uz/api/v1/clients/{client_id}/alert-label-mappings \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-d '{
"source_label": "instance",
"target_field": "host_address",
"transform": "strip_port",
"priority": 10
}'
Target Fields
| Target Field | Matches Against | Example |
|---|---|---|
client_slug | clients.slug | cluster=acme-prod maps to client acme-prod |
environment_slug | environments.slug | namespace=staging maps to environment staging |
host_address | hosts.ip_address or hosts.hostname | instance=10.0.1.5:9100 maps to host 10.0.1.5 |
team_name | teams.name | team=platform maps to team platform |
Transforms
| Transform | Behavior | Example |
|---|---|---|
exact | Use label value as-is | cluster=production matches production |
strip_port | Remove :port suffix | instance=10.0.1.5:9100 becomes 10.0.1.5 |
regex | Regex extraction | Extract hostname from FQDN patterns |
Common Mapping Configurations
Kubernetes cluster setup:
[
{"source_label": "cluster", "target_field": "client_slug", "transform": "exact", "priority": 10},
{"source_label": "namespace", "target_field": "environment_slug", "transform": "exact", "priority": 20},
{"source_label": "instance", "target_field": "host_address", "transform": "strip_port", "priority": 30},
{"source_label": "team", "target_field": "team_name", "transform": "exact", "priority": 40}
]
Bare-metal/VM setup:
[
{"source_label": "datacenter", "target_field": "environment_slug", "transform": "exact", "priority": 10},
{"source_label": "instance", "target_field": "host_address", "transform": "strip_port", "priority": 20}
]
If label resolution fails (no matching host, environment, or client), the alert is stored with host_id = NULL. It still appears in the global alerts list and client-scoped views.
Correlation
When the AlertWorker finishes processing an alert, it publishes a message to proxima.alerts.correlate. The CorrelationWorker gathers infrastructure context and writes it to the correlation JSONB column on the alert group.
What Gets Correlated
| Context | Source | Window |
|---|---|---|
| Metrics snapshot | VictoriaMetrics | ±30 minutes around first_fired_at |
| Recent changes | PostgreSQL (change_events) | PROXIMA_CORRELATION_CHANGE_WINDOW before the alert — 24 h by default. When that window is empty, falls back to the host's last-N changes, each labeled with its true offset. |
| Related alerts | PostgreSQL (alert_groups) | Firing alerts on same host/environment (limit 10) |
| Compliance score | PostgreSQL (control_evaluations) | Average across all frameworks at correlation time |
| Host snapshot | PostgreSQL (hosts) | Hostname, OS, architecture |
How Correlation Works
When the AlertWorker finishes processing an alert group, it publishes the alert_group_id to proxima.alerts.correlate. The CorrelationWorker picks it up and:
- Loads the alert group from PostgreSQL
- If
host_idis set (label resolution succeeded):- Resolves the VM tenant via
tenantCache(host → client →vm_account_id) - Queries VictoriaMetrics for up to 3 metrics chosen from the alert's own name and labels in a ±30-minute window around
first_fired_atwith 60-second step (see below) - Computes trends: separates data points into baseline (before alert) and at-alert (closest point to fire time). Trend =
spikeif at-alert > 1.5× baseline,dropif < 0.5× baseline,stableotherwise - Queries change events from the correlation window (24 h by default, limit 10) — includes agent-detected file changes, backend inventory diffs, and external webhook events (GitLab, GitHub, ArgoCD)
- Queries compliance score — averages across all frameworks for the host
- Captures host snapshot — hostname, OS, architecture
- Resolves the VM tenant via
- Queries related alerts — other firing alert groups on the same host or environment (limit 10, excludes self), with relative time offsets
- Stores the correlation JSON in
alert_groups.correlationand logs a"correlated"action to the audit trail
If any individual enrichment step fails (e.g., VM query timeout, change store error), it logs a warning and continues with the other steps. Partial correlation data is better than none.
Correlation JSON Structure
The correlation column on alert_groups contains this JSON structure:
{
"metrics_snapshot": {
"window": {
"from": "2026-03-15T10:16:00Z",
"to": "2026-03-15T11:16:00Z"
},
"series": {
"cpu_usage_ratio": {"baseline": 0.341, "at_alert": 0.872, "trend": "spike"},
"memory_used_ratio": {"baseline": 0.623, "at_alert": 0.910, "trend": "spike"},
"disk_used_ratio": {"baseline": 0.779, "at_alert": 0.785, "trend": "stable"}
}
},
"recent_changes": [
{
"id": "d4e5f6a7-...",
"source": "gitlab",
"event_type": "deployment",
"summary": "Merged MR !482: reduce memory limits",
"file_path": "/etc/kubernetes/manifests/api-server.yaml",
"severity": "high",
"created_at": "2026-03-15T10:28:00Z"
}
],
"related_alerts": [
{
"id": "a1b2c3d4-...",
"grouping_key": "DiskSpaceFull",
"status": "firing",
"severity": "P1",
"first_fired_at": "2026-03-15T10:34:00Z",
"relative_offset": "-12m"
}
],
"compliance_score": 0.72,
"host_snapshot": {
"id": "e3f4a5b6-...",
"hostname": "prod-k8s-01",
"os": "Ubuntu 24.04",
"arch": "x86_64"
}
}
Field reference:
| Field | Type | Description |
|---|---|---|
metrics_snapshot.window | object | Time range queried (from/to in RFC 3339) |
metrics_snapshot.series | map | Metric name → trend object. The (up to 3) metrics are selected per alert — see the keyword table below. Values are 0.0–1.0 ratios. |
metrics_snapshot.series.*.baseline | float | Average value of data points before first_fired_at |
metrics_snapshot.series.*.at_alert | float | Value of the data point closest to first_fired_at |
metrics_snapshot.series.*.trend | string | "spike" (above 1.5× baseline), "drop" (below 0.5× baseline), "stable", or "unknown" (the series has samples but none before first_fired_at, so there is no baseline; baseline is then a placeholder 0) |
metrics_snapshot.series.*.has_data | bool | false when the metric had no sample in the window at all; the other fields are then placeholders and the UI shows "no data" |
Metric selection is alert-aware. A disk, iowait, memory or database alert is investigated with the right series rather than always cpu/mem/disk. The worker scans the distinct alert names, the grouping key, and the common-labels blob for the first matching keyword:
| Keyword | Metrics queried (most relevant first) |
|---|---|
iowait | cpu_iowait_ratio, disk_used_ratio, cpu_usage_ratio |
disk | disk_used_ratio, disk_inodes_used_ratio, cpu_iowait_ratio |
inode | disk_inodes_used_ratio, disk_used_ratio, memory_used_ratio |
swap | memory_swap_used_ratio, memory_used_ratio, cpu_usage_ratio |
oom, memory, mem | memory_used_ratio, memory_swap_used_ratio, cpu_usage_ratio |
cpu, load | cpu_usage_ratio, cpu_iowait_ratio, memory_used_ratio |
latency | cpu_usage_ratio, cpu_iowait_ratio, disk_used_ratio |
postgres | pg_cache_hit_ratio, memory_used_ratio, cpu_usage_ratio |
database | pg_cache_hit_ratio, memory_used_ratio, disk_used_ratio |
The first matching keyword wins, and the result is bounded to ~3 metrics. cpu_usage_ratio,
memory_used_ratio, disk_used_ratio is only the no-match fallback, not the standard set.
| recent_changes | array | Up to 10 most recent change events in the correlation window (24 h by default) before the alert |
| recent_changes[].source | string | "agent", "gitlab", "github", "argocd" |
| recent_changes[].file_path | string | File path (if applicable, omitted otherwise) |
| related_alerts | array | Up to 10 other firing alert groups on the same host/environment |
| related_alerts[].relative_offset | string | Time offset from this alert's first_fired_at (e.g., "-12m" = fired 12 min before, "+2m" = fired 2 min after) |
| compliance_score | float | Average compliance score across all frameworks (0.0–1.0). null if no evaluations exist. |
| host_snapshot | object | Host info at correlation time. null if host_id was not resolved. |
All metric values in metrics_snapshot.series are 0.0–1.0 ratios (not percentages). The frontend multiplies by 100 for display. This matches the OpenMetrics naming convention used throughout the agent.
metrics_snapshotisnullwhen: the host has no VM data, the tenant cache is unavailable, or all metric queries failed.recent_changesandrelated_alertsare always arrays (empty[]when none found, nevernull).compliance_scoreisnullwhen no compliance evaluations exist for the host.host_snapshotisnullwhenhost_idwas not resolved by label mapping.
Alert Lifecycle
Alert groups follow a four-state lifecycle:
| Status | Description |
|---|---|
firing | One or more alerts are actively firing |
acknowledged | An operator has acknowledged the alert group |
resolved | Closed: every alert in the group resolved, a person resolved it, the stale-group reaper closed it, or the silence sweeper resolved a silence that had ended over 24h earlier |
silenced | Paging suppressed until a set time, when the silence sweeper ends it |
State Transitions
- Acknowledge: Sets the status to
acknowledged— fromfiringorsilenced— withacknowledged_byandacknowledged_at. Requiresalerts:writepermission. When the group's escalation is armed, every acknowledgement starts the "still on it?" loop: the acknowledger is asked to confirm, and if nobody confirms within the confirm window the group is auto-unacknowledged — it returns tofiringand escalates again (ack_auto_unackedon the timeline). Nothing else returns an acknowledged group tofiring: a further firing alert leaves it acknowledged. - Resolve: Manually resolves the alert group from any open status, with its firing alerts, and halts its escalation. Does not reach back to Alertmanager. A person's Resolve is an assertion and is written as given — only a source's resolve is corroborated first.
- A source's resolve is corroborated before it is written. When a delivery carrying no firing alert would close an open group, Console first looks for independent evidence that the entity is still alive — the host's check-in time, and the uptime monitors bound to that host. Evidence of trouble refuses the resolve: nothing is written, the group stays firing with its escalation untouched, and a
resolve_not_corroboratedrow records why. Anything else writes the resolve and records whether a signal confirmed it (resolve_verified) or nothing could speak (resolve_unverified, the common case). The check is bounded at 5s and fails open, so it can never hold a resolve back on a failure of its own. See Resolution Corroboration. - Stale-group reaper: a
firinggroup that nothing has updated for the reaper's TTL (PROXIMA_INCIDENT_STALE_TTL, default6h) is resolved by the system, with its firing alerts. It never touches an acknowledged or silenced group. A source that keeps re-sending a still-firing alert keeps its group fresh — which is why Alertmanager'srepeat_intervalmust be shorter than the TTL. Reaps are counted per source type onproxima_alert_group_reaped_by_source_total. - Silence: Sets
silenced_untilto a future timestamp and suppresses paging until then. The silence ends on its own at that time: within 30 seconds the silence sweeper returns the group to the status it had when it was silenced —acknowledgedif it was acknowledged, otherwisefiring— and paging resumes where it stopped (an armed escalation picks up on the timer's next tick; an acknowledged group's "still on it?" loop resumes). The status to return to is recorded asprior_statuson thesilencedtimeline entry; silencing an already-silenced group carries the earlier one forward, and a silence recorded before this rule existed returns tofiring. The timeline gets a system Silence ended entry saying which status the group went back to. - A silence that ended more than 24 hours before it was noticed resolves the group instead. This covers silences that never ended under earlier releases, and a sweeper that was down for a day: resuming them would page everyone at once for incidents that are days or months old. The group is resolved by the system with the same writes a manual resolve makes (its firing alerts resolved, its escalation halted), and the timeline says why. Nothing is lost: if the problem is still firing, its next alert opens a fresh group, which can page again.
proxima_alert_silences_ended_total{outcome="expired"|"resolved_stale"}counts both. - Unsilence: Ends the silence early by the same rule — back to
acknowledgedif it was acknowledged, otherwisefiring. On a group that is not silenced it clears any leftoversilenced_untiland changes no status. - Reopen: New firing alerts do not join a resolved group — the next firing for its scope opens a new group. A resolved group reopens in place only in a race: a delivery that had already found the group open writes its firing alert just after the resolve commits, and the recount, judging the status under its row lock, returns the group to
firingwith a new firing epoch,resolved_atcleared and its escalation state reset.
A resolution never creates an alert group
Only a delivery carrying at least one alert with status firing can create an alert group. A delivery with no firing alert — an Alertmanager batch of resolves, or a group-level recovery that carries no alerts at all — can only close a group:
| Delivery | Open group for (source, grouping key)? | Result |
|---|---|---|
Any alert firing | — | Group found or created, exactly as before (this is the paging path) |
| Resolves only | Yes | The group resolves — unchanged behavior |
| Resolves only | No | Dropped. Nothing to resolve. Logged at INFO, counted, and ACKed — not an error and never retried |
Dropping is the correct outcome, not a lost alert: the source is resolving something Proxima has already closed. It matches the contract every pager uses (PagerDuty's Events API accepts a resolve for an unknown dedup key and creates nothing).
Watch the drop rate with:
# Redundant resolves, by source type. A steady low rate is normal for a flapping source.
increase(proxima_alert_resolve_no_open_group_total[15m])
source_type="proxima_probe" is expected to be non-zero right nowUptime monitors gate the firing half of an emission on paging_enabled and never gate a
resolve, deliberately — a held resolve leaves a group open that the monitor's next outage
merges into, which would page nobody at all. So while paging ships off, every monitor
recovery publishes a resolve that finds no open group, gets dropped, and ticks this counter in
its own series. Roughly one tick per recovery is the healthy shadow-window shape. It only
becomes a question once a monitor has paging_enabled on and its resolves still find no
group.
Before this rule, a resolve for an already-closed group fell into the find-or-create path, matched nothing (its predicate excludes resolved groups), and inserted a new firing group with an invented P5 severity and zero alerts — which the recount then resolved in the same second. The result was a stream of empty "phantom" alert groups that had never fired.
A phantom could not page anyone directly — escalation only arms a group whose status is still firing, and the recount resolves it first. But it was not inert: a phantom still reached incident assignment, where it could open an incident it was then unable to close (it had no incident at ingest time, so the member-resolve hook never ran for it). Under ACTIVE grouping, a later real alert joining that scope is suppressed as a follower. So phantoms could not raise a page, but they could hold an incident scope open and suppress one.
Paging safety nets
Two mechanisms make sure a firing alert that should page does page, even when a step along the way fails.
The correlation watchdog
An accepted alert reaches escalation only through correlation, and the hand-off from the alert
worker to correlation is a NATS publish made after the alert is saved. If that publish failed, or
arming escalation failed inside correlation, the alert used to sit firing and page nobody until
the stale-group reaper closed it.
Correlation now records, per firing episode, that it reached a paging decision — armed, not
routed, paging off for the client, team in shadow, held by a maintenance window, or a duplicate
whose paging belongs to its incident's leader. A decision not to page counts: those groups are
decided too. The correlation watchdog looks for firing groups whose current episode has no
decision and re-sends their correlation. Re-sending is safe: correlation can run any number of
times, and escalation is armed only once per episode.
| Setting | Default | What it does |
|---|---|---|
PROXIMA_CORRELATION_WATCHDOG_INTERVAL | 60s | How often one replica sweeps. |
PROXIMA_CORRELATION_WATCHDOG_GRACE | 2m | How long an episode may stay undecided before it is re-sent, and the minimum gap between two re-sends of the same episode. |
PROXIMA_CORRELATION_WATCHDOG_MAX_RETRIES | 5 | Re-sends per episode before giving up. Must be at least 1. |
Worst case, a lost hand-off pages about GRACE + INTERVAL (≈3 minutes) late instead of never.
GRACE is measured from the group's last update, and every escalation step updates the group.
So when a re-arm fails on a group that is already
paging, the retry waits until the running chain pauses for GRACE; until then the group keeps
paging on its current chain.
At most 200 groups are re-sent per sweep; the rest follow on the next sweep. A re-send that fails
(NATS down) is not counted against the episode and ends that sweep, so an outage cannot use up
the budget. A new firing episode always starts with a fresh budget.
Watch two counters:
proxima_alert_correlation_recovered_total— re-sends. A few after a NATS blip is the watchdog working. A steady rate means hand-offs are being lost; each one is logged at WARN withalert_group_id,client_idandfiring_epoch.proxima_alert_correlation_unrecovered_total— episodes the watchdog gave up on afterMAX_RETRIESre-sends. An episode that was never armed will not page; one that was armed and only failed a later re-arm keeps paging on the chain it already had. Alert on any increase, and read the ERROR log beside it (watchdog retries exhausted) for the group.
Groups already firing at upgrade time that had armed escalation are marked decided by the
migration. Firing groups that were never armed — paging off, no matching route, shadow team —
start undecided, so the watchdog re-sends their correlation, and correlation then marks them
decided — normally on the first re-send. Expect a one-off bump of
proxima_alert_correlation_recovered_total after the deploy.
This can page. The re-sent correlation uses the client's configuration as it is now. A group
that did not page when it fired but would now — on-call turned on for the client since, a route
that now matches, a policy fixed since — arms and pages a few minutes after the upgrade, possibly
several at once. Before upgrading, count them with the query in the alert ingestion standard
(docs/standards/alert-ingestion.md, "Liveness and the paging decision", Rollout), and resolve
stale ones or warn the on-call teams.
A more severe alert re-arms escalation
Alertmanager groups alerts by its group_by labels, so a warning and a critical for the same
rule can land in one Console alert group. The group's severity rises to the most severe firing
alert, but its escalation used to stay on the route it was first armed on — a critical joining
a group armed on a P3 Telegram-only route never got the P1 voice chain.
Now, when correlation finds an armed group that has become strictly more urgent than the severity it was armed at and now matches a different route:
firinggroup — escalation is re-armed on the new route, once per raise. The old chain is fenced: its pending steps can no longer page. The new chain starts at its first step now. The timeline shows Escalation re-armed: severity raised P3 → P1 (the two tiers vary), andproxima_escalation_rearmed_total{to_severity}counts it.acknowledgedgroup — nothing is re-paged; someone already owns it. The timeline shows Severity raised P3 → P1 while acknowledged — not re-paged, once per episode and severity, andproxima_escalation_severity_raised_while_acked_totalcounts it.
A drop in severity never changes a running escalation, and a raise that still matches the same route changes nothing — it is the same chain. Groups armed before this release keep the old behaviour until they resolve: their escalation does not record the severity it was armed at, so there is nothing to compare a raise against, and they are never re-armed. The alert card in Telegram is not re-posted when the new route belongs to another team; the pages follow the new route's escalation policy.
REST API Reference
Alert Source Management
Requires alertsources:read, alertsources:write, or alertsources:delete permissions.
| Method | Path | Description |
|---|---|---|
| POST | /api/v1/clients/{id}/alert-sources | Create alert source (returns token once) |
| GET | /api/v1/clients/{id}/alert-sources | List alert sources |
| GET | /api/v1/clients/{id}/alert-sources/{id} | Get alert source |
| PUT | /api/v1/clients/{id}/alert-sources/{id} | Update alert source — name, source_type, config, enabled; every field optional, omitted fields unchanged |
| DELETE | /api/v1/clients/{id}/alert-sources/{id} | Delete alert source |
The same handlers are also mounted unscoped at /api/v1/alert-sources and
/api/v1/alert-sources/{sourceID}. Updating is a PUT (not PATCH); see
Changing a source's type for what happens when source_type changes.
GET /api/v1/alert-sources takes an optional client_id:
client_id | What is listed | Scope |
|---|---|---|
| Present | That one client, exactly as before | The handler re-checks Can(alertsources:read, client_id) itself — membership is not enough, and the middleware falls back to a union check on the path-client mount |
| Absent | Every client the caller may read sources on | ClientIDsWithPermission(alertsources:read) — not AllowedClientIDs: holding some other permission on a client is not permission to read that client's ingest-token health, refusal reasons and counters |
A super-admin's unscoped set means every client; a caller who holds the permission on no
client gets [], never everything. On the nested mount /api/v1/clients/{id}/alert-sources
the path {id} supplies the client when the query param is absent, so a URL that names a
client in its own path still lists exactly that client rather than quietly widening.
Alert Label Mappings
Requires alertsources:read, alertsources:write, or alertsources:delete permissions.
| Method | Path | Description |
|---|---|---|
| GET | /api/v1/clients/{id}/alert-label-mappings | List label mappings |
| POST | /api/v1/clients/{id}/alert-label-mappings | Create label mapping |
| PUT | /api/v1/clients/{id}/alert-label-mappings/{id} | Update label mapping |
| DELETE | /api/v1/clients/{id}/alert-label-mappings/{id} | Delete label mapping |
Webhook Receiver
Public endpoints. The two legacy doors authenticate via a token in the URL path; the generic
door is header-only — see below.
| Method | Path | Description |
|---|---|---|
| POST | /api/v1/alerts/webhook/alertmanager/{token} | Receive Alertmanager webhook (returns 202) |
| POST | /api/v1/alerts/webhook/grafana/{token} | Receive a Grafana 8.x legacy (panel) alerting webhook (returns 202) |
| POST | /api/v1/alerts/webhook/generic | Receive a payload from any sender that can POST JSON (returns 202). Authenticated by Authorization: Bearer <token>, never a URL token — full contract, worked examples and response codes in Generic Alert Webhook |
The two legacy paths resolve the source by token hash, so the path segment is cosmetic — the
source's stored source_type selects the parser. A 400 carries one of the
reason codes; on these two doors an unknown or disabled token is always
404 (never 401, so the endpoint does not confirm which tokens exist). The generic door
answers 401 for every authentication failure instead, with an identical body regardless of
reason — see Generic Alert Webhook for why. A
repeat delivery inside the replay guard's window (60s on alertmanager and generic, 24h on
grafana_legacy)
answers 200 {"status":"skipped"} without republishing.
Alert Browsing
Requires alerts:read permission.
| Method | Path | Description |
|---|---|---|
| GET | /api/v1/alerts | List alert groups (filterable) |
| GET | /api/v1/alerts/{id} | Alert group detail with correlation, alerts, and log |
| GET | /api/v1/alerts/counts | Alert counts by status (scoped to user's clients) — returns all four buckets: firing, acknowledged, resolved, silenced |
Query Parameters for /alerts
| Parameter | Default | Description |
|---|---|---|
page | 1 | Page number |
per_page | 20 | Items per page (max 100) |
client_id | — | Filter by client |
environment_id | — | Filter by environment |
host_id | — | Filter by host |
status | — | Filter by status: firing, acknowledged, resolved, silenced. Repeatable — pass ?status=firing&status=acknowledged to match any of several statuses (server-side status = ANY(...)). A single value still works for back-compat. |
severity | — | Filter by severity: P1, P2, P3, P4, P5, or unknown. The UI dropdown offers P1–P5 only; pass unknown via the API to find alerts whose wire severity could not be mapped. |
search | — | Free-text search (case-insensitive) over the alert name (common_labels.alertname) and the grouping key. Backed server-side. |
from | — | Start time (RFC3339) — matches first_fired_at >= from |
to | — | End time (RFC3339) — matches first_fired_at <= to |
Enriched List Fields
Every list item in GET /api/v1/alerts carries three read-time enrichment fields (computed by JOIN / COALESCE / correlated subquery — no schema change, no migration). They power the Client / Team / Users columns on the Alert Groups page:
| Field | Type | Meaning |
|---|---|---|
client_name | string | The owning client's display name (LEFT JOIN clients). |
team_name | string | The resolved page-owner team, in COALESCE precedence: the team-owned escalation policy → the frozen escalation_snapshot.team_id (survives a policy delete) → the client's covering team. Empty string on an unrouted alert with no covering team — an accepted data gap, not a bug. |
involved_users | array of {id, name} | The DISTINCT union of the users who touched the alert group's current firing episode: its acknowledger ∪ the users paged this episode (notification chains for this firing_epoch) ∪ the human audit actors (audit-log entries with actor_type=user). Always an array ([] when none), never null. Deleted or service-account ids simply drop out (no users row to match). |
Bulk Actions
Three bulk-mutation endpoints back the Alert Groups page's selection toolbar. Each takes a list of alert-group ids and applies the action to every one, enforcing alerts:write per id (route-level alerts:write plus a per-id tenant check). Ids the caller can't write, or that don't exist, are skipped, not failed — the batch returns 200 with a partial-success split.
| Method | Path | Description |
|---|---|---|
| POST | /api/v1/alerts/bulk/acknowledge | Acknowledge every id in ids |
| POST | /api/v1/alerts/bulk/resolve | Resolve every id in ids |
| POST | /api/v1/alerts/bulk/silence | Silence every id in ids until until (required, RFC3339) |
Request body:
{
"ids": ["a1b2c3d4-...", "e5f6a7b8-..."],
"until": "2026-03-16T08:00:00Z"
}
until is required only for bulk/silence. A single request carries at most 500 ids.
Response (200):
{
"data": {
"succeeded": ["a1b2c3d4-..."],
"failed": [
{ "id": "e5f6a7b8-...", "reason": "forbidden" }
]
}
}
A non-empty failed array is expected, not an error — it is how per-id RBAC surfaces. reason is one of not_found, forbidden, resolved, changed, or an internal error string. resolved means acknowledge or silence refused a group that is already resolved and changed nothing. changed means the group's status kept changing while the action was being applied, so nothing was written; retrying that id may succeed. The frontend renders forbidden, not_found and errors as "M skipped (no access)", resolved separately as "K already resolved", and changed as "K changed meanwhile — try again".
Alert Actions
Requires alerts:write permission. All actions are local to Proxima only — they do not reach back to Alertmanager or OnCall.
Acknowledge, Silence and Unsilence refuse a resolved group with 409 Conflict — "alert group is resolved" — and change nothing; Resolve stays idempotent. Each of the three is also a compare-and-set on the status it read: when another change lands between the read and the write — a person, the "still on it?" loop's auto-unacknowledge, the silence sweeper — nothing is written, the group is read again, and the action is decided again from what the group now is (a Silence records the status the group really had, so it ends by restoring that). After three such attempts the action answers 409 Conflict — "alert group changed, try again" — having changed nothing. The guard is in the database write, so a resolve that lands while the action is in flight still wins. Reviving a resolved group would make it the target of the next firing, which would then join it instead of opening a new group and page nobody.
Every path runs through the same service, and each tells its caller which refusal it was:
| Path | Group already resolved | Group kept changing |
|---|---|---|
| Web single action | 409 "alert group is resolved" — the alert detail page's Acknowledge and Silence toasts and the incident page's Acknowledge toast repeat it and refresh the page | 409 "alert group changed, try again" — repeated the same way |
| Web bulk action | the id under failed with reason: resolved | the id under failed with reason: changed |
| Telegram buttons (Ack, silences, Unsilence) | Already resolved | Changed meanwhile — try again |
| Voice keypad, press 1 (Asterisk) | "This alert has already been resolved." | "Sorry, that could not be done. Please use the console." |
| Voice keypad, press 1 (Twilio) | "This alert has already resolved. Goodbye." | "The alert changed while you pressed. Please try again or open Console." |
| Triage chat's acknowledge tool | an answer: already resolved, not acknowledged | an answer: changed while being acknowledged, not acknowledged |
Neither keypad refusal stamps ack_source, and a late press of either is counted refused.
| Method | Path | Description |
|---|---|---|
| POST | /api/v1/alerts/{id}/acknowledge | Acknowledge alert group |
| POST | /api/v1/alerts/{id}/resolve | Manually resolve alert group |
| POST | /api/v1/alerts/{id}/silence | Silence until specified timestamp |
| POST | /api/v1/alerts/{id}/unsilence | Remove silence |
Silence Request Body
{
"until": "2026-03-16T08:00:00Z"
}
Deduplication
Alerts are deduplicated using Alertmanager's fingerprint field:
- A partial unique index on
(alert_group_id, fingerprint) WHERE status = 'firing'enforces at most one active alert per fingerprint per group - If a firing alert with the same fingerprint already exists,
updated_atis refreshed (no duplicate created) - Resolved alerts free the fingerprint for future firing alerts
- Alert groups are capped at 1000 alerts each; additional alerts are rejected with a log warning
Permissions
| Permission | Scope | Description |
|---|---|---|
alerts:read | Browse | View alert groups, alert detail, correlation context |
alerts:write | Actions | Acknowledge, resolve, silence, unsilence alert groups |
alertsources:read | Config | View alert sources and label mappings |
alertsources:write | Config | Create and update alert sources and label mappings |
alertsources:delete | Config | Delete alert sources and label mappings |
All endpoints are scoped to the authenticated user's accessible clients. Super admins can see all alerts across all clients.
Frontend
Alert Groups Page
Navigate to On-Call → Alert Groups to view all alert groups. The page is designed for Grafana OnCall parity — a filter row, clickable stat cards, a bulk-action toolbar, and a rich table — so on-call responders can triage a wall of alerts the way they already do in OnCall.
Filter row
A single filter row drives the list; all filter state lives in the URL, so a filtered view is shareable and survives a reload:
- Client — narrow to one client (multi-tenant). Defaults to "All clients".
- Status chips — a multi-select row of togglable Firing / Acknowledged / Resolved / Silenced chips. Click to add or remove a status; the list shows the union of the selected statuses. Deselecting the last chip lands on "all statuses" (not a rebound to Firing). The chips stay in sync with the stat cards.
- Search — free-text, case-insensitive, matched server-side against the alert name and grouping key (debounced).
- Date-range — a preset picker (24h / 7d / 30d / custom). "Custom" reveals
From/Todatetime-localinputs. Matched againstfirst_fired_at. - Severity — P1–P5.
Stat cards
Four clickable cards — Firing / Acknowledged / Resolved / Silenced — show the live count per status (from GET /alerts/counts) and double as the primary status filter: clicking a card toggles that status through the same path as the chips, so a card and its chip are always in the same state. The active card is shown filled. This replaces the old plain-text count summary.
Bulk actions
Writers (alerts:write) get a selection checkbox column plus a header select-all. Selecting one or more rows raises a floating bulk toolbar with Resolve / Acknowledge / Silence over the whole selection (Silence opens an until-picker). The toolbar calls the bulk endpoints above.
Partial success is expected. Because each id is checked against alerts:write on its own tenant, a batch can partly succeed — the toast reads "N done, M skipped (no access)". That is per-id RBAC working, not an error.
Rich table
The table is a wide, expandable grid: checkbox · expand · ID · Severity · Status · Alert · Integration · Client · Team · Host · Users · Count · Age. It scrolls horizontally on narrow screens, and the Integration and Host columns collapse on smaller breakpoints. Notable columns:
- Integration — the alert source type (e.g.
alertmanager). - Client — the owning client (
client_name). - Team — the resolved page-owner team (
team_name). This can be blank (—) on an unrouted alert with no covering team — an accepted data gap, not a bug. - Users — an avatar cluster of the involved users (
involved_users): the acknowledger, everyone paged this firing episode, and any human who acted in the audit log — deduplicated. Avatars show initials with a deterministic per-user color; beyond three, the rest collapse into a+Nchip, and the whole cluster is labelled with every name for screen readers.
Other page behavior:
- Hybrid expandable table — click (or keyboard-activate) any row to expand inline and see the correlation preview, summary, and action buttons. The Investigate lifecycle lives in the expanded row and on the detail page.
- P1–P5 severity badges — color-coded priority levels (P1=Critical red, P2=High orange, P3=Medium amber, P4=Low blue, P5=Info gray).
- Sortable columns — Severity, Status, Age (server-side sort).
- 30-second polling — alerts and counts refresh automatically.
- Sidebar badge — red count of firing alerts on the On-Call navigation entry.
- In maintenance · not paged — shown on the row and the detail page while a maintenance window holds the alert's page; links to the window. The timeline row Not paged: in maintenance names the window; if the alert is still firing when the window ends it pages at once and the timeline records Maintenance ended: paging resumed.
Alert Detail Page
Click "View Full Context →" on any alert to open the detail page (/alerts/:id). Two-column layout:
Left column — Correlation Data:
- Metrics around the alert — the host's own series charted across the alert window (see below)
- Metric cards with the value at alert time, the baseline, and a SPIKE / DROP / STABLE / UNKNOWN badge. An
unknowncard says "no baseline — no samples before the alert fired" rather than showing a baseline of 0% - Recent changes with source badges (GitLab, Agent, ArgoCD) and time offsets
- Related alerts in Before/After card layout
- Host information bar
- Individual alerts in group table
Right sidebar — Activity Timeline:
- Unified chronological feed mixing all event types (newest first):
- Alert status changes (fired, acknowledged, resolved, silenced)
- Correlation events (config changes detected, related alerts)
- Operator notes
- Notes composer at bottom for adding investigation notes (Markdown supported)
- Notes are preserved as post-mortem evidence
Metrics around the alert
Above the metric cards, the page charts the metrics the correlation worker picked for this alert, using the host's real series rather than the three numbers in the snapshot. The charts use the incident style from the charts standard (direction C):
- Series: drawn in slate ink.
- Firing window: the part of each line between
first_fired_atandresolved_at(or now) is copper, over a faint copper band. - Baseline: a dotted line at the snapshot's pre-fire baseline, labelled "baseline 31%".
- Summary banner: one line naming the largest spike or drop, e.g. "CPU Usage rose to 92% from a baseline of 31% — firing since 14:05".
Copper marks the firing window, not a threshold. Alerts arrive by webhook and Console stores no alert rule or threshold, so the chart never draws a threshold line it would have to invent.
| Aspect | Behaviour |
|---|---|
| Who sees it | The alert must be resolved to a host, and the viewer needs metrics:read for the alert's project and environment. Otherwise the section is absent and only the metric cards show. |
| Window | Starts at the snapshot's window.from, or 30 minutes before first_fired_at if there is no snapshot. While the alert is open (firing, acknowledged or silenced), it ends at now and slides forward once a minute. Once resolved, it ends 10 minutes after resolved_at (or at the snapshot's end, if later) and no longer moves. Capped at 30 days. |
| Which metrics | Up to three: anomalous metrics first, largest change first, then the rest in snapshot order. |
| Disk and inode metrics | One line per mount, up to the five fullest (ranked by their latest value in the last 15 minutes of the window). The rest are counted as "+N more mounts", and a small legend names each line. |
| Baseline on multi-mount metrics | Not shown. The snapshot's baseline and at-alert value describe only one mount (the first series the metrics store returned), so a chart with several mounts gets no baseline line, no "at alert · baseline" subtitle, and no banner. The banner names the strongest change among single-series metrics instead, if there is one. |
| Markers | Other alerts on the host appear on the marker rail in copper; deploys and agent-offline events keep their usual style. This alert itself is left off, because the highlighted window already shows it. |
| No data | A metric with no samples in the window shows "No samples for this host in the alert window". If one mount fails to load, the others still show; only when every request for a metric fails does its card say "Failed to load metric data". If no metric has any sample, or the metrics request fails, the section is hidden and the metric cards remain. The page never shows an empty or NaN chart. |
The data comes from the existing host metric endpoints (GET /hosts/:id/metrics/:name/series and GET /hosts/:id/metrics/:name). No new API is involved.
Investigation Lifecycle
An on-demand L1 investigation can be started and tracked directly from the alert detail page — you no longer have to run it from the alerts-list expanded row. The Investigation section renders the full lifecycle in one slot:
- Not investigated → an Investigate button (shown only to viewers with
alerts:writeon the alert's client; read-only viewers see the state without the button). - Investigating… → clicking Investigate removes the button and shows an "Investigating…" status line until the verdict lands. This state is durable: it survives a page reload and is visible to other viewers looking at the same alert, because it is backed by a real server-side signal (
triage_in_progresson the alert detail) derived from the existing triage claim — not a local, per-browser flag. A crashed or abandoned run self-expires after the claim lease (30 minutes), so a stuck run never leaves a false "Investigating…". - Verdict → once the L1 analysis concludes, the section flips to the trust verdict block automatically (the page polls every 30 seconds).
- Already analyzed → if this firing episode was already analyzed, the section says so and offers a Re-run anyway button (with a confirmation, since it spends an LLM call).
The detail page and the alerts-list expanded row share the same triage logic, so the two views stay consistent — they differ only in what happens on a fresh trigger (the list row navigates to the linked incident; the detail page stays put and lets polling surface the verdict).
Admin Pages
- Alert Sources (
/oncall/alert-sources;/admin/alert-sourcesredirects there) — manage Alertmanager and Grafana-legacy webhook integrations across projects. Token-based authentication with one-time display. Each row also carries a per-source ingest health state (Producing / Rejecting / Not producing / No deliveries yet) with the last delivery and accepted timestamps, the accepted / rejected / unknown-severity counts, and the last refusal reason — see Per-source ingest health, including what that health state does and does not catch. See The Alert Sources page spans projects for the filter and the create flow. - Alert Label Mappings (
/admin/alert-label-mappings) — configure how Alertmanager labels map to Proxima entities (client, environment, host, team).
The Alert Sources page spans projects
The project selector is a filter, not a gate. The page lands on All projects and lists every source on every project the caller may read sources on, so the first thing it renders is data rather than an instruction to pick something. There is no state in which the page shows nothing but a prompt.
- A Project column, first and always rendered. It names each row's project in both states — filtered and unfiltered — so the column never appears and disappears under you, and a single-project view still confirms which project you are looking at.
- The search matches project too. It has always matched the project name; on a list that spans projects that finally means something. The placeholder says so: name, type or project.
- Switching the filter cannot serve stale rows. The filter travels in the query key, so the previous project's rows are never served from cache for another project.
- Create inherits the filter. With a project filtered, the create dialog does not ask the question a second time: it shows that project as a read-only row with a Change control that re-enables the picker — not a disabled dropdown, which reads as broken. Creating for a different project stays possible. From All projects, the dialog asks which project, and submit is blocked until one is chosen; that is the one case where the question is asked, because it is the one case where nothing answered it.
What is listed is bounded by permission, not by the selector: see
GET /api/v1/alert-sources for the scoping rules.
Host Detail — Alerts Tab
The host detail page includes an Alerts tab showing all alert groups for that specific host.
Permissions
| Permission | Access |
|---|---|
alerts:read | View alerts, alert detail, host alerts tab |
alerts:write | Acknowledge, resolve, silence alerts, add notes |
alertsources:read | View alert source and mapping admin pages |
alertsources:write | Create/edit alert sources and mappings |
alertsources:delete | Delete alert sources and mappings |
Incident Response
The full incident-response layer that originally sat behind Grafana OnCall now ships in Proxima Console as the L1 incident agent:
- Escalation policies — multi-step escalation with notify-users, round-robin queues, on-call-schedule resolution, and channel notifications (
escalation_store,escalation_timer_worker) - On-call schedules and rotations — per-client schedules that resolve who is on call at notification time (
oncall_store) - Telegram notifications — escalation steps fan notifications out to bound Telegram chats (
notificationworker) - Automated triage — the L1 agent triages incoming alert groups against an incident memory loop (
triageworker)
Console is mid-cutover from Grafana OnCall to this native stack; per-client cutover state is tracked in the backend.
What's Deferred
| Feature | Timeline |
|---|---|
| Internal alert sources (healthcheck failures, host offline) | Post-launch |
| Inhibition and silence rules (label matchers) | Post-launch |