Skip to main content

Resolution Corroboration

A source saying an alert resolved is evidence, not proof. The failure this check exists for is absence of data reading as recovery: an exporter dies, a target leaves service discovery, a rule stops being evaluated — and the condition may be worse than when it fired. Before the alert worker writes a resolve that would close an open alert group, internal/resolvecheck.Verifier looks for independent evidence that the entity is still alive. Positive evidence of trouble refuses the resolve; anything else lets it through.

What this is not — it never re-evaluates the alert's condition

Console cannot re-derive whether the condition still holds. Alertmanager sends labels, annotations and a generatorURL — never the rule expression or the threshold — so the only alerts whose condition Console could re-evaluate are the ones whose rule Console wrote itself. The check asks a different and answerable question: is the thing this alert is about still reachable? A resolve it accepts is not a statement that the problem is fixed.

Which resolves are checked​

Only a source's resolve, and only when it would close a group that is open right now. Six conditions must all hold, and each rules out a real delivery:

ConditionWhat it rules out
A verifier is wiredNothing: it is wired unconditionally in cmd/server/workers.go, with no ops-bot token involved
The client has not turned it offA client that opted out of corroboration — see Turning it off for one client. It is on by default, and this is the last condition checked, so it costs one client read only on a delivery that would otherwise be checked
The delivery carries no firing alertA firing delivery, and a batch whose envelope says "resolved" while carrying a firing member — that is a live incident, and it closes nothing
It is a group-level resolve, or it carries alertsAn empty delivery that would close nothing anyway
The group was not created by this deliveryA group that has only just opened
The group is not already resolvedA group somebody already closed

A resolve for a scope with no open group never reaches the check at all: it is dropped on the find-only path, as Alerting describes.

Three closes are deliberately not checked​

  • A person pressing Resolve. A human resolve is an assertion, not a report — someone is claiming the incident, and the system's job is to record that claim, not to argue with it. If the condition persists, the problem's next alert opens a new group.
  • The stale-group (TTL) reaper.
  • The silence sweeper resolving a silence that had ended more than 24 hours earlier.

The last two are Console's own system closes, each with its own timeline row saying so. Corroborating them would be Console checking its own arithmetic against a signal that has nothing to do with it.

Turning it off for one client​

The check is a per-client feature flag, resolve_check, in the client's features map. It defaults to true for every client, existing and new — the one default-on flag in domain.LaunchDefaults added since launch. That departs from the prevailing pattern deliberately: this is a safety behaviour that already runs for every client, and shipping it default-off would leave it dormant and unproven.

Set it to false and that client's source resolves are written unchecked: no verifier call, no verdict row on the timeline, nothing counted, and the delivery behaves exactly as it did before this check existed.

{ "features": { "resolve_check": false } }

It is written through the super-admin client update (PUT /api/v1/clients/{id}), which replaces the whole features map — send the client's existing features with this key changed, never this key alone. The admin Features editor does not render a toggle for it, but it preserves it: the editor merges its own checkboxes over the stored map, so a resolve_check set through the API survives unrelated feature edits.

A single stuck alert does not need this flag

Turning it off is a standing decision about a whole client, and it is not the remedy for one group that will not close. A person's Resolve in Console is never corroborated — a human resolve is an assertion, not a report — so a stuck alert is closed by one click, with no flag change and no deploy. Reach for the flag only when the check is working correctly and is unwanted for that client.

Why the flag exists at all is worth stating precisely, because it is not a safety valve for the verifier breaking. Every failure mode of the verifier already fails open toward accepting the resolve (see It fails open, by construction). The flag exists because the check can be correct and unwanted: a refused resolve keeps the group open and keeps paging whoever is on call, and a client may not want Console making that call for them.

If reading the flag fails, the check is skipped​

A database error reading the client's features skips the check and writes the resolve, logging at WARN. This is deliberately the opposite of the oncall_enabled gate on the paging path, which fails closed: that gate protects against a page being sent, so not knowing must mean not sending; this one protects against a resolve being withheld, so not knowing must mean not withholding. No infrastructure error may hold a resolve open.

A worker built with no feature seam at all runs the check for everyone — unwired means "no opinion", not "disabled". Whether the check exists at all is decided by the verifier's own wiring and nothing else.

Nothing counts a skip​

A client with the check disabled records no metric. proxima_alert_resolve_checks_total's label set is closed at confirmed|unverified|contradicted, and a fourth value would claim a check concluded something when no check ran. The skip is visible on the INFO log line beside the delivery, which names the alert group and the client.

The signals, and what they actually cover​

A verdict carries three signal slots, in a fixed order. Two of them are asked. One is inert.

SignalWhat is readTroubleAliveWhen it cannot speak
hosthosts.last_seen_at — the agent's check-in timeNo check-in for longer than 3 minutesChecked in within 3 minutesThe alert names no host
monitorThe enabled uptime monitors bound to the alert's host in its environment, newest result per monitor, ordered by monitor nameAny one of them is downOne is up and none is downThe alert names no host or no environment, no monitor is bound to it, or every result is older than 5 minutes
sourceNothing. It always reads unknownneverneveralways — see below

The host signal reads last_seen_at and deliberately not hosts.agent_status: that column is a stored verdict a sweeper refreshes on a timeout plus its own check interval, so it lags reality by minutes, and the live fact is the check-in time.

A monitor state older than 5 minutes is skipped entirely: a monitor whose probes stopped holds its last state forever, so without that window a monitor that went down and was then forgotten would contradict every resolve for that entity until somebody deleted it. pending, maintenance, paused and unknown are not health readings and count as neither — treating a maintenance window as evidence would put an operator's own admin action on the card as proof of an outage.

So in practice the check asks at most two signals, and only for an alert that names a host. An alert with no host has nothing to ask and resolves as unverified.

The source signal is inert, in both directions​

It consults nothing and always reads unknown. This is not a wiring gap to be closed by pointing it at a column:

  • The only clock available to it is alert_sources.last_delivery_at, and that column is written by the delivery being verified — the webhook handler stamps it on a deferred, best-effort goroutine while the resolve is written downstream by the alert worker.
  • If the stamp lands first, a fresh timestamp is the source vouching for itself. Confirming a resolve on that basis is the opposite of independent evidence.
  • If the worker wins the race, or the best-effort write simply failed, the value read is the previous delivery's — routinely hours old, because Alertmanager re-notifies a firing alert on repeat_interval (commonly 4h). Calling that trouble would refuse a resolve, and keep paging people, on the strength of a dropped telemetry write.

It is doubly inert for uptime-monitor alerts: RecordIngest's only caller is the alert webhook handler and probe emission never touches it, so a per-client proxima_probe source never stamps that column at all.

A real source-liveness signal needs a clock the delivery under verification does not write — the last time this source delivered something other than this alert, or a per-source heartbeat. Until one exists, read the third slot as "never asked", not as "asked and had nothing to say".

How the readings fold into one verdict​

Trouble outranks alive, alive outranks unknown, and unknown alone is unverified. The first trouble signal supplies the verdict's reason, and the order is host, then monitor — so when a host's check-in gap and a down monitor both report trouble, the card names the host's measurement.

The three outcomes​

OutcomeMeaningWhat happens to the resolve
confirmedA signal reported the entity alive, and none reported troubleWritten, exactly as before this check existed
unverifiedNothing could speak — no host on the alert, no monitor for the entity, a reader error, or the deadlineWritten anyway, and labelled honestly. Never a confident green
contradictedA signal positively reported troubleNot written. The group stays open with its escalation state untouched, and the card says why

unverified is the common outcome, and a high rate is not a fault. Alert→host resolution measured 5.1% fleet-wide, so most alerts have no host to ask about; proxima_alert_host_resolution_total is the instrument that explains this one's shape. A deployment where nearly every resolve is unverified is the expected shape, not a broken check.

What a refused resolve changes: one row, and nothing else​

A contradicted resolve writes nothing at all about the resolve. Specifically, none of this happens:

  • no member alert is closed;
  • the group row is untouched — status, resolved_at and severity all stay as they were;
  • the firing epoch is not bumped;
  • escalation keeps its snapshot, its fence token and its timer, so paging continues on its own terms;
  • no resolved or updated row is written for that delivery, and nothing is published to correlation.

What it does write is one timeline row (resolve_not_corroborated, actor_type system) and a request for the Telegram card to re-render. The delivery is then ACKed: it was handled, nothing was written, and there is nothing to retry — returning an error would only redeliver it to re-run the same check against the same contradicting signal.

Four things are deliberately dropped with that ACK

Because the delivery returns before step 5, none of the writes below it happen for that delivery: its labels and annotations are not refreshed onto the group, its updated audit row is not written, nothing is published to proxima.alerts.correlate for it, and the recount that refreshes alert_count and firing_count does not run. All four are the price of refusing the resolve before the first write rather than half-writing it and unwinding. None is retried — the delivery is ACKed, not NAKed — so the group keeps the previous delivery's labels and counts until the next delivery for this scope refreshes them.

Doing nothing was chosen over parking the group in a new status: that would need a status CHECK migration and would have to teach roughly a dozen predicates about it, three of which fail silently (the reaper never reaps such a group, escalation can never re-arm it, and the next delivery's recount rewrites it to resolved anyway).

One row per firing episode​

The three rows are written through the at-most-once fence keyed on (alert_group_id, action, firing_epoch), so a source re-sending its resolve on every evaluation leaves one row per episode, not one a minute. A re-fired episode gets its own row. The three actions carry no CHECK constraint and needed no migration.

The numbers, and where they are configured​

ConstantValueWhat it bounds
CheckTimeout5sThe whole check, all signals together — not per signal
HostAliveWithin3 minutesHow recently a host must have checked in to count as alive
MonitorFreshWithin5 minutesHow recently a monitor must have produced a result for its state to count as a reading

All three are Go constants in backend/internal/resolvecheck/verifier.go. They are not environment variables, not per-client settings and not tunable at runtime — a per-client verification window would be a knob nobody could reason about from a card that says an alert is still firing. Changing one is a code change and a deploy.

It fails open, by construction​

Verify never returns an error and never blocks past CheckTimeout, even against a reader that ignores its context. A reader error, a panic inside a reader, a recorder that panics, or the deadline all degrade to an unknown signal — never to trouble. Only positive evidence of trouble contradicts, so every failure mode of this subsystem makes a resolve more likely to be accepted, never less. A verifier that could hold a resolve open would keep paging people for something already fixed.

What the room sees​

OutcomeThe Telegram card
confirmedThe usual green 🟢 RESOLVED card, exactly as before
unverifiedThe same green card. Nothing marks it as unverified in chat — that distinction lives on the timeline in Console
contradictedThe card keeps its firing state and gains one line

The contradiction line renders directly under the scope line:

🔴 P1 · Replica lag on psql01
AcmeCorp / production / psql01
⚠️ STILL FIRING — at 03:14 UTC the source reported resolved, but node02 has not checked in for 4m12s
  • It is never shown on a resolved card. A card cannot say RESOLVED and STILL FIRING at once; an acknowledged or silenced card does show it, because those groups are still open.
  • Its reason always names a host's check-in gap or a named monitor — never a source, because the source signal is never asked.
  • It dates its measurement (at 03:14 UTC), in the card's own time zone — the one the silence end and the timeline's stamps use. The stamp comes from the refusal row, not from the moment the card was rendered, and it is there because the line is sticky: see Known limits.
  • The edit that delivers it is forced. The card syncer normally skips an edit when the tracked message already shows the group's state, and a refused resolve deliberately changes no state — so without forcing, this line would render and then be discarded. Forcing skips only that comparison, and only when the refusal row was actually written, so the room gets exactly one such edit per firing episode however often the source re-sends its resolve.
  • A card edit notifies nobody, so this is a silent correction to the message already in the room, not a new page. Paging is the notification chain, and a group whose resolve was refused simply keeps escalating on its own terms.
  • A refusal whose reason is missing or unreadable renders no line at all, rather than a warning with nothing behind it.

See Telegram Alert Card for the card's rules in full.

What Console shows​

The alert detail page's activity timeline renders all three rows, with the verifier's reason verbatim underneath each:

RowCopyTone
resolve_not_corroboratedThe source reported resolved; Console did not accept itWarning, labelled Resolve not accepted
resolve_verifiedResolved — confirmed by an independent checkNeutral
resolve_unverifiedResolved — source-reported, unverifiedNeutral

Neither accepted outcome is styled as a success: the group's own resolved row is written beside them and already carries the green, and a green resolve_unverified would contradict the sentence that created it. The refusal is a warning rather than a danger — nothing broke, the group is in exactly the state it was already in, and this is the row that explains why a group somebody expected to be closed is still on the board.

The row's metadata.signals array is deliberately not rendered. Its third entry is always the inert source signal, and printing it would invite exactly the misreading this page exists to prevent.

Metrics​

MetricLabelsWhat it counts
proxima_alert_resolve_checks_totaloutcome = confirmed | unverified | contradictedOne per delivery that would have closed an open group. Deliveries carrying a firing alert, and resolves for an already-closed group, are never checked and never counted — so this is not a count of resolve deliveries
proxima_alert_resolve_check_signal_errors_totalsignal = host | monitor | sourceReadings a check could not take: the reader errored, panicked, or the deadline was spent on it. A signal that is merely absent is not an error and is not counted. source can never occur while that signal reads nothing

Neither carries a group, client or source label; the identifiers are on the resolve check: … log lines beside each check. Both are per replica.

No alert rule ships for either, deliberately. Every failure mode of this check makes a resolve more likely to be accepted, so neither counter can indicate a page being held open. The pair is still worth reading together: a rising signal-error rate against a flat unverified rate is corroboration quietly going blind, which looks exactly like a healthy system from the outside.

# What checks are concluding.
sum by (outcome) (increase(proxima_alert_resolve_checks_total[1h]))

# "Could not ask" versus "nothing to ask" — only readable as a pair.
sum by (signal) (increase(proxima_alert_resolve_check_signal_errors_total[1h]))

Known limits​

  • A rule that genuinely stopped firing while its condition persists resolves as unverified. That is the common outcome, not an edge case: with no host on the alert there is nothing to ask, and the resolve is written. This check narrows the window in which absence-of-data closes an incident; it does not close it.
  • A contradicted resolve is retried only when the source sends another resolve. Nothing re-runs the check on a timer, and no background sweep revisits the refusal. A source that resolves once and never speaks again leaves the group open until a human or the stale-group reaper closes it — which is the intended direction of failure, but it does mean the group's ending depends on the source repeating itself.
  • The source signal answers nothing, so a source that has genuinely gone silent cannot be detected here.
  • An alert with no host is never corroborated, including every uptime-monitor alert: a probe alert cannot yet be mapped to its own monitor, and asking "is some monitor in this environment down?" would let an unrelated service refuse this alert's resolve. That coverage returns when a probe alert can be mapped to its own monitor through its grouping key — not by widening the scope of the read.
  • The contradiction line needs a card to exist. It renders on a Telegram card, and no card exists until paging is live for the alert's team — see the rollout preconditions under known limits. On a deployment with no Telegram cards the refusal is visible in Console and in the metrics only.
  • unverified is invisible in chat. The card of an unverified resolve is an ordinary green resolved card; only Console's timeline says the resolve was nobody's word but the source's.
  • The contradiction line is sticky for the whole firing episode, and its measurement ages. Because the refusal row is fenced to one per episode, the first refusal supplies the line's reason until that episode ends — through an acknowledgement, a silence, and any later refusal that would have given a different reason (a down monitor rather than a check-in gap). The measurement is never re-taken: if the host came back a minute after the refusal and the source never re-sent a resolve, the card still reads node02 has not checked in for 4m12s. This is why the line dates its measurement (at 03:14 UTC) rather than being hidden once it ages — suppressing it would need a rule for which later evidence cancels a refusal, and a card that quietly dropped the warning would be worse than one a reader can date. Console's timeline is where the full sequence of rows is read.
Not yet observed in production

This check is proven by unit tests and by tests against a real PostgreSQL. At the time of writing no production contradicted resolve has been observed, and the end-to-end path — a real source resolving an alert for a host that has stopped checking in, refused, with the line appearing on a real card — has not yet been exercised against live traffic.

  • Alerting — ingestion, the alert lifecycle, and why a resolution never creates a group.
  • Telegram Alert Card — the card the contradiction line renders on, and everything else it shows.
  • On-call Resolution Safety — the other half of the question: the gaps that stop a page reaching anyone in the first place.
  • Uptime Monitors — the probes behind the monitor signal.