Skip to main content

On-call Resolution Safety

Even a well-authored on-call setup can have gaps that prevent a page from reaching anyone: an alert matches no route, a schedule resolves to nobody, or a rotation is left empty. This page describes the four observability signals that surface these gaps — so operators can find and fix them before an incident proves the gap the hard way.

Observability only — these signals never page

All four gaps are captured as metrics + structured logs. None of them trigger a page. An unrouted alert is not the engine's job to escalate — it is a configuration problem to detect and fix. Wire the metrics below into Grafana alerts and you will find gaps ahead of a real incident.

G1 — Unrouted alert (default fallback)​

An alert that fires for a client with no matching escalation route is a silent loss: nobody is paged, no record is written, and the alert continues to fire indefinitely with no on-call response.

Console now captures every such alert with:

  • Metric: proxima_escalation_unrouted_total{severity} — incremented once per unrouted alert group. Labels the severity tier (P1–P5) so you can distinguish a missed P1 from a routine P5.
  • Log (WARN): structured, carries alert_group_id, client_id, grouping_key, and severity.

The alert itself is not paged — an unrouted alert means the routing table has no policy to send it to, so there is nowhere to page to. The intended response is to author the missing route.

# Any unrouted alert in the last 15 minutes — treat a spike as a routing gap.
increase(proxima_escalation_unrouted_total[15m]) > 0

To drill into which client is affected, filter logs by client_id:

./scripts/obs.sh logs '"escalation unrouted"' 1h

G2 — Reactive no-on-call signal​

An escalation policy can reach a NOTIFY_SCHEDULE step whose schedule resolves to nobody at that moment: the schedule may have no rotation covering the current time, or no user may be assigned to the active slot. Previously this was a silent WARN-skip; the step advanced without any signal.

Console now emits:

  • Metric: proxima_escalation_no_oncall_total{reason} — reason is one of:
    • missing_schedule — the schedule referenced by the step does not exist.
    • no_oncall — the schedule exists but resolves to nobody right now.
  • Log (ERROR): structured, carries schedule_id, team_id, alert_group_id, and client_id (plus error on a schedule-load failure) so you can tie the no-on-call event to the specific firing incident.

The step still advances after emitting the signal — the escalation does not stall. But the no-oncall event is now observable instead of silent.

# Any no-oncall resolution in the last 15 minutes — means a live page hit a gap.
increase(proxima_escalation_no_oncall_total[15m]) > 0

A no_oncall event during an active incident is high-priority: a page was attempted and nobody received it.

G3 — Schedule Auditor (proactive gap detection)​

The ScheduleAuditorWorker sweeps all live-team schedules ahead of time, looking for coverage gaps and empty rotations before they cause a missed page. It runs on a configurable interval and emits metrics and structured WARN logs for every issue found.

The auditor holds a Postgres advisory lock (one replica sweeps per tick) and is enabled automatically when the ops-bot token is configured — the same gate as the escalation timer worker.

Observability only

The auditor never pages and never pings the dead-man's-switch heartbeat. A coverage gap is a configuration problem, not an engine failure.

Metrics emitted​

MetricTypeMeaning
proxima_oncall_schedule_gap_total{schedule,team}counterA coverage gap was found in the look-ahead window for this schedule.
proxima_oncall_empty_rotation_total{schedule,team}counterA rotation in this schedule has no users assigned.
proxima_schedule_auditor_sweeps_total{result}counterSweep outcomes: clean | error.
proxima_schedule_auditor_last_success_timestamp_secondsgauge (unix s)Wall-clock of the last successful sweep. Staleness = auditor down.

The schedule and team labels carry the schedule id and team id (UUIDs), so you can identify the affected schedule directly from the metric.

Structured log fields (WARN)​

Each gap emits a structured WARN log with:

  • schedule_id — which schedule has the gap (UUID).
  • team_id — which team owns it (UUID).
  • gap_start, gap_end, leading_open, trailing_open — the uncovered window and whether it is unbounded at either edge (coverage gaps only).
  • rotation_id — the empty rotation (empty-rotation warnings only).

Query gaps in the look-ahead window:

./scripts/obs.sh logs '"schedule gap"' 24h
./scripts/obs.sh logs '"empty rotation"' 24h
# Any coverage gap found in the last sweep window.
increase(proxima_oncall_schedule_gap_total[1h]) > 0

# Any empty rotation found in the last sweep window.
increase(proxima_oncall_empty_rotation_total[1h]) > 0

# Auditor stopped running (staleness guard).
time() - proxima_schedule_auditor_last_success_timestamp_seconds > 7200

Environment variables​

All three knobs are optional — the defaults are safe for production:

VariableDefaultMeaning
PROXIMA_SCHEDULE_AUDITOR_INTERVAL1hHow often the auditor sweeps.
PROXIMA_SCHEDULE_AUDITOR_WINDOW336h (14 days)How far ahead the auditor looks for gaps.
PROXIMA_SCHEDULE_AUDITOR_GAP_TOLERANCE60sGaps shorter than this are ignored (handoff jitter / sub-minute transitions).

Querying live​

make obs-metrics PATTERN=proxima_oncall_schedule
make obs-metrics PATTERN=proxima_schedule_auditor

G4 — Schedule guards (configuration-time enforcement)​

Two guards prevent the most common misconfiguration classes at write time, before they cause a silent resolution failure.

by_day is required for daily and weekly rotations​

For frequency=daily or frequency=weekly, the by_day field is now required. This routes these rotations through the DST-correct masked resolver, which ensures handoffs stay at the correct local wall-clock time across daylight-saving transitions.

For a 24/7 rotation that covers every day, set by_day to all seven weekdays:

"by_day": ["MO", "TU", "WE", "TH", "FR", "SA", "SU"]

A backfill migration (000149) added an all-days mask to existing daily/weekly rotations that had no by_day set, so all pre-existing schedules are already compliant.

Attempting to create or update a daily/weekly rotation without by_day returns 400.

You cannot empty a rotation of an enabled schedule​

Removing all users from a rotation that belongs to an enabled schedule is now rejected with 409. An empty rotation resolves to nobody — the same state as the no_oncall signal — so the guard prevents this class of silent coverage gap from being authored.

To empty a rotation:

  1. Disable the schedule (PATCH /api/v1/oncall/schedules/{id} with "enabled": false).
  2. Remove the users from the rotation.
  3. Re-enable (or delete) the schedule as appropriate.

G5 — Dispatch delivery completeness (orphaned claim auditor)​

Every notification Console sends goes through an idempotency ledger: it claims the send slot (INSERT … status='sending'), calls the external channel (Telegram, voice), then marks the row terminal ('sent' or 'failed'). If the process crashes between the external call and the terminal write — a host OOM, a JVM kill, a kernel panic — the row is left permanently 'sending'. The next dispatch on the same key hits the ON CONFLICT DO NOTHING dedup fence and silently no-ops: the page is not re-sent, no error is logged, and the on-call engineer never hears about it. This is the orphaned dispatch claim problem.

What the auditor does​

The DispatchCompletenessAuditor (a background singleton, held on a Postgres advisory lock so exactly one replica sweeps per tick) periodically scans notification_dispatch for rows that are still status='sending' after a configurable cutoff age. For each orphan it:

  1. Marks the row terminal 'failed' — keeping it as the completeness record so the ledger is never left dangling.
  2. Emits metrics and a structured ERROR log (dispatch_id, alert_group_id, epoch, channel, target_ref, step, path, age_seconds) so the event is observable.
  3. Does not re-send.
No re-send — and why

A stuck 'sending' row is ambiguous: the crash may have happened after the external call succeeded — the Telegram message was delivered, the Twilio call went out — and marking it failed does not change that. Re-sending would double-page the on-call engineer for the same alert, which is worse than a late page.

The actual re-page guarantee lives elsewhere: the escalation ladder (P0-A) fires its REPEAT step on a live, fresh key that is not blocked by the orphan, and the notification chain worker (P0-B1) re-arms a user's next channel on a fresh step cursor. Both paths produce a new dispatch claim with a new key, so they are never blocked by an orphaned 'sending' row on the old key.

This replaces the previous ReapStuckSending reaper, which deleted orphaned rows with no alert — a silent discard that erased the evidence and produced no metric.

Metrics emitted​

MetricTypeMeaning
proxima_dispatch_orphans_total{channel,path}counterAn orphaned claim was found and marked terminal. channel is the delivery channel (telegram, voice); path identifies the ledger step origin (escalation, memory, chain, or initial).
proxima_dispatch_auditor_sweeps_total{result}counterSweep outcomes: clean (no orphans, no errors) | error (the sweep itself failed).
proxima_dispatch_auditor_last_success_timestamp_secondsgauge (unix s)Wall-clock of the last sweep that completed without error. Staleness means the auditor is down.

A rising proxima_dispatch_orphans_total rate is the key signal: it means the process is crashing between Send and Mark, and pages are losing their terminal write. The path label tells you which part of the paging stack the orphan came from.

# Any orphaned dispatch claim found in the last 15 minutes.
# A single orphan is a crash artifact. A sustained rate is a systemic delivery gap.
increase(proxima_dispatch_orphans_total[15m]) > 0

# Auditor stopped sweeping (staleness guard). Separate from findings.
time() - proxima_dispatch_auditor_last_success_timestamp_seconds > 300

Drill into which channel and path are affected:

make obs-metrics PATTERN=proxima_dispatch_orphans
make obs-metrics PATTERN=proxima_dispatch_auditor
./scripts/obs.sh logs '"orphaned dispatch claim"' 1h

Environment variables​

VariableDefaultMeaning
PROXIMA_DISPATCH_AUDITOR_INTERVAL1mHow often the auditor sweeps.
PROXIMA_DISPATCH_AUDITOR_CUTOFF5mA 'sending' row older than this is treated as an orphan. Must exceed the longest realistic Send latency.

The auditor is enabled under the same ops-bot-token gate as the escalation timer worker and notification chain worker. If Telegram paging is not configured, it is disabled.

Ack confirm-loop (after a page is acknowledged)​

The gaps above (G1–G5) all concern getting a page out to someone. This section covers the other side: what happens after a human acknowledges the page. An ack is a promise ("I've got this"), and a promise can go quiet — the acker gets pulled onto something else, goes off shift, or simply forgets. The ack confirm-loop makes an acked-but-unresolved incident that goes silent a loud, escalating signal instead of a silent assumption that "someone's on it."

This loop does page — unlike G1–G5

G1–G5 are observability-only and never page. The ack confirm-loop is different: on silence it re-escalates (re-pages the chain), and in its terminal state it keeps nagging the acker. It replaced the old silent ack_watch auto-un-ack, which quietly un-acked and re-paged the whole chain when a timer fired — never first asking the acker whether they were still on it.

What replaced the old behavior​

Previously, an ack with ack_timeout_minutes set entered an ack_watch state and, when the timer fired, the engine silently un-acked and re-paged the whole chain — it never asked the acker "still on it?" first. Worse, once the chain's repeat_count was exhausted the escalation halted with next_escalation_at = NULL: an acked, unresolved incident with nobody paged and nothing signalling it. Both are gone. An ack now converts the escalation timeline into a reminder loop that re-escalates only on silence and never ends in silence.

State machine​

An ack arms a small state machine on the escalation timer (the escalation_snapshot.mode column — no new table, no migration):

StateWhat happens
ack_reminderEvery ack arms this (including a policy whose ack_timeout_minutes is null). On the reminder cadence, a "still on it?" DM/badge appears and the state moves to ack_confirm.
ack_confirmA ⏰ Still on it? reminder has been sent; the loop waits for a confirm tap. If the acker confirms → back to ack_reminder (reminder re-armed, no re-page). If silent past the confirm window → un-ack + re-escalate (re-page the chain), bounded by repeat_count.
ack_abandonedRe-escalations spent and still unresolved. A never-silent terminal state: it keeps nagging the acker (🚨 Unresolved and unconfirmed — please resolve or hand off) and holds an alertable metric until the incident resolves. next_escalation_at is never NULL. A confirm tap from here pulls the incident back into the gentle ack_reminder loop.

Resolving the incident anywhere clears the loop.

Confirm affordances (web + Telegram parity)​

The acker confirms they're still on it from either channel — both call the same backend path (AlertService.ConfirmAck), which re-arms the reminder and rotates the escalation fence token so a concurrent stale silence-timeout can't also fire:

  • Telegram: the ✅ Still on it inline button on the reminder DM.
  • Web: POST /api/v1/alerts/{alertGroupID}/confirm (alerts:write, tenant-checked). The Alert and Incident detail pages render a "still on it?" confirm badge with a ✅ Still on it button (and a loud Abandoned banner in ack_abandoned). This is the load-bearing confirm surface.

Cadence env vars​

VariableDefaultMeaning
PROXIMA_ACK_REMINDER_DEFAULT30mReminder cadence used when the policy's ack_timeout_minutes is null (ack → first "still on it?").
PROXIMA_ACK_CONFIRM_WINDOW10mSilence window after a reminder is sent before it counts as no-response → re-escalate.
PROXIMA_ACK_ABANDONED_INTERVAL60mNag cadence while a group is in ack_abandoned (re-nudge + re-emit the metric).
ack_timeout_minutes is the per-policy reminder cadence

When a policy sets ack_timeout_minutes (5–240), that value is the reminder cadence for its groups (frozen at arm time). Its meaning shifted from "silently un-ack after N minutes" to "ping 'still on it?' after N minutes." A null value falls back to the global PROXIMA_ACK_REMINDER_DEFAULT.

Metrics to alert on​

MetricTypeMeaning
proxima_escalation_ack_abandoned_totalcounterAlertable — a live abandoned incident. Incremented on entry to ack_abandoned and on each nag tick. A nonzero rate means an acked incident has gone unresolved and unconfirmed right now.
proxima_escalation_ack_reescalations_totalcounterA silence re-escalation: the acker went quiet past the confirm window and the chain re-paged. A rising rate means acks are not being followed through.
proxima_escalation_ack_reminders_totalcounter"Still on it?" reminder DMs planned/sent.
proxima_escalation_ack_reminder_failures_totalcounterBest-effort reminder-send failures (self-healed next tick).
proxima_escalation_ack_confirms_total{source}counterConfirm taps, source ∈ {telegram, web}.
# A live abandoned incident — acked, unresolved, nobody confirmed, re-escalations spent.
# This is the loudest ack-loop signal: page on it.
increase(proxima_escalation_ack_abandoned_total[15m]) > 0

# Acks that went silent and forced a re-escalation — a page reached someone who then dropped it.
increase(proxima_escalation_ack_reescalations_total[15m]) > 0

Drill into the loop live:

make obs-metrics PATTERN=proxima_escalation_ack_
./scripts/obs.sh logs '"ack_abandoned"' 24h

Liveness — P0-A covers the loop​

The ack confirm-loop rides the same next_escalation_at timeline that the escalation dead-man's switch already watchdogs — its overdue sweep processes acknowledged rows through the shared plan body, so the ack-loop modes inherit that liveness guarantee. There is no separate ack auditor: an abandoned incident is a human-response failure (surfaced by proxima_escalation_ack_abandoned_total), while a wedged timer is an engine failure (surfaced by the dead-man's-switch heartbeat) — the two signals stay distinct.

G6 — Why nobody was paged (per-alert evidence on the timeline)​

G1–G5 are fleet-wide: they tell you that pages are going missing, not why this alert never woke anyone. That question — "why didn't I get paged?" — is the commonest one an on-call system has to answer, and it used to be unanswerable from the alert itself: a group whose page was suppressed looked exactly like a group that never fired.

Eight group-level decisions end (or, for maintenance, hold) a page before any target is even chosen. Each one now writes a row to the alert's own timeline (alert_group_log, actor_type='system'), visible on the alert detail page:

ActionWhat it provesWhat to do
not_routedThe client's enabled routes were read and none matched this alert. Nothing was armed.Author the missing route (same gap G1 counts).
l1_disabledThe client's paging feature flag — oncall_enabled — was read and is off. The reason token keeps the retired l1_agent name on purpose: historical rows in this table carry it and a partial unique index is built on the literal, so renaming it would orphan both. It has never reported the AI flag.Enable oncall_enabled, or accept that this client pages nobody.
suppressed_as_followerThe group joined an open incident as a duplicate member under active grouping and did not raise the incident's paged severity — page once per incident.Nothing: this is by design. The incident's leader was paged.
team_in_shadowThe frozen page-owner team's cutover flag says Console is not the pager for it.Expected before cutover; after cutover, check the team's escalation_live.
policy_not_entitledA route matched and its policy loaded, but the policy belongs to neither a team nor this alert's own client — paging it would wake another owner's engineers, so arming is refused.Poisoned route/policy data. Fix the route; any occurrence is an alarm (see proxima_escalation_cross_tenant_skip_total).
policy_invalidThe escalation policy behind the matched route cannot page anybody.Open the policy and fix its steps.
in_maintenanceA maintenance window covered the group when a page was about to go out (arm, re-page, re-triage, a follower member, or an escalation step). The row's at is <site>:<window id>:<occurrence start> and its metadata names the window and when it ends, so two windows covering one group write two rows. The page is held, not dropped: the release pages it at the window's end if it is still firing. Readiness-drill groups are never held.Nothing, if the work is planned. If not, end or narrow the window.
folded_into_maintenance_releaseMore than 10 held groups were still firing when one occurrence ended; one (the most severe that would actually page) paged and this one was listed on it. paging_group_id in the metadata names it. Paged individually if still firing and unacknowledged 30 minutes after that group is acknowledged, resolved or silenced — or 30 minutes after the release if that group never armed.Work the paging alert.

policy_invalid is written from three distinct sites, and the row's metadata says which one wrote it (three sites, four call sites: the plan site is reached from both the escalation timer's own tick and the dead-man's-switch auditor's recovery re-plan — the same planner, the same remedy):

  • {"at":"arm"} — the policy's steps do not decode, or name a step type this build cannot execute, so nothing was armed.
  • {"at":"dispatch"} — the frozen snapshot carries no page-owner team, which a team-owned policy always does.
  • {"at":"plan"} — the frozen snapshot no longer decodes when the timer re-plans it. This one is the loudest in consequence: the planner clears next_escalation_at, which permanently stops every future page for that group until something re-arms it (a fresh correlate pass). Before this evidence existed, that kill was completely silent.

in_maintenance uses the at metadata the same way, with the window appended: sites arm, repage, triage, follower, dispatch; folded_into_maintenance_release has site release. The metric's at label carries the site only.

What is deliberately NOT recorded

A decision that merely failed to determine something writes nothing. A client row that would not load, a route list that errored, a policy that would not load, a live-check that could not be answered, and a re-page severity test-and-set that errored all suppress the page too — but none of them proves a cause, and a reason the reader acts on had better be the real one. Those paths leave a WARN in the logs instead.

The same applies to an acknowledged or resolved group: somebody acted, the timeline already carries their acknowledged/resolved row, and calling that a paging failure would be a lie.

One more silence is deliberate but different: a build whose L1 notification dependencies are unwired records nothing, because that is a fleet-wide wiring fact, not a per-alert decision — a row saying so would be identical on every alert forever.

One row per reason, per firing episode​

An alert group is long-lived and reopens in place (a re-fire bumps its firing epoch and pages again). The evidence is fenced on (alert_group_id, action, firing_epoch, at), so:

  • a retry, a redelivered message, or an escalation timer re-planning the same group every tick leaves one row;
  • a re-fired episode that is suppressed again gets its own row, so "why didn't I get paged this time?" has an answer for the current episode rather than one stamped weeks ago; and
  • each emitting site of policy_invalid gets its own row. That last part is not symmetry: without it the first site of an episode masks every later one, and the reachable pair is dispatch → plan — so the row that would be suppressed is the one recording the permanent disarm, the loudest thing this evidence exists to surface.

policy_invalid is the only reason that names its site, which is not the same as being the only one with more than one emitter. l1_disabled is written from two places — the correlate gate and the escalation timer's dispatch-time gate — and both deliberately stamp no site, so the fence collapses them into one row per episode rather than one per emitter. That is the intended outcome: both emitters prove the identical fact (the flag was read and it is off) and send the reader to the same fix. A reason gets distinct at values only when its emitters would send the reader somewhere different.

The fence is belt and braces. A conditional insert (WHERE NOT EXISTS) answers the cases that actually occur, and migration 000187's partial unique index (uq_alert_group_log_no_page_reason) catches the one it cannot: two concurrent deciders under READ COMMITTED both observing an empty table. The index is partial to these six actions, because alert_group_log is legitimately append-only for everything else — correlated, analyzed and the ack confirm-loop actions repeat by design.

It fences; it never rejects. A unique violation is read by the writer as "already on the timeline" — the same answer the conditional insert would have given a moment later — so adding an index to a table on the paging path cannot turn a telemetry write into a failure. That is also why the action column still carries no CHECK constraint: a new reason must never need a migration, and a constraint would be one more way a paging-path write could fail. Migration 000187 extends the column's documented vocabulary instead.

The write can never cost a page​

Every one of these writers sits inside the escalation path, so the rule is that telemetry must never fail or delay a page. Errors are logged and dropped, panics are recovered, and the write runs on a context detached from the caller's cancellation with a 3-second deadline of its own. Per write that is exact: a wedged database costs an observation, never a lost page.

Per tick it is a budget rather than a guarantee: the escalation timer walks a batch of 50 rows sequentially (each row is either a planned dispatch or a permanent disarm), so a partial database wedge that leaves alert_groups healthy but stalls alert_group_log can add up to 50 x 3 s to one tick. Where that delay lands is deliberate:

  • the writes made during the dispatch loop (l1_disabled, team_in_shadow, policy_invalid at dispatch) delay the live groups planned behind the suppressed one in the same batch; while
  • the plan-time disarm evidence ({"at":"plan"}) is written after the whole dispatch loop, so it can never delay a page that is already due in that tick — only the next tick, because the ticker is serial.

The dead-man's-switch auditor writes that same {"at":"plan"} evidence for the disarms its recovery re-plan performs, and it gets its own placement and its own bound — the paragraph above describes the timer only. Its candidate list has no LIMIT (the sweep returns every overdue group, not a batch of 50), so there is no N to multiply, and its candidates are by definition groups whose page the timer already failed to deliver. Two things bound it: the writes are drained after the whole candidate loop and after the heartbeat decision, so they sit behind every recovery page in the sweep; and the drain as a whole is capped at 15 s (disarmEvidenceTickBudget), after which the remaining explanations are dropped with a WARN naming how many. What is left is a delay to the next sweep of 15 s plus one in-flight write, instead of an unbounded one. Those dropped rows are gone for good — the disarm cleared next_escalation_at, so the group is in no later sweep — which is the accepted cost of not making the recovery path slowest exactly when it is the only path left. The budget is only expected to be reachable when the evidence store is wedged, and that expectation is arithmetic about a single-row insert, not an observation: none of this has run against live traffic.

The permanent-disarm reason ({"at":"plan"}) is recorded after the planner's transaction commits, never inside it: an insert that failed inside that transaction would roll the disarm back and turn a telemetry fault into a state bug.

Metric​

MetricTypeMeaning
proxima_no_page_reason_write_failures_total{outcome}counterAn evidence write that did not land. outcome is error (the store returned an error) or panic (the store panicked and the recorder swallowed it so the page could proceed). The label is folded against those two constants, so a third value added later shows as unknown instead of forking the series.
# Alert timelines are missing their no-page reasons — the blindness is coming back.
increase(proxima_no_page_reason_write_failures_total[15m]) > 0
Not yet observed in production

This evidence is proven by unit tests, mutation testing and integration tests against a real PostgreSQL. At the time of writing no production occurrence of any of the six reasons has been observed, and the frontend renders these actions with the generic neutral timeline style (the raw action string, e.g. policy_not_entitled) rather than a human-readable label — a friendlier rendering is still owed.

G7 — Who we never even tried to reach (per-target evidence in the dispatch ledger)​

G6 explains a page that ended before any target was chosen. This one explains a page that ended after one was: the notification chain reached a real person, could not reach them on that channel, advanced, and left nothing behind. A chain that ran to completion having notified nobody was indistinguishable from one that notified everybody — including to the person it failed to reach, who could not learn any of this from their own alert.

Five per-target drops now write a terminal skipped row into the notification_dispatch ledger, carrying skip_reason:

skip_reasonWhat it provesWhat to do
no_telegramThe user record loaded and carries no usable Telegram chat id in its traits.Have them bind their Telegram account, or drop the telegram step from their chain.
no_phoneTheir on-call profile has no phone. The lookup succeeded and found none — a lookup that errored is retried, not skipped.Add the phone on /profile (it is user-global since D3).
unverifiedTheir phone is present and not verified, and this deployment runs PROXIMA_VOICE_VERIFIED_GATE=enforce.Have them verify the number, or accept that voice will not ring for them.
channel_unconfiguredThe contact resolved, but this build registered no implementation under the step's channel name.Configure the provider. A deployment with no voice provider drops the voice step of every chain this way.
page_render_emptyThe page renderer returned no error and an empty body, and the step's channel refuses to send an empty one. Only the Telegram branch checks this, because only Telegram refuses it; a render that errored is re-armed and retried, not skipped. This is a defensive guard: read zero rows as "not seen", never as "the renderer is healthy".Investigate the alert's rendered body — the person and their contact were both fine.

The row lands under the same ledger key the page itself would have used — same alert group, same firing epoch, same channel, same step — so the deliberate non-attempts sit in the same table as the deliveries. Read the caveats below before treating the result as a list of everyone who went un-paged:

SELECT channel, skip_reason, count(*)
FROM notification_dispatch
WHERE alert_group_id = $1 AND status = 'skipped'
GROUP BY 1, 2;

What a skipped row does and does not say​

  • It names the person, not the contact: target_ref carries the synthetic user:<user_id>. A skip has no destination — that is the point of the row — and for unverified persisting the number we refused to dial would be exactly the wrong thing to keep. The synthetic form also does not collide with either shape a real destination takes (a chat id is numeric, a phone is E.164), which is load-bearing: the ledger key includes target_ref, and a skip row occupying the key a later real send claims would make that send look like an already-claimed duplicate and drop the page. There is no CHECK and no validator on the column, but the separation is enforced at the one place it could be violated: Claim refuses a target_ref inside that reserved namespace outright. A future channel whose destination were itself user:<uuid>-shaped therefore fails loudly on its first send — NAK, retry, dead-letter with a log line — instead of having its pages silently dropped as duplicates.
  • unverified is silent on a default deployment. The gate ships as warn, which pages the unverified phone anyway; only enforce refuses it. Zero unverified rows therefore means "the gate is off" at least as often as it means "every phone is verified".
  • It is not a complete census of people who were not paged, for two separate reasons. First, a drop whose cause the branch cannot prove writes nothing at all — a contact lookup that errored is a transient retry, a render that errored is re-armed, and a chain step naming a channel this build has no contact lookup for (the API admits only telegram and voice, so that means hand-written data) gets a WARN and no row, because none of them tells an on-call engineer anything they would act on.
  • Second, and larger: a send that was attempted and FAILED leaves no row either. MarkFailed reclaims the claim key by deleting the ledger row, and it runs on every send error — including a permanent one, such as a Telegram bot the user has blocked or a number that is dead, which is arguably the commonest way a real person goes unpaged. The deletion is deliberate and is not going to change: keeping the row would hold the claim key and turn a retryable failure into a permanent silence, which is far worse than the reporting gap. So an absent row for a person means "delivered, or attempted and failed" — never "skipped", and never "paged". A failed row does not fill the gap either: MarkFailed never writes one, and the only rows in that status come from the completeness auditor sweeping a claim left orphaned in sending. A crash between the send and the ledger write is the headline case but not the only one — a MarkSent that merely errored, or a send that outran the sweep cutoff, strands an identical row with the process alive. And a failed row does not mean the send failed: it means delivery is ambiguous — the external call may well have succeeded — which is exactly why the auditor marks it terminal rather than re-sending and risking a double page.
  • The reason is first-write-wins, and it can go stale. Every reason computes the same ledger key for the same (group, epoch, channel, person, step), and the insert is ON CONFLICT DO NOTHING. A step whose lease expired and re-fired keeps the first attempt's reason even if the second attempt resolved differently — the operator added the phone in between, or the voice provider came up. The row says why that attempt was dropped, not why the person ended up un-paged. (This is the deliberate opposite of the call_record.not_connected_reason column, which is last-write-wins because it describes repeated origination attempts at one call, where a stale confident verdict would hide a worsening trunk.)
  • Because of that, a skipped row and a sent row can coexist for the same person and the same step: the retry that succeeded claims the real destination as its target_ref, so it does not conflict with the skip's synthetic one. Answer "was this person paged?" from the presence of the sent row, never from the presence of a skipped one.

skipped is terminal​

The dispatch completeness auditor sweeps this table for claims stuck in 'sending' and marks them terminally failed. A skipped row is written once, by one INSERT, and never transitions; the auditor's predicate selects only status = 'sending', so it never sees one.

The write can never cost a page​

The recorder sits inside the notification path, so the same rule applies as for G6, with one addition specific to a chain: it runs after the chain has already advanced. The advance is what releases that user's run to their next channel — the one that might still reach them — so a wedged ledger can never sit in front of the escalation the skip itself makes necessary. Beyond that, errors are logged and dropped, panics are recovered, and the write runs on a context detached from the worker's cancellation with a 3-second deadline of its own. The insert is at-most-once (ON CONFLICT DO NOTHING), so a re-fired lease or a redelivered claim leaves one row rather than a pile.

MetricTypeMeaning
proxima_dispatch_skip_write_failures_total{outcome}counterA per-target skip write that did not land. outcome is error or panic, folded the same way. Separate from the group-level proxima_no_page_reason_write_failures_total because the two writers fail for different reasons and are fixed in different places.
# Skipped targets are going unrecorded — the per-target blindness is coming back.
increase(proxima_dispatch_skip_write_failures_total[15m]) > 0
Not yet observed in production

Proven by unit tests, mutation testing and integration tests against a real PostgreSQL. At the time of writing no production skipped row has been observed, and none of this surfaces in the UI yet — the rows are readable by SQL only.

Rolling back migration 000186 destroys this evidence

The down migration deletes every skipped row: the status CHECK cannot narrow back to sending|sent|failed while rows hold a value it forbids. Rolling back therefore erases part of the answer to "why didn't I get paged?" for every alert already recorded. Copy the rows first if that answer still matters.

Summary: signal matrix​

GapWhen it firesMetric(s)Log levelAction
G1 Unrouted alertAlert fires with no matching routeproxima_escalation_unrouted_total{severity}WARNAuthor the missing escalation route
G2 No on-callNOTIFY_SCHEDULE step resolves to nobodyproxima_escalation_no_oncall_total{reason}ERRORFix the schedule or add an override
G3 Schedule gapAuditor finds uncovered window ahead of timeproxima_oncall_schedule_gap_total{schedule,team}WARNAdd coverage (rotation or override)
G3 Empty rotationAuditor finds a rotation with no usersproxima_oncall_empty_rotation_total{schedule,team}WARNAdd users to the rotation
G5 Orphaned dispatch claimClaim stuck 'sending' after crash between Send and Markproxima_dispatch_orphans_total{channel,path}ERRORInvestigate process-crash source; rising rate = delivery gap
Ack abandonedAcked incident goes silent past all re-escalations, still unresolvedproxima_escalation_ack_abandoned_totalERRORRe-claim (✅ Still on it) or resolve — a live abandoned incident
G6 Nobody was pagedA group-level decision ends a page before any target is chosentimeline row on the alert (not_routed, l1_disabled, suppressed_as_follower, team_in_shadow, policy_not_entitled, policy_invalid, in_maintenance, folded_into_maintenance_release) + proxima_no_page_reason_write_failures_total{outcome} for the writer itselfINFO/WARNOpen the alert's timeline: it names the cause and the fix
G7 A target was never attemptedA chain step drops one person on one channel (no Telegram id, no phone, unverified under enforce, no channel implementation in this build, or an empty rendered page body)notification_dispatch row with status='skipped' + skip_reason; proxima_dispatch_skip_write_failures_total{outcome} for the writer itselfWARNFix that person's contact, or configure the missing channel provider
G8 The call connected but the page did not informA voice page plays the generic prompt instead of the alert, never connects at all, or is acknowledged by button rather than by keypadcall_record columns tts_result / tts_fallback_reason / not_connected_reason / ack_source; proxima_voice_tts_render_total{result,reason} and proxima_voice_not_connected_total{reason,purpose}; proxima_call_record_evidence_write_failures_total{stamp,outcome} for the writers themselvesWARNSee Voice Paging: a rising fallback rate means pages are landing and conveying nothing, and trunk_rejected is an infrastructure incident rather than an on-call one
  • On-call & Escalation Management — schedules, rotations, policies, routes, covering teams, cutover.
  • Paging Control — how routes decide whether and whom to page.
  • Escalation Dead-Man's Switch — guard against a wedged escalation engine.
  • Notification Chains — per-user paging channel order and fallback.
  • Resolution Corroboration — the other direction: what Console does before it accepts a source's claim that an alert has ended.
  • Voice Paging — the third piece of "why didn't I get paged?": what the call_record evidence columns say once a target was chosen and dialled (what the callee actually heard, when they answered, how the acknowledgement arrived, and why a call never connected).