On-Call Readiness Drill
A paging engine can be perfectly healthy and still fail the one thing that matters: the human on the other end. The escalation dead-man's switch guards the engine — it pages you when the escalation worker stops advancing. The readiness drill guards the human — it periodically proves that whoever is on-call right now can actually be reached, before a real incident finds out they can't.
It is a human dead-man's switch: a background worker randomly voice-calls the current
on-call ("press 1 to confirm you're available") and sends a Telegram nudge if the call is
missed.
Everything below about the coverage-gap arm — the synthetic P1 alert group, arming it
through the team's coverage-role policy, PROXIMA_ONCALL_READINESS_COVERAGE_ROLE, and
the proxima_oncall_readiness_coverage_* metrics — describes a path that no longer runs.
EnableCoverageGap is not called in cmd/server/workers.go, the raiseCoverageGap seam is
nil, and escalatePass returns immediately.
What a missed drill does now: places the call, detects the miss, sends the one Telegram nudge DM, and records the miss on the readiness matrix. That is all. Nothing is created in the alerts list and no escalation chain is armed.
Why it was removed. The drill's window originally shaped only the deterministic target
instant while the due-test was now >= target, so an on-call who came on duty later in the
UTC day was called the moment they took the rotation — three engineers were phoned at 00:00
local. That is fixed (the window now gates the call), but while it was live the midnight calls
went unanswered, two consecutive misses is the coverage-gap trigger, and the bug manufactured
P1 pages about the pager itself. The drill is worth keeping; the synthetic page it could raise
is not.
The code and this documentation are kept because the arm is one wiring line away from
returning, and cmd/server/coverage_gap_unwired_test.go fails if that line comes back
by accident. Re-enabling it is a deliberate act that deletes that test. It also takes the
alert-group episode store (EnableCoverageGap(esc, groups, store.NewAlertGroupEpisodeStore(db), role)): the arm marks its synthetic episode decided so the correlation watchdog leaves it alone
— and the watchdog also skips readiness_drill groups outright — because a re-correlated drill
group would be planned onto the client's real escalation routes.
Historically the drill also pursued an unreachable on-call past the grace window: it paged
through the team's escalation policy for the coverage role, so a coverage gap reached a
live human before an incident did. It did not page the whole team: notify_team is not an
authorable step, and the team learned from the alert card the synthetic group posted. See
Authoring the coverage policy for who was on that chain and
why the first hop was a DM rather than a call.
The drill ships disabled. The worker is constructed only when the ops-bot token is set
(PROXIMA_OPS_TELEGRAM_BOT_TOKEN) and the master switch is on
(PROXIMA_ONCALL_READINESS_DRILL=on). With the switch off — the default — no worker
starts, no call is placed, and nothing pages. It is inert to production until you opt in.
The staged flow
Every drill walks a small state machine — pending → confirmed | missed → escalated —
that escalates only as far as it has to. A confirmation at any stage ends the drill.
- Random voice drill (call + one retry). The worker resolves who is on-call now and,
at a deterministic-random instant inside the allowed window — and only while the
clock is still inside that window (see the window is a gate, not just a
target) — places a
purpose=readinessvoice call over the same shared voice provider the alert path uses. The call plays the readiness prompt and waits for the operator to press1. There is one retry if the first call doesn't connect.- Press
1→ the drill is marked confirmed. Done. (Confirming availability is not an alert-ack — it needs noalerts:write; the acting user is the server-resolvedcall_record.contact_id, never a body claim.)
- Press
- Miss → one Telegram nudge + a grace window. If the call goes unconfirmed within the call + retry window, the drill is marked missed and the worker sends exactly one Telegram DM — "On-call readiness check — tap ✅ I'm available" — with an inline ✅ button. The button is token-fenced (like the P0-D ack reminder) and confirms only the tapping user's own drill. Tapping it → confirmed. Then a grace window (default 10 minutes) elapses.
- Still silent → the drill concludes
escalated. A page waits for a SECOND silence. If the grace window expires with no confirm — no keypress, no ✅ — the drill is stamped escalated: the on-call is marked readiness-degraded (so real alerts reroute past them immediately) and the Readiness page raises a finding naming them. Nobody's phone rings. Only when the next drill for the same person also fails to reach them does the worker arm a syntheticsource_type=readiness_drillfiring alert and drive it through the team's escalation policy whoserolematchesPROXIMA_ONCALL_READINESS_COVERAGE_ROLE. From there it is a normal page: it can be acknowledged, it covers up the chain, or it reaches the policy's loud no-coverage terminus — the whole team now knows this on-call is unreachable. See Why the first miss pages nobody.
role pages nobody, silentlyThe lookup is PoliciesForTeamRole(team, role) — it matches the role column, not the
policy's name. A team with a policy named …-critical but with role empty resolves
nothing, and the worker takes its documented Nil-skip path: the drill is stamped escalated,
proxima_oncall_readiness_coverage_skipped_total{reason=no_policy} is incremented, a WARN is
logged, and no page is armed. The drill row then carries escalated_at with a NULL
coverage_group_id, and it will not retry — that is deliberate, since a config gap must not
spin.
This was live in production until 2026-09-26: all eight policies had an empty role, so three
escalations in one week detected an unreachable on-call correctly and told nobody. Check it
with SELECT name, role FROM escalation_policies before trusting the chain, and never flip
the env var to a new role before every team has a policy carrying it — the flip re-creates
this exact hole for any team that lacks one.
Call-less confirms (a drill that never needed to fire). The worker skips the call and
records the bucket confirmed without ringing anyone if the on-call has already
proven reachable in this window — they confirmed a previous drill, or they acked a real
page since the window started (an on-call who just handled a live incident is
demonstrably available). The Readiness panel counts these as confirmed and says no call
placed in the day's detail, so one is never mistaken for a press 1.
Skipped drills (a drill that could not fire). When the on-call has no paging phone
there is nothing to call, and nothing will change inside the bucket — so the worker records
a skipped drill with skip_reason = no_phone, the same reason the notification chain
records when it skips a voice step. The strip shows it as an amber gap with the reason on
hover. Two rules govern it:
- A skip is not "cannot be paged". A missing number breaks the voice leg only; a real incident still reaches that person on Telegram, exactly as the chain treats it.
- A skip never sets or clears readiness-degraded. It proves nothing about reachability either way, so it is passed over when deciding whether someone is degraded. Otherwise, the day after a missed drill, a no-phone skip would become the "latest" drill and silently end a live reroute.
Failures that may resolve on their own — the phone lookup errors, no voice provider is
connected, the call record cannot be created, or the originate fails — are not recorded
as drills: a recorded drill consumes the day's bucket and would cancel the retry. They are
counted on proxima_oncall_readiness_fire_failures_total{reason} and retried on the next
tick. So a day that ends with no drill at all means nobody was on call, drills were off, or
every retry failed — that counter and the worker's logs say which.
Blast-radius bound. At most one drill per on-call per 24-hour bucket (a unique DB index on the bucket key), and the coverage-gap alert is deduped per bucket. A wedged or looping worker cannot storm an on-call with calls or the team with pages.
Why the first miss pages nobody
A single missed drill and two in a row are different facts about a rotation, and only the second one is an unmanned rotation. Someone who steps away from their phone for ten minutes has missed a drill; the rotation is still theirs and still answerable ten minutes later. Someone who misses two consecutive drills is not answering at all.
The first miss is not ignored — it is acted on without waking anyone:
- the on-call is marked readiness-degraded, so a real alert arriving now reroutes past them before it ever rings their phone (that is the protection people actually need, and it is automatic and immediate);
- the Readiness page raises a degradedEngineer finding naming them, which clears itself on their next confirm;
- the drill row concludes as
escalated, so it drops out of the escalation sweep and never retry-spins.
What the first miss does not do is ring a second engineer's phone at 04:00 about a colleague who was in the shower. That is how a drill becomes the thing everybody mutes — and a muted drill protects nobody.
"Consecutive" is counted in drills, not in hours. The gate reads the newest concluded drill
for the same person strictly before this one, and arms only if that one also failed to reach them
(missed or escalated). If someone misses today and is not on call tomorrow, the rotation is
not unmanned tomorrow, so nothing pages; the ladder resumes the next time a drill finds them
silent. A confirmed prior drill breaks the run even if it was a late off-phone confirm — they
demonstrably responded. A skipped prior drill is invisible to this read: it placed no call,
so it is evidence about our configuration (no paging phone on file) and about nothing else.
Counting it as a reach would break a genuine run of misses; counting it as a miss would page a
covering team because somebody's number is missing.
If the prior-drill read errors, the ladder arms — a WARN is logged and the page goes out. An
unreadable history must never become the reason a genuinely unmanned rotation goes unpaged. The
cost of failing open is an occasional page on a first miss, which is exactly the behaviour that
shipped before this gate existed; the cost of failing closed would be silence on the one path the
whole drill exists to catch. Withheld first-miss pages are counted on
proxima_oncall_readiness_first_miss_not_paged_total — that counter rising while
_coverage_gap_total stays flat is the gate working, not a fault.
Authoring the coverage policy
A coverage gap and an incident are different failures, so they deserve different chains. An incident means something is broken, get the responders. A coverage gap means nobody is holding the pager — an ownership problem, whose audience is whoever can reassign a rotation.
role is a CLOSED enum, enforced in Go (validPolicyRoles) and by a CHECK constraint, with a
unique index permitting at most one policy per (owner, role). Since migration 000241 it
admits three values: critical and default are the covering-team materialization targets (P1
routes and everything else), and coverage is this chain — never a materialization target.
Give the gap its own policy and point the env var at it:
kind: escalation_policy
version: v1alpha1
metadata:
name: team-1-coverage
spec:
owner_type: team
owner_id: TEAM 1
role: coverage
repeat_count: 2
ack_timeout_minutes: 15
steps:
- type: notify_users # whoever owns this rotation
user_ids: [[email protected]]
important: false # Telegram first; the phone rings 5 min later
- type: wait
wait_seconds: 1800
- type: notify_users # head of engineering
user_ids: [head-of-[email protected]]
important: true # Telegram + voice, no grace
- type: repeat
Then PROXIMA_ONCALL_READINESS_COVERAGE_ROLE=coverage, and create all the policies first —
the danger block above explains what a missing one costs.
Why not reuse critical or default. Both open with notify_schedule, which resolves
"whoever is on call now" — the very phone the drill just proved unreachable. Under critical,
where every step is important: true, that step places an immediate voice call to a dead number
and burns five minutes before reaching anyone who could reassign the rotation. default is
gentler (non-important, so a DM with the call five minutes behind it) but it is the
default-severity incident chain, so coverage and incident paging could never be tuned apart and
anyone editing a -default policy would have to know it serves two jobs. A dedicated role drops
notify_schedule altogether and keeps the two independent.
Why the important flag decides the volume. It is the only proportionality lever
available — there is no working-hours gate (StepNotifyIfTime has no case in the planner and
is not authorable). A non-important step's fallback chain is Telegram → wait 5m → voice; an
important one is Telegram → voice back to back. A missed drill means nothing is on fire,
and the soft-absence reroute already widens real incidents for
a degraded on-call — so the gap needs to be seen and fixed, not to wake anyone. A DM that
goes unread for five minutes and then rings the phone is the right proportion for the first
hop; important: true earns its place further down the chain, once time has passed with no
human responding. (The fallback chain applies only to a user who authored no notification chain
of their own; a user with one gets theirs.)
Escalate on persistence, not on the first hop. One unanswered drill is a phone on silent. Three days running, or nobody on the team reachable at all, is a management problem. Routing the first hop to an executive trains them to ignore it, and a 2am page for an event where nothing is broken destroys the signal's credibility within a month.
Why there is no peer step and no chat step. escalation.SupportedStepType allows exactly
five step types — wait, notify_users, notify_users_queue, notify_schedule, repeat.
notify_team and notify_channel are refused on write with ErrUnsupportedStep by every
path, the API and pc apply alike, even though the planner has cases for them. A static
notify_users list cannot express "the team except whoever is on call today", because the
on-call rotates. Peers learn from the alert card instead: the coverage gap is an ordinary
firing alert group, so it posts an internal card to the team's bound internal-audience chats
through the normal notification fan-out, with no policy step involved. That visibility depends
on oncall_enabled being on for the team's representative client.
A one-person team has no coverage policy that can work. If the team's only member is the engineer who missed, every step naming a team member pages them about themselves. Such a team needs a second member, or a coverage policy naming people outside it.
The window is a gate, not just a target
The window does two things, and for a year it only did one. It picks where in the day the deterministic target lands — and it must also gate whether a call may be placed at all right now. Without the second half, the due-test is "now ≥ target", which reads as due from the target until the end of the UTC day.
For an on-call who was already on duty at UTC midnight those are indistinguishable: their target is inside the window and the worker fires within a tick of it. They are not indistinguishable for an on-call who appears later in the UTC day.
Every Proxima rotation hands over at 00:00 Asia/Tashkent = 19:00 UTC. At that instant each
team's on-call changes, which mints a fresh bucket key (<user>:<schedule>:<utc-day>) for the
UTC day already in progress — and that bucket's target was computed inside the window,
already hours in the past. So the drill was due the moment the new person came on duty. On
2026-10-05 three engineers were voice-called at 00:00:28 local, within 60 ms of each
other, one per team. Over the preceding 14 days it fired 3 times, affecting 6 engineers
(3 on 10-05, 2 on 09-28, 1 on 09-25) — always at the local-midnight handover, but not every
night; what suppresses it on the other nights has not been established.
The due-test now requires now to be inside the window as well. The cost is deliberate: a
drill whose target passes while the worker is down, or whose on-call appears after the window
closes, is skipped rather than fired late. For a pager that is the right direction — a
missed drill costs nothing a human notices, a 00:00 call costs exactly the trust the drill
exists to measure.
Nobody goes undrilled for long. The midnight-handover on-call gets the next UTC day's
bucket, whose target lands inside the window — 10:00–18:00 local for the production
05:00-13:00 band — so they are drilled during their own working hours instead. One drill per
shift, in daylight.
The escalation is deliberately NOT gated. missPass/escalatePass act only on drills that
already exist, so gating creation bounds them too: the tail past the window's end is at most the
call window plus the grace (5m + 10m by default, so ≤ 18:15 local on the production band).
Gating it would mean the system discovers an unreachable on-call and says nothing, which is the
one thing this subsystem must never do.
A full-day window contains every instant, so a deployment that configures no window keeps the any-hour behaviour exactly. The gate can only ever subtract hours an operator has already said they do not want calls in.
Configuration
All knobs are environment variables (Twelve-Factor; defaults in code). Every one is inert
unless the ops-bot token is set and PROXIMA_ONCALL_READINESS_DRILL=on.
| Variable | Default | Meaning |
|---|---|---|
PROXIMA_ONCALL_READINESS_DRILL | off | The master switch. on enables the drill; any other value (incl. the default) keeps it disabled. |
PROXIMA_ONCALL_READINESS_INTERVAL | 1m | How often the worker sweeps (resolves who's on-call, checks whether a drill is due). |
PROXIMA_ONCALL_READINESS_WINDOW | "" (any hour) | An optional HH:MM-HH:MM (24-hour clock) window the drill's target instant is clamped into and outside which no drill is placed at all (both halves matter — see the window is a gate). The band is UTC — offsets from UTC midnight, not the server's local zone — because the drill's 24-hour bucket is UTC-aligned. Convert from your team's local time first: e.g. to keep drills within 08:00–22:00 in UTC+5 (Tashkent), set 03:00-17:00 (the UTC equivalent). Empty means any hour. A malformed value degrades to any-hour with a WARN rather than crashing startup. |
PROXIMA_ONCALL_READINESS_GRACE | 10m | The nudge→escalate grace window: after a missed call and the Telegram nudge, how long the on-call has to confirm before the team is paged. |
PROXIMA_ONCALL_READINESS_COVERAGE_ROLE | critical | Inert (arm unwired 2026-10-05). Formerly: which team escalation-policy role the coverage-gap page was armed through. |
The "is a drill due now?" decision is a pure function of (now, user, schedule, existing drills, window) using a stable hash — there is no live randomness — so the
target instant is deterministic and the same each sweep until the drill fires.
Observability
The drill exposes ten proxima_oncall_readiness_* metrics on /metrics that complete
the voice picture alongside proxima_voice_*: the voice metrics say whether the pager
can place a call; these say whether the human actually answered. The full table and the
two recommended alerts live in the observability standard —
docs/standards/observability.md
(On-call readiness-drill metrics). In short:
-
Alert on any increase in— do not build this monitor. With the arm unwired this counter can never move, so the alert would be permanently green and would read as "coverage is fine" rather than "nothing is checking". Watchproxima_oncall_readiness_coverage_gap_totalproxima_oncall_readiness_missed_totalinstead: a miss is now the last automated step, so it is the only signal that an on-call was unreachable. -
Alert on any increase in
proxima_oncall_readiness_coverage_skipped_total— a team's on-call config has a gap (no representative client, or no policy carrying the coverage role), so a drill wanted to page a coverage gap but couldn't. A silent hole in the safety net — fix the team's client assignment, or give it a policy whoserolematches PROXIMA_ONCALL_READINESS_COVERAGE_ROLE. -
proxima_oncall_readiness_first_miss_not_paged_totalis informational, not an alert. It counts drills that concludedescalatedwhile the coverage ladder was withheld because this was the on-call's first consecutive miss (why). It is the normal path for an isolated miss; it is its own counter rather than areasonon_coverage_skipped_totalprecisely because that one is alerted on — a deliberate withhold is not a config gap.
_fire_failures_total{reason} counts drills the worker could not place and will retry — a
sustained rate means drills are silently not happening, because these leave no drill row.
The _confirmed_total / _missed_total / _escalated outcome split and the
_sweep_last_success_timestamp liveness gauge round out the funnel. No PII is ever a label
or a log field — only user_id / schedule_id / team_id / drill_id and the bounded
outcome/reason labels appear; never the phone number or any keypress.
Ops prerequisite: provision the readiness prompt
The drill call plays a sound:proxima-readiness-prompt prompt (the "press 1 to confirm
you're available" audio). Like the phone-verification prompts,
this media must be provisioned into Asterisk's sounds directory before a live drill
can speak.
Provision it exactly the way the P2-C verify prompts were — copy the recorded audio into
the named-volume sounds directory with docker cp. Use the .alaw and .ulaw
codec-native formats, not .wav: format_wav.so is not loaded on the production
Asterisk box, so a .wav file will not play — Asterisk resolves sound:proxima-readiness-prompt
against the codec-native variants.
# From wherever the recorded prompt lives (both formats), into the Asterisk sounds volume:
docker cp proxima-readiness-prompt.alaw proxima-asterisk:/var/lib/asterisk/sounds/en/
docker cp proxima-readiness-prompt.ulaw proxima-asterisk:/var/lib/asterisk/sounds/en/
If the prompt is missing, the call still originates but has no audio to play, so a valid
on-call would be unable to hear the "press 1" instruction — provision it as part of
enabling the drill, alongside the verify prompts (proxima-verify-prompt,
proxima-verify-ok) and the alert/ack prompts.
Rollback note: purge readiness call records first
The migration that introduced the drill widened the call_record.purpose CHECK constraint
to allow 'readiness' (alongside 'alert' and 'verification'). Its down-migration
reverts that CHECK to ('alert','verification') — which will fail if any
call_record row has purpose='readiness'.
So in an environment where drills have already fired, a rollback must purge those rows first:
-- Run BEFORE applying the down-migration, or the constraint re-creation fails.
DELETE FROM call_record WHERE purpose = 'readiness';
In an environment where the drill was never enabled (the default), no readiness rows exist and the down-migration applies cleanly with no purge needed.
Soft-absence reroute (Phase 2)
Phase 1 stops at detection — it proves the on-call is unreachable and pages the team about the coverage gap. But that leaves a hole: a real incident that fires while the primary on-call is demonstrably unreachable would still be routed to that unreachable operator (and only reach anyone else on the normal escalation timer, minutes later).
Soft-absence reroute (P2-D Phase 2) closes it. When a real incident resolves to an on-call who is readiness-degraded, the fire path also pages the next escalation tier (typically the team lead) — at the first step, immediately, alongside the primary. The degraded engineer is not skipped; a second, reachable human is simply brought in at once instead of after the escalation ladder times out.
An on-call is readiness-degraded when their latest drill concluded they're unreachable
(escalated) — a skipped drill concludes nothing and is passed over — and no reachability signal has cleared it since — see
the auto-clear and manual clear below.
Safety guarantees
The reroute is deliberately conservative — it can only ever page more people, never fewer:
- Never drops the primary. The rostered on-call is always paged first, exactly as today. The widen is strictly additive — it adds the next tier on top; it never reroutes away from the primary or replaces them. If the primary is reachable after all, they still get the page.
- Fail-open in every branch. The degraded-check, the escalation-snapshot read, and the "is there a next tier?" lookup each fail open: if any of them errors or comes up empty, the failure is logged and swallowed and the incident is still paged and still escalates normally. A bug in the reroute path can only ever fail to widen — it can never mis-route, drop, or delay a real incident.
- OFF is byte-for-byte the old path. With
PROXIMA_ONCALL_REROUTE=off(the default), the degraded-check is never even consulted — the fire path is identical to the pre-Phase-2 behaviour. Inert to production until you opt in. - Double-page-safe. The next tier is armed with that tier's own importance flag, so when the escalation timer later advances to that step naturally, the paging engine's idempotency key coalesces the two rather than paging the tier twice.
Phase 2 rides on top of Phase 1. The reroute is enabled by PROXIMA_ONCALL_REROUTE=on
(default off), and it is doubly inert unless the Phase-1 drill is also on
(PROXIMA_ONCALL_READINESS_DRILL=on) — with no drills running, nobody is ever marked
degraded, so nothing widens.
Clearing the degraded state
Being marked degraded is not permanent — the moment the engineer proves reachable again, the reroute stops widening their incidents. There are two ways it clears:
- Auto-clear (the engineer proves reachable). Any reachability signal from the engineer
clears it automatically — a later drill confirm (press
1or the Telegram ✅), or an ack of a real page. An on-call who just acknowledged a live incident is demonstrably at their post, so the next incident routes to them normally with no widen. - Manual "Engineer is back." A team manager (or the engineer themselves) can clear the
degraded state immediately from that engineer's readiness page
(On-Call → Readiness → their name) — an Engineer is back button, backed by
POST /api/v1/oncall/readiness/{userID}/clear-degraded. This is useful when the engineer is reachable again but no drill or real page has happened to auto-clear it yet. The endpoint is self-or-team-manager gated and authorizes the exact drill it mutates (you can only clear a degraded state on a team you can write), so it can't be used to reach across tenants.
Observability
Two counters, both pairing with the Phase-1 proxima_oncall_readiness_* metrics:
proxima_oncall_reroute_widened_total{team}— a real incident was widened to the next tier. This is the meaningful signal. A sustained non-zero rate (especially on one team) = a chronically-unreachable on-call — a rota problem to fix, not a paging one.proxima_oncall_reroute_skipped_total{reason=no_next_tier}— the on-call was degraded but the policy had no next tier to widen to (the incident is still paged; this makes a degraded-but-un-widenable page visible).
The full table and the alert live in the observability standard —
docs/standards/observability.md
(On-call soft-absence reroute metrics).
Configuration
| Variable | Default | Meaning |
|---|---|---|
PROXIMA_ONCALL_REROUTE | off | The Phase-2 master switch. on widens a real incident to the next tier when the resolved on-call is readiness-degraded; any other value (incl. the default) keeps the fire path byte-for-byte identical to pre-Phase-2. Doubly inert unless PROXIMA_ONCALL_READINESS_DRILL=on too. |
The Readiness page: "can we be paged?"
/oncall/readiness in the console is not a drill log. It answers a single operational
question — can we be paged right now, and if not, what do I fix? — and it renders exactly
three things:
| Part | What it answers | Where the data comes from |
|---|---|---|
| One card per team | Who does a schedule step reach right now, until when (the end of that person's UNBROKEN stretch — see When "until" is not the next hand-off), and what does the drill record say about everyone rostered in the checked window? | GET /api/v1/oncall/now (per team) + GET /api/v1/oncall/schedules/{id}/calendar (per schedule, for the resolved roster and the gaps) + GET /api/v1/oncall/readiness |
| The drill record | How is every engineer doing on drills over a window you choose — one row each, one square per day | GET /api/v1/oncall/readiness?from=&to= — see The fleet drill record |
| What to fix | One list, ordered by how completely each item stops a page from landing | the three reads above plus GET /api/v1/oncall/readiness/coverage |
The drill record sits above the fix list: it is the evidence, and the fix list is the conclusion drawn from it.
The order of the fix list is the point. A client with no matchable route outranks three engineers missing a second channel, because the first means nobody is paged at all and the second means one leg of a chain is shorter. Nothing in the list is hardcoded: every item is produced only when the data says it is true, and an empty list means "nothing we can see is wrong" — which is what the caveat below exists to qualify.
Each part renders only from its own successful reads. A failed, forbidden or still-loading query degrades to not known here — never an error box, and never silence — so a caller who can read four teams gets four answers rather than one answer and four red boxes. The per-team and per-schedule reads have no fleet endpoint, so they are capped (12 teams, 8 schedules) and the page states how many it actually checked.
Three things are deliberately not on this page. The fourteen-day calendar strip (gap
detection stays and feeds a finding; the schedule page draws the calendar itself), the
read-only drill configuration (it lives at /oncall/settings with the rest of the
environment variables), and the per-engineer drill strip — which became its own page.
The standing caveat is an affordance, not a paragraph. The hedge that qualifies every verdict — this is what we can see, not a delivery guarantee — is reached from an ⓘ control on the verdict chip instead of being printed under all four states. The rule is unchanged and is not negotiable: a clear state is never green and never claims a page will land (Reading the verdict column). What changed is that a sentence rendered identically in every state is furniture, so the page's vertical space goes to data instead.
The fleet drill record
One row per engineer, one square per day, over a window you choose. It answers the question none of the older views did: not how is this person doing but how is the fleet doing — a person who never answers is a broken row, legible before a word is read.
What a square means
Six states, each carried by a glyph and a fill, never by colour alone, and each square
carries a full sentence as its title and aria-label, so there is no legend to consult
and the grid survives a greyscale screenshot:
| State | Glyph | What it says | Counted in answered? |
|---|---|---|---|
confirmed | ✓ | The on-call answered the call — the voice leg is proven working | yes |
confirmed_late | ~ | Confirmed, but the call went unanswered — the ✅ tap on the nudge DM arrived instead | no |
missed | ✗ | A call was placed and went unconfirmed within the call + retry window | yes |
escalated | ! | The grace window expired with no confirm. Incidents reroute past them; whether a covering engineer was also paged depends on whether this was their second consecutive miss — the square's sentence says which | yes |
skipped | – | A drill was due and no call was placed — today only no_phone, no number on file | no |
pending | … | In flight. Nothing has concluded yet | no |
none | (empty, dashed) | No drill that day. Almost always "not on call" | no |
The row's right-hand side reads confirmed / answerable as a fraction, the last confirm as
relative time, and the median time-to-confirm. The engineer on call right now is marked with
a left edge and the words on call now — worst-first sorting can otherwise bury the person
currently holding the pager several rows down.
Columns are UTC calendar days, because that is the unit a drill is bucketed in: bucketIndex
floors unix seconds by 86400, so "one drill per on-call per 24h" is a guarantee about UTC days.
Bucketing by the viewer's local day breaks the 1:1 correspondence between a column and a drill —
at UTC+5 a drill fired at 21:09Z falls into the next local day, so its real column renders empty
while its neighbour holds two. That was observed: an engineer with a confirmed drill on both 25
and 26 Sep UTC showed a blank 26th in Tashkent. A column that cannot correspond to a drill is the
wrong column.
Every drill timestamp on these pages is UTC too, and says so. The columns were made UTC on
2026-09-27 while the per-drill timestamps stayed on the viewer's clock, which put two units on one
screen: a drill fired 2026-09-27T19:00Z reads as "Sep 28, 12:00 AM" in Tashkent while its square
sits under the 27th. On 2026-09-28 a reader went looking for that confirmed drill under the 28th
and found an empty square — the one thing an empty square must never mean. Drill instants now
render as Sep 27, 2026, 19:00 UTC, with the suffix in the string rather than a footnote
elsewhere on the page.
The rule stops at the drill record. Schedule instants — when a shift ends, when cover resumes, how long a gap runs — stay in the viewer's clock, because those are moments in a human's day and someone asking "am I on call tonight?" means their own tonight.
Why skipped is excluded from the fraction. A skipped drill placed no call, so nobody
ignored anything — the voice leg had no number to ring. Counting it as answerable would
report an engineer with no paging phone as one who ignores pages, and those two findings
need opposite fixes: the first is add a number, the second is this person is not
reachable. For the same reason the square is drawn neutral, never in a failure colour.
Why confirmed_late is excluded. A nudge DM is only ever sent after a call is missed,
so a confirmed drill carrying a nudged_at timestamp is one whose call rang out. The person
is demonstrably alive; their phone leg demonstrably did not work. Production produced exactly
this on 2026-09-25 — a drill fired at 01:49, nudged at 01:50, escalated at 02:00, and
confirmed at 08:05, six hours and sixteen minutes later, by a tap, on a call nobody
picked up. It was recorded as a plain confirmed, indistinguishable from the four engineers
who pressed 1 within twenty-odd seconds. Had it been a real page, nobody was coming.
That is what this state prevents: a drill whose entire purpose is proving the phone works must not pass on evidence that the phone failed. The square sits in the warning family rather than the success one, and the row still shows the confirm under last confirmed — the honest pair being "last confirmed Thursday" beside "0/1", which reads as they responded, and not by answering the phone.
Why pending is excluded. An in-flight drill accuses nobody. Folding it into either
confirmed or missed would make the grid claim an outcome the data does not yet support —
the same error skipped exists to prevent — and the fraction would move on its own as the
drill concluded.
Two drills on one day (an engineer on call for two teams) collapse to the loudest
concluded outcome, ordered escalated > missed > confirmed_late > skipped > confirmed, with pending last — a
call placed and unanswered outranks a call never placed.
The square's sentence names every drill of that day, not just the decider.
Rows sort worst first: anyone who escalated in the window, then anyone who missed, then by oldest last-confirm. The person to chase is the top row.
The grid is folded from the drills that were returned — there is no roster input — so an
engineer who was never drilled in the window simply does not appear. An absent row is not
a clean record. "Rostered but never drilled" is a finding (it appears in the fix list and
as the undrilled state on the team card's roster line), never a grid row. Read the grid for
how the drilled engineers did, and the fix list for who was never asked.
The window
The window lives in the URL (?from=&to=), so "these three weeks" is a shareable link. The
page lands on 14 days; the picker offers 7d / 14d / 30d plus an absolute range, and the
arrows step a whole window back or forward, because the grid's unit is a day and its span
is the window.
That window becomes the endpoint's own from/to (RFC3339), and the page clamps exactly as
the endpoint clamps, so what is fetched and what is drawn can never disagree. See
The readiness read is windowed for the rules and what
happens at the edges.
An empty grid never means "everyone failed". A window with no drills at all renders one sentence naming the window instead of a grid; a failed drill read says the history could not be read, which is not a claim that no drills ran.
The readiness read is windowed
GET /api/v1/oncall/readiness is bounded by time, not by a row count:
| Rule | Behavior |
|---|---|
from, to | Optional RFC3339 bounds on fired_at, half-open [from, to) — a drill exactly at from is returned, one exactly at to is not, so adjacent windows tile without double-counting a drill |
| Omitted | Defaults to the last 30 days (to = now, from = to − 30d) |
from after to | 400. It is a caller bug, and an empty grid is the most alarming sentence this page can say — it must not be produced by swapped parameters |
| Wider than 90 days | Clamped, not refused: from moves forward to to − 90d. The recent end is kept, because the recent end is what the question is about. The page says so when it happens |
readinessMaxRows = 5000 | A safety cap on the response, not the window. Hitting it logs a WARN naming the window — the response is then missing the oldest drills of it |
The read used to return the newest 200 rows across every visible team with no date bound. That looked equivalent — at roughly one drill per on-call per day it reached back about three weeks — but it shrank as the fleet grew and, worse, it truncated unevenly: its oldest day returned some engineers and not others. A per-engineer-per-day grid draws a missing row as no drill, so a view whose entire job is showing who did not answer could invent a gap. The bound is now a window the reader chose and can widen, and the remaining cap is loud rather than silent.
The per-engineer page: /oncall/readiness/{userId}
Every engineer named on the readiness page links to their own drill history: every drill we
hold for them, newest first, with time to confirm per drill (confirmed_at − fired_at).
Voice and Telegram reachability chips come from GET /api/v1/users/{id}/oncall-profile,
which is self-or-admin — a caller who may not read it gets no chips rather than a guess.
There are no aggregate analytics on this page, on purpose. No average, no sparkline, no score, no percentage. With eleven drills across three days a "33% confirm rate" reads as an unreliable engineer when the truth is that the two drills dragging the fraction down are the two that placed no call at all, because there was no number on file. The per-drill duration is what makes the real finding legible — six hours and sixteen minutes against everyone else's twenty-two seconds — and averaging the two is exactly the operation that erases it. Aggregates become honest once there are weeks of history behind them.
There is still no per-user endpoint: this page reads the same fleet-wide
GET /api/v1/oncall/readiness and filters it client-side. It asks for an explicit 90-day
window — the clamp ceiling, not the overview's 14-day default — so "their record" means
their drills among every visible team's drills of the last 90 days. Two consequences worth
knowing: a drill older than 90 days is not shown here at all, and on a fleet busy enough to
hit readinessMaxRows = 5000 in 90 days the oldest of those drills are missing (the backend
logs a WARN when that happens). Lengthening it properly needs a per-user endpoint, not a page
change.
A skipped drill is a gap, not a failure, everywhere on both pages: its own count bucket
(never folded into missed), a neutral pill rather than a warning colour, and the
degraded verdict read from the newest non-skipped drill — the same rule the backend's
LatestDrillForUser uses, so a skip the day after an escalation cannot clear a badge the
backend is still acting on. A confirmed drill with call_placed: false is likewise not a
press-1: the on-call had already proved reachable that day, so no call fired, and the row
says so.
GET /api/v1/oncall/readiness/coverage
Returns one row per client the caller can see:
{
"data": [
{
"client_id": "…",
"client_name": "Acme",
"matchable_route_count": 0,
"has_default_route": false,
"oncall_enabled": false,
"l1_agent": false
}
]
}
Two properties are load-bearing:
matchable_route_countcounts what could actually arm a chain, not the route table — two filters, composed exactly asarmEscalationcomposes them:enabled = true OR is_default = true(matchingListEnabledRoutes), and then a route must name an escalation policy (matchingescalation.PagingRoutes). The name is the invariant —MatchRoutereturns a default route unconditionally, enabled or not, so a disabled default route is matchable while a disabled non-default one is not, which is why "enabled" was never the right word for it; and a notify-only route (no policy: it sends a Telegram message or a feed event and pages nobody) matches for notify while arming no chain, which is why "route" is not the right word either.has_default_routefollows the same rule: a notify-only catch-all is a catch-all for messages, not for pages. Counting the raw table — or only the first filter — would misreport exactly the configuration you are diagnosing.0means an incoming alert arms no chain and pages nobody.oncall_enabledis an independent gate. The paging predicate returns beforearmEscalationand the feature defaults to off, so a client with routes but paging off still pages nobody. This endpoint is the only place a super-admin can see the real per-client value —/auth/mehands them a wildcard "everything on" entry instead.l1_agentis the deprecated name for this same value, still served alongside it so a frontend mid-rollout keeps working; the two can never disagree — the handler fills both from one source. Readoncall_enabled. Note that the client's AI flag,ai_triage, is not a paging gate and is not reported here at all: a client with AI off pages normally.
Tenant-scoped in-handler: an unscoped super-admin sees every client; anyone else sees only
clients where they hold oncall:read, schedules:write or escalation:write. An empty
visible set returns an empty list, never all clients. The alert list's "captured, not
paged" banner reads the same endpoint, which is why it works on All clients and can
state the paging gate definitively.
When "until" is not the next hand-off
GET /api/v1/oncall/now reports two different instants about the person on call, and the
difference is the whole point:
shift_end— when the CURRENT shift ends, the next boundary the rotation produces.covered_until— when that person's unbroken stretch ends: consecutive shifts held by the same user, merged, stopping at the first hand-off to somebody else or the first unstaffed gap.
They differ whenever a rotation emits shifts more often than it changes people. A rotation with a
24-hour duration and every weekday in its by_day mask emits a shift per day while the person
advances per week, so seven consecutive shifts belong to one engineer and shift_end names
tonight's midnight seven times. Observed 2026-09-28: an engineer on call all week was shown as
"on call until Sep 29, 12:00 AM".
Anything phrased as "on call until" reads covered_until. A null means the stretch runs past
the window the endpoint materializes (14 days), so no end was observed — and then the UI says
nothing about when cover ends rather than falling back to shift_end, which is the understatement
this field exists to replace. An override that hands the pager to someone else mid-run ends the
stretch exactly as a rotation boundary does, because the merge reads the already-resolved shifts.
Reading the verdict column
Wherever a client's paging verdict is rendered — the alert list's "captured, not paged" banner, and the readiness page's fix list — the two directions are deliberately not symmetric:
Cannot be pagedis a fact. Each visible gate is sufficient on its own — no matchable route, oroncall_enabledoff, each independently means an incoming alert arms no chain.No blocker foundis not an outcome, which is why it is neutral rather than a green check and why it always ships with a caveat. There are roughly twenty gates between an alert arriving and a phone ringing; this endpoint sees two. The two that matter most right now are invisible to it and are both currently shut:escalation_livedefaults to false (000093_oncall_cutover.up.sql) and the escalation timer'sIsLivecheck suppresses every side effect for a team in shadow, and a schedule can resolve to nobody on call. A page that rendered a green "Can be paged" today would be giving the wrong answer for every configured client.
A third state sits between them: a client with routes but no catch-all (is_default)
route. Only a default route is matched unconditionally, so matchable_route_count > 0
without one proves merely that some incoming alerts match — a lone severity = P1 route
leaves everything else captured, not paged. That client is reported with its own sentence
rather than being rounded to either verdict.
Note what the clear copy claims and where it stops: it says an incoming alert matches a
route, never that a chain is armed. MatchRoute succeeding is provable from this
payload; armEscalation still runs GetPolicy, an owner-entitlement guard,
BuildSnapshot/ValidatePolicySteps and a marshal afterwards and bails on any of them,
and none of that is visible here.
It means no visible configuration gate is shut. It does not mean a page will land — see the cutover and on-call gates above.
See also
- Escalation Dead-Man's Switch — the engine guard (the readiness drill is the complementary human guard).
- Voice Paging — how the drill call is placed and driven (the same shared provider, ARI flow, and DTMF ack path), plus phone verification.
- On-call & Escalation Management — the schedules and escalation policies the drill resolves the on-call from and pages the coverage gap through.
Live validation
The worker's due-check and state machine are unit-tested, but the call, the keypress,
the nudge, and the coverage-gap page can only be proven on a real deployment. After merge
and deploy, set PROXIMA_ONCALL_READINESS_DRILL=on with a short
PROXIMA_ONCALL_READINESS_INTERVAL and a test schedule whose on-call is the tester, then
confirm: a readiness call places and "press 1" confirms (proxima_oncall_readiness_confirmed_total
ticks, last-confirmed updates); a forced miss delivers the Telegram nudge and the ✅
button confirms; and letting the grace expire arms a source_type=readiness_drill
coverage-gap alert that pages through the team's critical policy to its terminus. The
full co-driven checklist lives with the implementation plan in the repository.
Forcing a drill for a live test
The "is a drill due?" decision is deterministic per (user, schedule, UTC-day bucket): the
drill fires once at a hash-derived instant within the current UTC day. That makes an
on-demand test call non-obvious — two things to know:
- Deleting the drill row does not force a re-fire across a day boundary. The bucket key includes the UTC day, so a row deleted for an already-past bucket is moot after 00:00 UTC — the next day's bucket has its own (future) target.
- A past-band window forces an immediate fire. Set
PROXIMA_ONCALL_READINESS_WINDOWto a band that ends at or just before the current UTC time (for example, if it is06:34UTC, set00:00-06:34). The target clamps into that past band, sonow ≥ targetand the worker fires on its next sweep. The band is UTC (see the config table) — check the current UTC time, not local.
Procedure (Kubernetes / External Secrets):
# 1. Ensure the on-call for a test schedule is you, with a VERIFIED phone.
# 2. Provision the prompt sound on the Asterisk box if absent (see "Ops prerequisite"
# above) — otherwise the call connects but plays silence.
# 3. Set the force-window in the secret source (e.g. ESO), using a band ending at the
# current UTC minute: PROXIMA_ONCALL_READINESS_WINDOW=00:00-HH:MM (HH:MM = now, UTC)
# The env has no auto-reloader, so restart the backend to re-read it. `kubectl rollout
# restart` does NOT work on the Argo Rollout CRD — patch restartAt instead:
kubectl -n console-system patch rollout console-backend --type merge \
-p "{\"spec\":{\"restartAt\":\"$(date -u +%Y-%m-%dT%H:%M:%SZ)\"}}"
# 4. The drill fires within one interval. To test the MISS path, do not press 1 (hang up):
# the call goes terminal -> "readiness drill missed - nudge sent" -> Telegram nudge ->
# grace -> coverage-gap (or a no_client / no_policy skip for an unconfigured team).
# 5. AFTERWARD: remove PROXIMA_ONCALL_READINESS_WINDOW (back to any-hour) and restart again,
# so real drill timing is not skewed into a fixed band.
Validated end-to-end in production (2026-08-14): the confirm path (press 1 →
readiness_confirmed) and the miss → nudge → belated-confirm-cancels path (a ✅ tap during
grace stamps the drill confirmed, and the escalate pass — which acts only on missed rows —
skips it, so no team is spuriously paged).