Notification Chains
When an escalation step decides who to page (an on-call user, or a named user), a per-user notification chain decides how that one person is reached — which channels, in which order, with what delays between them. This guide covers the chain model, how to author a chain, the failure behavior, and the built-in fallback.
Notification chains are executed by the NotificationChainWorker, a Postgres-timer worker that runs alongside the escalation engine. It is enabled together with the rest of the paging stack (it requires the ops Telegram bot token). Voice legs additionally require Twilio to be configured; without Twilio a voice step is skipped (see Voice not configured).
The timed chain model
A chain is an ordered list of steps frozen at the moment the escalation fires for that user. There are two kinds of step:
- NOTIFY — page one channel:
telegramorvoice. - WAIT — pause for
wait_secondsbefore the next step.
When an escalation NOTIFY fires for a user, the escalation engine freezes that user's chain into a run keyed on (alert group, firing episode, user, important) and hands off; the chain worker then walks the frozen steps on its own timeline. Because the chain is frozen at fire time, editing a user's chain never disturbs an in-flight page — the change takes effect on the next episode.
A typical chain:
telegram # page Telegram immediately
WAIT 300s # give them 5 minutes to acknowledge
voice # if still firing, call their phone
The WAIT is honored: the voice leg does not fire until the wait has elapsed. If the user acknowledges (or the alert resolves) during the wait, the fence halts the chain and no further channels are paged.
Default vs important
Every user can have two chains:
- the default chain — used for a normal page.
- the important chain — used when the escalation step is marked
important. An important page typically rings the phone sooner (shorter or no WAIT), because a critical incident should not wait out an ack window.
The escalation step's important flag selects which of the user's two chains is frozen.
Authoring a chain (API-first)
Chains are authored through the API. A read-only view of your own chain also lives in the console under On-Call → My On-Call (/oncall/me), alongside your paging phone; a super-admin can edit any user's chain from that user's admin detail page.
Set a user's default and important chains with:
PUT /api/v1/users/{id}/notification-policy
The body carries the two ordered chains as step lists. Each step is either a NOTIFY step (notify_by: "telegram" | "voice") or a WAIT step (wait_seconds: <n>):
{
"default": [
{ "notify_by": "telegram" },
{ "wait_seconds": 300 },
{ "notify_by": "voice" }
],
"important": [
{ "notify_by": "telegram" },
{ "notify_by": "voice" }
]
}
- A NOTIFY step sets
notify_byand omitswait_seconds. - A WAIT step sets
wait_secondsand omitsnotify_by. - Read the exact request/response shape from the generated API types / Swagger; the fields above are the load-bearing ones.
Who may read and write a chain
The endpoint is self-or-admin, enforced in-handler (it was super-admin-only until the
/oncall/me work; leaving it there meant the one person a chain describes was the one person
who could not read it):
- Yourself — any authenticated user may read and replace their own chain. This branch never consults client scope, so an engineer with no client assignments still owns their chain.
- A
users:writeholder may read and replace the chain of a user who shares a client with them and is not a super admin. Holdingusers:writesomewhere is necessary but not sufficient; a colleague on the same client holding onlyusers:readis refused. - Users only. Service accounts — including a wildcard (
client_scope: ["*"]) API key holdingusers:write— are refused on both verbs. A machine identity does not get to decide which channels reach a human. (This is not a restriction the relaxation added: the previous super-admin gate already excluded API keys, because a service account is never a super admin.)
Writes are recorded in the audit log as notification_chain_updated against the user whose
chain changed, flagged self-service or not. Reads are not audited.
Limits
A chain is bounded on write, so it cannot be used to make a pager silently go quiet:
- at most 20 steps per chain;
- no single
wait_secondsover 14400 (4 hours); - and the total of a chain's waits may not exceed 14400 either — so the limit cannot be side-stepped by splitting one long wait into several legal ones.
These are pragmatic ceilings rather than engine-imposed ones: nothing in the escalation engine
forces a chain to finish within any window, and the point of the bound is simply that the
longest silence a chain can introduce is one an operator can notice and recover from. A
rejection returns 400 naming the exact limit it hit. The limits apply to writes only —
a chain already stored that exceeds them keeps running and is simply un-resavable until it is
brought back inside them.
Channel contacts
A NOTIFY step only pages if the user has a usable contact for that channel:
- telegram — resolved from the user's
telegram_idtrait. - voice — resolved from the user's global on-call paging phone (
user_oncall_profile, set on their profile page or viaPUT /api/v1/users/{id}/oncall-profile). It is user-global, not per-client — D3 moved it off the client roster. UnderPROXIMA_VOICE_VERIFIED_GATE=enforcean unverified number is skipped and the chain falls through to the user's other channels; the defaultwarnplaces the call anyway.
A step whose channel has no usable contact is treated as a permanent skip: the chain advances past it to the next step rather than stalling (see below).
Failure behavior: permanent vs transient
The whole reason delivery runs through the chain worker (rather than paging every channel inline) is robustness of the handoff between channels. Each NOTIFY step's send outcome is classified:
- Permanent failure — a dead/blocked contact, a
403, an unregistered channel, or an empty render. The chain advances to the next channel. A user whose Telegram is blocked still gets the voice call. - Transient failure — a temporary send error or a DB blip. The same step is re-armed with backoff, up to
PROXIMA_NOTIFICATION_TRANSIENT_MAX_ATTEMPTS(default 3) attempts, after which it is treated as terminal and the chain advances. - Sent — success; the chain advances to the next step.
Before chains, on-call delivery paged a user's channels inline and early-returned on the first channel error — so a dead Telegram contact silently dropped the page and the voice call never happened. The chain executor fixes this: a permanent failure advances instead of aborting.
The fence overrides everything: on every step the worker rechecks that the alert group is still firing. An acked or resolved group halts the whole episode's chains for every user — no further pages fire.
The fallback (user has no authored chain)
If a paged user has no authored notification policy, the engine synthesizes a timed default fallback so they are never silently un-paged:
- Default (non-important):
telegram → WAIT 300s → voice— page Telegram, give a 5-minute ack window, then call. - Important:
telegram → voice— page Telegram and call back-to-back, with no wait.
Authoring a chain for the user overrides the fallback entirely.
The default fallback rings phones — read this before assuming otherwise
The important flag on an escalation step controls how soon the phone rings, not whether it rings. Both built-in fallbacks end in voice. So a default catch-all route — the one that carries P2 and below — does phone people, five minutes behind the Telegram.
This is not theoretical. On 2026-09-30 a P2 on one client walked a -default escalation chain and phoned the on-call engineer, then the team lead, then the CTO, then the CEO, because none of them had authored a notification policy and the chain repeated twice. Routing was correct throughout; the phone calls came from this fallback.
Since then both fallbacks are configurable per deployment, so "may a non-critical alert ring a phone" is an environment decision rather than a compile-time constant:
| Env var | Built-in default | Meaning |
|---|---|---|
PROXIMA_ONCALL_DEFAULT_NOTIFY_CHAIN | telegram,wait:5m,voice | The chain for a non-important step — everything a default route sends. |
PROXIMA_ONCALL_IMPORTANT_NOTIFY_CHAIN | telegram,voice | The chain for an important step — what a critical route's P1 sends. |
Format: a comma-separated ordered list of telegram, voice, and wait:<duration> (anything time.ParseDuration accepts). To stop non-critical alerts ringing phones entirely:
PROXIMA_ONCALL_DEFAULT_NOTIFY_CHAIN=telegram
The two sides are independent — overriding the default chain leaves the important one at its built-in value, so quieting P2 does not quiet P1.
Empty means "use the built-in default", not "notify nobody." A chain that names no channel is a silent page, the worst failure this subsystem has, so it is deliberately not expressible: a value naming only waits is refused, and any malformed value logs a warning and falls back to the built-in chain. A typo here degrades to today's paging behavior, never to silence. If a step should page nobody, remove the step from the escalation policy instead.
Both keys appear on the read-only settings page, where an unset chain reads (built-in default).
Configuration
| Env var | Default | Meaning |
|---|---|---|
PROXIMA_NOTIFICATION_CHAIN_INTERVAL | 2s | How often the chain worker ticks. |
PROXIMA_NOTIFICATION_TRANSIENT_MAX_ATTEMPTS | 3 | Transient re-arm attempts per channel before advancing. |
The worker is gated on the ops Telegram bot token (the same gate as the escalation timer and auditor): with no Telegram channel there is nothing to page through, and the whole paging stack — including chains — is off.
Voice not configured
The voice channel is registered when either provider is configured:
- Asterisk —
PROXIMA_ASTERISK_ARI_URL,PROXIMA_ASTERISK_ARI_USER,PROXIMA_ASTERISK_ARI_PASSWORD - Twilio —
PROXIMA_TWILIO_ACCOUNT_SID,PROXIMA_TWILIO_AUTH_TOKEN,PROXIMA_VOICE_PUBLIC_BASE_URL
PROXIMA_VOICE_PROVIDER picks between them: asterisk, twilio, or empty to auto-select — which
prefers Asterisk when it is configured, falling back to Twilio.
Only when neither is configured does a voice NOTIFY step resolve no channel and get
permanently advanced past (a config gap, not a stuck chain) — the chain still completes, just
without the call.
Observability
| Metric | Labels | Meaning |
|---|---|---|
proxima_notification_chain_steps_total | channel | Successful chain sends per channel. |
proxima_notification_chain_retries_total | channel | Transient re-arms per channel. |
proxima_notification_delivery_failures_total | channel, kind | Delivery failures, split permanent vs transient. |
Live validation procedure
The chain worker only runs on a deployed backend (it needs Postgres, the ops bot, and — for voice — Twilio). Validate on a dev/staging deployment after the branch is deployed:
-
Author a timed chain via the API for a test user (as that user, or as a
users:writeholder who shares a client with them):PUT /api/v1/users/{testUserId}/notification-policy
{
"default": [ { "notify_by": "telegram" }, { "wait_seconds": 60 }, { "notify_by": "voice" } ],
"important": [ { "notify_by": "telegram" }, { "notify_by": "voice" } ]
}(Use a short
wait_secondslike60for the test so you don't wait 5 minutes.) Ensure the user has atelegram_idtrait and an on-call paging phone on their profile. -
Fire a test escalation that pages this user (e.g. a P1 alert routed to a policy whose step notifies the user / their schedule). Confirm the sequence with real timing:
- the Telegram page arrives first;
- after the WAIT elapses, the voice call is placed.
-
Blocked-Telegram case — repeat with a user whose Telegram contact is dead/blocked (or delete the
telegram_idtrait). Confirm the voice call still lands — the permanent Telegram failure must not drop the page. -
Check the metrics in VictoriaMetrics:
proxima_notification_chain_steps_total{channel="telegram"}and{channel="voice"}both increment;proxima_notification_delivery_failures_total{kind="permanent"}increments for the blocked-Telegram case, whilevoicestill succeeds.
-
Acknowledge the alert mid-chain and confirm the chain halts — no further pages.
Related
- On-call & Escalation Management — schedules, policies, routes, and who gets paged.
- Escalation dead-man's switch — the auditor that guards the paging engine against a wedged timer.