Paging Control
This page describes how Proxima Console decides whether, and whom, to page when an alert fires — covering the full control surface from the coarse per-client on/off switch down to the per-route severity and label matching logic.
How paging works: route-driven
Paging is route-driven. An alert pages a human if and only if an enabled escalation route for that client matches the alert. There is no hardcoded paging gate — no "only P1/P2 pages" rule burned into the engine. The routes and the policies they point at are the single paging control surface.
A route matches on any combination of:
- Severity (P1–P5, per-route) — a route fires for the severities it is authored for.
- Environment — a label/slug selector restricting matching to a specific environment.
- Host — a hostname or IP selector restricting matching to a specific host.
- Label selectors — arbitrary key/value label matchers on the alert.
- Default (catch-all) — a route marked
is_defaultmatches any alert not caught by a more-specific route above it. Position matters: routes are evaluated top to bottom, first match wins.
Because a route carries the severities it covers, you can page on any severity tier you choose — P3, P4, or a mixed set — simply by authoring a route that matches at that tier. Equally, a severity you don't route is never paged.
Safe default: no routes → no pages
A client with no routes never pages anyone. You must author at least one route — or assign a covering team that materializes the default pair of routes — before any paging fires. This is intentional: an unconfigured client is silent, never noisy.
Covering-team shortcut
Assigning a covering team materializes exactly two routes for the client automatically:
| Route | Severity | Target policy |
|---|---|---|
| Critical route | P1 | Team's critical-role policy |
| Default route | P2 | Team's default-role policy |
Covering teams page on P1 (critical policy) and P2 (default policy) only — P3–P5 are not paged via a covering team. There is no severity-null catch-all: if you need to page a covering team at a lower tier, author an explicit route for that severity.
These are source='covering_team' routes and are replaced atomically when you reassign the covering team. Your hand-authored manual routes are never touched by this operation.
Materialization runs on (re)assignment. Covering-team assignments made before this change still carry the old severity-null catch-all — reassign the covering team to re-materialize the bounded P1/P2 pair.
See On-call & Escalation Management — Assigning a covering team for the API and UI details.
Per-client on/off: oncall_enabled
The oncall_enabled feature flag is the coarse on/off switch for paging, and for paging only. When oncall_enabled is off:
- No routes are evaluated.
- No escalations are armed.
- No pages fire, regardless of what routes are authored.
Turn it on in the client's Features section (Client Settings, super-admin) under "On-call & paging" to enable paging for that client. The flag does not affect alert ingestion or correlation — alerts still flow in, are enriched and are grouped — it only gates paging.
It is not silent after the fact: when the gate refuses, it records a no-page event against the alert group whose reason is l1_disabled. That token keeps the old name on purpose — it is written into historical alert_group_log rows and a partial unique index is built on the literal — so grep for l1_disabled even though the flag it reports is now oncall_enabled. See Resolution Safety for the full no-page reason table.
oncall_enabled and ai_triage are two switches, not one
These two flags used to be one boolean called l1_agent. That boolean read as an AI switch and was labelled "L1 Agent", but it also decided whether anyone got paged — so a client that wanted paging without LLM spend could not have it, and turning the AI off turned the pager off with nothing but a l1_disabled row to say so.
They are now independent, and each states its own consequence:
| Flag | Off means | Does not affect |
|---|---|---|
oncall_enabled | Nobody is paged. No route is evaluated, no escalation is armed, no notification chain runs. | AI investigation — a non-paging client can still be auto-investigated. |
ai_triage | Nothing is investigated automatically. No auto-triage, no memory draft, no remediation proposal, no reactive /engage reply. Alerts still page exactly as before. | Paging — not one gate on the paging path reads it. |
Both default off. Neither is a precondition for the other, and no code path may make one speak for the other: the no-page record belongs to the paging gate alone, and triage-quality warnings belong to the AI gate alone.
A client whose stored features still hold only {"l1_agent": true} resolves both new flags from it — domain.ResolveFeatures carries a legacy alias. Migration 000230 backfills the stored rows and deliberately leaves the legacy keys in place so a rollback still resolves correctly. See Per-Client Feature Flags.
Auto-triage vs. manual triage: ai_triage_manual
When ai_triage is on, an alert can trigger an automatic AI root-cause analysis (RCA/triage). The triage mode is controlled by the ai_triage_manual feature flag — which changes when an investigation runs, never whether anyone is paged:
| Flag | Mode | Behaviour |
|---|---|---|
absent / false | Automatic | Auto-triage runs when the alert fires (P1 only — see below). |
true | Manual | No auto-triage. An operator triggers investigation on demand. |
Toggle in the client's Features section: the "Investigation mode: Automatic / Manual" switch appears under "AI investigation" and is visible only while ai_triage is on. ai_triage_manual keeps the same polarity as the l1_manual_triage it replaces — true still means manual.
Auto-triage is P1-only
Auto-triage (AI RCA) runs automatically for P1 alerts only. P1 is the only tier where the cost and latency of an LLM investigation is justified without an operator's explicit request.
For P2–P5, the paged engineer triggers triage on demand:
POST /api/v1/alerts/{alertGroupID}/triage
Requires alerts:write. Returns 202 with the alert-group and incident IDs. The on-demand path is identical to the auto path — it reuses the same triage worker and idempotency layer, so submitting it twice for the same alert group does nothing.
When ai_triage_manual is true, auto-triage is disabled for all severities — including P1. The POST /triage endpoint always works regardless of the flag. Paging is untouched either way.
Retired: Telegram-group broadcast for alert paging
The v0.1 Telegram-group alert broadcast — which posted alert notifications directly to client group chats — is retired. On-call engineers are now paged exclusively via their personal escalation notification chain (routes → escalation policy → per-user notification chain of telegram / voice steps).
Group posting was removed because it was uncontrolled: any firing alert would broadcast to every bound group, with no per-severity or per-route gating, no ack integration, and no delivery guarantee. The escalation chain model replaces this entirely, giving each on-call engineer a personal, ordered, retried notification sequence with a proper ack fence.
Group posting will return in a later phase as a controlled NOTIFY_CHANNEL escalation step type — an explicit, opt-in step in an escalation policy that broadcasts to a group when the operator deliberately includes it. Until then, group chats receive Jira/Service Desk notifications only, not alert pages.
Summary: the control surface at a glance
| Control | Where | Effect |
|---|---|---|
oncall_enabled feature flag | Client Settings → Features ("On-call & paging") | On/off for all paging for this client. Off = nobody is paged |
| Escalation routes | Admin → Escalation (routes table) | Which alerts page (severity + env/host/labels + order) |
| Escalation policies | Admin → Escalation (policy builder) | Who gets paged and in what order (steps, WAIT, REPEAT) |
| Covering team assignment | Client detail → Covering team | Auto-materializes P1 (critical) + P2 (default) routes into the team's policies; P3–P5 not paged |
ai_triage feature flag | Client Settings → Features ("AI investigation") | On/off for AI investigation. Off = alerts still page, nothing is auto-investigated |
ai_triage_manual flag | Client Settings → Features | Auto-triage off (on-demand only via POST /triage). Never affects paging |
| Per-user notification chain | Admin → Users → notification policy | How a paged individual is reached (telegram, voice, WAIT delays) |
Related
- On-call & Escalation Management — schedules, policies, routes, covering teams, cutover.
- Notification Chains — per-user paging channel order and fallback.
- Escalation Dead-Man's Switch — guard against a wedged escalation timer.
- Alerting & Correlation — alert ingestion, lifecycle, and correlation context.
- L1 Incident Agent — what
ai_triageturns on, and why it pages nobody. - Per-Client Feature Flags — every flag, its default, and the retired names.