Skip to main content

Paging Control

This page describes how Proxima Console decides whether, and whom, to page when an alert fires — covering the full control surface from the coarse per-client on/off switch down to the per-route severity and label matching logic.

How paging works: route-driven​

Paging is route-driven. An alert pages a human if and only if an enabled escalation route for that client matches the alert. There is no hardcoded paging gate — no "only P1/P2 pages" rule burned into the engine. The routes and the policies they point at are the single paging control surface.

A route matches on any combination of:

  • Severity (P1–P5, per-route) — a route fires for the severities it is authored for.
  • Environment — a label/slug selector restricting matching to a specific environment.
  • Host — a hostname or IP selector restricting matching to a specific host.
  • Label selectors — arbitrary key/value label matchers on the alert.
  • Default (catch-all) — a route marked is_default matches any alert not caught by a more-specific route above it. Position matters: routes are evaluated top to bottom, first match wins.

Because a route carries the severities it covers, you can page on any severity tier you choose — P3, P4, or a mixed set — simply by authoring a route that matches at that tier. Equally, a severity you don't route is never paged.

Safe default: no routes → no pages​

A client with no routes never pages anyone. You must author at least one route — or assign a covering team that materializes the default pair of routes — before any paging fires. This is intentional: an unconfigured client is silent, never noisy.

Covering-team shortcut​

Assigning a covering team materializes exactly two routes for the client automatically:

RouteSeverityTarget policy
Critical routeP1Team's critical-role policy
Default routeP2Team's default-role policy

Covering teams page on P1 (critical policy) and P2 (default policy) only — P3–P5 are not paged via a covering team. There is no severity-null catch-all: if you need to page a covering team at a lower tier, author an explicit route for that severity.

These are source='covering_team' routes and are replaced atomically when you reassign the covering team. Your hand-authored manual routes are never touched by this operation.

Reassign to pick up bounded routes

Materialization runs on (re)assignment. Covering-team assignments made before this change still carry the old severity-null catch-all — reassign the covering team to re-materialize the bounded P1/P2 pair.

See On-call & Escalation Management — Assigning a covering team for the API and UI details.

Per-client on/off: oncall_enabled​

The oncall_enabled feature flag is the coarse on/off switch for paging, and for paging only. When oncall_enabled is off:

  • No routes are evaluated.
  • No escalations are armed.
  • No pages fire, regardless of what routes are authored.

Turn it on in the client's Features section (Client Settings, super-admin) under "On-call & paging" to enable paging for that client. The flag does not affect alert ingestion or correlation — alerts still flow in, are enriched and are grouped — it only gates paging.

It is not silent after the fact: when the gate refuses, it records a no-page event against the alert group whose reason is l1_disabled. That token keeps the old name on purpose — it is written into historical alert_group_log rows and a partial unique index is built on the literal — so grep for l1_disabled even though the flag it reports is now oncall_enabled. See Resolution Safety for the full no-page reason table.

oncall_enabled and ai_triage are two switches, not one​

These two flags used to be one boolean called l1_agent. That boolean read as an AI switch and was labelled "L1 Agent", but it also decided whether anyone got paged — so a client that wanted paging without LLM spend could not have it, and turning the AI off turned the pager off with nothing but a l1_disabled row to say so.

They are now independent, and each states its own consequence:

FlagOff meansDoes not affect
oncall_enabledNobody is paged. No route is evaluated, no escalation is armed, no notification chain runs.AI investigation — a non-paging client can still be auto-investigated.
ai_triageNothing is investigated automatically. No auto-triage, no memory draft, no remediation proposal, no reactive /engage reply. Alerts still page exactly as before.Paging — not one gate on the paging path reads it.

Both default off. Neither is a precondition for the other, and no code path may make one speak for the other: the no-page record belongs to the paging gate alone, and triage-quality warnings belong to the AI gate alone.

No client's behaviour changed when this split shipped

A client whose stored features still hold only {"l1_agent": true} resolves both new flags from it — domain.ResolveFeatures carries a legacy alias. Migration 000230 backfills the stored rows and deliberately leaves the legacy keys in place so a rollback still resolves correctly. See Per-Client Feature Flags.

Auto-triage vs. manual triage: ai_triage_manual​

When ai_triage is on, an alert can trigger an automatic AI root-cause analysis (RCA/triage). The triage mode is controlled by the ai_triage_manual feature flag — which changes when an investigation runs, never whether anyone is paged:

FlagModeBehaviour
absent / falseAutomaticAuto-triage runs when the alert fires (P1 only — see below).
trueManualNo auto-triage. An operator triggers investigation on demand.

Toggle in the client's Features section: the "Investigation mode: Automatic / Manual" switch appears under "AI investigation" and is visible only while ai_triage is on. ai_triage_manual keeps the same polarity as the l1_manual_triage it replaces — true still means manual.

Auto-triage is P1-only​

Auto-triage (AI RCA) runs automatically for P1 alerts only. P1 is the only tier where the cost and latency of an LLM investigation is justified without an operator's explicit request.

For P2–P5, the paged engineer triggers triage on demand:

POST /api/v1/alerts/{alertGroupID}/triage

Requires alerts:write. Returns 202 with the alert-group and incident IDs. The on-demand path is identical to the auto path — it reuses the same triage worker and idempotency layer, so submitting it twice for the same alert group does nothing.

Manual mode unaffected by P1 auto-triage

When ai_triage_manual is true, auto-triage is disabled for all severities — including P1. The POST /triage endpoint always works regardless of the flag. Paging is untouched either way.

Retired: Telegram-group broadcast for alert paging​

The v0.1 Telegram-group alert broadcast — which posted alert notifications directly to client group chats — is retired. On-call engineers are now paged exclusively via their personal escalation notification chain (routes → escalation policy → per-user notification chain of telegram / voice steps).

Group posting was removed because it was uncontrolled: any firing alert would broadcast to every bound group, with no per-severity or per-route gating, no ack integration, and no delivery guarantee. The escalation chain model replaces this entirely, giving each on-call engineer a personal, ordered, retried notification sequence with a proper ack fence.

Group posting will return in a later phase as a controlled NOTIFY_CHANNEL escalation step type — an explicit, opt-in step in an escalation policy that broadcasts to a group when the operator deliberately includes it. Until then, group chats receive Jira/Service Desk notifications only, not alert pages.

Summary: the control surface at a glance​

ControlWhereEffect
oncall_enabled feature flagClient Settings → Features ("On-call & paging")On/off for all paging for this client. Off = nobody is paged
Escalation routesAdmin → Escalation (routes table)Which alerts page (severity + env/host/labels + order)
Escalation policiesAdmin → Escalation (policy builder)Who gets paged and in what order (steps, WAIT, REPEAT)
Covering team assignmentClient detail → Covering teamAuto-materializes P1 (critical) + P2 (default) routes into the team's policies; P3–P5 not paged
ai_triage feature flagClient Settings → Features ("AI investigation")On/off for AI investigation. Off = alerts still page, nothing is auto-investigated
ai_triage_manual flagClient Settings → FeaturesAuto-triage off (on-demand only via POST /triage). Never affects paging
Per-user notification chainAdmin → Users → notification policyHow a paged individual is reached (telegram, voice, WAIT delays)