Voice Trunk Health
The native on-call engine can page an operator by voice call over a self-hosted
Asterisk voice node with a SIP trunk (the P2 voice platform — see
infra/asterisk/README.md
for how the node is provisioned and hardened). Voice is one additional paging
channel alongside Telegram, not a replacement for it.
If that SIP trunk goes down — the upstream provider drops the registration, throttles the account after a rejection storm, or the voice node loses network — voice calls stop going through. Console runs a small trunk-health worker that polls the trunk over Asterisk's REST interface (ARI) and exports a single gauge so you can alert on it before an operator misses a call.
A down trunk means voice paging is degraded — it does not mean the on-call engine is dead. Telegram paging is unaffected and continues to fire. So this signal is observability-only: the worker never pages and is deliberately decoupled from the escalation dead-man's switch heartbeat (the same rule the schedule auditor follows). Treat a trunk-down alert as "fix the voice path soon", not "the pager is down".
The voice path ships disabled. The trunk-health worker is constructed only when
the Asterisk ARI endpoint is configured (PROXIMA_ASTERISK_ARI_URL /
PROXIMA_ASTERISK_ARI_USER / PROXIMA_ASTERISK_ARI_PASSWORD all set). With an empty ARI
config — the default, and the state in production today — the worker is not started and
proxima_voice_trunk_up is absent. That is expected and inert; there is nothing to
alert on until the voice node is stood up per the runbook.
How it works
Every poll interval (PROXIMA_VOICE_HEALTH_INTERVAL, default 30s) the worker asks ARI
for the state of the configured trunk endpoint (PROXIMA_ASTERISK_TRUNK, default
sarkor-trunk):
- Trunk online (registered/qualified) → set
proxima_voice_trunk_upto1and stampproxima_voice_trunk_last_ok_timestamp_seconds. - Trunk not online, or the probe errors (transport error, non-2xx, throttled) → set
proxima_voice_trunk_upto0.
To keep the logs quiet, the worker logs loudly exactly once per down-transition (a single ERROR — "voice trunk unreachable — voice paging degraded; Telegram unaffected") and once on recovery, never on every poll while it stays down.
Two carriers: per-path probing and the Trunk Health page
A deployment can configure several voice paths (PROXIMA_VOICE_PATHS), each a
carrier on an Asterisk box — in production, Sarkor on asterisk01 and Skyline on
asterisk02. The worker then probes every path independently, on that path's own
box, and writes each verdict to its own row. One carrier being down neither stops the
other being probed nor hides it.
proxima_voice_trunk_up carries a path attribute in that mode, which is what makes
the two SIPs distinguishable. Without it both paths write the same series and "Skyline up,
Sarkor down" reads as one gauge flapping between 1 and 0. A single-trunk deployment has
nothing to distinguish, so it records no attribute and its existing series is unchanged.
The Trunk Health page
On-Call → System → Trunk Health (/oncall/voice) is the operator view of the same
facts: one card per configured path with its probe verdict, its dial preference, when it
was last OK, and its last two failures. It is super-admin only — a voice path is
global infrastructure with no owning client, so no per-client permission fits it. The
trunk-down Telegram notice links straight to it.
Four distinctions the page is built around, and that matter whether you read them there or in the metrics:
-
A path that is up can still be unable to SPEAK. A page's audio is delivered to one machine, and playback names it only when the box the call is on has confirmed it. The audio delivery row answers that separately from the trunk's state.
A configured box with no confirmation reads unproven, not broken, because the row cannot tell two causes apart:
- The fetcher on that machine is not reporting under the name the backend expects, so every page it carries rings and then plays the generic prompt instead of the alert text. This has no other symptom anywhere — it is why the row exists.
- Nothing has had occasion to confirm. A confirmation is only written when a call needs rendered audio, and only a page does: readiness drills and verification calls play static prompts, so a fleet that has placed nothing but drills sits at "never confirmed" while being perfectly healthy.
So the row names both and asserts neither. To settle it, place a real page (or an alert-purpose test call) on that box and watch the row flip to "confirmed by this box" — a drill will never move it. See the prefetch interlock.
-
A probe error and a call error are different facts. A trunk can qualify fine and still refuse INVITEs, so the page shows the last failed probe and the last call that never connected in separate rows. A clean probe beside a recent call error is a carrier problem, not a monitoring gap.
-
A down path is still dialed. Whether a page is attempted on a path depends on that box's ARI event loop, not on this probe. The probe is evidence for a human; it is never a gate that could stop a page being placed. (The dial-time skip that does exist is liveness of the event loop, and it fails open: a path whose liveness is unknown is dialed.)
-
"Not probed" is not "down". A configured path with no health row means the worker has not reported on it yet — at startup, or because it is not running. Rendering that as down would send you to chase a carrier that is fine.
Is failover armed?
The page's headline answer comes from the backend, and it is deliberately strict: failover is armed only when at least two paths are configured and at least two are believed up. Two configured with one down is not redundancy — it is one carrier away from no phone paging at all. One path configured is reported as such ("a page has nowhere to go if this carrier fails"), and no path up says plainly that phone paging cannot be placed and that Telegram is unaffected.
Metrics reference
| Metric | Type | Meaning |
|---|---|---|
proxima_voice_trunk_up | gauge | 1 when the SIP trunk is reachable/qualified, 0 when down/throttled. == 0 = voice paging degraded. Carries a path attribute per configured voice path (absent in a single-trunk deployment). |
proxima_voice_trunk_last_ok_timestamp_seconds | gauge (unix s) | Wall-clock of the last successful trunk probe. Lets you see how long the trunk has been down. |
Query them live via the observability helper:
make obs-metrics PATTERN=proxima_voice_trunk
Recommended alert
Alert when the trunk stays down for a couple of minutes — long enough to ride out a single flapping probe, short enough to catch a real outage before it costs a page:
# Voice paging degraded: the SIP trunk has been down for > 2 minutes.
# Telegram paging is unaffected — this is a "fix soon", not a "pager is down".
proxima_voice_trunk_up == 0
With more than one path configured this expression fires per path, which is what you
want: the alert names the carrier that is down. The condition worth a louder alert is
every path being down — min(proxima_voice_trunk_up) == 0 only catches one of them,
while max(proxima_voice_trunk_up) == 0 catches the case where no carrier can place a
call at all.
Fire it after a for: 2m hold. Because voice is degraded-not-broken, route this at a
lower severity than the escalation dead-man's switch alerts — a warning to the voice
platform owner, not an out-of-band page. (The dead-man's switch, which does page
out-of-band, covers the case where the whole engine stops advancing escalations.)
If you want to catch the worker itself wedging (stopped polling), alert on staleness of the last-ok gauge instead — but only when you know the voice node is meant to be up:
# Optional: last successful probe is stale (worker stopped polling a live trunk).
time() - proxima_voice_trunk_last_ok_timestamp_seconds > 300
See also
infra/asterisk/README.md— provisioning, Sarkor trunk hardening, ARI lockdown, and the live-validation checklist for the voice node.- Escalation Dead-Man's Switch — the out-of-band guard that pages when the escalation engine stops (a different, higher-severity signal).
- On-call & Escalation Management — the paging engine voice sits alongside.