Skip to main content

Voice Paging

The native on-call engine can page an operator by voice call as one additional channel alongside Telegram. A voice call places an outbound phone call, speaks a short summary of the alert, and lets the operator acknowledge by pressing 1 — or escalate by pressing 2, "not me, get the next person now" — on the keypad. Voice is registered as a step in a notification chain; if it fails or the operator never picks up, the chain auto-advances to the next channel so an alert is never silently dropped.

This page covers how the call is placed and driven — the provider abstraction, the call/ack flow, and the text-to-speech (TTS) contract. For the health of the SIP trunk that carries these calls, see Voice Trunk Health.

Inert until a voice node is configured — production is unchanged

Voice paging ships behaviorally identical to before. No provider is selected until one is fully configured. The self-hosted Asterisk path is only built (and its ARI event loop only started) when the ARI config (PROXIMA_ASTERISK_ARI_URL / PROXIMA_ASTERISK_ARI_USER / PROXIMA_ASTERISK_ARI_PASSWORD) is set. With that config absent — the default, and the state in production today — the Asterisk controller never starts and voice stays on the existing Twilio path (or on no voice at all if Twilio is not configured either).

The provider model​

Voice paging is placed behind a small PhoneProvider abstraction so the same notification-chain step, the same ack authorization, and the same spoken text work regardless of who actually dials the phone:

  • asterisk — self-hosted Asterisk voice node driven over its REST interface (ARI), with the alert text spoken by the Proxima TTS service. This is the P2 voice platform.
  • twilio — the original hosted-TwiML path, retained as a config-selectable fallback. Nothing about the Twilio flow changed.

Both providers share one ack path — the same action.VoiceActor seam used by the Twilio HTTP handler: it resolves the called contact → its linked user → a per-client alerts:write RBAC check → the ack, and auto-advances the escalation on a failed call. There is no forked authorization: pressing 1 on an Asterisk call is authorized exactly as a Twilio keypad ack is.

Which provider is active (PROXIMA_VOICE_PROVIDER)​

PROXIMA_VOICE_PROVIDER selects the active provider. It defaults to auto (empty):

PROXIMA_VOICE_PROVIDERResult
(unset / empty) — autoPrefer Asterisk when its ARI config is set; else Twilio when configured; else voice is disabled. A deploy to a Twilio-only environment keeps voice on Twilio.
asteriskForce Asterisk. If its ARI config is not set, voice is disabled (an explicit-but-unconfigured selection never silently switches to Twilio).
twilioForce Twilio. If Twilio is not configured, voice is disabled.

On startup the backend logs the selection, so you can confirm what is live:

INFO voice channel active  component=backend provider=asterisk

or, when neither provider is configured:

INFO no voice provider configured; notification-chain voice steps will permanent-advance

(When Asterisk is built you'll also see INFO asterisk voice provider built … trunk=… app=….)

The ARI call flow​

When a notification chain reaches its voice step, the active provider is asked to originate one outbound call. On the Asterisk provider that call is driven by a small Stasis state machine over ARI:

  1. Originate — POST /ari/channels places the call to the operator's verified phone through the SIP trunk, handing it to the Stasis application. The returned ARI channel id is persisted to the call_record as its SID (the way a Twilio CallSid would be) so the ack can be attributed back to the call.
  2. Answer → play the summary — on StasisStart (the operator answers), the controller plays the TTS-rendered summary media on the channel (POST /ari/channels/{id}/play). What it names is a local sound file, sound:px-<hash>, and it names it only when the audio has been confirmed onto the box — see the prefetch and its interlock. Both the synthesis and the download onto the box were started when the call was placed, in parallel with the trunk setting the call up; see the render happens at originate.
  3. Gather the keypress — after playback finishes, the controller waits for a keypress. Pressing 1 (ChannelDtmfReceived) runs the shared ack path, plays a prompt saying what the press did, and hangs up once that prompt has finished playing (see the sign-off prompt). An acknowledgement that landed — or a press whose alert-store write failed after authorization, which is recorded as delivered (see known limits) — plays sound:proxima-ack ("Acknowledged. You are now handling this alert."); a press on an alert that had already resolved plays "This alert has already been resolved.", and a press the authorization path refused (an unknown call, an unlinked phone, a user without alerts:write, a missing alert group), or one whose acknowledge the alert service gave up on because the alert kept changing while it was applied, plays "Sorry, that could not be done. Please use the console." (Twilio speaks the same outcomes in words: "This alert has already resolved. Goodbye." and "The alert changed while you pressed. Please try again or open Console.") An acknowledgement that lands flips the alert to acknowledged on the keypress, not when the prompt ends. Pressing 2 runs the escalate path instead — the escalation timer is fast-forwarded, the responder hears which of five things actually happened, and the call hangs up with the alert still open. See pressing 2 to escalate. Any other key is ignored and the gather window keeps running.
  4. No answer → one gentle retry — if the call is never answered within the answer window (30s), the controller re-originates once (re-attaching the same call_record SID). If the retry also goes unanswered, the call is treated as failed and the escalation auto-advances to the next channel.
  5. Answered but no keypress → replay once, then end — if the operator answers but never presses a key within the gather window (12s), the summary is replayed once. If there is still no keypress, the call ends delivered-but-unacked — it does not advance on its own; the normal escalation ack-timeout takes over from there.

Every media-play failure is soft: a render or play error is logged and the call keeps going with the fallback prompt — a broken render service never silences a call.

ARI wire assumptions​

The ARI client makes a handful of concrete assumptions about Asterisk's REST/WebSocket interface. They are unit-tested against a fake, but the live box is the correctness gate (see the live-validation checklist).

One row of this table is no longer an assumption. It is a disproved claim, and it is left here struck through rather than deleted, because the header's own promise — the live box is the correctness gate — is precisely what collected the evidence, and a table that quietly loses its wrong rows teaches nobody that it can have them.

Was assumedWhat the live box didWhere it is written down now
Asterisk fetches an https:// media URI itself, via res_http_media_cacheRefused the play at submit, scheme is unsupported, on a real page to a real phone on 2026-08-25. The URL was reachable from the box — curl returned 200 and 44,418 bytes over verified TLS. The fetch was never attempted.Why sound: and not a URL

The current rows:

InteractionAssumption
OriginatePOST /ari/channels with query-param form; response body is {"id": "<channel-id>"}.
Play mediaPOST /ari/channels/{id}/play?media=<uri>. The value is always one of the six schemes — sound:px-<hash> on the confirmed path, sound:proxima-alert otherwise. Since 2026-09-08 the unconfirmed path no longer submits the renderer's https:// URI: doing so produced a silent page, because a submit ARI accepts and a submit ARI refuses fail in different places and only the second is visible here. See what a page actually submits. Nothing ever prefixes sound: onto an https:// URI — Asterisk would read the whole thing as a file name, find nothing, and play silence.
Media schemesres_stasis_playback accepts exactly six: sound:, recording:, number:, digits:, characters:, tone:. This is read off the binary, not assumed — the module carries those six prefixes and zero ast_bucket references. It is the one row in this table with a live counter-example behind it.
HangupDELETE /ari/channels/{id}; a 404 is treated as success (channel already gone).
Events WSwss://…/ari/events?app=<stasis-app>&api_key=<user>:<password> — ARI carries basic-auth as api_key=user:password on the WebSocket.
Event decodeEvents are keyed on type, channel.id, and (for DTMF) digit.
Playback lifecycleA play Asterisk accepts creates a playback and eventually emits PlaybackFinished; a play it refuses creates none and emits nothing. The call's gather phase — and therefore its ack-timeout — is armed only on that event, so a refused play with no local-prompt retry would leave the call open indefinitely. This is the one assumption in the table that a unit test cannot even model, because the fake would only be agreeing with itself; it is why the refused-play fallback exists and it needs the live box.
Late keypad deliveryA ChannelDtmfReceived frame Asterisk has already sent is still readable after the channel is destroyed. The main case does not need this to be true: the event loop is a synchronous reader, so a keypress that arrived while the loop was busy is already sitting in the receive buffer, and TCP does not retract sent frames — it is delivered when the reader gets to it, hangup or no hangup. The assumption is only load-bearing for a keypress Asterisk had not yet sent when the channel went away. If Asterisk drops those, the late-ack handling is simply inert for that case and the old behaviour (a discarded keypress, now counted) stands — it cannot make anything worse. Watch proxima_voice_late_dtmf_total{outcome="honored"} on the live box.

Event-WebSocket reliability​

The ARI event WebSocket is the connection the backend drives calls over, and an idle WebSocket can be dropped by an intermediary (or the box itself) without a clean close — so the backend would keep an open socket that silently receives nothing. If that happens mid-page the call answers but is never driven: no prompt, no ack, a silent drop. To detect it, the transport runs a client keepalive — it sends a WebSocket ping on an interval and enforces read/write deadlines, so a dead connection surfaces within seconds and is reconnected automatically. The knobs (PROXIMA_ASTERISK_WS_*, see Environment variables) tune the ping interval and the read/write timeouts.

Two backend metrics make this observable on /metrics: proxima_voice_ari_connected (gauge — 1 when the event WebSocket is up, 0 while disconnected/reconnecting) and proxima_voice_ari_reconnects_total (counter — reconnect churn). ari_connected is distinct from proxima_voice_trunk_up (that is the SIP trunk over REST, and can read 1 while the event socket is down): alert on a sustained ari_connected=0, because the backend cannot drive any call while it holds.

Multi-replica: leader-elected paging (P2-B)​

On 2+ backend replicas the ARI event socket has a co-location problem: if every pod ran the Stasis event loop they would replace each other's WebSocket, and each pod's call state is in-memory — so a call originated on pod A whose ARI events land on pod B is silently dropped (the reason Phase 1 was validated at a single replica). Phase 2 fixes this with a single elected leader. A session advisory-lock VoiceLeader (one pod holds the lock; Postgres auto-releases it when that pod dies) owns the ARI event WebSocket and the VoiceController, and all voice originations are routed to that leader over NATS: each originate site — the notification-chain voice step and the verify-phone handler — publishes a VoiceOriginateRequest (a call_record id plus the already-resolved phone/DID and purpose — no secrets, no verification code) to a VOICE JetStream stream on subject proxima.voice.originate, and the leader alone consumes it and places the call. Origination and its ARI events are therefore always on the same pod, so the "voice event for unknown channel" case drops to ~0.

The leader is stamped with proxima_voice_leader (1 on the leader, 0 elsewhere) — see the observability standard (Voice paging metrics) for the routing metrics and the alerts (sum(proxima_voice_leader) != 1 for split-brain or no leader; a persistent published-vs-consumed gap for a down consumer).

Failover is deliberately simple and backstopped. On a leader crash the lock frees and another replica acquires it within PROXIMA_VOICE_LEADER_ACQUIRE_INTERVAL (default 5s). Queued originations redeliver to the new leader and are idempotent (each call is stamped Nats-Msg-Id = call_record id, and the consumer skips a call_record that already has a SID), so a redelivery is never a double-page. A single in-flight call — already originated, awaiting DTMF — is lost when its leader dies; that is the one residual gap, and it is covered by the P0 backstop: the escalation engine auto-advances to the next channel and the dead-man's switch re-fires a missed page. Durable call state is intentionally out of scope — the pager stays correct without it.

Pressing 2 to escalate​

A voice page offers two keys. 1 acknowledges — "I have this". 2 escalates — "not me, get the next person now" — and fast-forwards the escalation timer instead of waiting out the ack timeout.

Pressing 2 does not acknowledge, and there is no undo

Two facts a responder needs before they press it:

It is not an acknowledgement. The alert stays open and stays firing. ack_source is left NULL, the alert group is not marked acknowledged, and nothing about the incident is recorded as handled. An escalate that quietly acked would be a silenced page: the chain it was asked to advance would stop, nobody else would be called, and the record would say a human took it.

A mis-pressed 2 cannot be recalled. By the time the prompt finishes, the next step has already fired. That is why the key is 2 and not something adjacent to 1, and why every prompt names both keys explicitly rather than leaving the option to be discovered.

What a responder hears, and what each one means​

The keypress is mapped to exactly what happened and never to a friendlier summary of it. Five outcomes, five prompts:

You hearWhat actually happenedWhat to do next
"Escalated."The escalation chain advanced. That, and nothing beyond it.Nothing. Note the "nothing beyond it" below.
"This alert has already been acknowledged."Somebody acknowledged it while your call was live. A healthy race — ack wins by construction.Nothing. It is somebody's.
"This alert has already been resolved."The incident is over, and you are being paged about it anyway.Worth reporting: a page for a closed alert is a routing or timing defect, not a nuisance.
"There is no further escalation for this alert."The ladder really is exhausted, or was never armed. Only those two verdicts reach this sentence — it is the one answer that tells a responder to stop worrying, so it is spoken only when stopping is right.Nothing further will happen on its own. If the alert still needs someone, act in the console.
"Sorry, that could not be done. Please use the console."Your press could not be acted on: an unknown call, a phone not linked to a Console user, a user without alerts:write on that client, a missing alert group, a store error — or an advance the engine declined because it had already moved (stale_step, race_lost) or because its snapshot is broken (corrupt_snapshot).Use the console. In the declined-advance cases the escalation may well be running without your press, so this is "look", not "nothing is happening".
This table used to say five verdicts collapsed into "no further escalation"

They did, and on 2026-09-12 a live page showed the cost. A responder pressed 2, the engine returned stale_step, the call said "There is no further escalation for this alert" — and the escalation timer paged the next person forty-five seconds later. The sentence asserted the opposite of what the system then did.

stale_step means the step fence rejected a late press because the snapshot had already moved forward. race_lost means the CAS lost to a timer that had already fired. In both, something is demonstrably in flight. corrupt_snapshot is an engine bug, which is no basis for any claim about what remains. None of the three is an exhausted policy, and only an exhausted policy earns a sentence that tells a responder the incident is closed. Those three now speak the failure prompt instead.

And the next day the stale_step itself turned out to be the deeper bug. Making the prompt honest did not make the button work — see why pressing 2 never advanced anything, below. stale_step should now be rare from a keypress rather than universal.

no_timer is left on the "no further" side deliberately even though it is ambiguous — the column is also NULL when the engine has already advanced the timer to now. The exhausted ladder is the common case by a wide margin, and sending every responder with a correctly-finished ladder to the console would trade a rare wrong reassurance for a frequent wrong instruction. The timeline row keeps all five verdicts apart regardless of which sentence was spoken.

"Escalated." promises the chain moved and nothing else. An earlier draft said "Escalated. The next responder is being contacted." That asserts something nothing at that instant can know: the next step may match no route, may find nobody on call, may be gated off by the client's oncall_enabled flag, or may ring a phone that is never answered. Over-claiming on a paging path is the failure this whole line of work exists to end, so the prompt stops at the one thing that is true.

"Acknowledged" and "resolved" are kept apart on purpose. Both are good news; they are not the same news, and the responder's next move differs. An alert that is over while somebody is still being paged about it is a different fact from one that has just been taken.

One keypress can mean more phones ringing, not fewer​

Escalation advances the step, not one target's slot. If the current step pages three people, the other two calls keep ringing while the next step also fires. That is additive by design — escalation exists to widen the net, and truncating live calls to narrow it again would be worse — but it means pressing 2 never reduces the number of phones involved.

Where the evidence lands​

Every press is recorded, including the ones that changed nothing — an escalate against a chain with no depth left is direct evidence that the policy has no depth left, which is exactly what an operator needs before the next incident.

  • call_record.digit records the key: 1 for an acknowledge, 2 for an escalate. It is the only column that says a key came back over the trunk on a call that escalated, since ack_source stays NULL.

  • alert_group_log, action escalation_requested carries what came of it: metadata->>'result' is the spoken outcome above, and metadata->>'advance_outcome' is the one that matters when the answer was "no further escalation" — it names which of the five engine verdicts it really was (not_armed, no_timer, corrupt_snapshot, stale_step, race_lost). failure_reason does the same job for the failure prompt. requested_by is the responder, and a voice row's actor_type is system.

    The row is not voice-only. Since the Telegram alert card gained ⚡ Escalate, a tap writes the same action with the same result and advance_outcome vocabularies, and three differences. metadata->>'source' is telegram; a row with no source is a voice press-2. actor_type is user, with actor_id and requested_by the linked Console user who tapped, because a Telegram row is only ever written for an authorized, linked person. And there is no ledger_step: a tap is not tied to a paging call, and a 0 there would read as the initial dispatch. The timeline tells the two apart by source, not by actor_type, so the one action still renders one way. A query meant to count keypresses must filter on metadata->>'source' IS NULL.

  • proxima_voice_escalate_requests_total{outcome} counts keypresses by what they were told. A rising no_further is the interesting one: responders are asking for help their policies cannot give. See the metric registry.

The paging-outcome aggregate keeps escalates out of the acknowledged bucket. The readiness page's answered_with_ack used to mean any digit, so a press-2 would have rendered as "answered and acknowledged on the keypad" — a page nobody acknowledged, displayed as one. Non-acknowledging keypresses now have their own bucket, answered_other_keypress. See ack_source — did the voice channel work?.

Both prompts name both keys, including the degraded one​

A page never speaks a dedup key. A generic sender has no alertname concept, so ingestion sets both fingerprint and alert_name to the sender's dedup key. Spoken, a key like lv1-resolve-check-20260917 is read by the TTS engine as a cardinal number — "twenty million two hundred sixty thousand nine hundred seventeen" — which tells a responder nothing, and because the name leads the page it spent the 125-character bound before the sentence that mattered. RenderSummary suppresses the name when it equals the fingerprint, which is the exact test for "this is an identity, not a name" rather than a guess at what a machine-looking name is: an Alertmanager alert carries a real name and a hash, which differ, so HighCPUUsage is still spoken. When the name is suppressed the summary leads instead — the same precedence the alert card uses for its title — and the host clause rides with whichever clause leads, so a page with no usable name still says where.

The spoken page reserves the keypad instruction inside its byte budget rather than appending it. Every rendered page ends with "Press 1 to acknowledge, 2 to escalate." — 38 bytes held back out of the 125-byte page bound before the alert summary is bounded into what remains. Reserved, not appended, and the difference decides who learns key 2 exists: the bound cuts from the end, so a sentence appended after the summary is the first thing a cut reaches. Measured on the page shapes the renderer actually produces, an appended instruction survives on a short page, is cut mid-clause on an ordinary one, and disappears entirely on a long one — which would make "was the responder told about key 2" a function of how long the alert name happens to be. Reserving inverts what truncation costs: the tail of the summary clause, which a responder can read in the console, and never the part that tells them what the keypad does. A summary that was cut also gets back the full stop the cut removed, so the synthesizer hears a sentence break instead of running the alert text and the instruction together as one clause.

What that costs, against the budget where it is actually paid. Against the warm render budget — 11.5 s, the 6.5 s measured originate→answer floor plus the 5 s answer-render timeout — every page shape stays inside, an ordinary page going 81 → 120 bytes and ~5.8 s → ~8.6 s, and the 125-byte worst case is unchanged. Against the cold budget — 5 s flat, when the answering call lands on the replica that never did the warm pre-render or after a leadership change — the ~39 extra bytes cost +1.2 s at the fastest measured prose rate and +2.8 s at the worst, which moves a band of pages across the line: a short 34-byte page goes 2.45 s → 5.26 s and now misses where it used to fit. A page that misses the cold budget does not lose its tail; it ends on the content-free proxima-alert prompt, and pages that previously had no dead air can have up to ~2.1 s of it. Watch it on proxima_voice_tts_render_total{result="fallback"}, which is where it is loud.

The trade is deliberate, and it is not paid by the ability to escalate — the fallback prompt names both keys itself, so the one thing the reservation guarantees survives exactly the case the reservation made more likely. What is lost in that band is the alert's own text, which a responder can read in the console. Shrinking the instruction so it fits everywhere was rejected: it would swap a bounded, visible degradation for a silent one where some responders are never told key 2 exists, and a mis-pressed 2 cannot be recalled.

The static fallback names both keys too, and that is not a nicety. When the render service is unavailable the call plays the pre-recorded sound:proxima-alert — "A Proxima alert is active. Press 1 to acknowledge. Press 2 to escalate to the next responder. Check the console for details." — and 2 still works: the DTMF arm does not care which audio was played. A responder who hears the degraded prompt and is never told the option exists cannot use it, in exactly the case where an escape hatch matters most. The fallback carries the longer wording because it is a pre-rendered file and its words cost no render budget.

The five escalate prompts are pre-rendered files too (proxima-escalated, proxima-already-acked, proxima-already-resolved, proxima-no-further-escalation, proxima-escalate-failed), committed in .wav, .alaw and .ulaw; the words live in infra/asterisk/sounds/prompts.tsv, which the renderer, the deploy and the tests all read.

What 2 does on a call that is not an alert page​

Nothing, and that is a property under test rather than an accident. A phone-verification call collects a code and a readiness drill takes 1 to confirm; both are routed before the alert branch, so a stray 2 there never reaches an escalation. A 2 on a drill falls through to the gather timer and the prompt replays.

A 2 that arrives after the call has already ended is recorded and not acted on. A late 1 is honoured — the ack path is idempotent and over-paging is the safe direction — but nobody is on the line to hear which of the five escalate outcomes occurred, and being told what actually happened is the whole point of the feature. Late presses are counted on proxima_voice_late_dtmf_total{outcome="discarded"}.

Not yet exercised on a live call

Nothing on this path has been validated against a real trunk. Nobody has heard the five prompts — they are verified by measurement, not by listening — and voice paging is dark for an unrelated reason (the SIP trunk is de-registered and the provider is blackholing the box). See live validation.

Phone verification (prove the phone rings — and acks)​

A paging phone is only useful if it actually rings and the operator's keypress registers. Phone verification proves both in one call. It deliberately reuses the same inband-DTMF path as an ack, so verifying a phone doubles as an ack smoke-test: a number cannot verify unless its keypresses surface to Asterisk exactly the way a real press 1 to acknowledge would. That directly de-risks the trunk's #1 historical gotcha (a carrier that answers but silently swallows DTMF).

The flow​

  1. Open your On-call contact section (your own Profile, or — for a users:write admin — a user's detail page). A set-but-unverified number shows a muted "Unverified phone — voice paging not guaranteed" banner with a Verify phone button.
  2. Click Verify phone. The backend generates a short-lived 4-digit code, stores only its hash (the plaintext is returned once, for on-screen display, and is never logged or persisted in plaintext), and places a purpose=verification call to your on-call phone over the same provider and caller-ID the alert path uses.
  3. A dialog opens showing the code prominently ("We're calling {phone}. Enter this code on the keypad"). Answer the call and key the code — the call plays a sound:proxima-verify-prompt prompt and collects the digits over the same DTMF collector the ack path uses.
  4. On a match (within the 5-minute expiry and attempt cap), phone_verified_at is stamped. The dialog — which polls the profile while open — flips to a "Phone verified" success state, the banner disappears, and a green "Verified" badge replaces the "Unverified" one.

The endpoint (POST /api/v1/users/{userID}/oncall-profile/verify-phone) is self-or-admin: you may verify your own phone; a users:write admin may verify anyone's (subject to the same tenant-isolation guards). It returns 409 when the profile has no phone, and 502 when no voice provider is configured (nothing is stored and no code is returned on failure).

Ops prerequisite (co-driven)

Verification places a real call, so it needs the same live stack as alert paging: a configured provider (once PROXIMA_VOICE_PROVIDER's provider is up, verification works through the shared provider handle) and the two verification prompts — sound:proxima-verify-prompt and sound:proxima-verify-ok — provisioned in Asterisk's sounds directory alongside the ack/alert prompts. In an environment with no provider configured (e.g. CI), the endpoint returns 502 by design and no call is placed.

The verified gate (PROXIMA_VOICE_VERIFIED_GATE)​

Once phones can be verified, the voice resolver can act on phone_verified_at. The gate ships non-regressing and is flipped only after a verification drive:

PROXIMA_VOICE_VERIFIED_GATEBehavior
warn — defaultAn unverified phone is still paged (exactly as before verification existed), and each such decision bumps the proxima_voice_unverified_page_total metric, logs a WARN, and shows the profile banner. Non-regressing — no page that would have gone out before is suppressed.
enforceAn unverified phone is skipped for voice; escalation falls through to the other channels in the chain (Telegram, etc.). The same metric is bumped (now counting suppressed pages).

Roll-out order matters. Flip to enforce only after a verification drive, and alert on proxima_voice_unverified_page_total first — under warn that counter is your exposure gauge (how many pages would be dropped the moment you enforce). Enforcing before your on-call roster has verified their phones would silently strip voice from whoever hasn't verified — the opposite of a trustworthy pager. The gate mirrors the retention-off-by-default and P0 observe-then-flip pattern: ship observable, watch the counter, then enforce.

TTS: what the callee actually hears​

A voice page speaks the alert. Until the render service was deployed, every voice page played one pre-recorded prompt: the phone rang, someone answered, and heard that something had fired but not what, which host, or how bad. The callee now hears the alert itself:

"Priority one alert. HighCPUUsage on db-prod-01. CPU has been above 95% for ten minutes."

The spoken text is built by a shared allowlist renderer (RenderSummary) that reads only the priority word, the first alert's name, the host, and a length-capped summary/description annotation — never raw labels or arbitrary annotations, so no secrets or PII are ever spoken.

Where synthesis runs​

Synthesis runs in proxima-tts, a separate service in the console-system namespace (Kokoro, ONNX runtime, the voice baked into the image). It is not a process on the Asterisk box, and it is no longer Piper. Two hops are involved and they are worth keeping apart:

  1. The backend POSTs the spoken text to the render service inside the cluster and gets back a complete, short-lived https:// media URI.
  2. A fetcher process on the Asterisk VM downloads that URI and writes the audio into Asterisk's own sounds directory, before the callee answers. Asterisk itself never sees the URL; it plays a local file. See the prefetch and its interlock.
POST {PROXIMA_VOICE_TTS_RENDER_URL}
Content-Type: application/json

{ "text": "<the spoken summary>" }

→ 200 OK
{ "media_uri": "https://tts-console.prxm.uz/audio/<capability-token>.wav" }

The URI is a capability: everything the fetch needs is signed into the token, it carries no client, alert or file identifier, and it expires. /render itself is ClusterIP-only — it takes a client's alert text, so it is never exposed outside the cluster; only /audio is. The only client that ever dereferences that URI is the fetcher on the voice VM — which is why the backend logs it redacted to scheme and host: the path segment is a bearer credential, and a log line is the one place it would outlive the fifteen minutes it was scoped to.

Why sound: and not a URL​

The obvious design is to hand ARI the render service's https:// URI and let Asterisk fetch it. Asterisk certified-22.8 cannot. res_stasis_playback accepts exactly six media schemes — sound:, recording:, number:, digits:, characters:, tone: — and holds zero references to ast_bucket, the mechanism that would perform an HTTP fetch. res_http_media_cache does register http and https bucket schemes and is loaded on that box, which is exactly why this looked safe on paper; nothing in the ARI playback path consults them.

That is read off the binary rather than inferred, and it was confirmed the expensive way — by a real page to a real phone on 2026-08-25:

ERROR res_stasis_playback.c:376 play_on_channel:
Attempted to play URI 'https://tts-console.prxm.uz/audio/….wav'
on channel 'PJSIP/sarkor-trunk-00000008' but scheme is unsupported

The box itself was fine. From the voice VM, curl on that exact URL returned 200, 44,418 bytes, TLS verified. The fetch was never attempted — the URI was refused at submit.

A refused play does not degrade. It hangs.

This is the finding worth carrying to any future design on this path, and it is not obvious. A play Asterisk refuses creates no playback object, so no PlaybackStarted and no PlaybackFinished ever arrive — and the gather phase, and with it the call's ack-timeout, is armed only on that event. The call sat open for 94 seconds while the callee heard nothing and eventually hung up, and the row recorded tts_result = rendered, because the render genuinely had succeeded.

So the rule this whole mechanism is built around is not "prefer a local file". It is: never name a file we have not been told exists. "It is probably there" is precisely the reasoning that produced those 94 seconds.

The prefetch and its interlock​

Asterisk plays sound:<name> out of its own sounds directory perfectly well — that is how the pre-recorded prompt has always worked. So the audio is delivered to the box instead of being fetched by it.

  ① render ──► proxima-tts ◄── ③ GET /audio/<token>.wav   (outbound from the VM, TLS)
▲ ▲
backend ──② "needed soon" ──► fetcher (on the VM, OUTBOUND ONLY)
▲ │ ④ write /opt/asterisk/sounds/px-<hash>.wav
└───────── ⑤ ready ───────────┘
backend ──⑥ ARI Play(sound:px-<hash>) ──► Asterisk ──► ☎

Pull, not push, and that is a security decision rather than a convenience. The pager box grows no inbound surface at all: no new port, no second TLS certificate, no firewall rule. The fetcher opens connections and never accepts one; it has no listening socket, no health endpoint and no metrics endpoint. Its liveness is a heartbeat file that the container health check re-reads by running the binary again with -check.

The cost is about a second of latency, inside a window that is already 6.5–11 s wide, because the work is queued at originate rather than at answer.

Three columns carry it across two machines​

ColumnWritten byMeaning
sound_namethe backend, at originateThe local file to play: px-<sha256(spoken text)[:16]>. A reservation, not a permission.
sound_urlthe backend, at originateThe render service's media_uri. Only the fetcher ever dereferences it — it is explicitly not a URI Asterisk can play.
sound_ready_ata fetcher, and nothing elseSome box has the bytes. Not on its own permission to name the file — see below.
call_record_sound_boxthe fetcher on one boxOne row per (call, box): this machine has the bytes. On a multi-box deployment this is what authorizes naming the file.

The rule is stated in sound_ready_at's own column COMMENT, so \d+ call_record teaches it without anyone opening a design document. None of the three carries a CHECK constraint, for the same reason none of the evidence columns does: these rows are written on the paging path, and a constraint violation there does not reject a bad value — it costs a page.

Content-addressed names, and what that buys​

The name is derived from the spoken text, so two callees paged about the same alert resolve the same name: one file, one download, one write. A redelivered origination (a leader failover mid-page) re-derives the same name and rewrites the row with the value it already had, rather than reading as a changed name — which matters, because changing the name clears sound_ready_at. A confirmation is about a file, not about a row: when the alert really has moved on, the text differs, the name differs, and dropping the old confirmation is correct, because the file on the box is no longer this page's audio.

Two texts that differ only past PROXIMA_VOICE_TTS_MAX_TEXT_CHARS produce two names for identical audio — a duplicate download and nothing else. The dangerous direction, one name standing for two different audios, is held off by the render service being content-addressed on sha256(voice_id \0 bounded_text), and it holds only while the voice and that character bound are fixed. Change either inside one 15-minute window and the same name could stand for two renderings of the same words; the consequence is cosmetic (never a wrong alert, because the words are what the name is derived from), which is why the fix is recorded rather than taken.

Two boxes: delivery and permission are per MACHINE​

A global "ready" flag and a per-machine file is a silent page

sound_ready_at answers "did some box get this file". A call answers on one box. With a single Asterisk machine those were the same sentence; box failover made them different, and on 2026-10-04 17:21:54 the difference was audible in production:

dial.c:                PJSIP/skyline-trunk-0000000e answered
file.c: WARNING File px-f02b1f583a4c4ae9 does not exist in any format
res_stasis_playback.c: WARNING Playback failed for sound:px-f02b1f583a4c4ae9
…11 seconds… Playing 'proxima-ack.alaw'

asterisk01's fetcher had confirmed that sound five seconds earlier. The call answered on asterisk02, which never downloaded it. The callee heard nothing where the alert text belonged — not the generic prompt, nothing — because ARI accepts the sound: scheme and the failure happens later, inside Asterisk, where the submit-time fallback cannot see it.

Two mechanisms turned that into a coin flip per page, and both are fixed:

  • Delivery. ListPendingSounds filtered sound_ready_at IS NULL, so the first box to confirm removed the row from the other box's work list — the second box never saw the work. The fetcher now names its machine (?box=), and the per-box query's "already done" means this box has it. Both boxes therefore receive every sound.
  • Permission. Playback now requires a call_record_sound_box row for the box the call is actually on, resolved from the path the call was placed on — so a call that failed over is checked against the machine it failed over to.

Every "cannot verify" answer plays the generic prompt: a store error, a path with no configured box, a channel with no recorded path. The asymmetry is deliberate — a false no costs a content-free page the callee can still act on, a false yes costs the eleven seconds above.

Box identity is explicit configuration on both sides, never derived: PROXIMA_VOICE_PATH_<NAME>_BOX on the backend and the box variable in the fetcher's own environment. It is not inferred from the path or trunk name because a machine can host more than one carrier (asterisk01 carries Skyline as well as Sarkor when Skyline's credentials are present), and an inferred box is the same unchecked premise that produced the incident.

A single-path deployment takes none of this: below two configured paths there is one machine, the per-box check does not run, and playback reads sound_ready_at exactly as it did. Failover is what introduced the distinction, so a deployment that cannot fail over sees no change.

The per-box rows are keyed per call, not per sound name. A name-keyed row would be cheaper and would dedupe across calls — and it would outlive its file, because the reaper below deletes audio after about an hour. It would then tell a later call with the same text that its box holds audio that was deleted: the exact "the file is probably there" failure this whole feature exists to remove, reached through our own cache. A per-call row is read within minutes of that call's creation, so a reaped file simply shows as pending again and is re-downloaded.

Where to look when a page rings but says nothing: On-Call → System → Trunk Health shows, per path, when that path's box last confirmed any audio. A configured box with no confirmation is the symptom of a box-name mismatch — the fetcher on that machine runs, polls and downloads happily, nothing errors, and every page it carries plays the generic prompt.

Reaped by age, never by size​

The fetcher deletes files older than an hour and matching px-[0-9a-f]{16}, and the second half of that is not tidiness: the committed prompts live in the same directory and are older than everything in it, so an age sweep without the name filter would delete them — turning every degraded page into a silent one, which is the worst trade this system can make. Names are content-addressed, so a reaped page file is simply re-fetched.

Time-based rather than size-based is deliberate: a size-based reaper deletes files exactly when the directory is largest, which is during a page storm, which is when calls are ringing. That converts the one accepted race in this design — a file reaped between confirmation and playback — from vanishingly unlikely into routine.

What a page actually submits today​

Three outcomes, and it is worth knowing which one a given call took:

Row stateWhat the backend submits to ARIWhat the callee hears
tts_result='rendered', sound_ready_at setsound:px-<hash> — one submitthe alert
tts_result='rendered', sound_ready_at NULLsound:proxima-alert — one submit, after a bounded waitthe generic prompt
tts_result='fallback'sound:proxima-alert — one submitthe generic prompt
The middle row used to say something false, and a real page paid for it

This table previously described the unconfirmed row as "the renderer's https:// URI, which Asterisk refuses, then sound:proxima-alert — two submits", and called that refusal "the degradation working". On 2026-09-08 a page reached a real phone and the callee heard nothing at all for the whole call — not the alert, not the generic prompt — while call_record recorded tts_result = rendered, because the render genuinely had succeeded.

What is certain. The value submitted was the render service's https:// URI; res_stasis_playback cannot decode it; the local-prompt fallback did not rescue the call. Those three follow from the code path and the outcome, independent of any log.

What the table claimed, and why it cannot be the whole story. If the play had simply been refused at submit, the fallback would have caught it synchronously and the callee would have heard the generic prompt — the documented August behaviour, above. They heard silence instead. Three readings survive that, and the evidence does not choose between them: ARI accepted the play and it failed later inside Asterisk where the backend cannot see it; the fallback did not fire; or the fallback fired and proxima-alert is not playable on that box — which produces exactly "silence followed by a normal conclusion", the signature described below for a sound: name with no file behind it. All three say the same thing about the design: the caller's view of a play is not the same as the play's fate, and the branch was relying on a refusal it is not entitled to assume.

The third reading is not closed by this fix, and it is worth ruling out

If the static prompt file is missing or in the wrong format on the voice VM, then every fallback path on this page is silent and always has been — the fix routes more calls to a prompt that says nothing. It is cheap to check and it is the first thing to check after a silent page:

ls -la /opt/asterisk/sounds/proxima-*        # on the voice VM
# and in the box's own log, the signature of a named-but-missing file:
grep -E "does not exist in any format|Playback failed for sound:" /var/log/asterisk/messages.log

A File proxima-alert does not exist in any format line is the whole answer.

Two things changed, and neither depends on which reading is right:

  • The unconfirmed path no longer hands ARI a URI it cannot decode. It plays the static local prompt. Every path now ends at one of the six schemes, so there is no longer a submit whose outcome has to be reasoned about at all.
  • The prefetch gets a bounded chance to win. On that call sound_ready_at landed after the play went out — a fifth of a second after, by the row's own timestamps. The answer path now waits for in-flight audio, in two stages, and re-reads the row while it waits — see the prefetch wait.

The regression test's fake ARI accepts every play, deliberately: a fake that refuses the URI proves the fallback works and says nothing about this defect. That is how the suite stayed green while production was silent. :::

Failure modes — all degrade, one does not​

What failsResult
the render servicenothing is queued → no sound_name → the generic prompt
the fetcher is down, wedged, or never deployedsound_ready_at stays NULL → the generic prompt
the download 404s, truncates, or the disk is fullnever confirmed → the generic prompt
the confirmation is refused or lostNULL → the generic prompt
the callee answers before the download finishesthe answer path waits up to 1 s; if the bytes land, the alert — otherwise NULL → the generic prompt
the file is reaped between confirmation and playbacksilence

The last row is the one real gap, and it is accepted rather than closed: it needs the file to vanish in the seconds between the fetcher confirming and the callee answering, against a one-hour horizon. It is also the reason the reaper is time-based, above.

That "one-hour horizon" was not true on every path, and a whole-branch review found where. Names are content-addressed, so a repeat page about the same alert — an escalation step, a REPEAT, a flapping group — finds the earlier page's file already on disk and confirms it without downloading. That shortcut accepted a file of any age below the horizon, so the headroom at the moment of confirmation was not an hour, it was horizon − age, and it could be zero: a file 59 minutes old could be confirmed and swept moments later, and the callee would get the silence this row describes. The fetcher now refreshes the file's mtime as part of re-confirming it, so the reaper measures last usefulness rather than first download and every confirmation — cached or fresh — starts with a full horizon. A file whose clock cannot be reset is not reused at all; it is re-downloaded, which resets it as a side effect. The sweep also runs at the top of a poll rather than deferred to the end, so it can never follow the confirmations it would invalidate.

The atomicity that keeps the rest of the table honest is in the fetcher: it downloads to <name>.wav.part — not a name Asterisk can resolve, so a partial download is unnameable by construction — then fsyncs the file, renames it, fsyncs the directory, and only then confirms. A truncated file that plays is worse than no file, because it looks like success.

The render happens at originate, not at answer​

Synthesis is kicked when the call is placed, in parallel with the trunk setting the call up, and it is detached — it cannot delay the dial. Originate-to-answer measured 6.5–11.0 s on four real production calls, which is the window the render is meant to run underneath. By the time the callee picks up, the audio is normally already rendered and the answer-time render is a cache hit.

The prefetch is queued in the same place, and only on success. When the warm render returns a URI, the originate path reserves this page's sound_name and records its sound_url — a reservation, never a permission (see the prefetch and its interlock). That write is strictly best-effort and detached from the dial exactly as the render is: losing it costs the page its alert text and nothing else. Losing it is counted, though, unlike its sibling evidence writes — proxima_call_record_evidence_write_failures_total{stamp="sound_name"} — because this is the one stamp whose loss somebody can hear.

If the render is still running when the call is answered, what happens depends on which of the two render replicas the answering call reaches, and both cases are worth knowing:

  • The replica that pre-rendered. The answering call joins the render already in progress rather than starting a second one (the service is content-addressed and renders each distinct page once). The cost is a few seconds of silence before the alert is spoken — not a lost page.
  • The other replica. The cache is local to each pod, so a cold replica starts the render from zero against the answer-time client's 5 s timeout. Short-spoken pages still make it; the slower shapes below do not, and the callee hears the pre-recorded prompt instead, recorded as tts_result = 'fallback' with reason = 'transport'. The Service sets sessionAffinity: ClientIP so that both renders normally come from the same backend pod and reach the same replica; a change of voice leadership between placing and answering the call defeats that, and such a page is a genuine cold miss.
The answer-time render blocks other calls on the same connection

Synthesis at answer time runs on the single ARI event connection, so while it waits, no other call's answer, playback-finished or keypad event is processed — in a fan-out that delays a second callee's audio and their acknowledgement by the same amount. The worst case is about 25 seconds, not five: the render can take up to 5 s, the playback request that follows it up to 10 s, and on the degraded path the local-prompt retry another 10 s. It takes a cache miss and a refused playback together to reach that, which is what the pre-render exists to prevent. Since 2026-09-08 the second half is rare again. The unconfirmed path used to submit the renderer's https:// URI on every successfully-rendered page, so the 10 s retry arm was armed on all of them; it now plays the local prompt directly — one submit, no retry — and only the cache miss remains. The unconfirmed path does add a bounded wait of up to 1 s for in-flight prefetched audio, which is an order of magnitude less than the arm it replaced. On the confirmed path there is one submit, no retry and no wait. If you see several callees on one alert reporting slow or laggy calls, look for a run of tts_result = 'fallback' first — the degraded path is also the slow one.

25 seconds is longer than the 12-second window a callee has to press 1, and that window is timed off the event loop, so it keeps running while the loop is busy. Another call can therefore be concluded delivered-unacked and hung up while its callee's keypress is still queued. That keypress is not discarded: when the loop catches up, a 1 for a call this pod concluded in the last minute still acknowledges (or, on a readiness drill, still confirms), and proxima_voice_late_dtmf_total{outcome="honored"} counts it. If that press acknowledges nothing — the shared authorization path declines it (an unknown call, an unlinked contact, a denied user), the alert had already resolved, or the alert kept changing while it was applied — it counts as refused instead, because the counter reports what the actor returned rather than that it was called. A keypress that never reached the actor at all is discarded, and one for a channel this pod never concluded is unattributable. None of the four is a silent drop. A run of honored means the ARI loop is being blocked long enough to outlast the ack window — read it beside proxima_voice_tts_render_total{result="fallback"}, because the degraded render path is what produces the long block. A late ack can arrive after the escalation ack-timeout has already advanced and paged the next person; that is over-paging, the safe direction, and the same thing a Telegram or web ack arriving at that moment would do.

The queued case rests on nothing exotic: the event loop is a synchronous reader, so that keypress was already sent by Asterisk and is sitting in the receive buffer unread. A keypress Asterisk had not yet sent when the channel was destroyed depends on whether it sends one at all — see "Late keypad delivery" in the ARI wire assumptions table, which names the live box as the gate. If it does not, that case is exactly as it was before, and counted.

How long a page takes to render depends far more on what it says than on how long it is. At the default 125-character budget, ordinary prose and hostnames render in about 4 s, a pod name or an IP in about 5.5 s, and a summary carrying UUIDs or image digests in 7–9 s, because identifiers are spoken character by character. PROXIMA_VOICE_TTS_MAX_TEXT_CHARS is a proxy for that time budget, not a bound on it; raising it lengthens exactly the pages that are already slowest, and no value of it makes every page fit — which is why the fallback below, not the budget, is what guarantees the call says something. Text over the limit is cut at a word boundary, never mid-hostname — a truncated hostname sounds real, is not, and sends someone to the wrong machine.

🔴 The no-answer retry reaches the carrier but may never ring​

onAnswerTimeout re-originates once when a callee does not pick up. On a live no-answer page on 2026-09-12 the retry went out and the phone never rang:

leg 1  20:11:49.464  Called sarkor-trunk/sip:+998…
20:11:51.785 is ringing <- the handset alerted
20:12:19.546 destroyed <- answerTimeout, 30s
leg 2 20:12:19.556 Called sarkor-trunk/sip:+998… <- retry, 10ms later
20:12:20.476 making progress
(no "is ringing" line) <- the handset was NEVER alerted
20:12:29.104 hangup, ~8.6s later

Leg 2 got early media (183 Session Progress) and died without a 180 Ringing. That is the shape of the carrier answering on the subscriber's behalf — an announcement, a silent reject, or throttling of a second call to a number that just went unanswered, which the Sarkor PoC already flagged: "post-rejection throttling can kill alerting mid-incident."

Why this matters more than it looks. A no-answer page is supposed to get two chances. If the second never rings, the pager has one, and the difference between "they missed it" and "nobody was reached" is exactly the difference the retry exists to close.

It is NOT diagnosed. Asterisk's default verbosity does not log the SIP responses, so the carrier's actual reply on leg 2 is unknown. To capture it, reproduce with PJSIP logging on:

# on the voice VM, BEFORE firing the test page
docker exec asterisk asterisk -rx "pjsip set logger on"
# fire a page and let it ring out past answerTimeout (30s) without answering
docker exec asterisk sh -c "grep -a 'SIP/2.0' /var/log/asterisk/full | tail -60"
docker exec asterisk asterisk -rx "pjsip set logger off" # it is verbose; turn it off after

Read leg 2's response to the second INVITE. A 183 with SDP and no 180 means the carrier is answering instead of the handset; a 486/603 means an explicit reject; a 503 points at throttling. Each implies a different fix — a longer gap before the retry, a fallback trunk, or accepting that this trunk gives one ring per page.

Why pressing 2 never advanced anything​

Until 2026-09-12, the escalate key had never once advanced an escalation — not intermittently, but on every ladder whose next step is a wait, which is the normal shape of one.

The keypress was routed through AdvanceNow, which is written for a failed voice page: it fences the callback to the step that page came from, so a late callback cannot "skip the live step's wait". For a human pressing escalate, skipping that wait is the entire point of the button — the fence is backwards for that caller.

It fired every time, because the engine moves faster than a phone rings. PlanStep increments Seq on every step and a notify step returns NextAt: now, so the engine steps onto the following WAIT immediately after dispatching the page:

18:45:21  step 1 (notify) dispatches the page   call_record.step = Seq 1
the engine steps straight onto the WAIT Seq 2
18:45:27 the responder answers, carrying Seq 1
18:45:44 the responder presses 2 -> fence: 2 != 1 -> stale_step, declined

A keypress four seconds after answering was read as a late callback.

The fix is a second door, not a weaker fence. AdvanceByRequest drops the step fence and keeps every guard that protects the group:

KeptWhy
not-armednothing is armed to move
not-firingan acked or resolved alert is nobody's to escalate
no-timeran exhausted ladder has nothing to pull earlier
corrupt-snapshotan engine bug is not a license to advance
token-fenced CAS, only-move-earliera responder can pull a step earlier, never past the end and never backwards

AdvanceNow keeps its fence for the failed-page path it was written for. A repeat press still advances at most once — that property rested on the token CAS all along, and the second press loses it to the first exactly as before.

The prefetch wait, and the eight-millisecond miss​

The answer path may wait for prefetched audio that has not landed yet, rather than playing the generic prompt the instant it finds nothing on the row.

The first version of this wait asked its question two seconds too early

Measured on a real page, 2026-09-11. Nothing was broken — every component did its job:

20.343  call placed; the pre-render starts and takes the single render slot
24.179 the callee answers, 3.8s in, while that render is still running
24.230 the answer path reads call_record — NO sound_name yet — and starts its own render
25.947 that render returns: it JOINED the pre-render as a single-flight follower, and SUCCEEDED
25.947 the gate consults the record read at 24.230, sees no name, and does not wait
25.955 the pre-render stamps sound_name — 8 MILLISECONDS too late
27.820 the fetcher confirms the audio on the box

The gate judged a copy of the row read before a 1.7-second render, during which the row changed. The callee heard a content-free prompt with their alert 1.9 seconds away.

A first diagnosis blamed render-slot contention and was wrong. TTS_MAX_CONCURRENT_RENDERS is 1 and the loser of a collision gets an immediate 503, so "the answer-time render was refused" looked obvious. Measuring the service settled it: two concurrent requests carrying identical text both returned 200 in 2.57 s, because the render service's own gate (services/tts/app/render_gate.py) already collapses same-key work into one flight — the followers wait on the leader and then read the cache. The probe that appeared to show contention used two different texts, which is the deliberate fail-fast path. The answer-time render was a follower, not a casualty.

The wait re-reads the record as it polls, and runs in two stages:

StageBudgetQuestion
1500 msHas a sound_name appeared at all? A name shows up within milliseconds of the pre-render finishing. If none appears, nothing is coming and the call must not pay for it.
22.5 sIs the audio on the box? Entered only once a name exists — a real promise that bytes are in flight. Covers the 1.9 s measured, with room.

The stage-2 deadline is taken from the moment the name lands, so a name that appears at the very end of stage 1 still gets its full budget rather than the remainder.

Two conditions keep this from becoming a tax:

  • It arms only when audio can exist — the render succeeded, or a name is already on the row (which means the pre-render succeeded even if this one did not, after a pod restart or a transient 5xx). A deployment with no render service, or with a dead fetcher, waits zero.
  • The gather-timeout replay never waits. By the time it fires the audio has been on the box for ten seconds or is never coming, and blocking the shared ARI read loop a second time buys nothing on a call already going unanswered.

The sign-off prompt, and why it was never heard​

Every keypress arm ends by playing a prompt that tells the caller what happened — proxima-ack for an acknowledgement (and, for a press of 1 that acknowledged nothing, the already-resolved or could-not-be-done prompt), one of four outcome prompts for a press-2, proxima-verify-ok for a successful phone verification — and then hanging up.

Until 2026-09-10 none of those prompts had ever been audible

Confirmed by ear on a real page: the responder pressed 1, the ack was recorded correctly, and they heard nothing back. The box's log, 91 ms apart with a ChannelHangupRequest between:

12:18:09.460 file.c: <PJSIP/...> Playing 'proxima-ack.ulaw'
12:18:09.551 res_stasis_playback.c: ...: Playback failed for sound:proxima-ack

An ARI play is asynchronous. Submitting it creates a playback; the audio streams afterwards. Every arm submitted the play and called finish(..., hangup=true) on the next line, so the hangup beat the audio every single time. The press-2 prompts are the sharpest loss: they are the only thing distinguishing advanced, already acked, already resolved and no further step for the person holding the phone, and all four sounded the same — like nothing.

Why it went unnoticed for a month. A code comment explained the failure as a missing file — true on 2026-08-04, when proxima-ack genuinely was absent. It has been on the box since 2026-08-12 and the prompt still did not play. The same Playback failed line has two causes, and the comment asserted the harmless one, so the log looked already explained.

The hangup now waits for PlaybackFinished, with three properties worth knowing:

PropertyWhy
A failed playback still concludes the callAsterisk emits PlaybackFinished for a failed playback too — with a playback id, which is what separates it from a refusal at submit. A missing prompt file costs the words, never the call.
A five-second grace timer bounds the waitA wedged playback, an event lost between pods or a leadership change mid-prompt must never hold a trunk leg open. The longest prompt is about two seconds; five is that plus slack, and it is a ceiling rather than a target.
The hangup happens exactly onceThe event and the timer are two paths to the same act, and Asterisk can deliver PlaybackFinished more than once across pods. One mutex-guarded claim decides it.

A play ARI refuses creates no playback, so no event will ever arrive and there is no audio to wait for — that case hangs up at once rather than sitting through the grace period.

The acknowledgement is not deferred, only the audio. The actor runs and the call concludes synchronously on the keypress; ended_at still excludes this wait. Deferring the ack instead would delay an acknowledgement by up to the grace period, which is how a page somebody already answered gets re-escalated. And the wait is event-driven rather than a sleep, because it would otherwise stall the shared ARI read loop for every other call in a fan-out — the same constraint that bounds the prefetch wait.

Failure is always safe​

Every failure path — a request that will not build, a transport error, a non-200 response, an undecodable body, or a 200 carrying no media_uri — returns PROXIMA_VOICE_TTS_FALLBACK_URI (default sound:proxima-alert) and the call still goes out.

The same applies one step later: if Asterisk refuses to play a media value, the controller plays the local prompt instead — once, never in a loop. That is not only about the audio. A refused play creates no playback, so no PlaybackFinished event arrives, so the call never advances to gathering and its ack-timeout is never armed; playing the local prompt restores that transition too. Since 2026-09-08 the unconfirmed page no longer reaches this arm at all: it plays the local prompt directly rather than submitting a URI and relying on the reaction.

The old silent path is closed. A much narrower one replaces it.

This page used to warn that Asterisk could accept a play and then fail to fetch the audio — an expired capability token, or audio evicted from the render service's cache — leaving the callee in silence while the row read rendered, with Console never learning of it.

That path no longer exists, because Asterisk no longer fetches anything. A download that 404s, expires or truncates now fails on the fetcher, before any call is involved: it simply does not confirm, sound_ready_at stays NULL, and the page falls back to the generic prompt. The failure moved from inside Asterisk, where nothing could see it, to a process that logs it.

What replaces it is the single row at the bottom of the failure-mode table: a file reaped between confirmation and playback. It needs the file to vanish in the seconds between the fetcher confirming and the callee answering, against a one-hour reaper horizon — which is why the reaper is time-based rather than size-based.

There is one thing to keep in mind while reading that as an improvement: naming an unconfirmed file would be worse than the URI ever was. sound: is a scheme ARI accepts, so a name with no file behind it is not refused at submit — the play call returns no error, the local-prompt fallback (which reacts only to that error) never runs, and the callee hears nothing where the prompt should have been.

The call itself does not hang, and that is now evidenced rather than hedged. The box's own log, messages.log:226-228 (2026-08-04), shows a missing sound: file producing File proxima-ack does not exist in any format → Unable to open proxima-ack (format (ulaw)) → res_stasis_playback.c: 1785844257.7: Playback failed for sound:proxima-ack. That last line carries a playback id, so a playback object existed and reached a terminal state — meaning PlaybackFinished was emitted and the gather timer armed. Compare the refusal at messages.log:1018 (2026-08-26): no id, no Playback failed line, no playback object at all. So the cost of naming an unconfirmed file is silence followed by a normal conclusion, not the 94-second hang. sound_ready_at is what keeps that branch unreachable.

The fallback is deliberately not the same kind of value as a rendered URI: it is a file already sitting on the Asterisk box and needs no network at all. That asymmetry is the whole reason it is acceptable to put synthesis across a network from the paging path.

A cluster outage degrades the message, never the call

If proxima-tts is scaled to zero, evicted, unreachable, or slow, the page still happens — the callee's phone still rings, they still hear the generic prompt, and pressing 1 still acknowledges. What is lost is the content of the page, and it is lost loudly: the call record stamps tts_result = 'fallback' with a reason, and proxima_voice_tts_render_total{result="fallback"} rises. A degraded page is acceptable; a delayed or missing one is not.

An empty alert text is refused, not spoken. If the call_record or the alert text will not resolve, the render service answers 400 rather than synthesizing silence, so the call falls back to the pre-recorded prompt with tts_fallback_reason = 'http_status'.

When no render service is configured​

Leaving PROXIMA_VOICE_TTS_RENDER_URL empty is a supported deployment, and it is what production ran before this shipped. Every alert call plays the pre-recorded prompt, and the evidence says so explicitly: tts_result = 'fallback' with tts_fallback_reason = 'configured', counted as proxima_voice_tts_render_total{result="not_configured"} rather than as a failure.

Read reason = configured as "this deployment has no render service" — a true statement about the deployment, not an incident. It is a separate value precisely so that a fallback-rate alert cannot fire forever on a site that runs without one. It is also the value to look for first when voice pages are conveying nothing: a run of configured is a missing configuration, while a run of transport or http_status is a broken or unreachable service.

What the call record records (paging evidence)​

A call_record used to carry one mutable status and nothing else about how the call went. A call that rang for forty seconds before pickup and one that went straight to voicemail both ended as completed; a call whose render service was down played a content-free prompt, the callee acknowledged it, and the record said the page succeeded — indistinguishable from a page that actually informed someone. Six columns now record what happened, so a degraded page is countable instead of invisible.

ColumnWhat it holds
answered_atWhen the callee picked up.
ended_atWhen the call left the ARI state machine, whatever the outcome.
tts_resultrendered = the render service returned a usable media URI — not proof the callee heard the alert, see below; fallback = they heard the generic prompt.
tts_fallback_reasonWhich render path produced the fallback. NULL whenever tts_result = 'rendered'.
ack_sourceHow the acknowledgement arrived: keypad, opsbot, web.
not_connected_reasonWhy an origination never reached a ringing phone.

Three further columns — sound_name, sound_url and sound_ready_at — sit on the same row but are control state, not evidence: they steer what a live call plays rather than record what it did. They are described with the mechanism that uses them, under the prefetch and its interlock. sound_ready_at is the one that reads as both, and it is the column to check beside tts_result when asking whether a page conveyed anything.

Every one is nullable and never backfilled. A call placed before this shipped reads NULL, because that observation genuinely was not made and a fabricated value would be worse than an honest gap. None of them carries a CHECK constraint, deliberately: these rows are written on the paging path, and a constraint is one more way a paging-path write can fail. The named Go types plus the column COMMENT are what bound the vocabulary.

Time-to-answer measures from placed, and includes trunk setup​

answered_at - created_at is the only time-to-answer this data supports, and it is not ring time. Ring state is not observable here. A channel does not enter the Asterisk Stasis application until it is answered, so the first event Proxima sees for a call is the one saying a human is already on the line; the ARI event decoder surfaces no channel state either. There is no ringing_at to measure from, and the column deliberately does not exist.

The interval therefore includes everything between the origination being accepted and the callee picking up: trunk setup, carrier routing, and the ring itself. Read it as "how long from us asking for the call to somebody answering it" — a useful number, and a different one from "how long did their phone ring". Do not present it as ring time, and do not compare it against a carrier's ring-time figures.

Call duration is ended_at - answered_at. Anchor durations on ended_at and never on updated_at, which the telemetry writes now touch and which is therefore no longer a business signal.

ended_at is sparse, not merely nullable: it is NULL wherever the pod that owned the call did not see the call end — cross-pod event delivery, a voice-leadership step-down or a restart mid-call, or a record whose call_sid was never attached. That is an observation that was never made, never a zero duration, which is the reading the first SQL aggregate written against it would otherwise fall into.

tts_result — did the page convey anything?​

rendered means the render service returned a usable media URI for this call. Read it as the backend was handed playable audio, and not as the callee heard the alert — the column cannot carry the stronger claim, and the reason it cannot changed in August 2026.

It is no longer the empty-text case: the render service now refuses empty text with a 400 rather than synthesizing silence, so a call_record or alert text that will not resolve lands as fallback / http_status. (The counter value proxima_voice_tts_render_total{result=rendered_no_text} predates that refusal and is unreachable against the current service.)

What rendered cannot see is the delivery, and since the sound prefetch those are genuinely two facts rather than one. rendered says the render service handed back playable audio. Whether that audio ever reached the Asterisk box is sound_ready_at, and the pair answers what neither column can alone:

tts_resultsound_ready_atReading
renderedsetthe alert was named and played — the feature working
renderedNULLrendered, never delivered: the bytes did not reach the box in time, so the generic prompt played instead
fallbacksetthe answer-time render failed but the prefetched audio played anyway — the callee heard their alert

No new tts_result value was minted for this, deliberately. That third row is why one could not have been: it needs the cross product of a render outcome and a delivery outcome, and folding delivery into a vocabulary that a metric, a column and an immutable migration comment all share would mean naming every combination. It also means this column can under-claim — a confirmed sound plays even when the answer-time render just failed, and the row then says fallback. Under-claiming sends an operator to look at a page that was in fact fine; the reverse would hide a page that was not. The second column is right there.

The one thing neither column proves is that Asterisk played what it was handed. sound_ready_at says the bytes were on the disk when the fetcher looked; the playback state is not decoded off the ARI event stream, which remains undone.

fallback means they heard PROXIMA_VOICE_TTS_FALLBACK_URI, the pre-recorded generic prompt: the phone rang, a human answered, and they learned only that something was wrong.

NULL is sparse, never "the render failed": the call never reached the summary (it was never answered, or it is a verification or readiness call, which are routed to a static prompt), the call_record would not resolve when the stamp ran, the stamp was lost (check proxima_call_record_evidence_write_failures_total{stamp=tts_result}), or the row pre-dates evidence collection.

tts_fallback_reason says which path caused it — request_build, transport, http_status, decode, empty_uri, plus configured when this deployment has no render URL at all, and unknown when a fallback arrived carrying no reason. configured is a true statement about the deployment, not a failure: it is separated out precisely so a fallback-rate alert does not fire forever on a site that deliberately runs without a render service. empty_uri means the service answered 200 with no media_uri, which includes a service still replying with the pre-August-2026 sound_id field — a backend/render-service version skew therefore lands here by name instead of resolving to silence.

The reason is NULL on every rendered call and non-empty on every fallback, so the claim holds in both directions:

-- exactly the calls whose answer-time render did not produce audio
SELECT * FROM call_record WHERE tts_fallback_reason IS NOT NULL;

That is no longer the same set as "the degraded calls", and the difference is the third row of the table above: a call whose answer-time render failed but whose prefetched audio played is in this result and was not degraded. Exclude it explicitly rather than assuming the old equivalence:

-- the calls that actually conveyed nothing
SELECT * FROM call_record
WHERE tts_fallback_reason IS NOT NULL AND sound_ready_at IS NULL;

The counter proxima_voice_tts_render_total{result,reason} is the fleet-wide view of the same fact, and it can legitimately exceed a count over this column: it counts render attempts (a gather timeout replays one) and it is not gated on the call_record resolving. It also declares rendered_no_text, a render that succeeded with an empty alert text — but that value is unreachable against the current render service for the reason given above, and is kept only so a historical series stays readable and so a service that stopped refusing empty text could not fold the case into rendered.

ack_source — did the voice channel work?​

How this one is written

Unlike its sibling stamps, ack_source is written synchronously on the single Asterisk ARI read loop when the ack arrives by keypad — the loop every in-flight call shares. It is bounded at 2 seconds so a slow database cannot stall it, and a failure or panic is recovered and counted rather than propagated. That ceiling is argued, not measured: nothing here has run against live traffic, so treat it as the shape of the guarantee rather than an observed latency.

Only keypad proves it did: we placed the call, they answered, they heard enough to act, and the trunk carried the digit back. opsbot is the Telegram ✅ button and web is the console — both confirm the same incident or drill without proving anything about the phone. Before this column the two were one record, so a fleet whose trunk had stopped working entirely would still show a wall of confirmed drills for as long as people kept tapping.

Three things to know before counting it:

  • It is a lower bound on keypad acks, by design. The stamp is first-write-wins, and in the order "keypress, then tap" the later keypress leaves no trace. The rule can only ever fail to acquire a keypad proof, never manufacture one — and under-claiming voice success is the only tolerable error direction for a column whose whole job is proving voice works.
  • NULL is sparse. It also covers a keypress from a contact who lacks alerts:write (which acknowledged nothing), a keypad escalate (a different action, and a deliberate non-acknowledgement — it leaves the alert open; see pressing 2 to escalate), a manager's clear, and a stamp that failed. It is never "nobody acknowledged". The keypad resolve this list used to name is gone: 2 resolved on the Twilio IVR until press-2 became escalate, and one digit must not mean two opposite things across providers.
  • web has no writer yet. A web ack acts on an alert group, and one page can ring several people, so there is no single call to attribute it to. The value is defined so the vocabulary lives in one place; its absence is not a defect.

not_connected_reason — was it the telephony or the human?​

status said no-answer or failed and nothing more, so "the telephony system refused us" and "the human did not pick up" shared one bucket — despite demanding completely different responses. On the Sarkor trunk that distinction has teeth: a refused call triggers post-rejection throttling that can disable outbound paging for minutes, mid-incident.

ValueWhat it proves
trunk_rejectedThe telephony system answered our origination request and refused it (a 5xx from ARI).
originate_failedThe generic bucket: the origination failed with no verdict we may attribute.
no_answer / busyCallee-side outcomes. Defined, with no writer today (see below).

Read trunk_rejected as "telephony refused", never "carrier refused": a 5xx cannot separate the carrier from our own trunk endpoint being unreachable, nor either of those from an Asterisk-internal fault. Everything unattributable stays generic on purpose — a reason invented one level finer than the evidence lands in a column that is trusted more than a log line, and sends an operator to the carrier over a Proxima misconfiguration, or the reverse.

The 5xx rule is expected, not measured

The recorded production failure this exists to catch is Allocation failed / invalid URI '<aor>', which ARI produces when the trunk AOR is Unavailable — but its HTTP status has never been observed. 500 is expected, not measured, and no live trunk has been reachable to check it. If that reply is in fact a 4xx the value is simply never written and the call lands in originate_failed: an under-count, never an invented trunk incident. Verify it on the first live trunk run.

Three more caveats worth knowing before writing a query:

  • It is not page-scoped. Alert pages, readiness drills and phone verification calls all route through the same consumer and are stamped identically, so a bare COUNT(*) WHERE not_connected_reason = 'trunk_rejected' can exceed the number of trunk-refused pages. Scope it with purpose = 'alert' for pages; unscoped it is a lower bound on trunk-refused calls.
  • It is last-write-wins, the deliberate opposite of the columns above it, because a failed origination is redelivered and each delivery is a new attempt with its own verdict. Freezing the first would report a transient blip from attempt 1 while attempts 2–5 were all being refused. The rule is strictly chronological, not "the most specific reason wins". Attaching a provider reference clears the column, so a redelivery that finally places the call does not leave it claiming the call never connected.
  • no_answer/busy are unreachable from this path by construction, not by omission: Originate returns before the callee's phone has said anything, so those outcomes arrive later as a call status on a call that was placed. A carrier SIP rejection arriving after the origination was accepted (403 Q.850 cause=21) is invisible here and ends as a no-answer.

These writes can never cost a call​

Every one of these columns is written by best-effort code sitting on the paging path, under one rule: telemetry must never fail or delay a page. The TTS and lifecycle stamps run in their own goroutines off the single ARI event loop — that loop carries the events of every live call, including the DTMF ack that ends the page in progress, so a synchronous write there would let a wedged database stall paging fleet-wide. Every write is bounded by a deadline, detached from its caller's cancellation (the ARI loop runs under the voice leader context and is cancelled on step-down, which is exactly when this evidence matters most), and its errors and panics are logged and dropped. An unknown call_sid is a lost observation, never an error.

A dropped failure is counted, which is what makes dropping it safe to read. Every one of these swallowed failures — a store error, and a panic recovered so it cannot kill the ARI loop or 500 an acknowledgement that already landed — increments proxima_call_record_evidence_write_failures_total{stamp,outcome}, where stamp names the column (tts_result · answered_at · ended_at · ack_source · not_connected_reason · sound_name · digit) and outcome is error or panic (folded against those two constants, so a third value added later would read as unknown rather than fork the series). Read it before reading any of the columns above, because every one of them is NULL when the write was lost and NULL here means not recorded, never no: a writer broken on every page produces exactly the same rows as a fleet where nothing degraded, no call was answered and no ack arrived by keypad. Zero is the expected steady state.

stamp=sound_name is the odd one out and worth reading differently from its five siblings: it is not a lost observation but a lost control write — the sound prefetch was never queued, so the audio never reaches the box, sound_ready_at stays NULL, and the callee hears the generic prompt instead of their alert. A non-zero rate there is a degraded-page rate, and it is the one stamp whose loss somebody can hear.

stamp=digit was the last of the seven to become countable, and the reason is worth a sentence rather than a silent fix. The write that records which key the callee pressed was three bare UpdateStatus calls with a discarded error and no recover, on a seam whose caller is the single ARI read loop — so a panicking store driver there did not lose an observation, it ended the backend process and every live call on the connection to write a digit. All three now go through one shared writer that is bounded, detached, recovered and counted like its six siblings.

The rule is enumerated and proven across all twelve evidence writers by go test ./... -run TestTelemetryFailureNeverStopsAPage. That sweep proves the writes are fail-open; it does not catch a re-ordering regression (a recovered writer never blocks), and the tests that do are named in its header.

What has and has not been observed in production

This block used to say that no voice trunk had ever been reachable and that none of these columns had been written by a real call. That is no longer true, and it is worth being precise about what changed rather than deleting the caution.

Observed on live calls: a real page on 2026-08-20 wrote ack_source = 'keypad' — the trunk dialled, a human answered, they pressed 1, and the digit came back. A second real page on 2026-08-25 wrote tts_result = 'rendered' and is the call that disproved the ARI media-URI assumption; the callee heard 94 seconds of nothing. Both of those are this column set doing its job: the second page looked successful in every field except the one that mattered, and finding that out is exactly what these columns exist for.

Not observed: every other value, every threshold, and the whole of the sound prefetch. The sound_* columns have never been written by a real call, no fetcher has ever run on the voice VM, and nothing on the prefetch path has been proven by anything other than unit tests, mutation testing and integration tests against a real PostgreSQL. See the sound-prefetch cut-over for what proving it requires.

Carrier failover​

A voice page can be re-placed on a second carrier when the first one could not deliver it.

Configure an ordered list of paths, most preferred first:

PROXIMA_VOICE_PATHS=skyline,sarkor

and, for each name, three more keys built from the prefix plus the upper-cased name:

Key (for path skyline)Meaning
…_SKYLINE_TRUNKthe PJSIP endpoint on the box, e.g. skyline-trunk
…_SKYLINE_SIP_DOMAINthe provider host for explicit-request-URI dialing
…_SKYLINE_STRIP_PLUSon for a provider that refuses a dialed +998…

The prefix is PROXIMA_VOICE_PATH_. Those three travel per path because the disagreements are per provider: Skyline answers a dialed +998… with 404 cause=1 while Sarkor has taken the + in production for months, and Sarkor wants bell.uz as its domain rather than its own host. One global flag would dial one of them wrongly on every failover.

What triggers a switch, and what deliberately does not​

The decision is made only from the Q.850 hangup cause, through classifyHangup:

VerdictCausesSwitches carrier?
telephony1, 2, 3, 20, 27, 28, 34, 38, 41, 42, 58, 63, 88, 102yes — nothing ever rang
callee-side16, 17, 18, 19, 21no — the call got there; the pager worked
unknown / absenteverything elseno — classifyHangup asserts nothing, and a switch is an assertion

Switching on a callee-side cause would re-ring, from a different number, a person who had just declined. That refusal is the point of the feature as much as the switch is.

At most one switch per call. A second telephony failure means the problem is not this carrier, and walking the list would turn one page into a dialing loop while the escalation step's own timer still has to run.

What it does not do​

  • It does not survive the loss of the box. Every path is a trunk on the box PROXIMA_ASTERISK_ARI_URL points at, because one controller owns one ARI connection. Box failover is a separate mechanism — see the failover design §6.
  • It does not catch post-rejection throttling. Sarkor keeps answering OPTIONS while refusing INVITEs, so no health probe sees it and no hangup cause distinguishes it reliably. The signal there is not_connected_reason accumulating across calls.

Unset is the one-path deployment​

With PROXIMA_VOICE_PATHS empty, a single path is synthesized from PROXIMA_ASTERISK_TRUNK / _SIP_DOMAIN / _TRUNK_STRIP_PLUS, dialing exactly as before and never failing over. A malformed list degrades to the same single path with a warning — a typo here must not be able to stop the pager dialing. A one-element list inherits those scalars, so naming the current carrier is a valid one-line migration; a list of two or more must name every trunk explicitly, because inheriting there would hand a path the wrong carrier and silently spend the one failover on it.

Both trunks must be dialable from the active box

Listing two paths is not enough. The box must actually be able to reach both carriers. As of 2026-10-01 voice runs on asterisk02, and Sarkor has not whitelisted its IP (82.115.50.173), so only Skyline is reachable there and failover is inert — correct, configured, and with nowhere to go. It becomes real when Sarkor whitelists that IP, or when the active box moves back to asterisk01, which has both.

Environment variables​

All of these default to keeping production on the existing behavior (Twilio / no voice) and only take effect once the Asterisk ARI config is present.

Provider selection & TTS​

VariableDefaultDescription
PROXIMA_VOICE_PROVIDER"" (auto)Active voice provider: asterisk, twilio, or empty for auto-select (Asterisk when its ARI config is set, else Twilio).
PROXIMA_VOICE_TTS_RENDER_URL""Full URL of the render service's /render endpoint, including the path — a cluster-internal Service, e.g. http://proxima-tts.console-system.svc.cluster.local/render — no port: the Service listens on 80 and forwards to the container's 8080, so a URL carrying :8080 connects to nothing and every render, warm and answer-time, falls back with reason = 'transport'. Empty → fallback-only: every voice page plays the pre-recorded prompt and conveys no alert detail, recorded as tts_fallback_reason = 'configured'.
PROXIMA_VOICE_TTS_MAX_TEXT_CHARS125How many characters of the summary are spoken. A proxy for a time budget, not a bound on one — see the render happens at originate. Values above 2000 are refused by the loader (the render service rejects the text outright above that, which would make every page fall back).
PROXIMA_VOICE_TTS_FALLBACK_URIsound:proxima-alertARI media URI played whenever a render fails, no render URL is set, or the prefetched sound is not confirmed. Must exist in Asterisk's sounds directory — the deploy ships it, in .wav, .alaw and .ulaw.
PROXIMA_VOICE_FETCHER_TOKEN""The only authenticator on the two sound-prefetch endpoints, and not a user session — the caller is a machine on the voice VM. Empty means the routes are not registered at all: absent rather than open, so the fetcher 404s, nothing is ever confirmed, and every page speaks the generic prompt. Set-but-unusable is refused at startup rather than at request time: fewer than 16 printable non-space characters, or a value that could not survive an Authorization header — surrounding whitespace (which Go's server strips before the constant-time compare can match) or a control character such as the trailing newline vault kv get -field=… > file produces (which cannot be transmitted at all). Each of those would otherwise start cleanly, mount the routes, report set on the settings page, and 401 the fetcher forever. Rotating it means updating the box at the same time; the gap between the two degrades pages rather than failing them.
PROXIMA_VOICE_VERIFIED_GATEwarnWhether voice paging requires a verified phone. warn (default, non-regressing) still pages unverified phones and bumps proxima_voice_unverified_page_total; enforce skips unverified phones (escalation falls through). Flip to enforce only after a verification drive — see the verified gate.

Asterisk / ARI (gates the whole Asterisk path)​

The Asterisk provider is built only when the three ARI credentials are set. These are defined and detailed for the trunk-health worker — see Voice Trunk Health; the same ARI endpoint drives paging:

VariableDefaultDescription
PROXIMA_ASTERISK_ARI_URL""Base URL of the Asterisk ARI endpoint (e.g. https://asterisk-vm:8089). Empty → Asterisk voice disabled.
PROXIMA_ASTERISK_ARI_USER""ARI user (basic auth; also carried as api_key=user:password on the event WebSocket).
PROXIMA_ASTERISK_ARI_PASSWORD""ARI password (secret).
PROXIMA_ASTERISK_ARI_APPproxima-voiceStasis application name the controller subscribes to and originates into.
PROXIMA_ASTERISK_TRUNKsarkor-trunkSIP trunk endpoint used to place outbound calls (PJSIP/<to>@<trunk>).
PROXIMA_ASTERISK_WS_PING_INTERVAL20sHow often the ARI event-WebSocket transport sends a keepalive ping to detect a silently-dead connection. A non-positive value falls back to the default.
PROXIMA_ASTERISK_WS_READ_TIMEOUT60sHow long the transport waits for the next event (or pong) before treating the connection as dead and reconnecting.
PROXIMA_ASTERISK_WS_WRITE_TIMEOUT10sHow long a single WebSocket write (e.g. a ping) may block before it is treated as failed.
PROXIMA_VOICE_LEADER_ACQUIRE_INTERVAL5sOn 2+ replicas, how often a non-leader tries to acquire the voice leader session advisory lock — the upper bound on failover latency after a leader dies (P2-B).

On the voice VM (the sound fetcher)​

The fetcher is configured entirely from its own environment, written by infra/asterisk/deploy-asterisk.sh into /opt/asterisk/fetcher.env (mode 0600, over ssh stdin — never in an argv, which ps would expose to every user on the box). It needs PROXIMA_API_URL and the same PROXIMA_VOICE_FETCHER_TOKEN the backend holds; the rest have working defaults. The full table, and what happens if the poll interval is lowered without raising the backend's rate limit, is in infra/asterisk/README.md.

Twilio (fallback path — unchanged)​

VariableDefaultDescription
PROXIMA_TWILIO_ACCOUNT_SID""Twilio account SID (gates the Twilio provider together with the auth token + public base URL).
PROXIMA_TWILIO_AUTH_TOKEN""Twilio auth token (secret).
PROXIMA_VOICE_PUBLIC_BASE_URL""Publicly reachable backend base URL Twilio calls back for TwiML + status (e.g. https://api-console.prxm.uz).
PROXIMA_VOICE_FROM_DEFAULT(env-specific)Default caller-ID E.164 when no country-specific sender matches.

Known limits​

  • A press of 1 whose alert-store write fails is still told "Acknowledged". When an authorized press reaches the alert service and the acknowledge fails with a store error — not one of the service's two refusals, already resolved or changed meanwhile — the actor returns done: the Asterisk IVR plays sound:proxima-ack, Twilio says "Acknowledged. Goodbye.", and ack_source records keypad, because that column records that the voice channel delivered the key, not that the alert store accepted it. The alert itself was not acknowledged: its status is unchanged, and the error is logged as voice ack failed. This is pre-existing and deliberately pinned by TestVoiceActor_AckStoreFailureStillSaysDone; telling the caller the truth on this path is a follow-up.

See also​

  • Voice Trunk Health — the observability-only worker that alerts when the SIP trunk carrying these calls goes down.
  • Resolution Safety — the other two pieces of "why didn't I get paged?": the group-level reasons a page ended before any target was chosen (G6) and the per-target drops a notification chain never attempted (G7). The call-record columns on this page are the third piece — what happened once a target was chosen and dialled.
  • Notification Chains — how the voice step slots into a per-user paging chain (and why a failed call advances).
  • Escalation Dead-Man's Switch — the out-of-band guard that pages when the escalation engine itself stops advancing.
  • infra/asterisk/README.md — provisioning, trunk hardening, and ARI lockdown.
  • services/tts/README.md and infra/k8s/tts/README.md — the render service itself and the Kubernetes objects it needs.
  • infra/asterisk/README.md — the voice VM: the sound fetcher, its configuration, the sounds mount, and the box-level live-validation checklist referenced from the cut-over above.
  • Rate limiting — why the two prefetch endpoints carry their own 300/min budget rather than the webhook one.

Live validation​

The voice controller's state machine is unit-tested, but DTMF surfacing, audio, and call timing can only be proven on a real box. The original live-validation checklist — configuring ARI, deploying the render service, firing a test escalation, and confirming the ack/retry/replay paths — lives in the repository at docs/superpowers/plans/2026-07-30-oncall-p2b-live-validation.md.

Nothing about the render service or the sound prefetch has been validated on a live call. No alert audio has ever been heard through a real phone, the Kubernetes manifests have never been applied to any cluster, and no fetcher has ever run on the voice VM. The one step of that older checklist that still matters most is its step 4: scale proxima-tts to zero, fire a page, and confirm the call still connects and plays the local prompt with tts_result = 'fallback' — that is what proves the cross-network dependency is safe.

The press-2 escalate cut-over​

Nothing on the escalate path has been exercised on a live call, and nobody has heard its prompts. The five prompt files are verified by measurement — sample rate, channel count, bit depth, no clipping, one byte per frame in the companded pair, and a decode-back that lands inside companding error of the committed WAV — which proves they are well-formed audio saying the text in prompts.tsv. It does not prove anyone can understand them over a phone.

One assumption behind them is reasoned and not observed: that Asterisk prefers the committed .alaw over the .wav on this trunk. It follows from pjsip.conf's allow=alaw,ulaw, and it has never been watched happening.

Voice paging is also dark for an unrelated reason: the SIP trunk is de-registered and the provider is blackholing the box, so no call of any kind is currently placeable. See Voice Trunk Health. Restoring the trunk is a prerequisite for everything below.

Step 1 is the one that can invalidate the whole feature

Nothing waits for the escalate prompt to finish playing. ari.Play is an asynchronous submit — ARI returns as soon as it accepts POST /channels/{id}/play — and the controller then issues finish(…, hangup=true) → DELETE /channels/{id} one database round trip later, milliseconds afterwards. The prompts are 0.81 s (proxima-escalated) to 3.14 s (proxima-escalate-failed) of speech. If Asterisk terminates playback on hangup, the responder hears none of it, and everything else on this checklist can pass while the feature conveys nothing.

This shape is inherited, not introduced — the ack tone, the readiness confirm and the "Your phone is verified. Goodbye." prompt all do the same and have shipped since August — and it is reasoned from the protocol, never observed, because no trunk has been reachable to test it. It is not fixed here on purpose: holding the channel until PlaybackFinished means decoding playback state off the ARI event stream and restructuring the hangup path, in the file that calls itself the riskiest code in voice paging, with no live Asterisk to validate the restructure against. It belongs in a change that can be validated — but it must be checked first, because this is the first feature whose entire value is the prompt being audible.

When a trunk is back, this is what proving it takes — with a human on the phone:

  1. Deploy the prompts to the box before rolling out the backend. Run infra/asterisk/deploy-asterisk.sh, and confirm all five escalate prompts are present in .wav, .alaw and .ulaw. ARI accepts a sound: play for a missing file at submit and fails it asynchronously, so a backend rollout ahead of the VM deploy produces silence-then-hangup with nothing reporting it — which is the exact failure mode proxima-escalate-failed exists to prevent, arriving through the deploy instead.
  2. Press 2 and listen for the whole prompt. Before anything else on this list: did the responder hear the complete sentence, or was the channel torn down mid-word? See the warning above. Record what happened either way — this is the observation the follow-up needs.
  3. Fire a synthetic P1 through a narrowly-scoped test route.
  4. Answer, press 2, and confirm the callee hears "Escalated." and the chain advances.
  5. Confirm the alert is still unacknowledged and ack_source is still NULL — this is the one that must not be skipped, because an escalate that acks is a silenced page.
  6. Confirm call_record.digit is '2' and the readiness page shows the call under answered_other_keypress, not under answered_with_ack.
  7. Open the alert's Activity timeline and confirm the Escalation requested row renders with its outcome and its engine: / reason: detail line — that row is the only place the five collapsed engine verdicts survive.
  8. Press 2 on a single-step policy and confirm the caller hears "there is no further escalation" — and that the alert_group_log row records not_armed or no_timer in metadata->>'advance_outcome' rather than nothing.
  9. Confirm the pre-rendered page ends by naming both keys, and that a page whose alert name is long still ends with the instruction rather than truncating it away.
  10. Revert the test route, policy, source and any cut-over flip.

The sound-prefetch cut-over​

Do these in order. Each step is either reversible on its own or leaves the system in the state it was already in — a page that cannot find confirmed audio speaks the generic prompt, which is exactly what production does today. Steps 1–5 are preparation and change nothing audible; step 6 is the first one a callee could notice.

Anything marked 🔴 has never been executed anywhere and is being run for the first time by whoever follows this list.

  1. Apply migration 000196. Three nullable columns on call_record, no CHECK, no backfill. Reversible: down 1 was proven to return the table byte-identically, including column comments and indexes. Nothing reads or writes the columns until step 4.
  2. Mint the fetcher token and put it on the backend as PROXIMA_VOICE_FETCHER_TOKEN (16+ printable non-space characters, no trailing newline — see the note in the env-var table above; the backend refuses to start on a token it knows cannot authenticate, which is the failure you want, and it is loud). Until this is set the two endpoints are not mounted at all.
  3. Confirm the endpoints answer. From anywhere that can reach the API: curl -H "Authorization: Bearer $TOKEN" https://<api>/api/v1/voice/pending-sounds returns {"data":[]}; the same call without the header returns 401; a POST .../nope/ready returns 404. A 404 on the list means the token is not set on the backend.
  4. 🔴 Deploy the voice VM with infra/asterisk/deploy-asterisk.sh, which now requires PROXIMA_API_URL and PROXIMA_VOICE_FETCHER_TOKEN and aborts before shipping anything without them. This run performs the sounds-preservation step for the first time. The image declares /var/lib/asterisk/sounds as its own volume, and the new bind mount would otherwise hide it — including proxima-ack.alaw/.ulaw, which exist only on that box (nothing in the repository renders a proxima-ack, and sound:proxima-ack is played on every DTMF ack). The step copies the container's en/ onto the host before the bind takes effect, is guarded so it cannot run twice, and refuses rather than guesses if the bind is already present without the sentinel file. Verify it worked before doing anything else — this is the one step whose failure mode is losing a file nobody can regenerate.
  5. 🔴 Confirm the box's own checklist, in infra/asterisk/README.md §"Live-validation checklist": the fetcher is polling and healthy, the sounds directory holds all four committed prompts in all three formats plus the preserved proxima-ack.*, and — item (e2) — the .alaw on the box is byte-identical to infra/asterisk/sounds/proxima-alert.alaw. (e2) is the one open question this branch could not answer. The trunk is allow=alaw,ulaw, so Asterisk should prefer a raw companded file over a .wav it would have to transcode — but that preference is reasoned from the codec config, not observed, because observing it means playing a file on a box we hold read-only access to. If the committed .wav turns out to win, the good voice is on disk and the robotic one is what plays, and the symptom is indistinguishable from the prompts never having been fixed.
  6. 🔴 Fire one real page on a test route and read the row. sound_ready_at non-NULL is the fetcher's confirmation; the callee hearing the alert text rather than "A Proxima alert is active" is the whole feature. Check the fetcher's log for the download and the confirm, and the backend's for voice sound queued for prefetch and a play of sound:px-….
  7. 🔴 Prove the degradation, not just the success. Stop the fetcher, fire another page, and confirm the callee still gets the generic prompt and can still ack by keypad. This is the property everything else is built on, and it is worth spending a call to see rather than trusting.
  8. Settle the render/fetch replica pinning flagged in infra/k8s/tts/README.md. The prefetch de-escalated this from a silent 100%-of-pages failure to a visible one — a fetch that 404s now leaves sound_ready_at NULL and the page falls back — but it did not remove it: if the gateway's pinned replica is not the one that rendered, every download 404s and every page still degrades. It is now something to watch a metric for rather than a precondition to the rollout.

And one teardown that is not optional. A temporary live-test configuration was left in place during this work so the fix could be re-tested immediately — one escalation route enabled and one team forced live rather than shadow. It is a database change, not a repository one, so nothing in a deploy removes it. Take it down once step 6 or 7 has told you what it was there to tell you: a team left live by accident pages real people.

Two things this list deliberately does not promise. It does not prove Asterisk played what it was handed — playback state is not decoded off the ARI event stream. And the answer-time render still runs on the confirmed path, so a slow render service can still cost up to five seconds of dead air on a call whose audio is already sitting on the disk; removing it is a recorded follow-up rather than a shipped behaviour, because it would have to invent what tts_result says about a call that never rendered.