Telegram Alert Card
When an alert group's paging arms, each Telegram group bound to its client for the internal audience gets one message for it — the alert card — edited in place as the incident moves from firing to acknowledged, silenced and resolved. A client-facing group gets its own much shorter status message for a P1 or P2, edited the same way. Nothing is posted by a backend without PROXIMA_OPS_TELEGRAM_BOT_TOKEN, for a client without the oncall_enabled feature, or for an alert whose paging does not arm for a live team: see when a card is posted and the rollout preconditions under known limits.
This page carries the reasoning behind the card, not only its shape. Several of its rules look improvable until you know why they are what they are, so read the relevant section before changing one.
A Telegram edit notifies nobody. The card is how a room sees an incident and who owns it; it is not how anyone is woken. Paging is the notification chain — voice and direct messages — and the card changes nothing about it.
It follows that an edit failure never fails or delays a page. The page has already gone out; the edit is telemetry about it. Every card update is requested asynchronously, and every edit failure is logged and counted and never reaches the page path: a transient one (a 429, a 5xx, a network error) is retried up to three times, and a permanent one is remembered for up to an hour so it is not repeated.
One message per alert group, edited
Before the card, every notification was a fresh message. That caused two problems directly: a flapping alert group posted fifty messages into the room, and nothing in the room ever showed that somebody had taken the incident. People muted the channel, and alerts scrolled past with nobody owning them.
So each chat carries exactly one message per alert group — telegram_message holds one row per (alert_group_id, chat_id):
- The card is posted when paging arms. When the correlation worker arms an alert group's escalation — a route matched it — for a page-owner team that is live, not in shadow, it publishes the dispatch, and the notification worker posts the card into each chat bound to the client for the internal audience whose binding does not mute alerts and that has no card for the group yet, recording its message id and the state it shows. An alert no route matches, or whose team is still in shadow (its chats may still be receiving another system's posts), posts none; neither does a client without the
oncall_enabledfeature, nor a backend withoutPROXIMA_OPS_TELEGRAM_BOT_TOKEN, which wires neither worker. The card follows paging, so the client's AI flag (ai_triage) has no bearing on it: a client that pages without investigating still gets cards. An armed alert can still page nobody — nobody on call, no verified contact — and its card posts anyway, so a card in the room is not proof that someone was paged. - Once a chat has the card, everything after that is an edit of the same message: a further firing, the page for the group arriving again (a redelivered dispatch, or a new firing epoch — whose dispatch-ledger key would otherwise allow a second post), and every state change.
- The room sees ownership without anyone announcing it:
🔴 P1 · Replica lag on psql01
🟠 ACKNOWLEDGED by Aziz · 4m after firing
🔇 SILENCED by Aziz until 09:00 UTC
🟢 RESOLVED by Aziz · 38m open
The component that edits cards, the Syncer, only ever edits; it never posts. Posting belongs to the notification worker, on the dispatch ledger's claim path. If the Syncer could post, resolving a group that was never paged into a room would drop a resolved card into a room that never saw the alert.
A further firing edits instead of posting — and that is not a missed page
An edit does not notify, so a group that keeps firing does not buzz the room again. That is deliberate: the group chat is the shared record. The notification chain is the pager, and it re-pages on its own terms, whatever the card does. A fresh message for every firing would be exactly the noise the card exists to remove.
A refire after resolution is a new alert group
A group that has resolved and then fires again almost always becomes a new alert group — Console's grouping only ever matches an open group (status <> 'resolved'). The new group posts its own card when its paging arms for a live team, as any group does, and the old card stays resolved. Flapping within a firing group is one message; a resolve-then-refire cycle is one card per episode.
Rarely, a resolved group reopens in place instead. It takes a race inside a fraction of a second: a delivery for the group has already found it open, and a person's resolve (or the silence sweeper's) commits before that delivery writes its firing alert. The recount, which judges the status under its row lock, then finds the group resolved with a firing member and reopens it under a new firing epoch, with its escalation state reset. The alert worker asks the card to follow, so the resolved card is edited back to firing — an edit, which notifies nobody — and paging can arm again under the new epoch. The room gets no new card for that episode.
A resolve the check refused says so, and the card stays firing
A resolve the source reports is corroborated before it is written: Console looks for independent evidence that the entity is still alive, and a signal that positively reports trouble refuses the resolve. The group then stays exactly as it was — still firing, still escalating — and the card gains one line, directly under the scope line:
🔴 P1 · Replica lag on psql01
AcmeCorp / production / psql01
⚠️ STILL FIRING — at 03:14 UTC the source reported resolved, but node02 has not checked in for 4m12s
- The reason names a host's check-in gap or a named monitor, never a source: the source signal is never asked.
- The measurement is dated, in the card's time zone — the same one the silence end and the timeline's stamps use, so every time on the card agrees. The stamp is read from the refusal row, because one refusal supplies this line for the rest of the firing episode.
- The line is never on a resolved card — one card cannot say
RESOLVEDandSTILL FIRINGat once — but an acknowledged or silenced card does show it, because those groups are still open. - It is an edit, so it notifies nobody. The card is the shared record; the refused resolve changes nothing about paging, which continues on the chain's own terms. The edit is forced past the usual "the message already shows this state" check — the refusal changes no state, so nothing else would send it — and only once per firing episode, however many times the source re-sends its resolve.
- A refusal whose reason is missing or unreadable renders no line, by the degradation rule below.
A resolve that is written turns the card green as usual, whether an independent signal confirmed it or nothing could speak. The difference between those two is kept on the alert's timeline in Console, not in the room. See Resolution Corroboration.
What the card shows
Four lines get read, in this order:
- What broke, in words — the
summaryannotation, written by a human for a human. Without one, the alert name; without that, no title at all. Never "untitled", never "unknown", and never the grouping key, which is a hash. - Whose and where — client / environment / host.
- Is it spreading — the incident's blast radius, and the chronic count (how often this alert has fired within the recurrence window). The blast radius appears only when it carries news: more than one host or alert, or something downstream. A single-host, single-alert incident is already named on the "whose and where" line above, and a line that only restates it is a line the timeline or the raw payload gives up under the message budget. Counts agree with their nouns — 1 host, not 1 hosts.
- What it probably is — the AI cause, in its existing trust framing (verified / unconfirmed / unverified). Never re-worded, never promoted.
The card has its own clock line, under the scope. ⏱ firing 11m while it is open, ⏱ responded 28s · firing 11m once somebody has it, ⏱ responded 28s · open 4m once it is closed. These are the numbers a responder is measured on, so they get a line rather than a clause trailing the state — and a firing card previously carried no duration at all, which meant the most useful live fact on it ("nobody has answered this for 11 minutes") was absent. The state header keeps only the state and its owner. A response time is shown only when the acknowledgement came after the firing it answers: acknowledged_at is never cleared, so a re-fired group still carries the previous episode's acknowledgement, and printing that against this episode's start would claim somebody answered an alert nobody has answered yet.
The card reads as three blocks, separated by blank lines: who/what/where (state, title, scope, clock), then the state warning if there is one, then the machine's opinion. One tap away, in expandable sections on the same message — never a second message — are the timeline and the raw payload. Grafana OnCall posts its escalation log as a separate message per alert group, which doubles the objects in a channel meant to be signal.
The degradation rule: a line appears only when it is real
Enrichment fails in production, today. This is a real log line:
l1: alert resolved to no host — triage will be blind (check alert_label_mapping)
If the card had discarded the raw payload and the enrichment came back empty, the engineer would have less than Grafana gave them. So:
- A line appears only when it is real. No resolved host means no host on the card — not "host: unknown". A chronic count below the threshold, or a lookup that failed, is no line — never a zero. A whitespace-only value counts as absent. In the raw payload, blank fields (empty strings,
null,{},[]) render no line;falseand0still do, because they are values. - The raw payload is always one tap down. It is the floor the rest of the card degrades to.
It is the same principle as alert ingestion never fabricating a severity: an absent value is shown as absent, never as a confident guess.
Labels and annotations are demoted, not dropped
The raw payload section renders the first alert's labels, then its annotations, then every other field the sender gave. Labels carry what Console cannot know — a runbook_url on the rule, a client's own team label, a ticket reference — so moving them off the card's face is fine, and removing them would not be.
The message budget
A card is kept within 4000 UTF-16 code units, counted on the rendered HTML, markup included. Telegram measures length in UTF-16, where an emoji is two units, and rejects a message over 4096 outright. When a card would be longer:
- The head (state, title, scope, AI cause) is sized first, but always leaves each non-empty section its minimal real form. Free text — the AI cause — gives way first; identity values (severity, title, scope, actor) are cut last, and a cut value keeps its prefix and ends in
…. - The raw payload's minimal body is reserved before the timeline is fitted, because it is the floor.
- The timeline drops its oldest entries and keeps the newest, which are the ones being read, with a line saying how many earlier entries were left out.
- The raw payload takes what remains and gives up its tail. A payload with no room at all renders a one-line "raw payload too large — open in Console" rather than overflowing the message.
Who did it: attribution from the audit log
"by name" appears only when the audit log names a person for the action the card shows: the newest acknowledged, silenced or resolved entry, with actor type user. The name is the Console user's display name — never the email, which in a group chat is a disclosure, and never the user id, which is noise.
A resolve the source reported, or the stale-group reaper made, names nobody: the log cannot tell a source confirming a person's resolve from an episode that closed on its own. acknowledged_by is never used for a silence or a resolve, because it still names the acknowledger after someone else acts.
Sanitising interpolated values
When the card's formatting was verified against the live Bot API, Telegram parsed a bot_command entity out of the test text. A hostname or label value containing / would render as a tappable command link, and a client-supplied label value is untrusted input.
The card is written in Telegram HTML, whose escaping rule is a closed set (&, <, >, and " inside an attribute) that does not change with context. Every interpolated value is escaped, and the values are guarded in two different ways:
- In a code span, where Telegram parses no entities at all: the hostname, and the bodies of the expandable timeline and raw payload sections.
- Everything else is bare text — the severity, the title, the client and environment names, the actor's name, the silence end, the AI cause and its "needs human" reason. These are escaped, and an invisible zero-width non-joiner is inserted after every
/and@, so a value cannot form a bot command or a mention. That is the only guard. Telegram still auto-detects links, e-mail addresses,#hashtagsand$cashtagsin them, so a client-written summary containinghttps://…renders as a tappable link on an internal card, and the invisible character is copied along with any text someone pastes from the card.
If Telegram ever refuses a card as unparseable, the send fails as permanent telegram_render_rejected, so a rendering bug fails fast instead of looking like a dead contact.
Controls, per state
| Card state | Controls |
|---|---|
| Firing | ✅ Ack · Resolve · ⚡ Escalate — 🔇 1h · 🔇 4h · 🔇 until 09:00 UTC — Open in Console ↗ |
| Acknowledged | Resolve — Open in Console ↗ |
| Silenced | 🔔 Unsilence · Resolve — Open in Console ↗ |
| Resolved | Open in Console ↗ |
A tap on Ack, a silence or Unsilence whose alert has already resolved — a card someone taps after the incident closed — changes nothing and is answered with Already resolved. A resolved group is never revived from chat: the next firing for the same alert opens a new group — which posts its own card if its paging arms for a live team — instead of joining a closed one that would page nobody.
A tap that races another change to the same alert — someone else acknowledging or silencing it, the "still on it?" loop un-acknowledging it, a silence ending — is decided again against the alert as it now is, so a silence records the status the alert really had and ends by restoring that. If the alert keeps changing through three attempts, nothing is written and the tap is answered Changed meanwhile — try again.
A card in a state the renderer does not recognise gets the firing keyboard. Firing is the state that keeps asking for attention; a closed-looking card nobody can act on is the failure to avoid, and every action is still authorized on the server.
Escalate is on the firing card only. "Not mine" means page the next person, and the chain advances only for a firing group. Moving an acknowledged incident onward would first need an un-acknowledge, which is deliberately not built.
Why there is no Unack
An un-acknowledge looks like wiring, and it is building: what it means is undecided. Does it re-arm the escalation timer, and from which step? What happens to the "still on it?" reminder loop that every acknowledgement arms? Who is paged next? No service method exists, and a button would have to invent answers to all three. Until they are decided, someone who acknowledged by mistake has no chat control for it.
Resolve is an assertion, and Silence protects its meaning
A group also resolves when the source says so. A person pressing Resolve while the condition persists will see the alert come back as a new alert group, with a new card once its paging arms. That is correct, and it is precisely why Silence must exist: without it, Resolve becomes the make-it-stop button, the resolution record stops meaning anything, and people are annoyed by alerts "returning" that were never fixed.
Silence is bounded in chat, indefinite only in Console
An alert silenced forever is an alert nobody sees again, chosen by a tired person on a phone. So chat offers only 1h, 4h and until 09:00 — the next 09:00 strictly after the tap, so a tap at 10:30 silences until tomorrow morning rather than until a time that has already passed. That 09:00 is in the server's time zone, and the button names it (🔇 until 09:00 UTC on a server running in UTC). An indefinite silence stays in Console, where it is a deliberate act with an audit trail.
The card shows the silence and when it is meant to end, in the same zone and naming it — until 16:00 UTC later the same day, until Mon 09:00 UTC on one of the next six days, and with the date otherwise — further out, or once the end has passed, where a bare weekday would read as next week. The timeline's stamps are in that zone too, named once on the section's title. A silence ends at that time. Within about 30 seconds — the sweeper's interval — the silence sweeper returns the group to the status it had when it was silenced — acknowledged, with its owner, if it was acknowledged, otherwise firing — so an armed escalation, or an acknowledgement's "still on it?" loop, picks up again on the timer's next tick, and the sweeper asks the card to follow. 🔔 Unsilence ends a silence early by the same rule. The status to go back to is recorded when the silence is set, never inferred from who once acknowledged the alert, and silencing an already-silenced group again (from Console; the silenced card offers no silence) keeps it.
A silence that had already ended more than 24 hours before anything noticed it — a silence from before silences ended on their own, or a sweeper that was down for a day — does not resume paging for a stale incident: the system resolves the group, and the card turns resolved, naming nobody. If the problem is still firing, its next alert opens a new group, which arms paging if a route matches and posts its own card if that paging arms for a live team.
The card's state follows the group's status and nothing else, because the card must say what the paging engine will do and the engine keys only on status. So in the seconds between a silence's end time and the sweep, the card still says silenced — which is still true: nothing has paged yet.
For the same reason, acknowledged_by never decides the state. It is never cleared: when the "still on it?" loop auto-unacknowledges a group nobody confirmed, or an acknowledged group is resolved and reopens in place, the group is firing again and the card shows firing, with the firing controls, even though the acknowledger's name is still on the row.
Why Open is a link button, not a Mini App button
Telegram allows web_app inline buttons only in private chats, and cards live in groups. So Open in Console ↗ is a plain link to the alert in Console, <frontend URL>/alerts/<alert group id>. A link button produces no callback and so no authorization check, which is acceptable only because it points at Console, which authenticates the viewer itself. Without a configured frontend URL there is no Open button, rather than a dead link.
Escalate tells the truth
An escalate tap is answered with what the chain actually did, the way a voice press-2 is, and the person who tapped sees it as a Telegram toast:
| Outcome | Toast |
|---|---|
| The chain advanced | Escalated to the next step |
| Nothing further — no chain armed, or the ladder is exhausted | Nothing further to escalate to |
| Someone had already acknowledged | Already acknowledged |
| The alert had already resolved | Already resolved |
| The alert is silenced — the chain is paused, not exhausted | Silenced — unsilence to resume escalation |
| The engine had already moved, or the advance failed | Couldn't escalate — try again or use Console |
Escalation is not enabled on this Console (the backend answers 501) | Not available — no retry hint, because a retry cannot enable it |
| A reply this ops bot build does not recognize | Request sent — check the alert in Console — it claims only that the request arrived |
"Escalated" promises only that the chain moved: the next step may find nobody on call or ring a phone nobody answers, and over-claiming on a paging path is the failure this work exists to end.
Every escalate tap that passes authorization — including one that changed nothing, which is evidence that the policy has no depth left — writes an escalation_requested entry on the alert's timeline, attributed to the linked Console user, with metadata.source = "telegram" and the same result and advance_outcome vocabulary as voice. A failed timeline write is logged and never fails the tap.
Authorization: chat membership is not authority
Anyone in a group can tap a button. Being in the group authorizes nothing:
- The Ops bot forwards the tap to
POST /api/v1/telegram/ops/callback, which accepts only the pinned Ops bot service account. - The tapping Telegram user is resolved to their linked Console user. An unlinked user is refused, and told to
/link. - In a group chat, the chat must be bound to the alert's client. A group not bound to it is refused, and the toast says so.
- That user must hold
alerts:writeon the alert's environment — the gate the web Acknowledge, Resolve, Silence and Unsilence actions apply. A client-wide grant covers every environment, an environment-scoped grant covers only its own, and an alert with no environment falls back to the client-level check.
In a private chat with the bot — where an escalation page and the "still on it?" nudge land — step 3 does not apply, and every other step does, in the same order, for every control. No chat is ever bound to a private chat, so a binding check there could only refuse every tap; and it would prove nothing if it passed, because the chat's one member is the user, and chat membership was never the authority — the user's own alerts:write always was. A chat counts as private only when its id is the tapping user's own Telegram id, which is how Telegram numbers a private chat; any other chat is checked as a group.
The action then goes through the same alert service the web UI uses, so it lands in the same audit log and moves the card the same way.
A reply to the card becomes a note
Replying to an internal alert card, in your own words, records a note on that alert's timeline, attributed to your linked Console user. There is no command and no button: the reply is the gesture. The bot answers with one line — 📝 noted — and nothing else; the note itself is read in Console.
It is authorized exactly as a button tap is, in the same order: the pinned Ops bot service account, your linked Console user, the chat bound to the alert's client (skipped in a private chat with the bot, where no chat is ever bound), and alerts:write on the alert's client and environment. The alert is taken from the tracked card the reply points at — never from the bot's request — so a reply can only ever write to the alert whose card is being replied to.
What does not become a note:
| The reply | What happens |
|---|---|
| To a client-facing status message | Nothing is recorded. It is answered exactly as a reply to any non-card message, and falls through to the AI Investigator — see the known limit below |
| To a superseded card — one the bot has since re-posted | Nothing is recorded. The old message resolves to no alert, and guessing "the newest card in this chat" would file someone's words against the wrong incident |
From an unlinked Telegram account, or one without alerts:write | Nothing is recorded, and the bot stays silent — it never posts a refusal into the room |
A command (/…) or an @mention of the bot | Never a note. An @mention is a question addressed to the bot, and belongs to the Investigator rather than to the record in the asker's name |
| Over 2000 runes | Refused rather than truncated — a half-note attributed to a responder would put words in their mouth. The bot answers once: 📝 Too long to record as a note — shorten it, or add it in Console. |
| To any other bot message | Reaches the AI Investigator, exactly as before notes existed |
A note is an ordinary note row in the alert's audit log with metadata.source = "telegram", so it sits beside the web UI's notes rather than in a channel of its own. Replies are redacted before they leave the bot, so a secret pasted into a reply does not reach the record.
The card is redrawn as soon as the note is recorded, so the author sees their own words land on the timeline rather than waiting for some later, unrelated event to redraw it. A note changes no alert state, so this edit is forced past the usual "the message already shows this state" check, exactly as a refused resolve's line is — without that, the note would render correctly into a message that is never sent, and would appear retroactively at the next acknowledgement or resolve. Unlike the refused resolve, a note is not once-per-episode: a responder may reply as often as they like. What bounds the edits is the Syncer, not the author — replies inside the 2-second debounce collapse into a single trailing redraw that re-reads the group, runs share the same four slots as every other card edit, a retry_after Telegram named is honored, and a message whose edit failed permanently is not attempted again. Those bounds are per backend replica — the debounce is kept in memory, as the reconciler's backoff is — so a burst of replies costs at most one edit per running replica rather than one overall; with single-digit replicas and a human doing the typing, that stays small. The redraw is best-effort and runs after the note has committed, so a Telegram problem can never turn into a failed note: the bot still answers 📝 noted, and the note is in Console either way.
Two audiences, two renderings
| Internal chat | Client-facing chat | |
|---|---|---|
| Receives | a card for each alert of the client whose paging arms for a live team, unless the binding mutes alerts | the same alerts, P1 and P2 only, unless the binding opts out |
| Shows | the full card | one curated sentence: investigating, identified or resolved |
| Controls | per state | never |
| Internal detail | yes | never — no hostname, severity, AI text, timeline, raw payload or names |
Both are edited in place through the incident. In the client vocabulary, acknowledged and silenced both read identified.
They are two separate render functions, not one function with an audience flag. What engineers see while triaging must never reach a client, and a flag is one wrong boolean away from sending it there on some message. A client renderer that takes the same card and ignores almost all of it makes the separation structural rather than a judgement call on every message. The customer copy exists in one place, used by the first post and by every edit, so the two cannot drift.
The stricter audience wins. If a chat is re-bound so that a message's recorded audience and the chat's current audience disagree, the message is rendered for the client. Internal detail is never edited into a message that went out client-facing, and a chat that is client-facing now never gets internal detail because of what its message once was.
A third audience, which is not a card
A personal page — the Telegram DM an escalation chain sends to one responder — is tracked in the same table, under a third audience (direct), and is edited in place by the same syncer. It is documented separately, in Personal Pages, because it is not a card: it goes to a private chat, it is reconciled without any binding, and it edits to a one-line outcome rather than to a card. Everything on this page about cards in bound chats — the two renderings above, the topic routing below — is about the two audiences in the table, not about pages.
Which forum topic it lands in
A card goes to the chat's alerts topic when the chat has one configured, and to the thread its binding names when it does not — and the client-facing status message follows the same rule, in its own chat, for the same reason: a customer's room is as likely to be a forum as a responder's, and a status message arriving in the ticket topic is the same problem in a politer room. Nothing is configured by default, so a chat nobody has touched keeps delivering exactly where it always did. Set a topic with /topic alerts, run inside the topic you want — see Forum topics: routing by kind.
The rule is one function (store.ResolveThreads), and the service-desk path that routes Jira tickets to a tasks topic calls the same one. That is deliberate, not incidental: two producers, each with its own copy of "configured topic, else binding thread, else group root", would drift, and the drift would show up as a chat whose cards and whose tickets disagree about where the chat's messages go — a routing bug nobody reports because each half looks correct on its own.
Both fan-outs resolve their topics in one query each, so a chat's page never pays a per-chat lookup: the internal card's lookup covers only the chats that will actually be posted to, and the client fan-out gets its own, after the internal card has gone out. Only the client fan-out resolves before its message is built — the internal card is rendered once per invocation, before the lookup, because that render is also what records the chronic and blast-radius metrics and must happen exactly once whether or not any chat needs a post. Either way the lookup is one query per fan-out and never one per chat. A lookup that fails degrades — every chat falls back to its binding thread and the failure is logged — because routing must never delay or cost a page.
What updates a card
| Trigger | What it covers |
|---|---|
| The alert service | Every human state change — acknowledge, resolve, silence, unsilence — from the web UI (single and bulk), Telegram buttons, the voice keypad and the triage chat's acknowledge tool. The request is made after the change commits, never on error, and no handler carries card code of its own. |
| The triage worker | A finished AI analysis, once it has been persisted — the "what it probably is" line, with whatever trust framing the analysis carries. Asked for after the write commits, on both persist paths (a completed investigation and a run stopped by its deadline or token budget, which still records a best-effort answer), and never on a failed one. |
| The alert worker | The source resolving a group — its last firing alert resolving, or a group-level resolve — and a group reopening in place. |
| The notification worker | The first post, when paging arms; a dispatch for a group whose card already exists — a redelivery, or a new firing epoch — which edits rather than posts; and every fresh post, which asks for one follow-up so an acknowledgement landing while the message was being posted still reaches it. |
| The silence sweeper | A silence reaching its end time — the group goes back to acknowledged or firing — and a silence that had ended more than 24 hours earlier, whose group the system resolves. It asks for the edit after each change commits, as the alert service does. |
| The card reconciler | Every 30 seconds, on every replica that has the ops bot token: a tracked message whose recorded state disagrees with its group's status — an internal card in a chat still bound to the group's client for the internal audience, a client-facing status message in a chat still bound client-facing to that client, on a group updated in the last 7 days, or a personal page, which has no binding and no recency bound — at most 200 groups a sweep, across all three, most recently changed first. It catches the changes nothing asks about — the "still on it?" loop auto-unacknowledging a group, the stale-group reaper resolving a group (it returns only a count), a request lost when a pod died (requests live in memory) — and an edit whose retries ran out. A group it selects is synced whole. A group it keeps selecting under the same status is requested again after a growing wait, up to an hour. |
An analysis appears on the card as soon as it is written, rather than waiting for the next unrelated event. It is the one update nobody asks for: it arrives unbidden, minutes into an incident, and it is what a responder is waiting on while nothing else is happening — which is exactly the state in which nothing else would have redrawn the card. Like a note, an analysis changes no alert state, so this edit is forced past the usual "the message already shows this state" check; without the force it would render correctly into a message that is never sent, and surface retroactively at the next acknowledgement, resolve or note. Measured before the fix: an analysis written at 14:59:38 reached the room at 15:06:12, when a human happened to write a note. Unlike a note, an analysis is once per firing episode — an episode is analyzed at most once, and only an explicit re-triage runs another — so it needs no volume argument beyond the Syncer's own. The request is best-effort and made after the analysis has committed: a Telegram problem can never delay or fail a triage run, and the analysis is in Console either way. The reconciler below is no floor under this edit, for the same reason it is none under a note — its staleness test derives from alert_group.status, which an analysis does not change — so the forced request is the only thing that carries it.
The event triggers make a transition immediate; the reconciler is the floor under them: a message it covers is requested again within one interval of disagreeing with its group — later, within its backoff, when the same group keeps disagreeing under the same status — and catches up unless Telegram refuses the edit. It is not a floor for a message whose chat has since been unbound or re-bound, nor for a client-facing message on a group older than 7 days — see known limits. Two replicas asking for the same update is harmless, because a message that already shows its state costs no Telegram call.
One Syncer per process
Every trigger above goes through one Syncer per backend process. Its concurrency cap, per-group coalescing, retry-after floor and memory of refused edits live in that instance, and they exist to protect the Ops bot token, which also delivers escalation pages. A second instance would double the cap on the same bot and split the caches. Without PROXIMA_OPS_TELEGRAM_BOT_TOKEN there is no Syncer, no notification worker and no reconciler: nothing posts a card, and nothing is edited.
How an edit is delivered
- An unchanged state costs no Telegram call.
telegram_message.staterecords what the message shows, and a request for a card that already shows its group's state makes no Telegram call (the run still reads the group and builds the card to find that out). This holds across pods and restarts, and it is what makes a storm cost one message and no edits. - The last state wins. Requests for one group are coalesced over a 2-second window: the first runs at once, and everything inside the window becomes one trailing run that re-reads the group. Nothing is dropped — acknowledge-then-resolve within two seconds ends on resolved.
- A 429 defers the edit. A transient failure (a 429, a 5xx, a transport or decode error) is retried up to three times per burst, backing off 2s, 4s and 8s with ±20% jitter, and never sooner than Telegram's
retry_after. After the third retry the burst gives up and logsedit retries exhausted; the card then waits for the group's next sync, or for the reconciler. - "Message is not modified" is success. The message already shows exactly this, so the state is recorded and nothing is retried.
- A permanent failure is remembered, not repeated. A deleted card, a bot removed from the chat, or any other refusal leaves the stored state unchanged — it never claims the message shows what it does not — and the same edit to the same message is not attempted again for an hour, unless the card's state or audience changes or the message is re-posted.
- At most four runs are in flight across all groups, because every run reads from the database pool and edits on the bot token that pages share. A run that cannot get a slot in time is retried, within the same three-retry budget.
When edits fail: what operators see
Two counters say how card edits are going; the log lines below say which card. Neither counter carries an alert group, chat or client, and no alert rule ships for them — they are collected only.
| Metric | What it counts |
|---|---|
proxima_telegram_card_edits_total{result} | One per tracked message a run tried to bring up to date, by what came of it: ok; not_modified (Telegram said the message already shows it, recorded as delivered); transient (a 429, a 5xx, a transport or decode failure — retried); permanent (not retried by the run: Telegram refused it for good, which is remembered for an hour, or the failure was unclassified); skipped_known_failure (no call made, because the same edit to the same message already failed permanently within that hour). A message that already shows its group's state is not an edit and is not counted. |
proxima_telegram_card_edit_retries_exhausted_total | One per burst of runs for one alert group that still needed a retry after its three retries — the edit retries exhausted line below. |
Both are per replica. A healthy system shows none of the WARN lines below.
| Log line | Level | What it means |
|---|---|---|
alert card sync: edit failed | WARN | Telegram refused an edit; the card in chat_id still shows the previous state. With retryable=true it will be retried. With retryable=false it is not retried, and a refusal Telegram classified as permanent is remembered until the state changes. Counted by proxima_telegram_card_edits_total as transient or permanent. error carries telegram_edit_transient or telegram_edit_permanent and Telegram's description. |
alert card sync: edit retries exhausted | WARN | A group's runs still needed a retry after the burst's three retries: its edits kept failing transiently, a read kept failing, or no run slot came free. Counted by proxima_telegram_card_edit_retries_exhausted_total. The card stays stale until the group's next event-driven sync, or the reconciler's next sweep — for an internal card in a chat still bound internal, or a client-facing message in a chat still bound client-facing on a group updated in the last 7 days. A steady stream of 429s means the 2-second window is too short for the traffic, which is a real finding. |
alert card sync: build failed | WARN | The group could not be loaded. Retried, unless the group no longer exists. |
alert card sync: list bindings failed, alert card sync: tracker read failed | WARN | A database read failed. Retried. |
alert card sync: state write failed | WARN | The edit landed but recording it failed, which costs one extra, harmless edit later. |
alert card sync: audience write failed | WARN | A message was re-rendered for the stricter, client audience, but recording that audience failed. The next sync edits it again, harmlessly, and still renders it for the client. |
alert card sync panicked, alert card async sync panicked | ERROR | A bug. It is contained and logged with its stack; the page and the action that asked are unaffected. |
card reconciler sweep failed; retrying next tick | WARN | The reconciler's query failed, so cards stay as they were for one more interval. |
internal card: Telegram refused the card's markup; posting the plain-text page instead | WARN | Notification worker. Telegram rejected the card's HTML as unparseable (telegram_render_rejected) — a rendering bug. The chat gets the plain-text page with Acknowledge and Resolve instead of nothing. That message is not tracked, so nothing edits it afterwards; the line carries alert_group_id so the card that failed can be rebuilt and fixed. |
internal card: the plain-text fallback failed too | WARN | The plain-text page failed as well. The dispatch is retried. |
internal card: build failed | WARN | Notification worker. The card for a new post could not be built. This is the one card failure that fails a dispatch: the page is retried and redelivered rather than silently never reaching the room. |
At DEBUG there are also edit retry scheduled, skipping an edit that already failed permanently, message already shows this state, and no run slot free before the deadline; retrying.
Known limits
-
Nothing posts until paging is live for the alert's team, and that takes rollout steps no deploy performs. A card needs all of:
PROXIMA_OPS_TELEGRAM_BOT_TOKENon the backend; the client'soncall_enabledfeature (the paging half of the retiredl1_agent); a chat bound with/bind <client-slug> --audience internalwhose binding does not mute alerts; an escalation route that matches the alert to a team-owned policy; and that team flipped live through the on-call cut-over (POST /api/v1/admin/oncall-cutover/{teamID}, super-admin, gated). Teams start in shadow, so a fresh deployment posts no cards.GET /api/v1/oncall/readiness/coveragereports, per client,matchable_route_count,oncall_enabledandlive_route_count— the matchable routes whose owning team is live — so a zerolive_route_countshows the flip has not taken for any route. It does not say whether a chat is bound, or whether a route matches a particular alert. -
A silence ends up to 30 seconds after its end time, the silence sweeper's interval, and a silence whose end had passed more than 24 hours before the sweeper noticed it resolves its group rather than resuming paging.
-
A silence with no end time, set before this release by a race, is not ended by the sweeper — unsilence it in Console. Before this release, a delivery arriving just as the sweeper ended a silence could write
silencedback with the end time already cleared. The sweeper finds silences by their end time, so such a group stays silenced — and pages nobody — until a person unsilences it. They are deliberately not swept: how many exist is unknown, a missing end time says nothing about how old the incident is, and restoring them could page for incidents that ended long ago. Finding and handling them is a named follow-up. -
A client-facing status message's floor covers groups updated in the last 7 days. The reconciler selects a client-facing message whose chat is still bound client-facing to the group's client and whose text disagrees with the group's status — after a reaper resolve, an auto-unacknowledge, a lost sync request or exhausted retries — but only for a group updated within the last seven days. Client messages tracked before anything edited them still hold the default investigating, including for groups resolved long ago; the bound keeps the first sweeps from editing that history into customers' chats. An older group's message is corrected only if the group changes again.
-
A message whose chat was unbound or re-bound is outside the floor. Neither the internal nor the client-facing selection picks a message in a chat that is no longer bound to the group's client, or is now bound for the other audience. A re-bound message is corrected at the group's next event-driven sync, which renders it for the stricter, client audience; an unbound one is never edited again.
-
Resolve-then-refire is one card per episode, because a refire after resolution is a new alert group, which posts its own card when its paging arms — except, rarely, a refire that races the resolve, which reopens the resolved group in place: its card is edited back to firing, which notifies nobody, and paging can arm again under a new firing epoch.
-
The card's times are in the server's time zone, not the responders'. "Until 09:00", the silence end and the timeline's stamps all use the backend process's local zone — UTC in a container that sets no
TZ— and the card names it, so a responder at UTC+5 who taps 🔇 until 09:00 UTC at 02:00 local is silenced until 14:00 local, as the button says. A configured ops time zone is a follow-up. -
A failed topic lookup is a log line and nothing else. When per-kind routing cannot read
telegram_chat_topic, every chat in that fan-out falls back to its binding thread and a warning is logged — the page is never delayed for it. But there is no counter on that branch, so from the outside a chat that has no topic configured and a chat whose topic table is unreadable look identical: both deliver to the binding thread, quietly. The alert path and the service-desk path log the same sentence deliberately, so one log query surfaces both; a metric on it is a named follow-up. -
The card metrics are collected only. No alert rule ships for
proxima_telegram_card_edits_totalorproxima_telegram_card_edit_retries_exhausted_total, so a card fleet going stale pages nobody; it shows on the counters and in the logs. -
A reconciler sweep shares the four run slots with human-driven edits. In a flood of stuck cards, an acknowledgement's card edit can wait behind a sweep. This delays the card only, never a page. The cost is bounded per group: a group the sweep keeps listing under the same status — typically a card Telegram refuses to edit for good — is requested again after 30 seconds, then 1, 2, 4 minutes and so on, up to an hour, when the Syncer would retry the refused edit anyway. A status change is requested at once, and a group that catches up is forgotten. The backoff is kept in memory on each replica.
-
An escalation page's card in a private chat is not edited. Only the notification worker's posts are tracked, so a card sent to a person by Telegram DM keeps the state and the buttons it had when the page went out. A tap on it after the group changed acts on the group as it now is, and is answered with what that tap meets — Already resolved, Changed meanwhile — try again, or the action's own toast.
-
Muting a binding stops new posts, not edits. A card already in a chat whose binding later mutes alerts still follows its group.
-
Rollout precondition: before this reaches a large deployment, run
EXPLAIN ANALYZEon the reconciler's query (ListStaleCardGroups) against a production-shaped copy.telegram_messagehas no retention, so it only grows, and the plan can drive a per-client scan. -
A reply to a superseded card records nothing.
telegram_messageholds one row per(alert group, chat)and a re-post overwrites the message id, so a reply to the message that card used to be resolves to no alert. Nothing is written and the reply falls through to the AI Investigator. The group whose card that message once was is not recoverable from the table, and attaching the words to the chat's newest card instead would file them against the wrong incident. -
Notes appear in Console, not in the Mini App, which has no alert surface at all — there is nothing there to show a note on. The alert detail page in Console is the only place a note is read back.
-
A note's words, and its author, appear on the card. The timeline prints
Jamshid Yerzakov: checked the replica, lag is real— a note reads as the person speaking, because once the words are on the card the word note adds nothing and the name is the part that matters. An author Console cannot resolve keepsnote (user): …, since a bareuser: …would read as somebody's username. The text is capped at 120 runes with…and its newlines folded to spaces; the full note is one tap away in Console. Both are load-bearing rather than tidiness: a note may be 2000 runes against a 4096-character message, and notes are the newest entries while the budget drops the oldest, so an uncapped note would evict the timeline, the raw payload and eventually the analysis — and because the budget trims whole lines, a note containing a newline could have half its sentence dropped on its own.Every other action prints as action (who) — who being the person's display name on a human row, resolved for every human on the timeline so an acknowledgement reads the same whether or not that person also left a note, and the bare actor type when no name resolves. Per line that is usually one audit entry, except that consecutive system entries sharing a second are folded onto one line (
created, correlated), because they are one event to a reader; the fold is capped at three actions so a burst cannot become one line the budget is unable to trim. Human entries never fold: two people acting inside one second is two facts, and hiding who did what is worse than a longer section. A folded line still counts for every entry behind it, so the "N earlier entries" marker stays true when the budget drops it. A refused resolve's reason is not in the timeline at all — the card carries it as its own⚠️ STILL FIRINGline. -
The timeline says who was paged, on what, and what came of it. Each attempted page writes a
page_sentrow, so the card answers the question a responder reading the room actually has — is anyone being called?11:28:10 paged Jamshid Yerzakov by phone
11:28:12 paged Jamshid Yerzakov on Telegram
11:32:33 page to Aziza R. by phone failed — voice_no_answer
11:32:35 page to Aziza R. on Telegram failed — telegram_forbiddenA failed page is as much a fact as a delivered one. Silence on the card otherwise reads as somebody is being called, which is the worst thing for a reader to believe wrongly.
The row names a PERSON, never a destination. The dispatch ledger's
target_refis a raw E.164 number or chat id, which the store refuses to let reach a log line; the internal card is read by everyone in a group chat, so the row carries the paged user's id and the card resolves it to a display name. An id that will not resolve drops the name rather than printing a raw uuid at somebody.A failure shows its classified code, not the provider's words.
dispatch.SendError.Codeis documented as a bounded, low-cardinality label that never carries a phone number or chat id — that is what the card shows. The provider's exact text routinely embeds the destination it failed to reach, so it is recorded in the row's metadata for Console, where the reader is already authenticated, and never rendered into a chat. An error the dispatch layer never classified yields no code at all: the card then says a page failed without claiming to know why, which is true.The row is append-only — two pages to the same person are two facts, and an at-most-once fence keyed on the episode would erase the escalation that is exactly what a reader is looking for. The ledger's own
Claimalready makes the send at-most-once. The chain's ledger step is deliberately not shown: it is an opaque key (a base plus a generation stride plus a cursor), not the policy step a reader would recognise, so printing it would invent a number nobody can act on. Writing the row is best-effort by contract — detached context, its own deadline, panic-caught — because telemetry must never delay a page. -
A reply in a client-facing group reaches nothing. The note path declines it — a customer's prose must not become an operator-attributed entry in the internal record — and it no longer falls through to the AI Investigator either: the engage path refuses a
client-audience chat ahead of every other question about it, so no 🤖 answer appears in a customer's own room. It used to, and was observed doing so in production. See A client group never gets one for why that refusal is keyed on the room's audience rather than on the per-chat switch. -
Not built: bot provisioning, Slack, maintenance windows, and routing by project rather than by client and environment.