Ops Bot
The Ops Bot (@proximaopsbot) is the operations-facing Telegram bot. Unlike the Service Desk Bot (which is per-user and ticket-focused), the Ops bot is group-based: you add it to a Telegram group, bind that group to a Console client or team, and the group then becomes a live operations surface — alerts, Jira ticket activity, and an on-demand AI Investigator, scoped to what that bind grants — one client, or the clients a team owns.
It authenticates to the Console backend with a service-account API key and acts on the same RBAC seam the web UI uses: every action a person takes (acknowledge, resolve, silence, escalate, ask the AI) is attributed to their linked Console account and gated by their permissions.
The Ops bot is being consolidated onto @proximaopsbot (the bot already present in the client/team groups). Its routing config is moving off the legacy n8n bot onto Console — see the Overview engineering note.
What it does
1. L1 alerts (page + act)
When an alert's paging arms — a route matched it, and the team that owns the page is live rather than in shadow — Console posts to the groups bound to that client, sending as the Ops bot. That needs the ops bot token (PROXIMA_OPS_TELEGRAM_BOT_TOKEN) on the backend and the client's oncall_enabled feature (paging), and teams start in shadow, so nothing posts until the alert's page-owner team is flipped live. An alert no route matches, or whose page-owner team is in shadow, posts none. Arming is not the same as reaching someone: an armed alert can still page nobody — nobody on call, no verified contact — and its card posts anyway. The message depends on the group's audience:
- Internal (responder) groups get the alert card — one message per alert group, edited in place from firing to resolved, with the controls for its state: Ack, Resolve, ⚡ Escalate, bounded 🔇 silences, 🔔 Unsilence and Open in Console ↗ — unless the binding mutes alerts. See Telegram Alert Card.
- Client groups get a severity-gated, status-only message — P1 and P2 only, unless the binding opts out; no internal detail, no action buttons.
A person paged by Telegram direct message gets the same card in that private chat, and its controls act there too — see Authorization. That page is not edited afterwards, so it keeps showing the alert as it was when the page went out.
Tapping a control applies the action in Console, attributed to your linked user and checked against your alerts:write permission on the alert's environment. The card and the client-facing status message are then edited in place as the incident progresses, whether the change came from Telegram, the web UI, the alert source or a silence reaching its end time.
2. Jira ticket activity
The Ops bot delivers Jira ticket activity — new tickets, comments, and status changes — to the right groups: the responsible team group and the client group. Which group receives what is driven entirely by the /bind config (client ↔ project labels ↔ group, team ↔ group), so onboarding a new client/team is just a /bind in their group.
3. Reactive AI Investigator
@mention the bot — or reply to one of its messages — in a bound group, e.g.:
@proximaopsbot is web-prod-01 healthy right now?
…and it runs a single read-only AI Investigator turn and posts a 🤖 answer. It has no write tools and its output is secret-sanitized.
Two different questions decide what it may look at, and running them together is the easiest way to misread this feature:
| Decides | Answer | |
|---|---|---|
| Room scope | what this chat may ask about at all | the bound client (a client bind), or every client the bound team owns (a team bind) — either way, only those with ai_triage on |
| Turn scope | what one answer reads | exactly one client, every time |
The room
A chat bound with /bind <client-slug> has a room of one: that client, as it always had.
A chat bound with /bind --team <name> has a room of that team's clients — the team's client list, filtered to those whose ai_triage flag is on. Nothing is granted to the chat separately: the room is the set the team already owns, so widening or narrowing what a chat may ask about is done by changing the team's clients, in one place, where it is already audited.
A chat carrying both binds is resolved by the team, not by the union of the two — a union produces a set you cannot predict from either row alone. If the group bind names a client the bound team does not own, Console logs a warning and still uses the team's set; that combination is a configuration mistake, not a grant.
Which switch governs the chat
There are two reactive_enabled columns, and which one is read depends on whether the chat has a group binding, not on where its room came from.
| The chat has… | Its room is | Its switch is | Ships |
|---|---|---|---|
| a client bind only | that client | telegram_group_binding.reactive_enabled | on (true) |
| a team bind only | the team's clients | telegram_team_binding.reactive_enabled | off (false) |
| both | the team's clients | telegram_group_binding.reactive_enabled | on (true) |
The third row is the one to read twice. A chat carrying both bindings takes its room from the team and its switch from the group: the team column is never consulted, because the handler reads it only when there is no group binding to read. So the widened room arrives already switched on, under a column that has been true since the chat was bound.
This is not gated behind the new column. The one existing internal chat bound to both a client and a team goes from asking about one client to asking about its team's clients with no switch flipped and nothing to flip — in practice from the moment the ops bot finishes rolling, since until then the asker gate silences it entirely (see the deploy note below). If you are debugging why such a chat is or is not answering, telegram_team_binding.reactive_enabled is the wrong column — it is not read for that chat at all.
To silence a dual-bound chat, clear telegram_group_binding.reactive_enabled. To narrow what it may ask about, change the team's clients.
telegram_team_binding.reactive_enabled defaults to false — deliberately the opposite of the group binding's true — so a chat bound to a team and nothing else has no Investigator until someone turns it on, and applying the migration starts one in no such chat. There is no command and no admin route for it:
UPDATE telegram_team_binding SET reactive_enabled = true WHERE chat_id = -1001234567890;
A re-bind does not reset it. /bind --team "Other Team" on a chat that already has a team binding overwrites the team, thread and mention handle, and deliberately leaves reactive_enabled alone — the same rule that keeps a re-bind from silently undoing an admin's notify and language choices. The consequence is worth saying out loud: a chat whose Investigator you turned on keeps it on, now pointed at the new team's client set. Clear the column in the same breath if that is not what you meant.
One answer reads one client
The model never picks. A component that chooses what it may read is not a boundary, so the choice is made in Go — before any auth context is built, before any client is loaded, before a token is spent. In order:
- A room of one — nothing to disambiguate.
- The alert you replied to, when that alert's client is in the room. An alert that reached the chat by a forward or a re-post, whose client the room does not hold, is not an anchor: it is dropped, and the turn does not read it either.
- One in-scope client named in the message — whole-word, case-insensitive, matched against the client's name or slug. Multi-word names need every word, so "how's inter today?" does not select Inter Nation, and non-Latin names match (an ASCII-only tokenizer made a Cyrillic client name unmatchable forever).
- Otherwise it asks which one. Two in-scope clients named in one message is ambiguity, not a pick: resolving it by first match would make a tenancy boundary a coin flip.
The question it asks back — "I can look at Inter Nation or Proxima — which one?" — is a fixed string built in Go, never model output, because it is the one reply given before anybody has been authorized for anything. It costs nothing: no client load, no LLM call.
That reply lists every client in the room, and it is produced before the asker is authorized for any of them — gate 6 has not run yet. So any linked, active Console user in a team-bound chat can learn the team's full client list with one unanchored message, including one who holds no permission on a single client in it and will be refused an answer about every one of them.
It is bounded by the identity gate, which runs earlier: an unlinked member, a deactivated account and a stranger who was added to the group all get silence, not the list. So this is not "anyone in the room" — it is "anyone in the room who is somebody in Console". That follows from granting the room those clients at all, and it is written here rather than left to be discovered.
Who may ask
Being in the room is not permission. Two gates apply to the person, and both apply in client-bound chats too — this is not a team-chat feature:
- Identity — the asker must be a linked, active Console user (
/link <code>). Disabling someone's Console account silences their Investigator immediately, in every chat at once, with nobody touching Telegram. That is the offboarding hook, and it is why the bot keeps no identity table of its own. - Authorization — that user must hold
alerts:read,hosts:readandlogs:readon the turn's client. That is exactly the set the Investigator reads with: answering someone short of it would launder privilege through the bot — it would tell them things their own role forbids. It is checked against the turn's client, never the room, so a responder who may read one of a team's clients is not answered about the others merely because the team owns them.
The permission check runs through the same access resolution the web UI uses, so a super-admin (who holds no per-client rows at all) and an environment-scoped responder are both judged correctly rather than being refused by a raw lookup. It is not environment-precise, though, and nothing on this path is: the check is ac.Can(perm, clientID) — the same deliberately coarse predicate /bind --team is measured with — so a responder holding the read set on one environment of a client authorizes a turn that may read every environment of that client.
That is deliberate and it is intra-tenant: the turn is still pinned to a single client, and this is strictly tighter than what it replaced (a chat, with no question asked about the person at all). But if you hand out environment scope expecting it to be a boundary, the Investigator is not one — grant it to people you would let read the whole client.
A client group never gets one
A client-audience chat is refused first, ahead of every other question about the chat, and no switch changes that. The Investigator does not summarize — it reasons aloud about the client's own infrastructure: which hosts it read, what it ruled out, what it suspects. In a responder chat that thinking-out-loud is the product. In the customer's room it is a half-finished diagnosis of their own estate, in our voice, with no operator in between.
When the bot says nothing
Every refusal is silent — engaged=false, nothing posted — so from inside the chat they are indistinguishable. The backend can tell them apart. This is the order the gates run in, so it is the order to check:
| Check | Where |
|---|---|
Is it a client-audience group? | nothing below matters — refusing is the design |
| Is the reactive switch on? | Whichever column the chat's bindings select — see the table above. A chat with a group binding is governed by telegram_group_binding.reactive_enabled even when its room comes from a team; only a team-bind-only chat reads telegram_team_binding.reactive_enabled. Reading the wrong one is the likeliest wrong turn on this whole list. |
| Is the asker a linked, active Console user? | users.traits.telegram_id, users.status |
Is ai_triage on for anything in the room? | Client Settings → Features → AI investigation. It filters the room, so a team whose clients all have it off has an empty room and answers nothing. |
Does the asker hold alerts:read + hosts:read + logs:read on the one client this answer reads? | their roles |
ai_triage is then read a second time, on that same client, immediately before the LLM call — the room is resolved once per turn, but the gate protecting spend belongs next to the spend, so a flag flipped off mid-turn still costs nothing.
Two Info lines from the api.telegram_engage component name the two refusals a rollout produces most:
| Log field | What it means | Who fixes it |
|---|---|---|
refusal=no_asker_id | the request carried no telegram_user_id — the ops bot has not been upgraded, so every chat is quiet at once | whoever sequences the deploy |
refusal=asker_unlinked | that person is not a linked, active Console user; everyone else in the room is unaffected | that person, with /link <code> |
Neither line carries the raw Telegram user id: users.traits.telegram_id IS NULL answers "who is unlinked" better than a log ever will, and the log's job is which refusal, in which room.
The identity gate is fail-closed on purpose — tolerating a missing telegram_user_id would leave a permanent bypass any caller could trigger by omitting the field — so a backend that reads the field while the bot still omits it refuses every turn with refusal=no_asker_id.
There is nothing to sequence. The two images ship together: CI rewrites every console-system image reference to the same CI_COMMIT_SHA in one argocd commit, the ops bot's among them, and docker:opsbot always rebuilds on main and is one of the deploy job's needs — so a failed bot build blocks the bump rather than publishing a tag with no image behind it.
What remains is a rollout window, not an ordering decision. One commit updates both images, but they do not begin serving at the same instant: the backend ships as an Argo Rollouts blue-green, so its traffic moves when the rollout is promoted rather than draining gradually, while the ops bot's pods are replaced on their own schedule. While a new backend is answering and an old bot is still sending the old payload, every turn in every chat is refused and the room is quiet.
So a room widens when the BOT finishes, not when the backend does. A promoted backend with an unrolled bot is silence, not a wider room — gate 3 refuses before room scope is ever resolved, so nothing about the team's clients has been looked at yet.
refusal=no_asker_id around a deploy is that window. If it persists, the bot side has not finished — a rollout still awaiting promotion, or pods that never came up. This page deliberately does not say which: the ops bot's own rollout shape is defined in the argocd-infra repository, not in this one, so check it there rather than assuming either.
A team whose every client has ai_triage off has an empty room, and an empty room answers nothing. That is the flag working, not a fault — and it is the one case where the bot is quiet for everybody in the chat regardless of who they are or what they hold.
4. Replies to an alert card become notes
Reply to an internal alert card in your own words and the bot records your reply as a note on that alert's timeline, attributed to your linked Console user. No command, no button — the reply is the gesture. The bot acknowledges it with an ✍️ reaction on your own message rather than a new message, because these groups also carry pages and every extra line pushes a live alert further up the scrollback; the note itself is read in Console, on the alert's detail page.
If the reaction cannot be set — reactions restricted in that chat, the message too old, the bot lacking rights — the bot falls back to posting 📝 noted instead. The fallback is not optional: a note that landed with no acknowledgement at all looks exactly like a note that did not land, and you would reasonably retype it.
Telegram accepts reactions only from a fixed emoji set, and 📝 is not in it — setMessageReaction answers REACTION_INVALID for anything outside the set. Built from 📝, the reaction would fail every single time and silently fall back to the text receipt: the feature would look implemented and change nothing. A test pins our emoji against that set.
A reply to any other message of the bot's still reaches the AI Investigator exactly as before, and so does a reply the backend does not record — a message that is not a tracked card, or a client-facing status message. A recorded note never also spends an AI turn, and an AI answer is never also filed as a note.
- @mentions and
/commandsare never notes, even when they are replies: "@proximaopsbot what caused this?" is a question for the Investigator, not a note in the asker's name. - A refused reply is silent. An unlinked account, or a linked user without permission on this alert, gets no note and no message in the room.
- A reply over 2000 runes is refused, not truncated, and told so once: 📝 Too long to record as a note — shorten it, or add it in Console.
- Replies are redacted before they leave the bot, so a secret pasted into a reply does not reach the record.
A reply to a client-facing status message records no note — customer prose must not become an operator-attributed entry in the internal record. It no longer reaches the Investigator either: a client-audience chat is refused before any other question about the chat is asked, so a customer replying to a status message gets no 🤖 answer in their own room. That used to happen, and was observed in production; see A client group never gets one for why the refusal is audience-based rather than a per-chat switch.
Audiences: internal vs client
Every group is bound at one audience, which decides what it receives:
--audience | The group is… | Receives |
|---|---|---|
internal | a ProximaOps responder group | the alert card + its controls |
client | a customer-facing group | severity-gated, status-only (no buttons, no internal detail) |
client is the safe default, and it is the one the bind paths apply. /bind refuses a request that names no audience at all; the admin bind API and config-as-code both default an omitted audience to client — customer-facing, no buttons. (The column still defaults to internal, from migration 000086; no write path relies on it, and nothing should be written that does.)
Promoting an existing client group to internal requires super-admin — it would start pushing that chat full detail + action buttons, a privilege escalation of the chat itself. The admin API and pc apply require super-admin to bind a chat fresh as internal too; /bind gates only the promotion, having no existing row to compare a first bind against. Audience is a statement about who is in the room, not a switch for what the bot may do — see Two bindings, two questions.
Commands
Run these inside the group you want to configure (the bot is group-only for binds) — with one
exception: /start is the paging-verification command and only counts in a private chat.
| Command | What it does |
|---|---|
/start vfy_<token> | In a private chat only. Redeems a paging-verification deep link, proving this bot can DM you — the fact traits.telegram_id can never establish. See Personal Pages. Sent anywhere but a private chat it verifies nothing and says so; sent bare, it replies with help. |
/link <code> | Links your Telegram account to your Console user (generate the code in the Console profile). Required before you can bind or act on alerts. |
/bind <client-slug> --audience internal|client [--team <name>] | Binds this group to a client at the given audience. Captures the forum topic it was run in. |
/bind --team <team-name> [--mention <@handle>] | Binds this group to a cross-client team (no client). --mention sets the team-lead handle pinged on status notifications. The chat's Investigator can then cover every client that team owns. If the chat has no group binding it stays off until telegram_team_binding.reactive_enabled is set; if it already has one, that binding's switch keeps governing and the wider room is live immediately. |
/topic alerts|tasks|uptime | Sends that kind of message to the forum topic you run it in. See Forum topics: routing by kind. |
| (inline buttons) | The alert card's controls: Ack, Resolve, ⚡ Escalate, 🔇 1h / 4h / until 09:00, 🔔 Unsilence. |
| (reply to an alert card) | Records your reply as a note on that alert's timeline, attributed to your linked user. Acknowledged with an ✍️ reaction on your message (falls back to 📝 noted if the reaction is refused). |
| (@mention / any other reply) | Triggers the reactive AI Investigator — if the chat's reactive switch is on, and you are a linked, active Console user with alerts:read + hosts:read + logs:read on the client in question. |
Forum topics: routing by kind
The audience decides which chats a message reaches. For the kinds below, a kind decides where inside one chat it lands, when that chat is a forum with topics. Without it, a group that wants both alerts and ticket activity gets them in one thread: pages and Jira comments interleaved in the room people are meant to read during an incident.
Not every message Console sends is routed by kind. Per-kind routing applies to the L1 alert card, the client-facing status message, service-desk traffic, and uptime notifications — the producers listed in the table below. Uptime-monitor notifications are the one producer with a longer fallback: rather than a topic or the root, they follow their own ladder (explicit thread → uptime topic → alerts topic → binding thread → root — see Telegram targets). Everything else still goes where it always went: the daily briefing and other broadcasts (internal/broadcast) use the binding's own thread, and the resolution-memory draft card (worker/memory.go) does the same — note that this last one goes to the same internal responder chats the alert card goes to, so a chat with an alerts topic sees the card in the topic and the draft that follows it in the root. Configuring an alerts topic does not move the briefing or the memory draft, so a chat that has one can still see those arrive elsewhere. Bringing them in is a code change per producer, not a configuration one.
There are three kinds:
| Kind | What it carries |
|---|---|
alerts | the alert card in an internal group, and the client-facing status message in a client group |
tasks | service-desk traffic — new Jira tickets, comments, and status changes |
uptime | a monitor's or service's down/up message to the client's uptime targets. A chat with no uptime topic uses its alerts topic |
The vocabulary lives in Console's Go code, not in the database — telegram_chat_topic has no CHECK on kind — so adding another kind later is a code change and a line in this table, not a migration.
Setting a topic
Run the command inside the forum topic you want:
/topic alerts
The bot answers 📍 alerts will post in this topic, and from then on this chat's alert cards land there.
The thread comes from the message, never from an argument. That is the whole point of running it inside the topic: there is no id to type, so there is no id to get wrong, and the command cannot be aimed at a topic you are not looking at. It is the same mechanism /bind already uses to capture a topic.
- Run in the group root it is refused (with the usage text), because "post in the root" is expressed by having no route at all, not by a route pointing at the root.
- A kind the build does not know is refused rather than stored, so a typo cannot create a row nothing will ever read.
- Re-running it in another topic moves the kind; there is one route per chat per kind, and the newer one replaces it.
- A failure is answered in the room. A
/topicis a person standing there waiting, and silence would read as success — the one outcome worth preventing, since they would walk away believing the routing is live.
Routes are stored per chat, not per binding. A chat can carry a client binding and a team binding at once, and a forum topic is a property of the chat itself — keying it to one binding would hide the routing from the other.
Setting a topic from Console
The same routes are readable and writable over the admin API, for a UI that would rather offer a menu than ask an operator to go type in Telegram:
| Endpoint | What it does |
|---|---|
GET /api/v1/admin/telegram/groups/{chat_id}/topics | the topics the bot has staged for this chat (the menu) and the routing configured now (the state), in one call |
PUT /api/v1/admin/telegram/groups/{chat_id}/topics/{kind} | points one kind at one topic |
DELETE /api/v1/admin/telegram/groups/{chat_id}/topics/{kind} | removes the route |
PUT accepts only a thread the bot has actually seen in that chat. An unstaged id is refused, because a message sent to a topic that does not exist fails silently in the room — the operator would see a saved setting and no messages.
What happens when a kind has no topic
Nothing moves. The rule, in order:
- the chat's configured topic for that kind, if it has one;
- otherwise the thread the chat's binding already names (what
/bindcaptured); - otherwise the group root.
So a chat nobody has configured delivers exactly where it delivered before topics existed, and removing a route is how you undo one — the kind falls back to the binding's thread. This is one function in Console (store.ResolveThreads), used by both the alert path and the service-desk path, deliberately: the fallback expressed twice would eventually be two different fallbacks.
It also means the feature ships inert. A deploy creates an empty table, and an empty table is indistinguishable from the feature not existing — pinned by an acceptance test that resolves every chat through both the current resolver and the pre-topics one and requires them to agree.
The per-chat mute switches stay exactly what they were, checked independently of routing and before it. Routing only decides where inside a chat a message goes; muting decides whether the chat is messaged at all. A binding with notify_jira = false resolves no task target, and configuring a tasks topic on it does not bring one back.
Which switch mutes which message depends on the audience, and kind alerts spans both:
| Message | Muted by |
|---|---|
the alert card in an internal group | notify_alerts |
the client-facing status message in a client group | client_notify |
service-desk traffic (kind tasks) | notify_jira |
So notify_alerts = false silences a responder group's cards and nothing else — it does not silence the status messages a client group receives, even though both travel under kind alerts. Muting a client group is client_notify = false.
A failed topic lookup degrades rather than failing: the message goes to the binding's thread and Console logs a warning. Routing must never cost a page.
How to onboard a group
- Add
@proximaopsbotto the Telegram group (and, if the group uses topics, run the next step inside the topic you want notifications to land in). - Make sure your Telegram account is linked:
/link <code>(generate the code in the Console profile → Telegram). - Bind the group:
- Client's status group →
/bind it-unisoft --audience client - ProximaOps responder group →
/bind it-unisoft --audience internal - Cross-client team group →
/bind --team "Infra Team" --mention @lead
- Client's status group →
- (Optional, forums only) Split the traffic by topic: run
/topic alertsinside your alerts topic and/topic tasksinside your tickets topic. Skip this and everything keeps landing in the thread/bindcaptured — see Forum topics: routing by kind. - Done — the group now receives the matching alerts and Jira activity.
- (Responder groups only) Turn the Investigator on if you want it. A group with a client binding already has it — and keeps it, on that binding's switch, even if you later add a team bind. A group bound to a team and nothing else ships with it off and needs
telegram_team_binding.reactive_enabledset. Then @mention the bot to ask — see Reactive AI Investigator for what it may read and who it will answer, and Which switch governs the chat if you are unsure which applies.
Authorization
Defense-in-depth (all enforced server-side, never trusting the bot's request body):
-
Service-account pin — the endpoints accept only the pinned Ops-bot service account.
-
Linked user — the acting human is resolved from your Telegram link; unlinked users are rejected.
-
Client bind (
/bind) — requirestelegram_ops:writeand that your linked account has access to the target client. Re-pointing a chat to a different client, or promotingclient→internal, requires elevated access (super-admin / access to the current owner). -
Team bind (
/bind --team) — a team group has no client tenant boundary, so it requirestelegram_ops:writeheld somewhere: super-admin, or the permission on at least one client. That bar is lower than "global" suggests and the page says "somewhere" on purpose —ac.Canis deliberately coarse, so an environment-scoped grant on a single environment of a single client satisfies it. Whether a chat with no tenant deserves a stricter rule is a tracked follow-up, not a property of any one command. -
Topic routing (
/topic) — exactly the team-bind bar above,telegram_ops:writeheld somewhere, caveat and all, and for the same reason (the chat may carry no client at all) — literally the same function, shared on purpose so the two cannot drift apart. And, when the chat does have a client binding, access to that client. The client check is conditional rather than always-on precisely because topic routing exists to serve team chats too, and an unconditional client check would refuse/topicin the chats the feature was built for. Being in the group authorizes nothing: this endpoint writes the row the paging path reads, and a mis-setalertstopic sends a client's cards into a thread their responders are not watching. -
Alerts — every alert card control requires
alerts:writeon the alert's client and environment for your linked account, the same gate as the web UI. Where the tap happens decides one more check:- In a group chat, the group must also be bound to the alert's client. Anyone in a group can tap, so the binding keeps a chat bound to one customer from acting on another customer's alert. A tap in a group that is not bound to the alert's client is refused with Couldn't apply — chat is not bound to this alert's client.
- In your private chat with the bot (escalation pages, the "still on it?" nudge), the tap is authorized by your own permissions alone. No chat is ever bound to a private chat, so a binding check there would refuse every tap while proving nothing: being in a chat was never the authority, your RBAC is. A chat is private only when its id is your own Telegram id; anything else is checked as a group.
Only an unlinked account is told to
/link; a linked user without permission is told so. -
Notes — a reply recorded as a note runs the same gates in the same order as a button tap: the pinned service account, your linked user, the chat bound to the alert's client (skipped only in your private chat with the bot), and
alerts:writeon the alert's client and environment. The alert comes from the tracked card you replied to, never from the bot's request body, so a reply can only write to that alert. A refusal here is silent in the room — the bot posts nothing. -
AI — the reactive Investigator is read-only, and gated on four things, not one. In the order they are checked:
- The room's audience — a
client-audience chat is refused outright, before anything else about the chat is considered. - The chat's reactive switch —
reactive_enabled, on whichever binding the chat actually has. A chat with a group binding is governed by the group binding's copy, which shipstrue— and that includes a dual-bound chat whose room comes from its team. Only a chat bound to a team and nothing else is governed by the team binding's copy, which shipsfalse. See Which switch governs the chat; picking the wrong one of these two is the most common way to misread this list. - The asker's identity — a linked, active Console user. Not "somebody in the group": a room is not an access grant, and deactivating a Console account silences that person's Investigator in every chat at once.
- The asker's own permissions —
alerts:read,hosts:readandlogs:readon the single client this answer reads, resolved through the same seam the web UI uses. Checked per turn, not per room, so a team's chat cannot answer someone about a client they may not read.
ai_triageis woven through the middle rather than tacked on the end. It filters the room before the permission gate — a client with it off is not in the room, so it is never the turn's client and never reaches step 4 — and it is then read a second time on the turn's client, immediately before the LLM call, so a flag flipped off mid-turn still costs nothing. It is the only gate that is checked twice, and it is not a substitute for any of the four above.Every one of these refusals is silent; see When the bot says nothing for how to tell them apart from the logs.
- The room's audience — a
Deployment & identity
The Ops bot is a Go service (telegram/opsbot/) deployed to Kubernetes (console-system), long-polling Telegram and calling Console over CONSOLE_API_URL with CONSOLE_API_KEY. Its bot token is OPS_TELEGRAM_BOT_TOKEN; the backend sends as PROXIMA_OPS_TELEGRAM_BOT_TOKEN. Because this bot is the one that DMs per-user pages, it is also the bot paging verification proves a private chat with — so the backend needs PROXIMA_OPS_TELEGRAM_BOT_USERNAME (to build the t.me/…?start=vfy_… deep link) and PROXIMA_OPS_BOT_SA_ID (redemption is pinned to this bot's service account). See Deployment.
A reply tells the Investigator which alert you mean
The bot treats an @mention on an alert card as a question, not a note — someone asking "what caused this?" wants an answer, not their words filed on the record. That routing was only half a feature while the Investigator never learned what this was: the engage turn carried the chat and the question, so the model knew only which client was asking and had to guess among their open alerts.
A reply now carries reply_to_message_id, and the backend resolves it through the same LookupByMessage the note path uses. When it resolves to a tracked internal card, the Investigator is given the alert as context — its group id, severity, status and title — so it can call its read-only tools on the right alert.
- Internal cards only. A reply to a client-facing status message carries no subject, for the reason the note path refuses one: the internal record and a customer's room are deliberately different surfaces, and a reply there must not silently pull an internal alert into the model's context. (Such a chat is refused before this point anyway — the audience gate runs first.)
- In the room's scope only. A tracked card whose client is not in the chat's room scope is dropped rather than followed. A re-bind, or an
ai_triageflip that drops a client out of the room, can leave an old card sitting in a chat the room no longer covers, and narrating one client's alert inside another client's turn is exactly what turn scope exists to prevent. - Best-effort. A message that is not a reply, an untracked message, and a failed group load all answer without a subject rather than failing the turn. An answer with less context beats no answer.
- The subject line is sanitized like any other text that reaches the model, and it is stated as context rather than as an instruction.
A reply is a chat turn, not an investigation
The engage path runs the AI Investigator, and it used to run it with triage's budget: MaxLoops: 20 — the same as a full root-cause analysis — no token budget, and the expert model tier. Measured in production, one "@bot what caused this?" made eleven LLM calls over 82 seconds while the ops bot's HTTP client gave up at ten. The answer was produced, returned 200, and thrown away; the room saw nothing.
That was only visible once a reply carried its alert: with no subject the model had nothing to investigate, returned a canned greeting in 6–9 seconds, and slipped under the timeout by accident.
A reply is now bounded three ways, because each bounds something different:
| Bound | Value | Why |
|---|---|---|
engageMaxLoops | 4 | round trips — enough to read the alert, check a thing or two, answer |
engageTokenBudget | 15 000 | spend — one pass can pull a large tool result, so a loop count is not a spend bound |
engageRunTimeout | 45s | wall clock — the runaway did not stop when the caller hung up |
It also runs on the standard tier (deepseek) rather than the expert tier (sonnet): a responder asking a question mid-incident wants an answer in seconds, and the deep-analysis model belongs to triage, which has the time and the mandate for it.
On the bot side, one HTTP client with a single 10-second timeout served every call — a note that answers in 26ms and an LLM turn alike. Calls now carry their own deadline: 10s for the database round trips (bind, link, discover, note, callback) so the bot still notices a sick backend quickly, and 60s for engage alone, comfortably outside the backend's own 45s bound so the bot hears an answer or an error rather than silence.