Skip to main content

Notification Routing — Data Model & Onboarding

Console is the single source of truth for client notification routing: which Telegram group a Jira/alert event reaches, whether it's sent, and under what policy. This page documents the consolidated data model, the runtime routing chain, how a client is onboarded, and how to automate that onboarding (e.g. from a Temporal workflow).

Why this exists — consolidation​

Historically a single client's "who / where / whether" was smeared across four systems:

SystemHeldStatus
n8n jira_bot_dbproject → Telegram chat routingbeing retired → Console
temporal jira_sync_config_kvnotify modes + review/autocomplete policybeing retired → Console
JSM (Jira)org ↔ service-desk structureupstream master (read)
CDM (ClickHouse)whether the customer is activeupstream master (read)
ad-hoc /bind Telegram commandper-group routingreplaced by the admin UI

Console consolidates the operational layer (routing + policy) into a real client record with an admin UI, an audit trail, and one internal API the temporal-workers read. It does not absorb the upstream masters — it reads the facts it needs from JSM and CDM and owns the routing/policy that results.

The consolidated model​

Two bindings, two questions​

A chat can carry two bindings, in two tables, and they answer different questions. Reading one as the other is how a room ends up doing something nobody chose.

TableAnswers
Group bindingtelegram_group_bindingWhich customer is this room about?
Team bindingtelegram_team_bindingWhich of our teams owns this room?

chat_id is the primary key of both, so a chat has at most one of each — and may have either, both, or neither. Neither is not a broken state: it is a chat Console routes nothing to.

What the group binding gives the room. Its client_id is the tenant boundary — the one client whose traffic may land here. What the room actually receives is decided by its audience: a client room gets the severity-gated, status-only message (P1/P2) and the client half of service-desk routing; an internal room gets the full alert card with Ack, Resolve, ⚡ Escalate and the silence controls. Note that the Jira client route is keyed on audience = 'client' (clientTargetsQuery below), so changing the audience does not only add detail — it moves the chat in or out of the client ticket fan-out.

What the team binding gives the room. Jira traffic whose component names that team, delivered to the chat's tasks topic, plus the team-lead mention_handle. Since the Investigator's team-scope work it also decides the room — the set of clients the AI may discuss at all, which is the team's team_clients filtered to those with ai_triage on. It grants no tenant of its own; the team's client list is the grant, audited where it already lives.

audience is a statement of fact, not a feature switch​

audience exists only on the group binding, and is internal or client. It is not a capability dial. The only test is:

Is the customer in this room?

If a customer, a customer's engineer, or anyone outside ProximaOps can read the chat, the answer is client — whatever you want the bot to do in it. Marking a room internal does not remove them from it. It only stops Console refusing them.

That is the whole point, and it is why this page states it rather than leaving it to the column comment. Everything keyed on audience is Console declining to say something in front of a customer. Flipping the field changes nothing about who is there; it changes whether Console still believes they are.

What promoting a room to internal actually grants, stated as a cost so it is not a casual act:

  • the full alert card — hosts, severity, firing entities, the internal timeline — where the room saw a two-line status message for a P1 or P2 only;
  • action buttons, visible to everyone in the chat. A tap is still checked against the tapper's own linked Console account, so a customer cannot act; they can see that the controls are there, and read every state they produce;
  • an AI Investigator that reasons aloud about that client's infrastructure — which hosts it read, what it ruled out, what it still suspects. In a responder room that thinking-out-loud is the product. In the customer's own room it is a half-finished diagnosis of their estate, in our voice, with no operator in between. That is not hypothetical: it happened in production and was fixed on 2026-09-24 (d336eb39) by refusing a client room ahead of every other question about the chat. Console now declines it — but nothing stops an operator marking a customer's room internal because they want the AI in it, which puts it straight back.

Promotion is super-admin, deliberately: ordinary access to the client is not a sufficient bar, because every ProximaOps user scoped to that client has it. The admin bind API and config-as-code (pc apply) go further and require super-admin to bind a chat fresh as internal too; the ops bot's /bind gates only the promotion of an existing client binding, having nothing to compare a first bind against.

The rule is a test, not a prohibition. A room named after a client that in fact contains only our engineers is correctly internal — internal-infra is exactly that. What decides it is who is in the room, never what the room is called.

When a chat carries both​

This combination is real and it has surprised two readers, so it is worth a paragraph rather than a footnote. Verified against production on 2026-09-24: chat -1002673453812 carries a group binding to client Proxima at audience internal (hand-authored, not from a file) and a team binding to TEAM 3.

Its two properties come from two different rows:

  • its room — which clients the Investigator may discuss — comes from the team: TEAM 3 owns Inter Nation, Proxima and Ucode, and Ucode is filtered out because its ai_triage is off, leaving {Proxima, Inter Nation}. The team wins outright; the union with the group binding's client was rejected because it produces a set you cannot predict from either row alone;
  • its on/off switch comes from the group binding, because the handler reads the team's reactive_enabled only when there is no group binding to read — and the group column has shipped true since the chat was bound.

So a dual-bound chat gets the wider room already switched on, and telegram_team_binding.reactive_enabled is the wrong column to look at when debugging it. The full table, the refusal order and how to tell the silences apart live in Which switch governs the chat and When the bot says nothing — not repeated here, because a security gate chain written down twice becomes two gate chains.

Console tables (what Console owns)​

TablePurposeKey columnsPopulated from
clientsthe customer entityid, name, slugcreated (CDM-active gated)
service_desk_org_mappingslabel → client, + the JSM org linkclient_id, jira_organization_id, jira_organization_name, project_label (UNIQUE), service_desk_id, UNIQUE(client_id,jira_organization_id)JSM org + the n8n project_label
telegram_group_bindingclient → chat; Jira routing reads only audience = 'client' (see Two bindings for what the other audience is for)client_id, chat_id, thread_id, audience, notify_jira, language, topic_namethe n8n chat_id (via backfill / discovery+bind)
telegram_team_bindinginternal team → chatteam_id, chat_id, mention_handle, reactive_enabled (ships false; read only for a chat with no group binding — see which switch governs), …team routing
telegram_org_overrideper-SUP-org chat overridesup_org_id, client_id, chat_idn8n org_client_overrides
telegram_chat_topicper-kind forum-topic route for a chatchat_id, kind (alerts|tasks), thread_id, set_by, set_at/topic <kind> in Telegram, or the admin topics API
client_notification_settingsper-client review/autocomplete policyclient_id, review_notify_mode, autocomplete_mode, autocomplete_inactive_days, …admin UI
notification_global_settingsglobal pipeline mode switcheskey (pxd_*_mode), valuesuper-admin UI
service_desk_id

There is currently one JSM service desk (id = 1, "Support"/SUP); every org lives under it, so service_desk_id = '1' for every mapping. If a second service desk is ever added, the onboarding must capture it per-org.

Upstream sources (read, never owned)​

  • CDM (ClickHouse) — cdm.jira_customers (company_name, status). The active gate: only customers with status = 'Active' are onboarded. CDM uses its own customer_key (CDM-NN); there is no key join to JSM org ids, so CDM is matched by company name (or a human-confirmed mapping for renamed brands). CDM data is never stored in Console — it only decides who gets a client.
  • JSM (Jira) — GET {site}/rest/servicedeskapi/organization (id, name) and /servicedesk (the service desks). Source of jira_organization_id, jira_organization_name, service_desk_id.
  • n8n jira_bot_db (project_mappings, team_settings, org_client_overrides) — the project_label → chat_id routing, imported by the backfill tool. Being retired once routing is fully on Console.

Runtime routing chain​

When a Jira/PXD event fires, the temporal-workers ask the Console resolver (POST /api/v1/internal/telegram/notify-targets) for the targets, sending the ticket's SUP org id (customfield_10002) and its labels:

The resolver finds the client's service_desk_org_mappings row primarily by jira_organization_id = sup_org_id (exact), falling back to a label-strip match (strip a trailing -<n>-<date> off both the ticket labels and som.project_label, then compare) for tickets that carry no org id. It then returns that client's chat from telegram_group_binding (dropping bindings with notify_jira = false, applying any telegram_org_override), and finally resolves each target's thread per kind: the request's scope (comment / status / new_issue) maps to kind tasks, and a chat with a tasks row in telegram_chat_topic gets that row's thread_id instead of the one its binding names. An empty or unrecognised scope — which is what every caller sent before per-kind routing shipped — skips that step entirely and keeps the binding's thread. Only the thread_id value changes; the set of chats resolved is identical for every scope. The worker sends via @proximaopsbot. See Ops Bot for the send path and Forum topics: routing by kind for how a topic is configured.

The fallback lives in one function

store.ResolveThreads(topics, chatID, bindingThread) — configured topic → binding thread → group root — is the only expression of that rule, and the alert-dispatch path calls the same one. It is not a SQL COALESCE in the queries above, deliberately: the topic table is read separately, in one batched lookup per fan-out, so the mute filters (notify_jira = true) stay inner joins. An outer join for topics over a query whose job is muting is how a mute switch silently stops muting. A failed topic lookup degrades to binding threads with a warning; routing never costs a notification.

project_label is a routing key — do not reshape it casually

The label-strip fallback makes service_desk_org_mappings.project_label load-bearing for routing, not an inert attribute. On 2026-08-16 the temporal reconciler backfill rewrote labels to a per-org-unique slug-type-n-date-<jsm_org_id> form (needed to satisfy UNIQUE(project_label) when one client owns several orgs). The extra segment shifted the strip boundary — udevs-sup-1-130922-1081 reduces to udevs-sup-1, but the ticket's udevs-sup-1-130922 reduces to udevs-sup — so the match failed and client notifications silently resolved zero targets for every migrated client (comment / status / new-issue / weekly review). The primary jira_organization_id = sup_org_id match above was introduced (eef7e3f8) precisely so routing identity lives on the org id, not the mutable label. Any future change to the label's shape must be checked against clientTargetsQuery in backend/internal/store/telegram_notify_target_store.go before shipping.

Onboarding a client (manual flow)​

This is the flow used to populate the active customers. Each step has a clear source:

  1. Gate on CDM — only status = 'Active' customers (skip Passive / In analysis / not-in-CDM). Renamed brands (the routing/JSM name differs from the CDM legal name) need a human-confirmed mapping.
  2. Resolve the JSM org — jira_organization_id, jira_organization_name, service_desk_id (=1).
  3. Create the client (name, slug) — idempotent on slug.
  4. Create the service_desk_org_mappings row — (client_id, jira_organization_id, jira_organization_name, project_label, service_desk_id). One client can own several orgs (e.g. a SUP + an FP org), but only one project_label per (client, org) (the UNIQUE(client_id, jira_organization_id) constraint). Extra labels for the same (client, org) keep routing via the Console-first / n8n fallback.
  5. Create the telegram_group_binding — client → chat_id (read from n8n), audience = 'client'. The backfill-telegram-routing tool does this in bulk once the mappings exist.

The fallback for un-onboarded / collision routes​

CONSOLE_ROUTING_SOURCE=console resolves via Console only; a label with no service_desk_org_mappings row (inactive customer, or a (client, org) collision) returns no target. The worker's empty-result fallback routes those via jira_bot_db instead of dropping them — parity until they're onboarded or n8n is retired.

Automating onboarding via Temporal​

The manual flow above maps cleanly onto a Temporal workflow. Recommended shape:

Activities & APIs (all idempotent — check-then-create or ON CONFLICT):

ActivitySource / callIdempotency key
CdmActiveCheckClickHouse cdm.jira_customers— (read)
ResolveJsmOrgGET /rest/servicedeskapi/organization— (read)
EnsureConsoleClientPOST /api/v1/clients (super-admin)slug
EnsureOrgMappingPOST /api/v1/service-desk/admin/organizationsproject_label (UNIQUE) + (client, org)
EnsureTelegramBindingPOST /api/v1/admin/telegram/groups/{chat_id}/bind (see Telegram Groups admin)chat_id

Design notes for the automation:

  • Idempotent end-to-end — every write is upsert/ON CONFLICT DO NOTHING, so re-runs and retries are safe. This is what lets a Temporal workflow own it.
  • CDM gate is mandatory — never create a client for a non-Active customer; that's the rule that keeps the consolidated view clean.
  • Renamed brands — the JSM/routing name (Royal Taxi) often differs from the CDM legal name (You Cloud), and there's no key join. Keep a small override map (project/org → CDM customer) the workflow consults; surface unmatched-but-active candidates for human confirmation rather than guessing.
  • The (client, org) constraint — if a client has several projects under one JSM org, only the first maps; the rest are intentionally left to the fallback (or relax UNIQUE(client_id, jira_organization_id) in a migration after confirming the JSM ticket-sync tolerates multiple labels per client-org).
  • Auth — Console admin endpoints are super-admin / clients:write; the internal resolver + GetConfig are service-account-pinned (telegram_ops:read). A Temporal worker SA needs the right scopes per call.
  • Bot identity — sending only needs the bot token; binding a chat requires @proximaopsbot to already be in that group (it is, for any chat n8n already routes to). Do not move the bot webhook as part of onboarding — see the cutover runbook.

Reference — key facts​

  • Resolver: POST /api/v1/internal/telegram/notify-targets → {targets:[{chat_id, thread_id, audience, mention, team_name, language, topic_name}]}.
  • Config: GET /api/v1/internal/notification-config?key= (modes + autocomplete_org_thresholds), X-API-Key, telegram_ops:read.
  • service_desk_id is always 1 (single SUP service desk today).
  • CDM is a gate, not storage — matched by company name; renamed brands need a confirmed override.
  • Backfill tool: backend/cmd/backfill-telegram-routing imports n8n routing into Console bindings (dry-run by default; -apply, -create-teams).
  • The flip + cutover sequence lives in the Console-flip cutover runbook (temporal-workers repo).