Producing Notification Events
This page is the contract for writing into the notification feed. If you are about to call Publish, read all of it — one of these rules is a security rule, and the code will not stop you breaking it.
The feed is generic on purpose: any subsystem can publish into it, and nothing in the schema knows what a "good" notification is. That generality is what makes it useful and it is also what makes it fragile. A generic feed with no written boundary decays into a task queue one well-meaning producer at a time. This page is the boundary.
hosts:readThe feed is read under a single permission, hosts:read, for every event in it. There is no per-event permission check anywhere — not in the store, not in the handler, not in the bell.
So the content ceiling is not "what my subsystem is about", it is "what a hosts:read holder is already entitled to see". The seeded Viewer role holds exactly hosts:read, metrics:read, healthcheck:read and assets:read — and nothing else. Publish a compliance finding, a rollout result, an audit fact or an alert detail today and a Viewer reads it in their bell while being refused it on every page in the product.
Before the first event that is not safe under hosts:read ships, a kind → permission map must be built and consulted at read time. That is a backend change to notification_event_handler.go and the store's read path, not a note in a ticket. Do it first, and land it before the producer.
The one call
The whole producer-facing API is one method:
// backend/internal/store/notification_event_store.go
Publish(ctx context.Context, in domain.PublishEventInput) error
// backend/internal/domain/notification_event.go
type PublishEventInput struct {
Kind string
ClientID *uuid.UUID
EnvironmentID *uuid.UUID
SubjectType string
SubjectID string
Title string
Body string
Link string
DedupKey string
}
Construct the store with store.NewNotificationEventStore(db) and depend on it through a one-method interface declared at your own package, the way the retention worker does:
// feedPublisher is the ONE method this producer needs. Declaring the narrow
// interface here rather than taking store.NotificationEventStore keeps the read
// surface (List, UnreadCount, AdvanceWatermark) and the expiry (DeleteOlderThan)
// out of reach of a later edit.
type feedPublisher interface {
Publish(ctx context.Context, in domain.PublishEventInput) error
}
Wire the producer up in backend/cmd/server/workers.go, beside the other background workers.
Publishing is a background-worker job. Nothing reachable from an HTTP request may write or expire an event — the API handler holds an interface that deliberately has neither Publish nor DeleteOlderThan on it. Keep it that way: a feed a request can write to is a feed a request can be made to spam.
Publish returns nil both when the row was inserted and when it was deduplicated away. That is deliberate (see the dedup key) and there is no way to tell the two apart. If your producer needs to know, it is asking the feed a question the feed does not answer.
Rule 1: the scope pair decides who sees it
Every event carries a scope pair — ClientID and EnvironmentID — and that pair, intersected with the reader's grants at read time, is the entire access decision. Nothing else on the row is consulted. SubjectType / SubjectID are deep-link decoration and are never resolved to a tenant.
That is the design, and the rejected alternative is worth knowing so nobody re-proposes it: a polymorphic event that names an entity (host, monitor, runbook_execution) and lets the reader resolve its tenancy would need a resolution path per entity type, and each one is a place a tenant guard can be missed. Two of the three cross-tenant defects this codebase has recorded came from exactly that shape. One pair, resolved one way, is the whole point.
How to fill it in
| Your fact is about | ClientID | EnvironmentID |
|---|---|---|
| One environment of one project | that project | that environment |
| A whole project, or something spanning several of its environments | that project | nil |
| The platform, with no tenant at all (a global catalog sync failing) | nil | nil |
Choosing between the first two rows is a containment decision, not a formatting one, and the safe-looking default is the wide one. environment_id is the feed's environment boundary, and it is the same boundary the rest of Console already enforces: the Versions page narrows an environment-scoped reader with AllowedEnvironmentIDs(hosts:read, client), so someone granted only staging never sees production's version rows there. Leave EnvironmentID nil on a fact that is really about production, and it lands in that reader's bell — content they are refused on the page that owns it. That is rule zero at the top of this page, breached without breaking any other rule on it — which is why this one is spelled out.
So the rule runs in both directions:
- If the fact is about one environment, tag it.
nilis not the cautious choice; it is the widest one available inside the project. - If the fact genuinely spans environments, or is not about environments at all, leave it
nil. That is what makes it a project-wide fact, and it correctly reaches environment-scoped readers too.
What you must not do is fan one genuinely project-wide fact out into one event per environment to "be safe". You have not narrowed anything — you have tripled the noise, and each copy needs its own dedup key or two of the three silently vanish.
One refinement worth catching early: a count or a host list in the Body is itself environment-crossing information. "3 hosts still run it" over a project-wide event tells a staging-only reader about production's two. If the detail you want to publish is per-environment, then what you have is several per-environment facts rather than one project-wide fact — tag each one and give it its own numbers. That is not fanning out; that is the events actually being different.
Neither choice can leak across projects — the project half is checked independently of the environment half — but within a project, the environment half is the only thing standing between one environment's reader and another's news.
Most producers start from a host. A host carries environment_id; the project comes from the environment:
SELECT e.client_id, h.environment_id FROM hosts h JOIN environments e ON e.id = h.environment_id WHERE h.id = $1
Three ways to get the pair wrong
The first is refused by the database; the other two are not, so read them once:
EnvironmentIDset withClientIDnil — rejected.CHECK (client_id IS NOT NULL OR environment_id IS NULL), andPublishreturns an error namingnotification_events_env_requires_client. It is a constraint rather than advice because the failure it prevents is invisible: such a row is not "an environment event", it is a global event, which is super-admin-only — so the producer has not narrowed its audience, it has emptied it, whilePublishreturnsnil. If you set an environment, set its project too.- An environment that does not belong to the project you named. Nothing checks the pair is coherent — this one really is on you. The event reaches that project's project-wide readers and none of its environment-scoped ones, for reasons nobody will be able to reconstruct.
- A project or environment id that does not exist. Both columns are foreign keys with
ON DELETE CASCADE, soPublishreturns an error rather than inserting — and, separately, deleting a project deletes its events. The feed is a notice, never the record; if the fact must outlive the project, it does not belong here alone.
Rule 2: the dedup key is the announcement's identity
DedupKey is unique per project — the constraint is UNIQUE NULLS NOT DISTINCT (client_id, dedup_key), so two projects can safely use the same key, and two global events cannot (the NULLS NOT DISTINCT is what makes dedup exist at all for the events every reader sees).
Publish uses ON CONFLICT DO NOTHING. Dedup therefore lives in the database, not in your worker's memory, and that is the point: it means "run it again" is the recovery strategy for every failure mode — a crash between deciding to publish and publishing, a retried tick, a redeployed pod, a backfill over a window you have already covered. Your producer never has to remember what it has already said, and never has to be exactly-once.
Building one
Key on the thing you are announcing, at the granularity at which you want it said once:
versions.eol:nginx:1.24 one announcement per project, per product, per version
versions.eol:nginx:1.24:2026-Q3 the same, re-announced once a quarter
compliance.drop:pci-dss:2026-09 one per framework per month
-
Never put
time.Now(), a run id, a batch id or a random UUID in it. That does not "make it unique", it disables dedup — and the producer then re-announces the same fact on every tick, forever, into a panel with no per-item dismiss. -
Never put the project id in it. Uniqueness is already per project; adding it is harmless but tells the next reader the scoping works the other way round.
-
Do put a period in it if you genuinely want the thing re-announced — a quarter, a month, the version that changed. That is a deliberate re-announcement, not a workaround.
-
It cannot be empty, and there is no way to opt out of dedup.
CHECK (dedup_key <> '')rejects the empty string, andPublishreturns an error namingnotification_events_dedup_key_not_empty. The constraint exists because an unsetDedupKeyinserts'', which collides with every other empty key in that project — the first such event lands, every later one is silently discarded, andPublishreturnsnilfor all of them. That is a producer which appears to work and publishes exactly one notification ever, and it is now impossible.NULLis not an escape hatch either: the column isNOT NULL, and even if it were not, the index isNULLS NOT DISTINCT, so two NULL-keyed events in one project would collide exactly as two empty ones do. Every event carries a key.
Rule 3: events, not state
A row says something happened at a moment. It is never updated, never re-derived, and there is no UPDATE path in the store at all.
If the condition later clears, you publish a second event saying so — you do not go back and change the first, and you do not delete it. A feed that tracked live conditions would be a second alerting system standing beside a mature one, with overlapping semantics, a different UI and no escalation; operators end up trusting neither.
Practically: write the title in the past tense about a specific moment. "nginx 1.24 reached end of life" ages correctly. "nginx is out of date" does not — it is a claim about now, published once, and read a week later.
Field reference
| Field | Type | Required | Notes |
|---|---|---|---|
Kind | VARCHAR(64) | yes | Rendering only, never severity. See registering a kind. |
ClientID | *uuid.UUID | — | See Rule 1. nil means global, super-admin-only. Required by a CHECK whenever EnvironmentID is set. |
EnvironmentID | *uuid.UUID | — | Legal only alongside a ClientID; the pair is enforced by a CHECK. |
SubjectType | VARCHAR(32) | no | Deep-link decoration. Never used for access. |
SubjectID | VARCHAR(128) | no | Same. |
Title | TEXT | yes | One line, past tense, written for a human. Nothing generates it for you. |
Body | TEXT | no | One or two lines of detail. |
Link | TEXT | no | In-app absolute path only — /versions?product=nginx. The bell refuses to link anything else (a protocol-relative //host or a javascript: URL renders as plain text), so an external URL here is silently inert. |
DedupKey | TEXT | yes | See Rule 2. Non-empty, enforced by a CHECK. |
The VARCHAR limits are enforced by Postgres: overrun one and Publish returns an error rather than truncating.
A worked example
A version-currency producer announcing an end-of-life product, once per project per version:
// backend/internal/worker/version_notify.go
// feedPublisher is the one method this worker needs; see above.
type feedPublisher interface {
Publish(ctx context.Context, in domain.PublishEventInput) error
}
// announceEOL publishes one end-of-life notice for one product in ONE
// environment. eolFact is produced per (product, version, environment)
// precisely so that it can be: see the scope pair below.
//
// Safe to call on every tick -- the dedup key is that same triple, so the
// second and every later call for the same environment is a no-op in the
// database.
func (w *VersionNotifyWorker) announceEOL(ctx context.Context, f eolFact) error {
in := domain.PublishEventInput{
Kind: "versions.eol",
// BOTH halves of the scope pair, because this fact IS about one
// environment. Leaving EnvironmentID nil here would put production's
// end-of-life software in a staging-only reader's bell -- content the
// Versions page itself refuses them (Rule 1).
ClientID: &f.ClientID,
EnvironmentID: &f.EnvironmentID,
// Decoration for the link. Never consulted for access.
SubjectType: "product",
SubjectID: f.Slug,
Title: fmt.Sprintf("%s %s reached end of life", f.Product, f.Version),
// THIS environment's host count, which is what makes it publishable at
// all: an estate-wide count on a project-wide event would tell that same
// staging-only reader how many production hosts are affected.
Body: fmt.Sprintf("%d host(s) in %s still run it. %s is the oldest supported release.",
f.HostCount, f.EnvironmentName, f.SupportedVersion),
Link: "/versions?product=" + url.QueryEscape(f.Slug),
// The environment is IN the key. Without it the first environment's
// notice consumes every other environment's copy within the project --
// silently, since dedup is per (client_id, dedup_key) and Publish
// returns nil either way.
DedupKey: fmt.Sprintf("versions.eol:%s:%s:%s", f.Slug, f.Version, f.EnvironmentID),
}
if err := w.feed.Publish(ctx, in); err != nil {
return fmt.Errorf("announce eol for %s %s in %s: %w", f.Slug, f.Version, f.EnvironmentID, err)
}
return nil
}
A genuinely project-wide fact takes the same shape with both environment lines removed — EnvironmentID left nil and the environment dropped from the dedup key. Use that only when the fact really has no environment (a catalog sync failing for the whole project), and remember that its Body may then carry nothing environment-specific: no per-environment counts, no host names.
Why this producer is allowed to exist today: version-currency content is derived from host inventory, and the Versions page is gated on hosts:read — the feed's own permission — so nothing it says is outside the permission ceiling rule zero sets. That is one of the two boundaries, not both: the Versions page also narrows an environment-scoped reader to their own environments, and nothing about hosts:read carries that. Environment containment is Rule 1's job, and it is the producer's to get right, which is why the example above tags every event.
Failure, retries and tests
A failed Publish should be logged and should not abort the run. There is no batching and no transaction — one call, one row — so the next tick re-publishes everything the dedup index has not already absorbed. Aborting a sweep on the first failed notice loses every notice behind it and gains nothing, precisely because re-running is free.
Test a producer by faking the one-method interface and asserting the PublishEventInput it builds. The two fields worth pinning are the scope pair and the dedup key, because those are the two that fail silently in production — a wrong scope shows the event to the wrong set of people, and a wrong dedup key either spams or publishes once and never again, and neither returns an error. You do not need to re-test containment itself: that lives in the store and is covered by notification_event_store_integration_test.go.
What fits
- A rollout finished — with a link to the rollout.
- A compliance score dropped — once the permission map exists; see rule zero.
- Paging was enabled on a monitor — a configuration change worth knowing about, that nobody watches a page for.
- A global catalog sync failed —
ClientID: nil, super-admin-only, because there is no tenant it belongs to.
The common shape: something changed, nobody was watching the page it changed on, and knowing about it next time you open Console is enough.
What does not fit, and what to do instead
This is the section that earns the page.
Alerts
An alert has acknowledge, resolve, silence and escalate. Piping alerts into the feed builds a second alerting system beside a mature one, with different semantics, a different UI and none of the guarantees — and operators end up trusting neither. Worse, the feed looks like it has been read when the badge clears, which is exactly the wrong signal about an alert.
Instead: publish "this happened" if it is genuinely news, and link to the alert. Let the alert own its own lifecycle.
Approvals, and anything actionable with its own lifecycle
A remediation waiting for approval, a ticket needing triage, a request to confirm something. The feed would show it as read while nobody acted on it, because opening a panel is not doing a job. There is no state here to record that someone took it, and no way to take it back if they did not.
Instead: the producer owns that state, on its own surface, with its own list and its own permissions. The feed may announce that the thing exists.
Anything addressed to one person
There is no recipient field and there will not be one. Producers supply a scope, and the readers are whoever holds grants inside it — which is not a set the producer knows, and is not a set that stays still.
Instead: if you must reach a named person, that is the paging chain, an email, or Telegram — surfaces built for delivery to an individual, with the delivery guarantees that implies. The feed has none.
Anything time-critical
The bell refreshes once a minute, pauses in a background tab, and is not looked at by anybody who is not already in Console. It is not a delivery mechanism and it makes no delivery promise at all.
Instead: alerting and on-call.
Three things that will never be added
Not "not yet" — these are closed, and each is closed for the same reason: it is a lever a producer would pull to outrank the others, and the moment one producer pulls it every other producer must.
- A severity or priority field. The first producer to mark itself
criticalforces everyone else to, and within two releases every notification is critical and the field means nothing. It would also make the feed look like a pager while backing none of a pager's promises, which is worse than having no feed. - Per-item done state. The affordance for "I handled this" — which the feed cannot verify, cannot re-open, and cannot show to anybody else. The surface that owns the work owns that state.
- A way to address an individual. See above. A recipient turns a scope-derived feed into a delivery system, and every property that makes read-time containment safe depends on there being no stored audience.
If you find yourself needing one of these, you are building something other than a feed. That is fine — build it somewhere else.
Registering a kind in the frontend
Kind drives rendering and nothing else. Add one line to frontend/src/components/notifications/event-kinds.ts:
export const EVENT_KINDS: Readonly<Record<string, EventKindDescriptor>> = {
"versions.eol": { label: "Version currency", icon: PackageX },
};
The label is a category — "Version currency" — never the event's title. The title always comes from the event, because the producer wrote it for a human to read in exactly that position.
Conventions:
- Lowercase, dot-separated,
subsystem.event. Max 64 characters. - A kind is a wire value between a backend release and a frontend release, so treat it as permanent: renaming one orphans every event already published under the old name into the fallback below.
A producer may ship before the frontend knows its kind, and that is a supported state, not a bug. An unregistered kind degrades to a generic bell icon and the label "Notification", with the event's own title and body rendered normally. That degradation is what makes a rolling deploy safe and what stops one backend release from blanking every reader's panel — so ship the backend whenever you like, and add the line when convenient.
Shipped kinds
Every kind actually publishing today, and the permission a reader needs to see it. A kind with no entry in kindPermission (backend/internal/api/notification_event_handler.go) reads under the blanket hosts:read permission described in rule zero, above — this table is only the kinds that read under something narrower instead.
| Kind | kindPermission | What it means |
|---|---|---|
monitor.state_changed | monitors:read | A monitor's confirmed up/down transition — a fact, not a live condition; see the alert it should never be confused with. |
monitor.state_changed is the first kind this feed has ever carried that is not safe under hosts:read alone — the first (and, today, only) entry in kindPermission. A kind with an entry there is governed exclusively by that permission's own scope, not the blanket hosts:read scope as well: the store's read query branches on kind, so a monitors:read holder sees monitor.state_changed events even without hosts:read, and a hosts:read holder without monitors:read does not see them at all. The two permissions do not imply each other.
See Notify routing for what actually publishes this kind, and why it fires independent of a monitor's paging_enabled.
Operating a producer
- Events are deleted after 30 days (
PROXIMA_NOTIFICATION_RETENTION_DAYS, default 30; zero or negative disables the sweep entirely and the table then grows without bound). A dedup key survives only as long as the row does — so a producer whose key has no period in it will re-announce a still-true fact about 30 days later. That is usually right. If it is not, put the period in the key. - The sweep is metered.
proxima_retention_rows_pruned_total{table="notification_events"}counts what went, andproxima_notification_retention_last_success_timestamp_secondssays when the sweep last completed. You need both: the counter records nothing for a sweep that deleted zero rows, which at a 30-day window is the normal tick, so a flat counter cannot distinguish "swept, nothing aged out" from "the sweep stopped". The gauge is what answers that. - There is no per-tenant retention and no kill-switch. The window is the feed's contract, and every fact in it is a copy of something the owning surface keeps for as long as it deserves.
Before you merge
- Everything I publish is safe to show to a holder of
hosts:readalone — or thekind→ permission map exists and I am in it. - Every event names a project, or is deliberately global (
ClientID: nil) and super-admin-only. - If my fact is about one environment, I tagged it. Leaving
EnvironmentIDnil shows it to readers scoped to other environments of the same project — who are refused that content on the page that owns it. - If I set an environment, I set its project too (a
CHECKenforces this), and the two actually belong together (nothing enforces that). - The dedup key derives from the fact rather than the run, and contains no timestamp or random value. (Non-empty is enforced by a
CHECK; meaningful is not.) - Running my producer twice publishes one notification.
- Titles read as things that happened, not as claims about now.
-
Linkis an in-app absolute path, or empty. - I publish from a background worker, never from a request handler.
- No severity, no priority, no recipient, no per-item state anywhere in my design.