Skip to main content

Uptime Monitors

An uptime monitor is a synthetic check that Proxima Console owns end to end: you describe what "healthy" means for an endpoint — a URL, a TCP port, a DNS answer, a certificate, or a heartbeat that something else must send — and Console keeps the definition, runs the check from its probe locations, and records the current state and the history of every state change.

They matter because they are Console's first first-party alert source. Every other actionable signal in the platform arrives from someone else's system — a Grafana or Alertmanager webhook, a pull source. Monitors are the first thing Console observes itself.

Monitors can page now — and nothing does, by default

The probe fleet, quorum, the state machine, the availability history and alert emission are all live: a monitor really does go up, pending and down, the transitions are recorded with the locations that confirmed them, and a confirmed -> down publishes an alert on the same ingest path a webhook uses.

That publish is behind a gate chain, and the first gate is paging_enabled, which ships false on every monitor. So the shipped default is a shadow window: everything runs except the page, and proxima_probe_would_page_total records what would have woken somebody. Turning it on is a per-monitor configuration change, and the flag is only the first of at least five preconditions that no deploy step performs — two of which are ordinary alerting gates that also ship off. Read the rollout procedure before flipping anything.

Monitors can also be grouped into a service that is judged and paged as one thing, so a whole service failing wakes somebody once rather than once per member — see Monitor groups. A group's paging_enabled is its own flag and ships false too, and a group may hold it only while none of its members does.

Also not yet shipped: certificate-expiry notices. Maintenance windows have shipped — see Maintenance windows and the maintenance state.

What a monitor is​

A monitor belongs to exactly one client and one of that client's environments, and its name is unique per client (a reused name is a 409, not a silent overwrite). Beyond that it is a definition plus a set of knobs.

FieldDefaultFloor / ruleWhat it means
name—required, unique per clientHow the monitor is identified everywhere.
kind—one of http, tcp, dns, tls, pushWhich check to run. See below.
target—required except for pushWhat to check. A push monitor has no target at all — it is checked in to.
config{}free-form JSON objectPer-kind options. See below.
interval_seconds60≥ 20How often the check runs.
retry_interval_seconds20≥ 10The faster cadence used while a check is failing but retries are not exhausted.
timeout_seconds10≥ 1, and < interval_secondsHow long one check may take. A timeout at or above the interval stacks probes on top of each other.
max_retries2≥ 0How many consecutive failures a location tolerates before it counts as failing. max_retries is retries, so a location must fail max_retries + 1 checks in a row.
min_confirmations2 (1 for push)≥ 1How many probe locations must agree before a state flip. A location is a vantage point, so this counts places, never machines — see Locations are vantage points.
location_selector{"kind":"public"}public or privateWhich probe locations run the check.
severityP2P1–P5The tier a firing monitor pages at.
upside_downfalse—Inverts the verdict: a failing check means healthy. For things that are supposed to be unreachable.
enabledtrue—Whether the monitor is in rotation at all. Pause/resume flips this.
paging_enabledfalse—See Paging.
host_id, asset_id, cluster_idnullmust belong to the same clientOptional links to the infrastructure the monitor watches.

Two rules about severity are worth stating explicitly, because they differ from the rest of the platform:

  • The set is exactly P1–P5. There is deliberately no unknown tier, even though alerts, alert groups and incidents accept one. unknown is an ingestion tier — the honest answer when an external source sends no usable label (see Alerting). A monitor's severity is typed by a human in Console's own form, so there is no unlabelled source to represent, and allowing unknown would only let someone author a monitor that pages at an untriageable tier.
  • Severity is what the emitted alert carries into route matching, exactly as any other alert's is. It is still passed through the canonicaliser on the way out rather than trusted verbatim, so a row that somehow escaped both the domain validator and the CHECK constraint pages at a visibly untriageable unknown instead of at a plausible tier nobody observed.

Check kinds and their config keys​

config is stored as an opaque JSON object; the backend validates nothing inside it, and the prober is the first thing that reads it. A typo in a config key is therefore not rejected — it is ignored by the executor that would have used it.

Kindtarget looks likeconfig keys the executor reads
httphttps://api.example.com/healthaccepted_status_codes, method, headers, body, keyword, keyword_inverted, json_query, max_redirects, ignore_tls
tcpdb.internal:5432none — an explicit port is required, because a raw TCP check has no conventional port to assume
dnsexample.comrecord_type (default A; A/AAAA/CNAME/TXT/MX/NS), resolver, expected
tlsapi.example.com:443none — port defaults to 443
push(none — must be empty)none

The create/edit form offers inputs for the http and dns keys only, because those are the only kinds with options an executor reads. The API accepts any JSON object in config, so a caller may author additional keys directly. A PATCH that includes config replaces it wholesale; one that omits the key leaves the stored object alone — unless the PATCH changes kind, in which case config is reset to {}, because the previous kind's keys no longer mean anything and the prober would otherwise read them.

This kind + config split is the reason a new check kind does not need a migration: adding one means a new kind value and a prober function, not a new column.

HTTP checks in detail​

  • accepted_status_codes takes ranges, e.g. ["200-299", "418"]. An unparseable entry fails the check as an internal error rather than falling back to the default: a typo'd range that reported healthy is the one outcome nobody investigates.
  • max_redirects defaults to 10. An explicit 0 means "do not follow redirects" — the redirect response itself is then judged against accepted_status_codes, so a 302 you accept is a pass.
  • ignore_tls skips certificate chain verification. It does not skip dates: an expired or not-yet-valid certificate still fails the check, because skipping verification is about trusting the chain, not about pretending a date has passed.
  • keyword / keyword_inverted assert on the response body: the check fails when keyword is absent (or, inverted, when it is present).
  • Every check dials a fresh connection (keep-alives are disabled). A synthetic check measures a connection being made, and a pooled connection would skip the very thing being observed.
  • For an https:// target the leaf certificate's expiry is recorded as a side effect, so an HTTP monitor also answers "when does this certificate run out?" without a second monitor.

The json_query contract​

json_query narrows where keyword is looked for. Without it, the subject of the keyword match is the whole response body; with it, the subject is the value at that path. There is one "contains" rule, so both behave the same way.

  • The path is dot-separated, and a numeric segment indexes an array: checks.1.name. Bracket form is accepted and identical: checks[1].name ≡ checks.1.name.
  • The resolved value is rendered as text before matching: a string as itself, a number in its shortest decimal form, a boolean as true/false, and an object or array as its compact JSON.
  • A path that does not resolve is a content failure in its own right (error class keyword), whether the body is not JSON, a field is missing, an index is out of range, the path descends into a scalar, or the value is null. A monitor asserting on a field that is not there is not healthy.
  • json_query on its own — with no keyword — asserts only that the path resolves.
{
"json_query": "status",
"keyword": "ok"
}

against {"status":"ok","checks":[{"name":"db","ok":true}]} passes, because the value at status contains ok. Change the body to {"status":"degraded"} and the check fails with keyword, naming the path in the error.

Push monitors​

A push monitor inverts the direction: instead of Console checking something, something checks in to Console on a schedule, and silence is the failure.

  • It takes no target (a target is a 400).
  • Its min_confirmations must be 1 — it is confirmed by its own check-in, not by probe locations.
  • Creating one mints a heartbeat token, returned in the create response as push_token.
The push token is shown exactly once

Only a hash of the token is stored. The plaintext exists only in the response that minted it — the kind: "push" create, or an edit that converts an existing monitor into a push monitor. No later read returns it: every other response carries has_push_token (a boolean) and no token.

pc apply is a third surface for the same rule rather than an exception to it: an apply that mints prints the plaintext once, on the line under applied, and no pc get — including -o yaml — ever returns it. See Config as code, which also covers the one case neither surface's wording above names: a push monitor that somehow has no stored hash gets a working token from its next write.

The form shows the token — and the two check-in URLs built from it — in a panel that stays open until you acknowledge it, precisely because closing the dialog would discard an unrecoverable credential. If it is lost, convert the monitor away from push and back to mint a new one — which invalidates the old token.

Checking in​

Two public, token-authenticated URLs, shown exactly once alongside the plaintext token (see the warning above):

GET /api/v1/monitors/heartbeat/{token}        -- "I ran, and I'm healthy"
GET /api/v1/monitors/heartbeat/{token}/fail -- "I ran, and I failed"

GET, not POST, and the token is in the path rather than a header — the whole point is that the integration on the customer's side is curl <url> at the end of a script, with no flags to spare. Both return 200 with {"status":"ok"} on any accepted call; an unknown or malformed token answers 404, never 401, so a bare string that matches nothing reveals nothing about whether tokens of that shape exist.

The failure route confirms down immediately, without waiting out the grace window below — a caller that has already told us the answer is not made to wait for a timer to agree.

The grace window​

A push monitor's grace window is:

grace = interval_seconds + max_retries × retry_interval_seconds

the same fields — and the same "N consecutive failures at cadence" shape — every probed monitor already configures, reused rather than given a new section in the form. Go quiet for longer than this window and a background sweep confirms down, exactly as a confirmed probe failure would: same state machine, same timeline event, same alert-emission gate chain. The create/edit form computes and displays this window live as you type.

A fresh monitor's window is measured from monitors.created_at, not from now — there is no earlier "last result" to measure from. A push monitor authored today and never called correctly will not show as down until its first grace window elapses: a down verdict before the customer's script has ever run once would be inventing an outage nothing observed.

Who confirmed it​

The timeline's confirmed_by names the heartbeat call itself (push) for an ordinary check-in or an explicit failure report, and the sweep's own finding (silence) when nothing arrived in time — so "Confirmed by push" and "Confirmed by silence" are both real, honest answers to "who saw this?", at the same one-witness scale a probe monitor answers it at several.

Probe locations (pops)​

A probe location is a vantage point on the network — eu-central, tashkent, a client's own datacentre — and it may be backed by one prober machine or several. A ProximaOps public location serves every tenant; a private one is deployed inside a single client's network.

The registry holds one row per prober, not one row per location. Several rows sharing a code are one location.

ColumnMeaning
codeStable short code (eu-central) naming the vantage point. Several prober rows may share it. It becomes a metric label and the value written into monitor_events.confirmed_by, so it is curated and effectively permanent — renaming one orphans every existing series.
kindpublic (must carry no client) or private (must carry one). Enforced by a CHECK constraint per row. Rows sharing a code must agree — see below.
capabilitiesWhat this prober can actually do, e.g. {"icmp": false}. A prober is capable unless explicitly told otherwise: an absent key means "yes". A location advertises a capability only if every prober behind it has it.
enabledfalse drains that prober: its results are refused and it stops counting either way — it can neither keep its location alive nor veto what its location can do.
agent_idThe enrolled probe agent behind this row. Still unique: one agent is exactly one prober at exactly one location, which is how the result path resolves a pop's reporting location from the agent id in the NATS subject.

Locations are vantage points, not machines​

Adding probers to a location increases resilience. It does not increase quorum.

This is the single most likely misunderstanding of the whole feature, and the reason the model was changed.

Three probers in Frankfurt are one place on the network. They share transit, they share a datacentre, and they fail together. If each cast its own vote, a Frankfurt transit failure would confirm a monitor down with what reads on the timeline as agreement across three independent locations — and adding a machine for redundancy would quietly weaken the bar it was meant to strengthen, because min_confirmations: 3 would suddenly be reachable from one building.

So a location votes once, however many machines back it. Machines are how a vantage point survives losing one; they are not evidence about the target. If you want a higher bar, add a location, not a machine.

The aggregation rule: any prober succeeding means the location is up.

That rule was not chosen because it is optimistic. It is the only rule where adding or removing a machine cannot change that location's vote for identical reality:

  • Majority breaks on a two-machine location (there is no majority of two) and shifts every time the deployment count changes — so scaling a pop from 2 to 3 machines silently re-decides outages.
  • All must succeed makes redundancy increase false outages: every extra machine is another thing whose local fault can drag the location down, so the more resilient you make a pop the more it lies.
  • Any succeeds is stable under deployment count. One machine reaching the target proves the target was reachable from that vantage point; another machine beside it failing is that machine or its local transit, and crediting it to the customer's target would be a fabrication.

Two more rules fall out of it:

  • A location's failure streak is the min of its probers' streaks, derived at evaluation and never stored. A location has been failing only as long as every machine there has been failing — if any prober succeeded recently, the location's streak is zero. Taking the maximum would let the sickest machine drive the location's time axis.
  • A location is capable of a check kind only if every prober there is. A location where only some machines can run the check would be probed by all of them and fail forever on the ones that cannot — a permanent down vote manufactured out of a configuration mismatch.

The stability is in the vote, not in the clock. The min rule makes the time axis sensitive to the deployment count in a way the vote is not: adding a machine to a location that is already failing gives that location a prober with a streak of one, so the location's streak resets and confirmation is delayed by up to max_retries intervals. Removing the least-failing machine moves it the other way. Both are transient, happen once, and settle within max_retries rounds — and the alternative (taking the max) would hand the sickest machine at a pop permanent control of the whole location's clock. So: adding a machine cannot change what a location says about identical reality; it can move when a confirmation lands.

Freshness is applied per prober, before the fold. A prober's row is updated in place and never reaped, so a machine that died at noon still holds a row saying "healthy" months later. Folding first would let that dead row win the any-succeeds rule permanently while borrowing a live sibling's timestamp to look fresh — the location would vote up, and fresh, forever. Dropping stale rows first means a location whose probers are all stale produces no vote at all, which is the honest answer: it knows nothing about right now, and that is not the same as healthy.

Rows sharing a code must agree on kind and client_id

Postgres cannot enforce that without a trigger, so nothing at the schema level stops you from giving a private prober the same code as a public pop. What happens instead is that eligibility requires every row at a code to match the monitor's selector, so a mismatched row takes the whole code out of the denominator rather than letting whichever row was read first decide.

That is fail-safe, not free: one mis-provisioned private row sharing a public pop's code removes that pop from every tenant's public monitors. It is reported by proxima_probe_location_disagreement_total{location}, which is why that counter exists.

Deploying a pop​

Two steps, and they are separate on purpose: installing enrolls an agent, registering points a location at it. A prober that is enrolled but unregistered runs, polls, and is given nothing to check.

  1. Install the agent as a prober. Four equivalent ways:

    The console's Add host wizard is the easiest — pick Probe (uptime pop) as the agent type and it generates the command with the data dir and volume already set.

    # A VM or bare metal
    curl -fsSL https://install.prxm.uz/agent.sh | sudo bash -s -- \
    --backend-url https://api-console.prxm.uz \
    --install-token <install token> \
    --agent-type probe \
    --data-dir /var/lib/proxima-probe
    # Docker. The data dir and the named volume are BOTH required: the image default lives
    # under /var/lib, which the non-root user cannot create, and a data dir with nothing
    # mounted at it re-enrolls on every recreate and orphans the previous agent.
    docker run -d --name proxima-probe \
    -e PROXIMA_BACKEND_URL=https://api-console.prxm.uz \
    -e PROXIMA_INSTALL_TOKEN=<install token> \
    -e PROXIMA_AGENT_TYPE=probe \
    -e PROXIMA_AGENT_DATA_DIR=/data \
    -v proxima-probe-data:/data \
    --restart unless-stopped \
    registry.prxm.uz/proxima/console/agent:latest
    # Kubernetes — the agent chart's prober section
    helm upgrade --install proxima-agent charts/proxima-agent \
    --set backendUrl=https://api-console.prxm.uz \
    --set installToken.value=<install token> \
    --set prober.enabled=true \
    --set prober.locationCode=eu-central

    On Kubernetes, a prober that serves more than one tenant must also set prober.defaultResolver to a public resolver, which the chart passes on as PROXIMA_PROBE_DEFAULT_RESOLVER; see the deployment requirement below.

    The separate data dir matters on a machine that already runs a host agent: sharing one makes whichever starts second load the other's state and adopt its identity.

    agent_type is part of the enrolled identity and is fixed at enrollment — changing it later means re-enrolling. The client and environment slugs are only the enrollment credential's scope: a public pop is not that client's pop, it probes monitors for every tenant whose selector resolves to it.

  2. Register the prober at Uptime → Probe locations, or through the API:

    curl -X POST https://api-console.prxm.uz/api/v1/probe-locations \
    -H "Authorization: Bearer <super admin token>" \
    -H 'Content-Type: application/json' \
    -d '{"agent_id":"<agent id>","code":"eu-central","display_name":"Frankfurt","kind":"public"}'

    For a private pop, kind is private and client_id is that client's id. Every verb is super admin only, reads included — probe locations are ProximaOps infrastructure rather than client data, and the store will not filter private rows by tenant.

    To take a pop out of rotation without retiring it, drain it from the page or PATCH /api/v1/probe-locations/{id} with {"enabled": false}. DELETE revokes the prober's credentials permanently and then removes the location — only a fresh enrollment brings that machine back.

    This used to be a direct INSERT

    Before the probe-location API existed, registering a pop meant writing to probe_locations by hand against production — which is how all three pilot pops were created. Raw SQL still works but bypasses the agreement check below, so prefer the API or the page.

Until the row exists the pop is refused with not a probe location and retries every 15 seconds — which is exactly what a pop whose registration was revoked also sees.

To add a second machine to an existing location, register a second prober with the same code — from the page, the API, or a row. agent_id differs; code, kind and client_id are identical. Nothing else changes: no monitor is edited, no quorum bar moves, and every monitor selecting that location keeps confirming on exactly the same number of vantage points it did before. What changes is that the location now survives losing a machine.

The API refuses a code whose rows would disagree on kind or client_id with a 409 naming the conflict, and the page warns you before you submit. That check is the reason to prefer either over SQL: a disagreeing row does not fail loudly — it takes the whole location out of every monitor's quorum denominator, and across tenants, because one client's private row under a public POP's code removes that POP from every tenant's public monitors.

-- Still worth running against a fleet that predates the API, or one edited by hand.
-- BOTH fields matter: two
-- PRIVATE rows for different clients agree on kind and are still a disagreement, because
-- eligibility matches on client_id too.
SELECT code,
array_agg(DISTINCT kind) AS kinds,
array_agg(DISTINCT client_id::text) AS clients,
count(*) AS probers
FROM probe_locations WHERE enabled GROUP BY code
HAVING count(DISTINCT kind) > 1 OR count(DISTINCT client_id) > 1;

To give a monitor a higher bar, add a new code in a genuinely different place and raise min_confirmations. To make an existing place harder to lose, add a row under the code it already has.

What a prober can and cannot do​

A probe agent runs only the transport (with its disk buffer), the credential renewer, the command dispatcher, the assignment client and the check scheduler. No inventory, no processes, no metrics, no logs, no file watch, no terminal, no runbooks, no SSH server, no self-update handler.

That is a security boundary, not tidiness. A prober's NATS JWT carries no tenant-scoped subject at all — the whole proxima.{client}.{env}.{agent}.> grant is withheld — because a public pop serves every tenant and that grant is precisely what a compromised pop would use to publish forged inventory or metrics into a client's namespace. What it holds instead is identity-only:

DirectionSubjectPurpose
Publishproxima.probe.results.{agent_id}Check results
Publishproxima.system.probe.assign.{agent_id}"What should I be checking?" request-reply
Publish/Subscribethe shared identity subjects (proxima.system.commands.{agent_id}, file transfer, credential renewal)Without these a pop could not be updated and would fall off the fleet at its first TTL

The reporting location is derived from the subject's agent id, never from the payload — wire.Result has no location field, and a smuggled one is ignored. A compromised pop can lie about what it saw; it cannot claim to be a different pop.

What a prober will refuse to dial​

Monitor targets are authored by tenants, and a monitor with the default selector runs on the shared public fleet. So where a probe may connect is a security boundary. Without one, any customer could aim the fleet at addresses that are not theirs and read the answer back from the result:

  • the network a pop itself sits in
  • the cloud instance-metadata endpoint of the machine it runs on
  • another customer behind the same carrier

Every check kind reports something worth having:

  • tcp says whether a port answered, which makes it a port scanner.
  • tls returns the peer certificate's issuer, subject and DNS names, which discloses internal certificates.
  • http reports keyword matches against request headers the tenant chose, which makes it an oracle.

So every connection a probe makes first passes a guard, agent/internal/probe/addrguard.

The check runs on the resolved address, not on the target string. The guard is the dialer's Control hook, which Go calls once for each candidate address, after name resolution and before connect(). A hostname that resolves into refused space is refused exactly as its literal address would be. So is every hop of an HTTP redirect chain, and so is a resolver named by hostname. A check on the string could not do this. A name resolves wherever its owner points it, and it can answer differently for the check than for the connection that follows.

What is judged is the connection, not the lookup:

  • http, tcp and tls: the target's hostname is resolved through the prober's own system resolver as usual. The connection to each resulting address is checked.
  • dns: the connection to config.resolver is checked. The name being looked up is never connected to, so it is not judged. A dns check with no resolver asks the prober's default resolver, PROXIMA_PROBE_DEFAULT_RESOLVER, when one is set, and that connection is checked in exactly the same way. With no default resolver, it uses the prober's own system resolver. That resolver is part of the pop's configuration rather than an address a tenant typed, so the connection to it is not checked. On a host running systemd-resolved it is 127.0.0.53.

Deployment requirement: a prober that serves more than one tenant must set PROXIMA_PROBE_DEFAULT_RESOLVER to a public resolver, and must use a system resolver that answers only public names. A dns check's answers are visible to the monitor's owner, so the resolver a check asks matters:

  • PROXIMA_PROBE_DEFAULT_RESOLVER is the resolver a dns check with no resolver asks. Write it like a monitor's resolver: 1.1.1.1, 1.1.1.1:53 or [2606:4700:4700::1111]:53. The guard applies to it. The prober refuses to start on a value it cannot use: one it cannot parse, a port that is not a number from 1 to 65535, or an IP address its dial policy refuses, such as a private address on a prober that has not opted into private targets, or loopback on any prober. A hostname is judged by the guard when a check dials it, like any other hostname.
  • When it is unset, those checks are answered by the prober's own system resolver. A prober that has not opted into private targets logs a warning at startup when this setting is unset: dns checks that name no resolver use this host's system resolver.
  • The system resolver still matters with it set. http, tcp and tls targets are resolved through the system resolver, and a refusal names the address a hostname resolved to. So a shared prober must not use a system resolver that answers private zones, such as a Kubernetes cluster's DNS service.

Always refused, on every prober​

No setting changes these.

ClassAddressesWhy
Loopback127.0.0.0/8, ::1The prober's own services, never a customer's.
Link-local169.254.0.0/16, fe80::/10169.254.169.254 is cloud instance metadata: the credential endpoint of whatever machine the prober runs on.
Unspecified0.0.0.0, ::connect() to 0.0.0.0 reaches loopback on Linux.
Multicast224.0.0.0/4, ff00::/8Nothing a monitor could be for. Link-local-scope multicast such as 224.0.0.5 is reported as link-local, and ff01:: as interface-local multicast. Either way it is refused.
Limited broadcast255.255.255.255No check can ever succeed against it, and a dns check naming it as resolver would send its query to every host on the prober's own network segment.
Zoned IPv6any address with a zone, e.g. fe80::1%eth0A zone selects an interface on the prober host. It matters only for an address that is not globally scoped, such as a link-local one. On a public address it would let a tenant pick the probe's egress interface. The address is refused outright, not stripped of its zone.
IPv6 prefixes that embed IPv4IPv4-compatible ::/96, IPv4-translated ::ffff:0:0:0/96, NAT64 local-use 64:ff9b:1::/48, 6to4 2002::/16, Teredo 2001::/32Each is a way of writing an IPv4 address as IPv6, and the guard does not judge the IPv4 address inside. Because the guard judges the address a target resolves to, a prober whose DNS64 uses one of these prefixes refuses every IPv4-only target; see the warning about local-use NAT64 below.

Two IPv4-in-IPv6 forms are judged, not refused. In both, the IPv4 address sits at a fixed position:

  • IPv4-mapped ::ffff:0:0/96, such as ::ffff:10.0.0.1, is judged as the IPv4 address it carries, under every rule on this page. Do not confuse it with IPv4-translated ::ffff:0:a.b.c.d, which is refused outright.
  • The NAT64 well-known prefix 64:ff9b::/96 is not refused as a whole. The guard takes out its last 32 bits, the embedded IPv4 address, and judges that address under the same policy. The reason is DNS64: on an IPv6-only prober whose DNS64 uses the well-known prefix, every IPv4-only target resolves into 64:ff9b::/96, so refusing the prefix would fail every such monitor. So 64:ff9b::8.8.8.8 is allowed, while 64:ff9b::127.0.0.1 is refused as loopback.

Refused by default: private address space​

This class covers:

  • RFC1918: 10.0.0.0/8, 172.16.0.0/12 and 192.168.0.0/16
  • unique-local fc00::/7
  • carrier-grade NAT shared address space 100.64.0.0/10 (RFC 6598)
  • reserved 240.0.0.0/4 (RFC 1112), apart from 255.255.255.255, which is always refused as the limited broadcast address
  • benchmarking 198.18.0.0/15 (RFC 2544)
  • the "this network" block 0.0.0.0/8 (RFC 1122), apart from 0.0.0.0 itself, which is always refused as unspecified

One setting refuses or allows all of them together.

None of these is a public monitoring target, and from a shared public pop none of them belongs to the customer who named it. They lead into the network the pop itself sits in or, for 100.64.0.0/10, into the carrier NAT of whoever else uses that provider. From a prober inside the customer's own network, they are exactly what the prober is there to monitor. 100.64.0.0/10 is the pod and node CIDR that EKS-style clusters hand out, and some clusters use 240.0.0.0/4 for pods as well. That is why this class can be opted into and the one above cannot.

The refusal names the class it found. 100.64.0.1, for example, is reported as carrier-grade NAT shared address space, not as private space, so an operator does not go looking for an RFC1918 address that is not there.

The opt-in is PROXIMA_PROBE_ALLOW_PRIVATE_TARGETS=true. The value 1 also works. Any other value, TRUE included, leaves it off. Set it on a prober you run inside the network it monitors, and never on a shared public pop. Nothing enforces that, because the backend cannot see a prober's environment, so the rule falls to whoever deploys the pop. The value is read once at startup, so a change takes a restart. The prober then logs probe dial policy installed with allow_private_targets, which is how to confirm the setting took effect.

None of the install paths in Deploying a pop sets it, the console's Add host wizard included. Add it yourself:

# A VM or bare metal. The installer writes the proxima-agent systemd unit and has no flag for
# this, so add a drop-in beside it rather than editing the unit it owns.
sudo mkdir -p /etc/systemd/system/proxima-agent.service.d
printf '[Service]\nEnvironment=PROXIMA_PROBE_ALLOW_PRIVATE_TARGETS=true\n' \
| sudo tee /etc/systemd/system/proxima-agent.service.d/probe-private-targets.conf
sudo systemctl daemon-reload && sudo systemctl restart proxima-agent
# Docker: one more -e on the run command.
docker run ... -e PROXIMA_PROBE_ALLOW_PRIVATE_TARGETS=true ...
# Kubernetes: the agent chart's prober.extraEnv. The value must be the STRING "true"; an
# unquoted YAML boolean is not a valid container env value.
prober:
extraEnv:
- name: PROXIMA_PROBE_ALLOW_PRIVATE_TARGETS
value: "true"

What is not refused​

The guard refuses the classes above and nothing else. Some special-purpose address space is therefore allowed. This list is not exhaustive:

  • 192.0.0.0/24, the IETF protocol assignments block. It holds globally reachable anycast services, such as 192.0.0.9 (PCP) and 192.0.0.10 (TURN), so refusing the block would break legitimate monitors.
  • fec0::/10, the deprecated IPv6 site-local prefix.
  • 100::/64, the IPv6 discard-only prefix.

The IPv4 entry is allowed in the same way when it is embedded in the NAT64 well-known prefix.

No address-class check can refuse IPv4 carried under a prefix the network operator chose. That covers a NAT64 network-specific prefix (RFC 6052 §2.2), 6rd (RFC 5969) and an ISATAP interface identifier (RFC 5214). None of them has a fixed prefix or a fixed position for the IPv4 address. Such a prefix only exists where the prober's own network runs a translator, so it is a property of where the pop runs, not of what a tenant typed.

An IPv6-only pop behind local-use NAT64 cannot probe IPv4-only targets

Some probers sit on an IPv6-only network and reach IPv4 through DNS64 with the local-use NAT64 prefix 64:ff9b:1::/48 (RFC 8215). Such a prober refuses every IPv4-only target, and no setting changes that. RFC 8215 lets the operator choose where the IPv4 address sits inside that /48. The guard therefore cannot extract the address, so it refuses the whole prefix. A pop in that position needs native IPv4, or a translator that uses the well-known prefix 64:ff9b::/96.

Probes do not use HTTP proxies​

Probes ignore HTTP_PROXY, HTTPS_PROXY and NO_PROXY. Through a proxy, the prober connects to the proxy and the proxy connects to the target. The only address a connect-time check could ever see is the proxy's. A pop with a proxy set for its own egress would therefore carry every http monitor straight past the guard, into private space and to 169.254.169.254 alike. It is also the wrong measurement: a location is a vantage point, and a check relayed through a proxy measures the proxy's network. tcp, tls and dns checks never used a proxy.

A prober whose only way out is an HTTP proxy cannot run http or https checks, and the failures do not say why. Each check fails with whatever the blocked direct connection produces:

  • timeout where egress is silently dropped
  • connection_refused or internal where egress is rejected

Nothing in the error mentions a proxy. No setting turns proxies back on for probes.

How a refusal shows up​

A refused connection fails the check like any other failure. It is not skipped and not reported as healthy. It is that prober's failing observation, folded into its location like any other: under the any-prober-succeeds rule, one prober refusing while another prober at the same location succeeds does not fail that location. There is no dedicated error class for it:

  • http, tcp and tls checks report internal; dns checks report dns.
  • The error message names the refused address and its class. For private space it reads probe target address class is refused: 10.0.0.5 is private address space, and this prober is not permitted to reach it. A monitor for an internal target needs a prober inside that network. The message does not name PROXIMA_PROBE_ALLOW_PRIVATE_TARGETS.
  • The message is stored as that prober's last_error. When the monitor moves to pending or down, the monitor's last_error is copied from a fresh failing location, and a location that confirmed the change is preferred.

Refused earlier, at create and edit time​

When a monitor is created or edited, the backend runs a smaller version of this check. A monitor that obviously cannot be probed gets a 400 naming the problem, instead of failing on every probe forever. That check is a convenience and not the boundary. The prober's guard is the authority. Passing the backend check says nothing about whether a probe will be allowed to dial.

  • What it judges. It judges only the address a check dials, and only when that address is an IP literal. For http, tcp and tls that is the target. For dns it is config.resolver; a dns check's target is only looked up, so it is never judged.
  • What it skips. It never judges hostnames, because the backend is not on the prober's network. The prober refuses a hostname only when it resolves into refused space. The backend also only recognises canonical literals: 127.0.0.1 is judged, while spellings such as 127.1 or 2130706433 are treated as hostnames and left to the prober.
  • On every location, public or private, it refuses loopback, link-local, unspecified and multicast literals and the limited broadcast address 255.255.255.255, e.g. target "http://127.0.0.1/health" is a loopback address, which no probe location may probe, public or private. It also refuses any literal carrying an IPv6 zone, such as [fd00::1%eth0]:443, before looking at its class, and the message names the zone.
  • On a public location, it also refuses RFC1918, fc00::/7 and 100.64.0.0/10, and the message points at a private probe location. This includes a monitor with no location_selector, and one whose selector cannot be read.
  • On a private location, it accepts those three ranges, because it cannot see whether that location's prober opted in. A prober that did not opt in refuses them at dial time.
  • What it does not refuse. None of the IPv6 prefixes that embed IPv4 is refused, and neither are 240.0.0.0/4 apart from 255.255.255.255, 198.18.0.0/15, or 0.0.0.0/8 apart from 0.0.0.0. Only the prober refuses those.
  • Edits. An edit is judged on the monitor as it will be stored. Changing only the selector, or only the target, is still caught.

Assignment is replication, not sharding​

Every location a monitor's location_selector resolves to probes that monitor. Quorum depends on several locations checking the same target, so sharding monitors across pops would make confirmation impossible: the selector chooses which pops, and then all of them probe.

Each assignment carries a generation. A result produced against a definition that has since changed is dropped and counted (proxima_probe_results_rejected_total{reason="stale_generation"}), never folded into current state.

A pop fetches its set at startup, retrying every 15 seconds until it succeeds, and refreshes it every 150 seconds thereafter. A probe_reassign command exists to nudge a pop immediately, but nothing publishes it yet — so a newly created or edited monitor starts being probed within about two and a half minutes, not instantly. For the same reason, pausing or editing a monitor produces a short burst of rejected results while the pop finishes working from its old set — counted as no_longer_assigned and stale_generation, deliberately never as unassigned, which stays a forgery signal.

A prober's credentials live on its poll; its quorum seat does not

A prober cannot heartbeat: the heartbeat subject is tenant-scoped and it holds no tenant grant. So agents.last_seen_at is advanced by the assignment poll (and by each accepted result), and a pop that is idle — assigned nothing at all — still proves it is alive every 150 seconds, half the 300-second stale timeout, so a single lost poll cannot make a live pop look dead. Without that, stale detection would mark a healthy pop inactive. That no longer costs the pop its credentials (an inactive agent renews; only a revoked one is refused), but it would report a live pop as gone.

There are two liveness clocks and they are not interchangeable. agents.last_seen_at is credential liveness, refreshed by the poll as above. probe_locations.last_result_at is quorum liveness — a seat in the denominator — and it is refreshed only by an accepted result. Polling proves a pop is running; only a result proves it can still probe, so a pop that polls happily while its egress is blocked loses its quorum seat while keeping its credentials. The window before that happens scales with the monitor's interval (max(5 minutes, 3 × (interval + 30s))), which keeps it wider than the freshness window at every interval: a pop whose vote still counts must still hold its seat, or losing both at once would lower the bar a transition has to clear.

Cadence, retries and quorum​

  • A monitor is checked every interval_seconds from every location assigned to it. A check that exceeds timeout_seconds counts as a failure.
  • The first check of a monitor at a pop is offset deterministically by a hash of (monitor, location), so a restart lands in the same slot instead of stampeding the target, and two locations never hit one target on the same second.
  • While a location's check is failing, that location tightens to retry_interval_seconds.
  • A 429 Too Many Requests is up, marked rate-limited, and that location slows down. An http check that gets 429 is a success, whatever accepted_status_codes says (slowing down is politeness, not a verdict), and its keyword check is skipped — a rate-limit page never carries the keyword. It never counts toward down and never pages. The prober then backs off that monitor at that location only: to the site's Retry-After when it sent a usable one, otherwise doubling from twice the interval (60 s → 120 s → 240 s), always kept within [interval_seconds, 240 s]. A zero, negative, unparseable or huge Retry-After is clamped, never trusted, and an interval of 240 s or more is never stretched. The first result that is not a 429 ends the backoff: success goes back to interval_seconds, failure tightens to retry_interval_seconds as always. Editing the monitor restarts its loop and resets the backoff.
  • Accepted risk: a site that answers everything with 429 (a misconfigured WAF) shows up while its real users are refused too. A per-monitor "treat 429 as down" switch is not offered yet.

Before either axis is judged, each location's probers are folded into one vote by the rule above: fresh rows only, up if any of them is up, streak the min across them. Everything below therefore speaks in locations, and the number of machines behind each is invisible to it by construction.

Quorum has two axes, and both must be satisfied before a monitor goes down:

  • Time — a location must have failed max_retries + 1 checks in a row. Its failure run is its own; the per-location time axis is what stops one flapping pop from driving a monitor-wide counter.
  • Space — at least min_confirmations locations must agree.

Time alone pages when a single network path is bad. Space alone pages on one simultaneous blip. Requiring both is what makes a down worth waking someone for.

Recovery is symmetric in space and asymmetric in time: min_confirmations locations reporting success is enough to go back up, with no retry count — a monitor that is working again should say so at once.

A split fleet that satisfies both a down-quorum and an up-quorum resolves to down: it cleared the strictly harder bar and it is the actionable half. Preferring up would let a half-broken target hide behind whichever locations still reach it.

Freshness: silence is not evidence​

A location's vote counts toward current state only if its last check ran within interval_seconds + 30s. A pop has to miss a whole round and the grace before its opinion is discarded.

A location backing off from a 429 is fresh for as long as it said it would be. While backing off, the prober stamps each result with next_check_seconds — when it will run that monitor next — and that prober's row stays fresh for max(interval_seconds + 30s, next_check_seconds + 30s), capped at 270 s (the 240 s backoff cap plus the grace). The widening belongs to the row and nothing else:

  • next_check_seconds is counted from the check's own start (CheckedAt), not from when the prober is next free to run it: the prober's timer for that next check does not start until this check has finished running and been handed to the publisher. A slow check therefore eats directly into the 30 s grace — the row's real next report can land up to roughly the check's own duration later than next_check_seconds + 30s alone implies, since timeout_seconds (which bounds a check's duration) is only required to stay under interval_seconds, not under the grace.
  • Another prober in the same location, or any other location's row, is still judged under interval_seconds + 30s. The monitor-wide window is never widened, because that would keep a dead prober's months-old "healthy" row voting — the exact thing judging each row before the fold exists to prevent.
  • The folded vote carries the widest window among the rows that survived into it, and quorum re-checks the vote (both the ballot count and the agreed failure rounds) under that same window, so a backed-off location keeps its seat and its ballot.
  • The liveness window is unchanged. 270 s is under its 300 s floor, which is why the cap is 240 s and not the 300 s other products use; TestRateLimitBackoffCapStaysInsideTheLivenessFloor fails if the cap is ever raised past it.
  • Every result rewrites the hint, so the first result that is not backing off (a success, or a failure) narrows the row straight back. A site that goes down during a backoff is reported by that location at its next check — at most 4 minutes later — while locations that are not rate-limited keep their normal cadence and detect it on time.
  • A hint the backend cannot use (not an integer, outside [1, 240], or on a result that is not a rate-limited success) is dropped and counted as proxima_probe_results_corrected_total{field="next_check_out_of_range"}; the row is then judged under the ordinary window, and the verdict is never touched.

Upside-down monitors and 429 (accepted, not special-cased). For an inverted monitor a 429 is a probe success, so the post-polarity verdict is unhealthy — the site is up. The prober still backs off, so that row stays fresh for up to 270 s, and detecting the site going down (the inverted monitor's recovery) is delayed by the backoff at that location. Inverted monitors against rate-limiting sites are rare enough that this is documented rather than special-cased.

During the agent rollout the fleet can be mixed: a prober still on the old agent calls a 429 down (and tightens to the retry interval), a new one calls it up. The location fold and quorum resolve that exactly as they resolve any split. The backend ships first, so a new prober's next_check_seconds is always understood.

If no vote is fresh, the state does not move at all. The system knows nothing about right now; reading that as recovery would close a live incident, and reading it as failure would page for our outage rather than the customer's.

The window also matters because a pop that lost its NATS connection buffers results to disk and replays them on reconnect, carrying their true (old) checked_at. Those results contribute history but do not vote on current state.

Partial failure is a signal about the pop, never a state​

When a minority of locations report failure while the quorum still sees the monitor as up, that is recorded as proxima_probe_partial_failure_total{location} and logged — and the monitor's state does not move. One pop failing while the others succeed is almost never the target; it is that pop or its transit, and the client cannot act on it.

Degraded quorum​

The bar actually applied is min(min_confirmations, locations that can vote right now) — where "can vote" means the location is assigned by the selector, capable of this check kind on every enabled prober behind it, and has at least one enabled prober reporting inside the liveness window — floored at 1.

Those two quantifiers differ on purpose. A location stays live while any of its machines is reporting, because evicting a seat that still has a working prober behind it would lower the bar. A location is eligible only if every one of its machines matches the selector and can run the kind, because a location that fails permanently on half its machines would manufacture an outage.

A monitor that used to confirm may now report QuorumDegraded instead

This is the fix working, and it will look like a regression the first time you see it.

Before the vantage-point model, three machines in one place cast three votes, so a monitor with min_confirmations: 3 could reach quorum from that one place. It now gets one vote from there, sits below its configured bar, and reports QuorumDegraded — while min_confirmations: 3 on a two-location fleet simply cannot be met by two locations.

Nothing is being hidden: the monitor still transitions at the reduced bar and says loudly that it did. But if a monitor stopped confirming after this shipped, the cause is almost certainly that its old confirmations came from prober redundancy rather than from independent vantage points. The fix is to add a location, or to lower min_confirmations to what the fleet can honestly supply.

When that is lower than the configured min_confirmations, the quorum is degraded: Console keeps working at reduced confidence and says so. Failing closed would hide real outages and failing open would manufacture them, so neither is acceptable; the only honest answer is to carry on and be loud about it.

You see it in three places:

  • The timeline entry carries the arithmetic: confirmed down by 1 of 1 available locations (quorum 1); QUORUM DEGRADED — 2 configured.
  • The monitor detail page shows a Quorum degraded badge beside the confirming count, with a line saying the state change was confirmed by fewer locations than the monitor requires.
  • proxima_probe_quorum_degraded{kind} counts the monitors this backend process last judged below their configured quorum.

What it means for you: the state below is real — it was confirmed — but it was confirmed by a smaller set than you asked for, so it carries the confidence of that set. Check that every location the monitor selects is deployed, enabled, capable of the kind, and reporting. A monitor whose min_confirmations exceeds the pops it can ever reach would otherwise sit looking healthy through the very outage it was created to catch.

proxima_probe_quorum_degraded is a PER-REPLICA gauge

Each backend replica only knows the monitors it evaluated. A monitor judged by two replicas is counted twice by sum(); one judged by a single replica is missed by max() over the others.

Alert on the existence of the condition, never on the number:

sum(proxima_probe_quorum_degraded) > 0

The magnitude lies somewhere between the true count and N times it. It is a trend, not a census — do not put it on a board as "monitors degraded".

The write path enforces the other half of this: creating or editing a monitor whose min_confirmations exceeds the number of distinct location codes its selector resolves to (counting only the ones capable of that kind, and excluding any code whose rows disagree on kind/client_id) is a 400 naming both numbers. It counts vantage points, not rows, for exactly the reason the numerator does — a ceiling drawn from machines would admit a min_confirmations no fold could ever reach. That check only fires when the selector resolves to at least one location; with an empty fleet every value is accepted, because an empty fleet is a deployment-wide fact rather than a property of the monitor being written.

The state machine​

A monitor's observed state is one row, updated in place, plus an append-only timeline of transitions.

StateMeaning
unknownNever checked. A monitor no location has reported on: newly created, or selecting locations that do not exist. Not a health verdict.
upA quorum of locations reported success.
pendingFailing, retries not yet exhausted — a real, unconfirmed outage. This is the notification-suppression state: it is what stops a single bad round from paging, and a monitor that recovers from here was never an incident.
downFailing past max_retries, confirmed by quorum.
maintenanceA maintenance window covers the monitor. Entered whatever the quorum says and sticky while covered — a probe result can never leave it, because that would re-arm paging inside a window someone opened on purpose. The monitor keeps checking. When coverage ends it leaves through the current votes: confirmed up → up; failing → pending; nothing decided → unknown. paused wins over it. Emission and notify refuse with in_maintenance.
pausedThe monitor is administratively disabled. Written by the prober path when results arrive for a disabled monitor, so a paused monitor does not sit reading down forever.

pending and unknown are never conflated. One is "something is wrong and being confirmed", the other is "nothing has ever looked". The console gives them different colours as well as different words, and a third treatment to a monitor whose state row is missing entirely — that is a storage anomaly, not a status, and filing it under unknown would hide it.

Leaving maintenance while still failing pages on the next result, not after a fresh retry cycle. A location's consecutive-failure count is never frozen by a window — only the monitor's state is — so failures during the window keep counting toward max_retries. A target that was broken all through the window usually already satisfies confirmation when coverage lifts: the monitor stops at pending for one result (leaving a window is never itself the page) and the next failing result confirms down and pages. This is deliberate — something still broken when planned work ends should page at once, as an alert held by the window does — and its cost is that a site which recovers 30 seconds after the window can page once.

A push monitor enters and leaves maintenance at its next evaluation. It has no probe stream to re-judge it, so the state moves when a heartbeat arrives or when the missed-heartbeat sweep finds it overdue. Until then the list shows the state it had; its pages are held regardless, because the emission gate asks the maintenance index itself rather than reading the state.

Monitor groups have no maintenance state. A window covering a group (directly, or through its client or environment) holds the group's page at the same gate; a member in maintenance is simply not measured, so the group's state stays the verdict of its measured members. Label rules do not reach groups on the emission path — group tags are not in the in-memory index.

Rate limited is not a state. A monitor whose target answers 429 stays up: there is no rate-limited state, no monitor_events row and no notification. The API's rate_limit object and the console's amber chip are the only places the slowdown shows (see API).

Alongside the state, the row carries since, last_change_at, last_result_at, consecutive_failures (the quorum-agreed number, not the worst pop's), confirming_locations, last_error, and alert_group_id — which is written by nothing and is not how a recovery finds the group to close. See The grouping key; the column is retained, with a comment saying so, because a column a spec references and nothing populates is how the next person gets misled.

Every transition appends a row to the event timeline: from-state, to-state, timestamp, the locations that confirmed it, and a reason carrying the arithmetic. The detail page renders it newest-first and names the locations, because "who saw this?" is the first question of any incident.

The timeline is not a dump of raw columns. From/to states render with the same chip vocabulary as the current-state card and the monitor list (MONITOR_STATE_META) rather than as raw lowercase strings, so a monitor never reads differently on two surfaces. Each row also states how long the previous state held — the first question of any incident is "how long was it down?", and the arithmetic is free: the previous row's at is already on the page. The oldest row loaded states no duration, deliberately: the timeline may continue past what was fetched, and inventing a number against data never loaded would assert something never measured.

A degraded quorum is a field, not a parse. GET .../events computes quorum_degraded server-side (uptime.ReasonIsQuorumDegraded), from the SAME marker quorumNote writes into reason — one parse, in the package that owns the text, rather than each caller re-deriving the fact independently. Before this, the console's own frontend matched reason.includes("QUORUM DEGRADED") on its own, so a wording change to the backend's reason text could silently stop the badge from rendering, with no test in either codebase positioned to catch it — the write and the read lived in different languages and never had to agree by construction. Do not re-derive this from reason a second time anywhere; read the field.

A truncated timeline says so. GET .../events is capped at 200 (limit, default 50), and meta.total carries the monitor's true transition count. A flapping monitor's timeline used to simply stop at the cap with nothing indicating more transitions existed — the same defect class as a paginated list whose header understates its own count. The detail page renders "Showing the N most recent of TOTAL transitions" exactly when meta.total > meta.per_page, and nothing when the timeline is not truncated.

confirming_locations is the set behind the last transition, not a live tally, and it is legitimately empty for a state no quorum confirmed (pending, paused).

Pause and resume touch enabled only. They do not rewrite the observed state: the state row records what a probe saw, and a configuration change must not fabricate a transition. The monitor moves to paused when the prober path next sees a result for it. Both are idempotent.

The vote ledger​

Each prober's latest opinion is one row in monitor_location_state, updated in place — one row per (monitor, prober), never an append-only history. A location is a vantage point that may be backed by several machines, so the row is keyed by the agent that cast the vote; keying it by the location code would make a second machine in one location silently overwrite the first. Raw results stay out of Postgres (five pops × a thousand monitors on a 60-second interval is ~7,000 rows a minute); the ledger is five thousand rows in total.

It has to be durable rather than in-process because several backend replicas share one durable JetStream consumer: results for one monitor land on different replicas, and an in-memory tally would give each replica a partial view of the fleet, so a quorum would never assemble.

A vote is only accepted if it is newer than the one it replaces, which makes a JetStream redelivery idempotent and stops a buffer replay rewinding a fresher vote.

Deploying migration 000190 produces an error burst. Expect it; do not roll back on it.

000190 drops and recreates monitor_location_state to re-key it by prober. The drop itself is benign — votes are ephemeral, every prober rewrites its row each interval, and NextState returns early on zero fresh votes, so a monitor's state freezes rather than flipping while the table refills. monitor_events, the durable record, is untouched.

The blue/green window is the part nobody sees coming. Old and new pods share one durable consumer, and the old backend's upsert names ON CONFLICT (monitor_id, location) — a constraint that no longer exists. So every probe result that lands on an old pod fails to persist and NAKs, for as long as both versions are serving. proxima_probe_results_rejected_total{reason="persist_error"} climbs and the ingest-failure log fills up.

Nothing is lost and nothing is corrupted: the retry budget is patient by design (MaxDeliver 10 spread over ~11 minutes, so a redelivery lands on a promoted pod), and any result that arrives after the freshness window contributes history without voting on current state. Nothing pages, because a monitor with no fresh votes cannot transition.

Incident timeline​

monitor_events is a flat log — every transition, one row each. GET /monitors/{id}/events reads it that way, newest first. GET /monitors/{id}/incidents reads the SAME rows folded into incidents instead: a maximal run starting the first time the monitor moves away from up and ending the next time it returns to up. No schema change and no new table — uptime.FoldIncidents walks the existing rows read-time, oldest-first, the direction they actually happened in.

An incident's fields:

  • entered_state is the worst state the run reached: down if any transition confirmed it, else pending for a run that recovered before ever confirming down (a blip, not an outage).
  • ended_at is null while the incident is still open — never a guessed "now". A UI renders a live counter for that case, not a static duration.
  • quorum_degraded is the OR across every transition in the run, not just its opening or closing one — one degraded confirmation anywhere in the incident marks the whole thing.
  • id is the id of the incident's OPENING transition, not a value of its own. Incidents are never persisted, so this is what gives one a stable identity — a future feature (a postmortem, an announcement) could reference it without this endpoint minting new rows for something that does not exist as a row.
  • transitions carries every row in the run, oldest first — the same detail /events shows, so a UI's drill-down never re-derives anything the flat timeline already computed.
  • Each transition's error is the incident's cause: the check error behind it, as monitor_state.last_error held it at that moment (a fresh error from a confirming location, else any fresh failing location's). The state row's copy is overwritten by the next check, so this is the only durable record of why. null means "not recorded" — the transition predates the column (migration 000233) — and is kept distinct from "", "no failing location reported error text". The console shows the error behind the transition into down as the incident's cause.

Paging is cursor-based, not total-based. Counting a monitor's whole incident history honestly means walking it, which gets expensive for an old or flappy monitor — so this endpoint answers has_more/next_before instead of a total, and pages by time (?before=, an RFC3339 instant) the way an operator actually browses history, not by page number. next_before is the earliest incident actually returned; passing it back resumes exactly there, whether the previous page stopped because it hit the requested limit or because its raw batch (monitorEventsMaxLimit transitions) ran out. An incident whose true start lies on an earlier, not-yet-fetched page is never shown with a fabricated start — it is dropped from the page that would have cut it in half, and has_more says so.

In the console​

The monitor page's Incidents card groups incidents by the calendar day they started on (Today / Yesterday / a full date), newest first. Each row is a shield (red while ongoing, green once resolved), "Down" or "Failing", the cause in monospace (or "Cause not recorded" for an incident from before causes were kept), the time range, and the length or "Ongoing". A row opens the incident's own page — /uptime/:monitorId/incidents/:incidentId — with the cause, when it started and how long it lasted, which locations confirmed it, the request that was checked, and the timeline told in sentences ("Confirmed down by 2 locations", the check error under it, the state machine's own reason in muted text). A service has the same list and page, naming its failing monitors where a monitor has its cause. There is no incident-by-id endpoint: the page walks the same paged /incidents list until it finds the id (at most ten pages) and otherwise says the incident is not in the recent history.

While an incident is open, the Incidents tile carries an "Ongoing incident ›" link to it, and every incident in the chosen window is an attention band on the response-time chart (see docs/standards/charts.md — windows are bands) whose tooltip names the cause and which opens the incident on click. "Load earlier" replaces a static truncation notice.

Monitor groups​

A twelve-endpoint service going down produces twelve pages. A monitor group is the fix: a named service — a set of one client's monitors, in one of that client's environments — judged as a single thing and paged once, in a page that still names the members that failed.

Groups are shipped end to end. Six REST endpoints author them, group_id on the monitor bodies fills them, the console drives both, and a group reaching down emits through the gate chain. monitor_groups.paging_enabled ships false on every group exactly as a monitor's does, so nothing pages until somebody flips it — and flipping it has preconditions no deploy performs.

Most of the rules below were bought with somebody else's outage

Grouping has a well-documented failure mode, and the products that have shipped it disagree about how to avoid it. Datadog's composite monitors, CloudWatch's composite alarms and Checkly's groups converge on the aggregate alerts and the members do not — but leave it to operator discipline or to inheritance defaults. Uptime Kuma models a group as a monitor and enforces nothing, and its issue tracker is the price list.

What Console doesWhose bug report it comes from
A group is never counted as a monitorKuma's dashboard reads "3 down" for two monitors and their group (#4597)
Paging authority is exclusive, refused at authoring timeKuma's still-open request for a "primary monitor" that silences its children (#3548)
all_children_down is not offeredKuma's all_down/any_down request, which only makes sense where children still notify (#5936)
A group overrides exactly one thing, and never silentlyCheckly moved group-level overrides to optional-and-off in V2 after group settings overwrote check settings (terraform-provider-checkly#332)
Group pause is deferred rather than half-builtKuma's group pause does not reach its children (#7341)

The one thing a group takes over is paging authority, through an explicit act that errors on conflict. Interval, retries, timeout, severity, location_selector, upside_down and every other monitor setting stay per-monitor and are never inherited, overridden or cascaded.

What a group is​

  • One client and one environment. monitor_groups.client_id and monitor_groups.environment_id are both NOT NULL, and a monitor may only join a group in its own client and its own environment — monitors_group_scope_fkey, a composite key over (group_id, client_id, environment_id), so the rule also holds for bulk updates and hand-run SQL.

    That is not tidiness. A first-party alert source carries its entity ids rather than describing them in labels, and that is what makes an entity-scoped escalation route match a probe alert at all (see Entity-scoped routes match). A group spanning environments could carry no environment_id, and would silently lose route matching for every client whose routes are environment-scoped. A group carries no host_id, because a group spans hosts by construction.

  • Not a monitor, and never counted as one. monitor_groups is its own table. A group is never probed, never appears in monitors, and appears in no list count, section header, stat card, firing_count, coverage ratio or metric that counts monitors. This is written down as an invariant rather than as a layout preference because it is the single most-reported defect of the alternative: modelling the group as a monitor is exactly what makes a dashboard report "3 down" when it is two monitors and their group.

  • Created empty, starting at unknown. There is no members field on the create; a group is filled from the monitor side. A group with zero members is legitimate — authored before the monitors that will fill it — and can never page.

  • One level: groups do not nest, and monitors.group_id is the whole of membership. No join table, no group-side write.

  • A group is called a service in the console. "Group" is what the table is; "service" is what it means.

Derived state: the fold, and what is left out of it​

A group holds no opinion of its own. Its state is the worst state among the members that are actually being measured, re-folded from scratch every time something re-derives it, and nothing else ever writes it.

"Measured" is an allow-list: a member counts only if it is enabled and its state is one of up, pending, down.

  1. Any measured member down → the group is down.
  2. Else any measured member pending → pending.
  3. Else at least one member measured → up.
  4. Else — nothing measured at all — unknown. Never up.

Everything else is excluded from the verdict and counted:

  • A paused member is excluded. It is not being measured; calling it down is a lie and calling it up is a bigger one. enabled is read separately from the state row on purpose: pause touches enabled only, so a monitor paused while it was down may still carry state = down until a trailing result lands — and whether one lands depends on timing. A fold that judged on the state alone would hold a whole group down, non-deterministically, on a member nobody is measuring.
  • An unknown member (never checked) is excluded. It is not evidence about right now.
  • So is anything else — maintenance, or any state a future build introduces. That is why the rule is an allow-list rather than "not paused and not unknown". Under a deny-list an unrecognised state counts as measured, contributes neither a down nor a pending, and the fold lands on up: a green service asserted from a member nobody looked at. On the five states the fleet actually produces the two rules are identical; they differ on exactly one thing, and it is that where a deny-list would say up, the allow-list says unknown.

The exclusion is paging-neutral, which is what makes it a display-honesty rule rather than a change to who gets woken: excluding a member can only ever remove a down from the fold, never add one.

A verdict computed from a partial sample says that it is partial. The state row carries member_count, measured_count, down_count and excluded_count; each timeline event carries the members that were down and the members that were excluded, by name as well as id, so the history stays readable after a rename or a delete; and the console renders a dashed "Partial — 2 of 5 not measured" badge rather than a clean verdict. That is the posture has_data and QuorumDegraded already take in this subsystem.

Two columns behave the way their monitor-side counterparts do. monitor_group_state.state can hold paused and nothing can reach it — group pause is deferred, and a group whose members are all paused derives unknown. monitor_group_state.alert_group_id is reserved, written by nothing and read by nothing, exactly as monitor_state.alert_group_id is; see The grouping key.

Paging: authority is exclusive, and it is enforced by refusal​

A group may hold paging_enabled. While it does, no member of it may — and the API refuses, at authoring time, every path that could create the pair:

The editThe refusal
Enabling paging_enabled on a group while any member holds its own400 naming those members, in the message and again in error.details.members
Enabling paging_enabled on a monitor whose group holds it400 naming the group
Moving a monitor that holds paging_enabled into a group that holds it400 naming both sides
Refusal at authoring time, never suppression at runtime — this is the load-bearing decision of the whole feature

The obvious design is the other one: let both flags stand and de-duplicate at emission, muting the members whenever the group pages. It is rejected, and not on taste.

Runtime suppression is a new silence mechanism, and a silence mechanism can fail. Suppose the members were muted to protect an operator from twelve pages, and the group's one page were then lost downstream — no matching escalation route, oncall_enabled off, the page-owner team still in shadow, a publish the broker refused. The outage would reach nobody, while every member that would have paged was deliberately silenced on the strength of a page that never arrived. That is strictly worse than having no groups at all, and it is worse in the one direction this subsystem refuses to move.

Refusal has no such failure mode, because there is nothing to un-suppress. The group emitter reads the group's flag and the member emitter reads the member's, and neither consults the other — so a page can only ever be lost the way it is lost today, through the existing, already-instrumented chain under Turning paging on. Grouping adds no new way to be silent.

The price is that the conflicting configuration stays reachable by routes no handler sees — two concurrent PATCHes each reading a world in which they are allowed, or one hand-run UPDATE monitors SET paging_enabled = true — so it is counted rather than prevented. proxima_monitor_group_authority_conflicts_total is the only thing anywhere that reports a group and a member both paging, and it never self-heals: it ticks every sweep until a human turns off one of the two flags.

Authority is only ever authored, never inherited. Moving a monitor out of a paging group — or deleting the group, which detaches its members — leaves that monitor's own paging_enabled exactly where it was, which is off. A service that paged as one thing therefore stops paging at all until somebody re-authors it. That is the honest outcome: the alternative is a page for one endpoint in the name of a service nobody re-authored.

Why all_children_down is not offered​

The paging rule is any_child_down, and in v1 it is the only one. Members failing after the first update the existing alert rather than opening a new one, so twelve failing endpoints stay one page that names twelve endpoints.

all_children_down is deliberately absent rather than merely unbuilt. The request it comes from (uptime-kuma#5936) is written for a world where children still notify, and there it is an additive summary sitting on top of per-child pages — strictly more signal. Under exclusive authority it inverts and becomes subtractive: the members are silent by construction, so "page only when all of them are down" means one member going down pages nobody at all. Copying the knob without copying its context imports a silent-outage footgun. If it is ever added it needs a withheld-page counter of its own and an unmistakable warning in the form.

The gate chain, and the resolve that is never gated​

A group page passes the same gate chain a monitor's does — paging_enabled on the group row, then ProbeSanityGuard tier 2, in that order — and it is the same code rather than a resemblance. The gate chain documents the chain and its counters; Monitor groups page as one thing documents which group transitions reach it.

Two properties are worth carrying here:

  • Tier 2 is re-checked at emission rather than inferred from the member states. A bug shipped to every prober fails every member of every group at once, and the derived page is that bug's loudest and most wrong expression: "everything in Payments is down".
  • Resolves are never gated. A held resolve is not a stale row somebody tidies up later. The find-or-create path matches status <> 'resolved', so the next genuine outage merges into the group that was never closed; the firing epoch bumps only on a reopen, which a group that never resolved can never have; and escalation arms only while the escalation snapshot is null, which the first outage already filled. A held resolve therefore makes the next outage silent, with paging_enabled on and the operator believing they are covered.

Alert identity​

  • The grouping key is proxima_probe_group:<group_id>, derived from the group's immutable UUID, so no edit an operator can make — the name, the membership, the severity, the flag itself — can re-key an open incident. It cannot collide with a member's proxima_probe:<monitor_id>.
  • alertname is ProximaMonitorGroupDown, deliberately distinct from a member's ProbeMonitorDown: a group page and a member page are different events with different remedies, and filtering or silencing one must not catch the other.
  • The resolve is per-fingerprint, never the group-level resolve flag — the group-resolve shape pairs with an empty alert list, and an empty list rewrites the group's severity to unknown at the instant it resolves.
  • Severity is the group's own, defaulting to P2 as a monitor's does. A group page is bigger, but defaulting to P1 would invent urgency the operator did not author.
  • The page publishes through the client's existing proxima_probe alert source. A group is not a new source type.

Evaluation, and the auditor that exists because a missed evaluation is silent​

Evaluation is edge-triggered: every write that can change the fold asks for a re-derivation — a member's confirmed transition, a pause, a resume, a delete, a move between groups, a group edit. No polling, no drift window. When a group's state is re-derived carries the full table and the six deliberate properties of those calls.

Edge-triggering has exactly one failure mode: a mutation path that changes the fold and does not ask. That failure is completely silent — the group keeps its last derived state forever, the UI renders it with the same confidence as a fresh one, nothing errors, nothing retries, no log line is written anywhere, and the page that should have gone out is simply never sent. So it is paired with a dead-man's switch, the pattern ProbeLocationAuditor, DispatchCompletenessAuditor and the escalation auditor already follow: MonitorGroupAuditor re-runs the real evaluation for every group in the fleet every five minutes and counts what moved.

A non-zero proxima_monitor_group_drift_total is a bug to fix at a mutation path, not a threshold to widen. correction="state" means a group's derived state was flatly wrong — a page or a resolve was missed, and the sweep has just now sent it, late. correction="counts" means the state was right and the evidence beside it was stale: the outage paged, and the operator was shown the wrong number of failing members. The correction is the cheap half. A corrected drift that is not counted looks exactly like a system with no drift, which is why the counter — not the repair — is the part that gets the missing evaluation call written.

Availability history​

Raw probe results go to VictoriaMetrics, in the monitor's own client's tenant namespace, resolved server-side from the monitor row:

SeriesLabelsValue
probe_successmonitor_id, location, kind1 / 0 — the probe's verdict, before upside_down
probe_success_maintenancemonitor_id, location, kind1 / 0 — as probe_success, for a result taken while the monitor was in maintenance
probe_duration_secondsmonitor_id, location, kindcheck duration
probe_phase_duration_secondsmonitor_id, location, kind, phasehow long one phase of the check took: dns, connect, tls, wait (time to first byte), transfer. Written only for a phase the check measured
probe_ssl_cert_expiry_secondsmonitor_idcertificate expiry (no location label — a certificate is a property of the target, so every pop sees the same value)

upside_down is applied by the reader, not baked into the series: a metric called probe_success must record whether the probe succeeded.

Rate-limited stretches in the availability figures​

A backed-off location checks up to 4× less often, and the figures are not re-weighted for it. Three effects follow:

  • Check-share figures are skewed. uptime_percent and total_checks are a share of checks, so a rate-limited stretch contributes fewer samples and is under-weighted against everything else in the window: 50 minutes rate-limited followed by 10 minutes hard down reads well below the true 83%.
  • The availability grid thins out. Buckets are check counts too, so at small bucket widths a backed-off stretch renders as sparse no-data cells beside a green monitor carrying the rate-limit chip.
  • Confirmation is later. The time-based availability table counts confirmed-down time, and confirmation from a backed-off location can be up to 4 minutes late. A real outage that begins inside a rate-limited stretch is under-reported by that delay.

The first two push a figure down and the third pushes it up, so in a window that mixes rate limiting and a real outage both kinds of figure move, in opposite-looking directions.

Maintenance in the availability figures​

While a monitor is in maintenance the ingest writes its success sample to probe_success_maintenance (same labels, same cardinality) instead of probe_success. Every check-share figure — uptime_percent, total_checks, the grid — therefore leaves maintenance checks out: they are a share of checks, and a period with maintenance has fewer checks behind its percentage. A grid bucket whose every check was a maintenance check is drawn in the maintenance colour (maintenance_checks on each bucket, on the detail bar, the list grid and a service's folded bar). The time-based table has a Maintenance column and leaves maintenance time out of both downtime and the measured length, so a run that entered maintenance and was still broken when it ended does not count its maintenance hours as downtime. Latency and phase series are unchanged. Push monitors have no probe_success series and write nothing here.

Latency phases​

A prober times each phase of an http, tcp or tls check and reports it on the result as timings, a schema-first field (api/proto/probe/v1/probe_timings.proto) in milliseconds:

  • dns — resolving the target's name, from the start of the lookup to a successful resolution. Absent for an IP target.
  • connect — the TCP connect, from a hop's first attempt to its first successful one. A Happy Eyeballs address family that failed first still counts as connecting. Absent for a dns check.
  • tls — the TLS handshake. Absent for plain http:// and for tcp checks.
  • wait — from the connection being ready to the first response byte. This includes writing the request, body and all, so a large request body reads as waiting, not as its own phase.
  • transfer — reading the final response's body only, from its first byte to the end of the read, capped at 1 MiB. dns, connect, tls and wait accumulate across every hop of a redirect chain (each hop dials afresh, and each hop's own wait is added in), but transfer never does: an earlier hop's body is discarded by the HTTP client without being read, and that time is in the check's duration but belongs to no phase.
  • A phase that did not happen is absent, not zero. An IP target has no DNS lookup. http:// and tcp have no handshake. tcp and tls make no request, so they have no wait or transfer. A dns check reports no phases at all. A present 0 means the phase ran and took no measurable time — it is a different fact from "not attempted."
  • Timings ride only on a check whose phases all finished: a response was read (status and keyword failures included), a tcp connect succeeded, or a TLS handshake completed. A transport failure carries none, because half a breakdown under a timeout would not add up to its duration.
  • The backend bounds every timing against the check's own duration before storing it. A single phase that is negative, non-finite, or longer than the whole check is dropped alone (proxima_probe_results_corrected_total{field="timings_out_of_range"}) — noise, one odd phase on an otherwise sane result. If the phases left over still sum past the check's duration, none of them can be true and the whole set is dropped instead (field="timings_self_contradictory") — this is the signature of a prober build reporting in the wrong unit (microseconds as milliseconds), not ordinary noise. A block that fails to decode at all drops every phase (field="timings_malformed"). None of the three ever touches the verdict.

GET /monitors/{id}/uptime attaches each step's breakdown to its latency point as phases, in latency and in latency_by_location, from avg by (phase) and avg by (location, phase) on the same grid as the duration line. A failing phase query degrades rather than failing the request: phases is simply omitted from the points it would have filled, a WARN is logged, and the rest of the response (uptime %, counts, the duration line itself) is unaffected — every other read behind this endpoint still answers a 502 on failure, exactly as before. The console draws the breakdown as a stacked area. A window that opens before its probes recorded phases says when they start. A monitor with no phase data draws the single response-time line, and the all-locations view notes "Breakdown from N of M locations" when only some locations report phases anywhere in the chart's window — N and M are counted once over the whole window, not per step, so the note does not flicker between steps as different locations happen to report a phase.

The measured total stays visible over the stack. A transport failure — a timeout, a refused connection — carries a real duration_ms (the check still took however long it took to fail) but no phase timings at all ("Timings ride only on a check whose phases all finished", above). Drawing only the five phase bands would therefore show that check as the previous healthy check's stack carried forward, or as a bare gap — either reads as the monitor being healthiest exactly during the incident. The chart draws the measured total as a solid neutral outline line ("Total response time" in the legend) over the phase stack: it normally sits flush on top of the stack, and during a timeout it spikes above it, leaving the gap between the line and the stack top visibly unattributed. The tooltip's Total row is this measured duration, never the sum of the phases, and adds a "Not attributed to a phase: N ms" line whenever the two differ by more than max(20%, 50ms) — see docs/standards/charts.md, "Stacked area".

Rollout: a prober reports phases once it runs an agent that measures them. It gets that agent through the fleet self-update rollout, after the backend is already live — no backend or API change is needed for a location to start reporting. MinSupportedAgentVersion is unchanged: an older prober is fully supported and simply reports none. While locations differ, "All locations" averages phases over the locations that report them, while the headline averages the duration over all of them.

GET /api/v1/monitors/{monitorID}/uptime?window=&buckets= computes availability from those series: uptime percentage, the check counts behind it, the mean duration, a per-location breakdown, a latency sparkline, and the per-location availability bar described below. window accepts Go duration syntax plus a whole-number d/w suffix, bounded to [5m, 30d], default 24h. buckets is the bar's target segment count, bounded to [10, 200], default 48 — see Segment count follows the container. Both are rejected with a 400 rather than a silent clamp when out of range or unparseable.

No data is not zero percent

A monitor nothing has checked answers 200 with has_data: false and uptime_percent: null. A 0 means every check in the window failed.

Clients must not default the null to 0 — that renders an unmonitored service as a failing one. The console guards on has_data and says "no probe data yet" instead of drawing a chart, and it keeps a third answer separate again: if the metrics backend is unreachable the panel says the history could not be read, because that is Console's outage and not a statement about the monitor.

The availability bar​

The same endpoint returns history: one bar per location, as a contiguous grid of equal buckets covering the window, oldest first. Each bucket carries total_checks and healthy_checks — plus checked_at in the one case it can be answered, see Reading one segment; bucket_seconds is the width of one segment.

Per location is the point. monitor_events carries no location, so a bar built from the monitor's state transitions can show that it went down but never that one vantage point disagreed with the rest — which is the most useful thing a multi-location quorum has to say, and the first thing worth knowing during an incident.

The console draws one bar that follows the page's location picker. "All locations" adds each location's checks per bucket rather than averaging percentages or taking the worst, so a bucket where one location failed while the others passed still classifies as partly failing, and a location that reported nothing adds nothing; hovering it names what each location saw ("eu-west 27/30 · us-east 30/30"). Picking a location draws only that location's bar, from the same per-location history this endpoint returns.

Four states, from the two counts:

BucketMeaning
healthy_checks == total_checks (and total > 0)every check in the bucket found it healthy
0 < healthy_checks < total_checkssome checks failed — a flapping bucket, not a clean verdict
healthy_checks == 0 (and total > 0)every check failed
total_checks == 0nothing was measured
An unmeasured bucket is not a healthy one

Buckets nothing landed in are transmitted, with total_checks: 0 — never omitted. A segment the server leaves out is a segment the client has to infer from the time grid, and an inferred gap gets drawn in whatever color is nearest to hand. That is how a period during which the probes themselves were down ends up rendered as a period during which the service was fine.

The check order matters for the same reason: healthy_checks >= total_checks is trivially true when both are zero, so "not measured" has to be decided before "fully healthy". The console renders it as its own fourth state, a hatched grey cell with its own legend entry, and its screen-reader summary counts unmeasured buckets out loud — a bar that is forty grey segments and eight green ones is not a healthy bar.

bucket_seconds is floored at the monitor's own check interval. A bucket finer than the interval is empty by construction — no check could have landed in it — and a comb of empty segments reads as an outage rather than as the sampling artifact it is. A bar therefore often carries fewer segments than it was asked for — see Segment count follows the container — and that is the honest shape: fewer, wider segments say less, and everything they say is true.

Reading one segment​

A single line above the bars says what the cursor is over: the location, the segment's whole range — both ends, from at and bucket_seconds — and the counts behind it. The line is always rendered at a fixed height, so hovering never reflows the bar out from under the pointer that summoned it, and it is one line for every bar rather than one per row, because "what am I pointing at?" has a single answer at a time.

The range is the part the native tooltip never carried. A segment is a bucket, not a moment: on a 24h window a bar of half-hour segments reads as 48 checks unless the width is spelled out. Every segment also keeps its title attribute — slower, but it survives in a DOM inspector and without the JS that drives the readout.

The legend and the cells never rely on colour alone. Each legend entry is a swatch, an icon and its words: a check mark for "All checks healthy", a warning triangle for "Some checks failed", a crossed circle for "All checks failed" and a dashed circle for "Not measured". A not-measured cell is hatched as well as grey.

Keyboard. Each bar is one tab stop, and focus lands on the newest segment. The arrow keys move one segment at a time, Home and End jump to the ends, and nothing wraps. Tab moves on to the next bar or out of the panel, so focus is never trapped and the location picker stays in the normal tab order.

A focused segment drives the same readout line as hover. Its accessible name starts with the state word, e.g. "Some checks failed, 00:30–01:00 — 3 of 6 checks healthy", and its title is the same text. The focus ring stays visible in Windows forced-colours mode.

checked_at — the exact time, when there is one

When a bucket holds exactly one check, the response carries checked_at: the real timestamp of that check, read from the probe series with MetricsQL's tlast_over_time rather than inferred from the grid. The readout then shows the minute instead of the range, because "sometime in this half hour" can span both sides of a deploy — and the moment a lone check landed is exactly what an operator is squinting at the segment to find out.

The field is absent for every other bucket, and that is the contract rather than an omission: a bucket nothing landed in has no time to report, and a bucket holding four checks has four. Never substitute the bucket's own at for it — at is a grid slot up to bucket_seconds wide, not an observation. A client that finds it missing on a single-check bucket (a page cached before the field shipped) falls back to the range.

Choosing the window​

The panel offers four windows — 1h, 24h, 7d, 30d — and every part of it follows the choice: the uptime figure, the check counts, the mean response time, the latency chart, the per-location breakdown and the bar itself. The endpoint accepts any window in [5m, 30d]; the console picks four rather than exposing a free-text duration, because the value goes straight into a request that answers 400 for anything out of range.

The choice lives in the URL, as ?window=, not in component state. A reload, a bookmark and a link pasted into an incident channel all land on the period the operator was actually reading — which is the whole reason it is a URL parameter, since "look at this 7d bar" is not a shareable claim if the link opens on 24h. Anything the parameter names that is not one of the four — absent, empty, hand-edited, truncated, or left over from a future set of presets — resolves to 24h and is never forwarded: a panel that cannot load, plus a 400 in the log, is a worse answer than the window nobody asked for.

The panel still labels itself from the response's echoed window_seconds, never from the preset that was clicked. The two are normally the same, and keeping the label on the answer rather than the request is what guarantees the figure and the words describing it cannot drift apart. It follows that a panel with no answer yet carries no window words at all rather than the most likely one: with four selectable windows, a caption assumed while the request is in flight would be wrong three times out of four.

Segment count follows the container​

?buckets= is how many segments the bar aims for, bounded to [10, 200], default 48; like window, an out-of-range or unparseable value is a 400 rather than a silent clamp. It exists because a fixed 48 renders one width on a phone and a very different one on a wide desktop — 5.7px versus 26.7px per segment on the two real screens this was measured against — cramped in one place and wasting the space it had in the other.

The console measures its own container rather than picking a number: the bar's first render asks for the default (there is nothing to measure before something has drawn), reports back the width its segments actually have once they exist, and re-queries with the corrected count if that changes the answer — a resize that rounds to the same segment count changes nothing. bucket_seconds is still floored at the monitor's own check interval on top of whatever count was requested — a bucket finer than the interval is empty by construction, and a bar therefore often carries fewer segments than either the default or the measured target, which is the honest shape: fewer, wider segments say less, and everything they say is true.

Two things the bar deliberately does not encode:

  • Why a check failed. error_class is stored on the result and on monitor_location_state, but it is not a label on probe_success — adding one there would fork the existing series and flatline every current query and alert. Coloring by failure reason needs a separate counter, and it could never retro-fill the history that already exists.
  • pending. Uptime Kuma's yellow means "failing but under max_retries", which is a monitor state. This bar's amber means "some checks in this bucket failed", which is a measurement. They are different claims and only the second one can be made per location.

Monitoring the pops themselves​

This is the part to read before you rely on a pop.

A prober ships no metrics and no logs. Both subjects are tenant-scoped and a probe agent holds no tenant grant, so nothing about the pop's own health leaves the pop host — not its buffer evictions, not its publish failures, not its check errors beyond what a result carries. Console sees a pop only through what arrives: assignment polls and results.

The consequence is blunt: the backend cannot distinguish "the pop is not probing" from "the pop's results are not arriving." Both look like silence, and silence is also what a pop with nothing assigned looks like. What you can watch from the backend side:

SignalWhat a change means
proxima_probe_results_total{location} falling to zero while other locations keep incrementingThat pop stopped producing, or its results stopped arriving.
every location going quiet at onceThe ingest path, not the fleet.
agents.last_seen_at going staleThe pop is not even polling — it is down, or it cannot reach NATS at all.
probe_locations.last_result_at going stale while agents.last_seen_at is freshThe pop is alive and polling but nothing it probes succeeds in reaching the wire — blocked egress, a wedged scheduler. It keeps its credentials and loses its quorum seat.
proxima_probe_results_rejected_total{reason="unassigned"} non-zeroA security signal, not a tuning one: a pop reported on a monitor that was not assigned to it.
proxima_probe_partial_failure_total{location} climbing for one locationThat pop or its transit is unhealthy, not the targets.

To see inside a pop, log in to its host and read the agent's own journal. Fixing that — a self-monitoring channel a prober is actually allowed to use — is deliberately deferred. What is no longer deferred is turning a silent pop into a signal: that is the fleet auditor below.

The fleet auditor — making silent vote-loss loud​

ProbeLocationAuditor sweeps the registry (advisory-lock singleton, so a finding is counted once per fleet rather than once per replica) and reports three things:

CounterWhat it means
proxima_probe_prober_stale_total{location}One prober machine reported nothing inside the audit window. Its siblings may still be keeping the location's vote alive — which is correct, and completely invisible without this.
proxima_probe_location_dark_total{location}No enabled prober at that location reported. The vantage point has left every quorum denominator: effective_min falls, and the monitors it watched begin confirming on fewer locations than their operator asked for.
proxima_probe_location_disagreement_total{location}Probers sharing one code disagree on kind or client_id, so that code is eligible for no monitor in any tenant until it is fixed.

Plus its own proof of life — proxima_probe_location_auditor_sweeps_total{result} and proxima_probe_location_auditor_last_success_timestamp_seconds. The three counters only ever go up, so on their own they cannot tell "healthy fleet" from "auditor blind since Tuesday". Alert on max() of the timestamp across replicas: only the replica that won the lock stamps it, so min() reads as a permanently dead auditor.

A location drained on purpose — every row enabled = false — is never counted as dark. Deliberate is not dark.

Two known limits, both loud rather than silent

It fires early above a ~270-second monitor interval. The audit window is derived from the quorum's own seat window (LivenessWindow(0) = 5 minutes) so the auditor can never be quieter than the denominator it reports on. But LivenessWindow scales with each monitor's interval and a single fleet-wide window cannot, so above roughly 270 seconds the auditor leads the denominator: a pop legitimately reporting once per long interval reads as stale every sweep while it is still seated and still healthy. Scaling per location needs the intervals of the monitors assigned to it — a read the auditor does not make — so per-location scaling is a follow-up.

A location with zero assignments reads as permanently dark. An idle pop probes nothing, so it produces no results, so last_result_at never advances. That is genuinely indistinguishable from a dead one on the signal the auditor has, and erring loud is the right direction — but on a fleet with unassigned pops, expect a standing location_dark line that is telling you about your selectors, not your pops.

The independence guards​

Quorum's whole premise is that location failures are independent: several vantage points agreeing is evidence about the target precisely because they do not share a failure mode. ProbeSanityGuard is what watches that premise, in two tiers. Neither tier is a singleton — its output is in-process state the result worker in the same replica reads on every result, so every replica must reach its own verdict. Alert on max() across replicas, never min().

Tier 1 — a broken location​

Detects: a location failing most of what it probes inside the window. A pop that has lost its own transit fails nearly everything at once, which looks exactly like hundreds of simultaneous customer outages.

Does: suppresses that location's votes. It loses its denominator seat and its ballot together, so every monitor it watched is now judged on fewer independent vantage points — a reduction in the evidence, which is why it is a gauge (proxima_probe_location_suppressed{location}, 1 while suppressed and 0 once counting again) rather than a log line.

Declines to act, visibly, in two cases, both counted by proxima_probe_sanity_withheld_total{location,reason}:

  • single_tenant_failure — the location met the failure bar but every failing monitor belongs to one tenant. That is a customer outage, not broken transit, and suppressing on it would take a shared pop away from every other tenant on the strength of one client's bad day. This is the guard working, and the counter records the false page it prevented.
  • single_tenant_fleet — every monitor that pop probes belongs to one tenant, so the transit-versus-customer question cannot be answered there at all and that pop can never be suppressed, however broken it is. A private pop is this shape by construction. This is not "the pop is fine", and must never be read as it.

Monitors in maintenance are left out of the failure rates (FailureRatesByLocation), so a client's planned outage cannot look like a broken location. That only shrinks the sample — the guard's documented safe direction — and a dip in the sample during a large window is expected.

Tier 2 — a broken fleet​

Detects: most locations failing at once across unrelated monitors. Tier 1 cannot see this failure, because it is the one that breaks the premise hardest: a bug shipped to every prober fails all locations together, so each one looks individually broken, every monitor confirms down on what reads as unanimous agreement between independent vantage points, and no surviving location is left to disagree.

Does: holds every synthetic page, for every tenant at once — proxima_probe_fleet_suppressed goes to 1.

It suppresses emission, never state. Monitors still transition, monitor_events rows are still written, the availability history is still true. Only the wake-someone-up action is held. A guard that froze state to protect the pager would corrupt the record an operator reads during the incident review.

Its numerator is tier 1's own verdict rather than a re-derivation, so the two tiers cannot contradict each other about a row they both looked at in the same sweep. A location tier 1 judged single_tenant_failure is not counted toward halting every tenant's pages — otherwise one client who dominates most pops would take the whole product's paging down when their datacentre went dark. A single_tenant_fleet location is counted, because undecidable is not exonerated and a private-pop fleet must stay trippable.

Tier 2 cannot tell a fleet bug from a genuine internet-scale outage

Both look identical: everything failing everywhere. Tier 2 declines to page in both cases, which is the wrong answer for the second one — and is still the right trade.

A real internet-scale outage reaches you through every other channel there is: your customers, your own dashboards, the news. A prober bug reaches you as every client's monitors paging at 3am for nothing, and it destroys trust in the pager permanently the first time it happens.

So tier 2 is deliberately biased toward the failure that is recoverable. When it trips, the state and the history are still true — read them.

Both tiers can be completely inert, and say so​

Both tiers refuse to judge a sample too small to be evidence. Both refusals are correct. Both also mean a safety guard can be entirely unable to fire while looking configured and healthy — and on a thin pilot fleet that is the likely state, not an edge case. proxima_probe_guard_inert{tier} is 1 while a tier cannot fire whatever happens, and the routes into it are:

TierInert when
locationNo location the guard can see has a verdict for at least MinSample monitors (default 5), or PROXIMA_PROBE_LOCATION_THRESHOLD is set above 1.0, which is the documented way to switch the tier off.
fleetThe fleet has fewer than FleetMinLocations locations (default 3), or PROXIMA_PROBE_FLEET_THRESHOLD is set above 1.0.

Two shapes on the current pilot fleet are worth naming outright, because both look like a working guard:

  • Tier 1 is two gates deep. A location must reach MinSample and its failures must span at least two tenants. On a thin or effectively single-tenant fleet it may never fire at all — the inertness gauge covers the first gate and proxima_probe_sanity_withheld_total covers the second, so between them the reason is always visible. Neither one alone tells you.
  • Tier 2 on a 3-pop fleet needs 3 of 3. With the default FleetThreshold of 0.75, two of three failing is 0.667 and does not trip. So the tier has exactly one trippable state — unanimity — and at two pops it is inert entirely, because the fleet is below FleetMinLocations.
Nothing in this repository alerts on proxima_probe_guard_inert. Add the alert.

The signal exists and is exported; no rule consumes it. Without one, an inert guard is indistinguishable from a healthy one on every dashboard, which is the exact confusion the gauge was built to end.

# A sanity-guard tier that cannot fire at all. Warning, not critical:
# on a thin fleet it is expected — it just must never be a surprise.
max by (tier) (proxima_probe_guard_inert) == 1

Pair it with the guard's proof of life, which answers a different question (blind because dead, rather than idle because unjudgeable):

# The guard has stopped sweeping in some replica. min(), not max():
# one replica with a frozen suppression set judges every result it handles on stale verdicts.
time() - min(proxima_probe_sanity_last_success_timestamp_seconds) > 600

Tuning​

Every threshold is an explicitly labelled guess. There are three pilot pops and no incident history, so nothing here has been fitted to anything; the shadow window is where the real numbers come from. All six are environment knobs for exactly that reason — a guess that needs a deploy to correct is a guess you are stuck with during the incident that proves it wrong.

VariableDefaultMeaning
PROXIMA_PROBE_SANITY_INTERVAL1mHow often the guard re-judges the fleet. Detection latency is the window plus up to one interval.
PROXIMA_PROBE_SANITY_WINDOW5mHow far back failure rates are drawn from. Derived from the quorum's own liveness floor so the guard and the quorum cannot disagree about which results describe the present; overriding it breaks that derivation, which the guard warns about at startup rather than refusing.
PROXIMA_PROBE_LOCATION_THRESHOLD0.8Fraction of a location's monitors failing before tier 1 judges its transit broken. Compared with >=.
PROXIMA_PROBE_LOCATION_MIN_SAMPLE5Monitors a location must have a verdict for before tier 1 can judge it at all. A pop with two assignments, both failing, is 100% and means nothing.
PROXIMA_PROBE_FLEET_THRESHOLD0.75Fraction of the fleet's locations failing at once before tier 2 holds every page. A 4-pop fleet trips at 3; a 3-pop fleet trips only at 3.
PROXIMA_PROBE_FLEET_MIN_LOCATIONS3Locations the fleet must have before "most locations are failing" is worth acting on.

Zero is not "off". Every knob is read through a positive-only loader: a non-positive value is refused, recorded as rejected, warned about, and the default stands — so an operator who explicitly disabled a check cannot have it silently re-enabled behind their back. To disable a tier, set its threshold above 1.0: a fraction of a whole can never reach it, the intent is legible in the environment, and the value survives to the settings page instead of reading as "unset".

Paging​

A confirmed transition publishes to proxima.alerts.ingest — the same envelope an Alertmanager webhook produces, deliberately reusing the consumer's own type so producer and consumer cannot drift. A probe alert is an ordinary alert whose source happens to be us.

Which transitions reach the pager​

TransitionEmission
-> downFires.
anything leaving down (down -> up, down -> paused, down -> maintenance)Resolves. down -> paused resolves too, because a paused monitor is never probed again, so no later transition could ever close its group. down -> maintenance resolves for the same shape of reason: the open "down" has to close so that a still-broken confirmation after the window is announced — and pages — against a clean slate. monitor_events keeps the truth (down -> maintenance, never down -> up).
-> pendingNothing. pending is the notification-suppression state; paging there would make max_retries meaningless.
everything elseNothing.

The grouping key, and why a resolve always finds its group​

The group identity is proxima_probe:<monitor_id> — derived from the monitor's immutable UUID. Repeated downs update one alert group rather than spawning many, and a recovery publishes the identical key, which closes exactly the group the down opened.

There is no scheme left to drift and nothing an operator can edit — name, target, severity, the paging flag — that can re-key an open incident. (monitor_state.alert_group_id is not the mechanism and is written by nothing: the group's id is minted by Postgres inside the alert worker after the publish, and the ingest payload carries no group-id field at all. The column is retained with a comment saying so.)

The resolve is per-fingerprint, the Alertmanager shape — a probe group has exactly one member alert with a stable fingerprint, so the exact member is closed by name rather than by a group-level recovery flag.

The gate chain​

The gate is a chain, not a flag, and the order is the order an operator would ask the questions in. Each half reports under its own counter, so "why didn't this page?" is answered by a label value rather than by inference.

  1. monitors.paging_enabled — is this monitor supposed to page at all?
  2. A maintenance window covers the monitor (in_maintenance) — asked after the monitor's own flag, so the shipped default keeps reading paging_disabled, and before tier 2, because planned work is the likelier explanation. The same gate holds the Telegram notify path.
  3. ProbeSanityGuard tier 2 — is the fleet currently trustworthy enough to page from?
MetricMeaning
proxima_probe_would_page_total{kind}Counted before the chain, always. Every confirmed down-transition that would have woken somebody. This is the shadow window's entire output and what turns tier 2's threshold from a guess into evidence.
proxima_probe_paged_total{kind}The publish the broker accepted.
proxima_probe_page_suppressed_total{kind,reason}The chain refused. reason is paging_disabled, in_maintenance or fleet_suppressed.
The gate holds the FIRING half only. A resolve is never held.

Both gates exist to stop somebody being woken up, and a resolve can only ever close something — so holding one protects nobody and costs a great deal.

A held resolve is not a stale row an operator can tidy up later. The find-or-create path matches status <> 'resolved', so the monitor's next outage merges into the group that was never closed; the firing epoch bumps only on a reopen, which a group that never resolved can never have; and escalation arms only while the escalation snapshot is null, which the first outage already filled. The next genuine outage would therefore page nobody, while paging_enabled was on and the operator believed they were covered. Tier 2 tripping mid-outage is not an edge case — it is the thing tier 2 exists to do — so that silence would be routine.

A non-zero proxima_alert_resolve_no_open_group_total{source_type="proxima_probe"} is EXPECTED, not a fault

Because resolves are deliberately ungated, every recovery publishes while paging is off. With no firing alert in the payload the ingest worker takes the find-only path, drops the delivery, and ticks that counter. During the shadow window it should be roughly one tick per monitor recovery, in its own source_type series, polluting nobody else's signal.

Read it as "the shadow window is working". It becomes worth investigating only once paging_enabled is on for a monitor and its resolves still find no group.

Monitor groups page as one thing​

Monitor groups has the model — what a group is, how its state is folded from its members, and why paging authority is exclusive. This section is the emission half: which group transitions reach the pager, what the page carries, and what re-derives a group's state in the first place.

Nothing about a group reaches a human by default

Every gate on a group page defaults off — paging_enabled on the group row, tier 2, a matching escalation route, oncall_enabled, a live page-owner team. See the rollout note under Turning paging on.

It passes the same gate chain, and that is not a resemblance — it is the same function. paging_enabled (on the group row) then tier 2, in that order, with resolves ungated. A group page is the loudest possible expression of a fleet-wide prober bug — every member of every group fails at once and the derived page says "everything in Payments is down" — which is exactly the case tier 2 exists for, so it is checked again here rather than assumed from the member states.

TransitionEmission
-> downFires. From up, pending or unknown alike.
anything leaving downResolves — including down -> unknown, which is what a group reaches when every member has been paused or removed. Nothing will ever be measured for it again, so no later transition could close its alert; leaving it firing would page indefinitely about a service Proxima has stopped watching.
-> pendingNothing. The notification-suppression state, as for a monitor.
everything elseNothing.

down -> pending needs two members to reach at all — a monitor never goes down -> pending itself, so "the only failing member starts retrying" cannot produce it. The real route is A down + B pending (group down), then A recovers and the group derives pending from B. Resolving there can flap: a resolve, then a fresh page if B fails too. That is the accepted cost, because holding the alert open through pending is worse twice over — a later recovery to up is no longer an edge leaving down, so it closes nothing and the group stays open forever; and while it stays open its escalation snapshot is non-null, so the next pending -> down re-fires into a group that arms no escalation. A flap is noise; that is silence.

Alert identity

  • Grouping key proxima_probe_group:<group_id>, from the group's immutable UUID — so no edit an operator can make (the name, the membership, the severity, the paging flag itself) can re-key an open incident. It cannot collide with a member's proxima_probe:<monitor_id>.
  • alertname is ProximaMonitorGroupDown, deliberately distinct from ProbeMonitorDown: a group page and a member page are different events with different remedies, and an operator filtering or silencing one must not catch the other.
  • The resolve is per-fingerprint, never the group-level flag — same reason as for a monitor: the group-resolve shape pairs with an empty alert list, and an empty list rewrites the group's severity to unknown at the instant it resolves.
  • Severity is the group's own, defaulting to P2 as a monitor's does. A group page is bigger, but defaulting to P1 would invent urgency the operator did not author.
  • The page publishes through the same proxima_probe alert source row as its members. A group is not a new source type.

What the payload carries

The group's environment_id is carried outright (the column is NOT NULL), so an environment-scoped escalation route matches without the client authoring any label mapping. There is no host_id — a group spans hosts by construction, so there is no id to carry, and a host-scoped route should not match a page about a service.

The summary names the failing members: "one page instead of twelve" is only worth having if the one page still says which twelve endpoints are failing. The list is bounded at ten names with the remainder counted (and 7 more) — truncating here, where the remainder is visible, beats truncating downstream in a paging channel, where it happens silently.

When a group's state is re-derived

A group holds no opinion of its own: its state is folded from its members every time something re-derives it, and nothing else ever writes it. So every write that can change the fold has to ask for a re-derivation, and the ones that are easy to forget are the ones no probe result will ever follow:

What happenedWhat re-derives the group
A member's probe transition is confirmedThe result worker, after that member's own page
A member is paused or resumedThe pause/resume endpoint — a paused member is excluded from the fold, and a paused monitor is never probed again
A member is deletedThe delete endpoint — the group it left; nothing in the database still connects the two
A member moves between groupsThe update endpoint, for both groups — the one it left is the one that gets forgotten, and it can sit down because of a member that is no longer in it
A group is created or editedThe group endpoints — neither changes the fold (membership is the monitor's own group_id), but evaluation is edge-triggered, so an edit is a free moment to make the state beside the form real
A group is deletedNo re-derivation — its state row and timeline cascaded and its members were detached, so there is no group left to derive. If it was down at the time, its resolve is published before the delete by ResolveOnDelete, since nothing afterward could ever find its way back to a row that no longer exists — see the grouping key

Six properties of those calls are deliberate:

  • The group evaluation runs after the member's own page, never before. Deriving a group reads its whole membership under a per-group lock, so ahead of the page it would put a lock wait between an operator and the alert about their own monitor.
  • A failed evaluation never fails the thing that triggered it. The probe result is already recorded and its verdict already durable, and the API write has already been committed and answered; failing either would trade a stale group state for a redelivery storm or for a false error about a change that succeeded. The stale state is visible and self-correcting — the next member transition re-derives it — and the drift auditor exists for the case where no next transition comes.
  • Both evaluations are bounded, and the bound is a liveness guard rather than a nicety: the per-group lock is a Postgres advisory lock that waits indefinitely, and the probe-result consumer dispatches from a single goroutine — so an unbounded wait there would not make one replica slow at ingesting results, it would stop it ingesting them at all, for every tenant, until the process restarted. The API path gets 15 seconds (a caller is waiting) and the ingest path 60 (nothing is).
  • The API path flushes the response before it evaluates. Without that, "after the response" is true of the code and false of the client: the JSON sits in the server's write buffer until the handler returns, so the caller would wait out the whole evaluation.
  • Tier 2 gates the API path too. The fleet verdict is in-process state built asynchronously when the backend connects to NATS — after the router exists — so the API's evaluator holds a write-once handle to it rather than the guard itself, and reads the verdict lazily at emit time. An unset handle reports "nothing suppressed", which is fail-open on purpose: a wiring mistake that silenced every page in the fleet would be worse than the fleet-wide prober bug tier 2 exists for, and it would be invisible.
  • The ungrouped majority pays nothing. A monitor with no group returns before any lock is taken and before the store is touched, which matters because every probe result in the fleet reaches that check.
MetricLabelsMeaning
proxima_monitor_group_would_page_total—Counted before the chain. The shadow window's output for groups; against proxima_probe_would_page_total it is what says how many member pages one group page replaced.
proxima_monitor_group_paged_total—The publish the broker accepted.
proxima_monitor_group_page_suppressed_totalreasonpaging_disabled or fleet_suppressed — the same two the monitor chain reports.
proxima_monitor_group_transitions_totalto_stateApplied transitions, including the edges that page nobody.

None of them carries a group_id, for the reason none of the probe counters carries a monitor_id, and it bites harder: groups are cheap to create. Which group did or did not page is in the log line and on the alert.

Authoring a group​

Six endpoints, mounted twice — at the top level with ?client_id=, and nested under /api/v1/clients/{id}/monitor-groups — exactly as the monitor routes are. The path parameter is {groupID} and never {id}, because under the nested mount id is already the client.

MethodPathNotes
GET/api/v1/monitor-groups?client_id=…The client's services with their derived state and counts. client_id is required; this endpoint never lists across tenants. Not paginated — a client has tens of services, not thousands.
POST/api/v1/monitor-groups201. The group is created empty and starts at unknown — there is no members field; fill it from the monitor routes (Joining and leaving a group).
GET/api/v1/monitor-groups/{groupID}The group, its state, and its members with each member's own state and enabled flag.
PATCH/api/v1/monitor-groups/{groupID}Partial: environment_id, name, severity, paging_enabled. There is deliberately no client_id — a group cannot be re-homed to another tenant.
DELETE/api/v1/monitor-groups/{groupID}204.
GET/api/v1/monitor-groups/{groupID}/eventsThe service's incident timeline, newest first, limit capped at 200.

Every route is gated on the monitor permissions — monitors:read, monitors:write, and monitors:delete for the delete — at the route and again in the handler against the loaded row's client. A group is not authority an operator holds separately from the monitors it folds; a second permission would let a role edit services it cannot see inside.

Four behaviours are worth knowing before you use them:

  • The name is trimmed before it is stored. UNIQUE (client_id, name) would otherwise make Payments and Payments two different services with one name between them, and the duplicate row would come with no explanation.

  • paging_enabled is exclusive with its members'. Turning it on while any member holds its own is a 400 whose body names those members, in the message and again in error.details.members with their ids. A member named there may be paused — pausing does not surrender paging authority, and a resume would restore the double-paging the rule prevents. Every other edit, a rename included, is never refused by this rule: the check reads the group as it will be after the write, so a group whose paging is off asks nothing. Turning paging off is likewise never refused — it is the remediation the refusal recommends.

  • Moving a populated group's environment_id is refused, not cascaded. The members' composite foreign key names the group's environment, so the move would orphan them; Postgres declines it and the API answers 400 telling you to move or detach the members first. An empty group may be re-homed — that is the legitimate case, fixing a group authored against the wrong environment before any monitor joined it. Renaming a populated group is unaffected.

  • Deleting a group detaches its members; it never deletes them. Each monitor's group_id goes null and each keeps its own paging_enabled exactly where it was — which is off — so leaving a group can never silently grant a monitor the authority the group held. A service that was paging as one thing therefore stops paging at all until its monitors are given their own flag or joined to another group.

    Deleting a group that is down publishes its resolve first, because nothing afterward could: the grouping key is derived from a UUID that stops belonging to anything the instant the row is gone. DeleteMonitorGroup calls ResolveOnDelete before the row goes, so the alert closes rather than firing forever with no key left to close it under.

current_state on the list and state on the detail are both nullable, and the null is not the unknown state: unknown means the group was evaluated and nothing was measured, while null means there is no monitor_group_state row at all — a broken invariant, since the row is seeded in the same transaction as the group. The two must render differently.

A group create or edit re-derives the group after the write commits, never inside a transaction with it. That is not a stylistic preference: the evaluation takes a per-group advisory lock on its own pooled connection and does its store work on others, so a caller holding a transaction across it holds a connection outside the evaluator's slot cap, and enough of those exhaust the pool while the capped holders wait for connections that will never free. A delete re-derives nothing — the state row and timeline cascaded, the members are detached, and there is no longer a group to derive.

Joining and leaving a group​

Membership is the monitor's own group_id, and it moves through the monitor routes. There is no join table and no group-side write: a group is created empty, and POST /api/v1/monitors and PATCH /api/v1/monitors/{monitorID} are what put a monitor in one or take it out.

# Join, or move between services.
curl -X PATCH .../api/v1/monitors/$MONITOR -d '{"group_id":"'$GROUP'"}'
# Leave. An explicit null is the ONLY way out; omitting the key changes nothing.
curl -X PATCH .../api/v1/monitors/$MONITOR -d '{"group_id":null}'

Five rules, and each of them is a refusal you will meet rather than a note:

  • group_id is three-state on update, like host_id/asset_id/cluster_id — and here the distinction is load-bearing rather than a convenience. A uuid joins or moves, an explicit null leaves, and an omitted key leaves membership exactly where it was. The update path replaces the whole row, so a client that sent membership only when it changed — or a handler that read it from the body alone — would silently detach a monitor from its service on every rename, pause or threshold edit, with a 200 and no trace.
  • A monitor may only belong to a group in its own client and its own environment. That is monitors_group_scope_fkey, a composite key over (group_id, client_id, environment_id), so it also holds for bulk updates and hand-run SQL. The API proves it before the write and answers 400 naming the group; a group in another tenant gets the same "group_id does not exist" a non-existent one does, so the endpoint cannot be used to probe which group ids exist elsewhere.
  • Re-homing a grouped monitor to another environment is refused, not cascaded. The monitor's environment is half of that composite key, so the move would orphan the membership. The 400 says to detach it from the group first — or send {"environment_id": …, "group_id": null} in one request, which is the way through. An ungrouped monitor moves freely.
  • paging_enabled is exclusive with the group's. Turning a member's own flag on while its group holds paging_enabled is a 400 naming the group; moving a monitor that holds its own into a group that holds it is a 400 naming both sides, with the monitor in error.details.members. Turning the monitor's flag off is never refused — it is the remediation those refusals recommend, and a check that read the stored row instead of the post-edit one would refuse exactly that, trapping the operator between an error and its own fix.
  • Leaving a paging group leaves the monitor's own paging_enabled off. Authority is only ever authored, never inherited — the same rule the ON DELETE SET NULL behind a group delete obeys. A service that paged as one thing therefore stops paging at all until somebody says otherwise, which is the honest outcome: the alternative is a page for one endpoint in the name of a service nobody re-authored.

A monitor create or update re-derives every group it touched, after the write commits — and a move touches two. The group a monitor left is the one that gets forgotten, and it can sit down because of a member that is no longer in it, forever: no probe result for the moved monitor will ever re-derive a group it is not part of.

The drift auditor — the backstop for a mutation path that forgets​

Edge-triggered evaluation has exactly one failure mode, and the table above is the list of places it can happen: a write that changes the fold and does not ask for a re-derivation. That failure is completely silent. The group keeps its last derived state forever, the UI renders it with the same confidence as a fresh one, no error is raised, nothing is retried, and no log line is written anywhere — from the code's point of view nothing happened, because nothing did. The page that should have gone out is simply never sent, and the first anybody hears of it is a customer.

MonitorGroupAuditor re-runs the real evaluation for every group in the fleet every five minutes and counts what moved. Advisory-lock singleton, so a finding is counted once per fleet and the fleet is re-evaluated by one replica rather than by all of them at the same instant.

CounterWhat it means
proxima_monitor_group_drift_total{correction="state"}A group's derived state was flatly wrong. A page or a resolve was missed, and the sweep has just now sent it, late.
proxima_monitor_group_drift_total{correction="counts"}The state was right and the evidence beside it was stale — the outage paged, but the operator was shown the old number of failing members.
proxima_monitor_group_authority_conflicts_totalA group holds paging_enabled beside a member that holds it too. See below.

Plus its own proof of life — proxima_monitor_group_auditor_sweeps_total{result} (clean, partial, error) and proxima_monitor_group_auditor_last_success_timestamp_seconds. Alert on max() of the timestamp across replicas, the opposite of the sanity guard's rule: only the replica that won the lock stamps it, so min() reads as a permanently dead auditor.

Four rules cover it, and none is redundant: sustained drift, an authority conflict, sustained partial sweeps, and the pulse going stale. partial needs its own rule precisely because it stamps the pulse — the auditor did run — so it is invisible to the stalled rule forever; error needs none, because it withholds the pulse and the stalled rule catches it. Three of the four use increase(...[1h]) with a threshold above zero rather than rate(...) > 0, and the reason is worth carrying to any new rule in this family: on a 1h lookback a single increment holds a rate above zero for a full hour, so rate > 0 with any shorter for: fires on one occurrence and is a level check wearing a rate's clothes.

The auditor sweeps once at startup, before its first tick. Otherwise a fresh replica leaves the liveness gauge absent for a full interval, and the rule watching that gauge is the only thing between a deleted line at the composition root and complete silence. It is safe under a rollout because the startup sweep takes the same fleet-wide advisory lock: every new replica asks, one wins, the rest return without touching the database — one extra sweep per deploy, not one per pod.

Six things about it are deliberate and worth knowing before changing it:

  • It corrects through MonitorGroupEvaluator.Evaluate and never writes monitor_group_state itself. The per-group advisory lock spans the whole derive-and-write because the read is what goes stale; a sweep with its own write path would re-open exactly the interleaving that lock exists to close — a stale verdict landing last, closing a live outage and publishing a resolve for an incident that is still happening — and it would do so on a timer, across every group in the fleet, with nobody watching.
  • A corrected drift that is not counted looks exactly like a system with no drift. The correction is the cheap part. The counter is the part that gets the missing evaluation call written, which is why a sustained non-zero rate is a bug to fix at the mutation path and not a threshold to widen.
  • A correction is a page, so a backend whose group evaluator has no emitter refuses to start. That wiring failure is the one nothing else in the family can see: an auditor whose evaluator lost its alert publisher repairs the state row, ticks drift_total, sweeps clean and stamps its pulse, so every rule reads healthy while the page it just discovered is quietly eaten. buildWorkers therefore refuses outright — the condition is unreachable from any configuration, so it fires on a source edit and nothing else. Read what a refusal actually does before relying on it: it is not a crash. The HTTP server keeps serving and /readyz does not check worker startup, so the pod stays Ready and a rollout promotes it — with no NATS-driven workers at all. That is a worse outage than the one being prevented, which is why the check sits above every worker start and must stay unreachable from configuration.
  • The auditor also reports it, and that report is unreachable today. Every start logs emitter_wired=yes|no|unknown, with no repeated as a WARN (monitor group auditor has no emitter). In this deployment the boot refusal above stands in front of the auditor's only construction site, so no cannot happen — it is kept for a second construction site that does not exist yet and would otherwise arrive with no check at all, not as a live second line of defense. The field is named emitter_wired and not can_page deliberately: all it can see is that there is somewhere to publish. Whether a page reaches a human is the paging_enabled/tier-2/route/oncall_enabled/team-live chain, none of which is visible from there and all of which default to off.
  • Each group's re-evaluation is bounded at 60 seconds — per group, not per sweep — and one group's failure never abandons the rest. pg_advisory_xact_lock waits forever, so an unbounded call would not make the sweep slow, it would end it, and every group ordered after the stalled one would stop being audited with nothing but a flat liveness gauge to say so. One budget for the whole sweep has the same shape at a larger scale: everything after the budget expires fails, under exactly the conditions (a fleet-wide outage queueing evaluations) that argued for the looser bound in the first place.
  • A group whose state cannot be READ is still re-evaluated. Only the reporting is lost — the correction, and the page riding on it, is the backstop the worker exists for, and a transient read error must not cost it. The sweep says so (partial) rather than silently reporting clean for a group it could not audit.

partial exists as a third sweep result because "the auditor is alive" and "something in the fleet is broken" are different questions: a sweep that re-evaluated every group but one still stamps the liveness gauge, so a single permanently broken group cannot make a live auditor look dead forever.

The drift counter has a floor, and it errs loud

A member that transitions between the auditor's before-read and its re-evaluation is picked up by that re-evaluation and counted as drift — even though the edge-triggered path was about to handle it, and (being idempotent) handled it correctly. The window is milliseconds per group, so this is rare rather than routine, but it is not zero. That is why the alert is increase(...[1h]) > 2 — a threshold above the floor — and not rate(...[1h]) > 0: on a 1h lookback a single increment keeps a rate above zero for a full hour, so rate > 0 with any shorter for: is a level check wearing a rate's clothes. Every correction, including a single one, still reaches the log; only the paging threshold sits above the floor.

Closing the window entirely would mean holding the group's lock across the whole audit — making the auditor's read the authoritative one, the exact thing it must not do.

The conflict nothing else can see: a group and a member both paging​

A group holding paging_enabled while one of its members holds it too double-pages every outage of that service, indefinitely. The auditor counts it as proxima_monitor_group_authority_conflicts_total, and it is the only thing anywhere that reports it. That is not redundancy with the API's refusal — it covers what that refusal structurally cannot:

  • Two concurrent PATCHes. Exclusivity is enforced by read-then-write across two independent requests — enable on the group, enable on a member — so each reads a world in which it is allowed, each passes its check, and both commit. Neither request did anything wrong.
  • Any path a handler never sees. Migration 000195 put group membership scope in the schema precisely because "a constraint also holds for bulk updates and hand-run SQL". Paging exclusivity has no such constraint, so one UPDATE monitors SET paging_enabled = true creates a conflicting pair silently.
  • A paused member. PagingMembers deliberately counts a member that is paused, because pausing does not surrender paging authority and a resume is one click away; the auditor's query counts it for the same reason, so the two ends of the rule cannot disagree.

And nothing re-checks at emission either, by design: "nothing is suppressed at runtime, so there is no new silence mechanism that can fail" — the group emitter reads the group's own flag and the member emitter reads the member's, and neither consults the other. That is the right call (a page must never depend on a row nobody paging is looking at) and it is precisely why the conflict has to be counted somewhere.

The auditor counts it and does not repair it. Repairing would mean deciding which side loses its page, and choosing wrong silences the only alert somebody is relying on. So unlike drift, a non-zero value here never self-heals: it ticks once per conflicting group on every sweep until a human turns off one of the two flags. The counter carries no group id; the ERROR log monitor group and member both hold paging authority names group_id, client_id, the group name and how many members are involved.

The alert source row​

source_type='proxima_probe' is carried the way every producer carries it — on the alert_sources row the envelope's source_id names — so each client needs one. It is created on demand by the result worker's first emission, not seeded, because clients are created continuously and a seed would only cover the ones that existed on deploy day.

The row is deliberately not authorable, not deletable and not editable through the API. Deleting it would stop every monitor in that tenant from paging until the next emission recreated it; re-pointing its source_type would break emission for that client permanently, because the reserved token would be held by a row the create path can no longer match.

The probe source will read "Never delivered" in the Alert Sources UI, forever

Per-source ingest health is written by RecordIngest, whose only caller is the alert webhook handler. Probe emission never touches that handler, so last_delivery_at, last_accepted_at and the per-source counters stay null and zero on this row however many pages it has carried.

A second surprise follows from it: the row appears on its client's first monitor recovery even with paging off, because resolves are ungated. So the first sign a client has probe monitors at all may be a permanently-"Never delivered" source row nobody created. Use proxima_probe_paged_total and the ingest-side counters for probe health, not this page's status column.

Turning paging on​

paging_enabled ships false on every monitor. The switch is in the monitor form (gated on monitors:write for that client) and the API accepts it in a create or update body.

At least five things must be true before the flip, and no deploy performs any of them.

Two of the five have nothing to do with monitors: a probe alert is an ordinary alert, so it passes through the same gates every alert does, and each of those gates ships off. They are listed here because an operator can satisfy the three monitor-specific ones, see nothing wrong anywhere, and still be paged by nothing at all.

Four of the five have one endpoint that answers them together

GET /api/v1/oncall/readiness/coverage returns, per client the caller can see: matchable_route_count (the worker's predicate, not the management route list, which counts disabled routes too — and only routes that name an escalation policy, since a notify-only route pages nobody), has_default_route, oncall_enabled (the resolved per-client paging flag, also served under the deprecated key l1_agent with the identical value), and live_route_count (how many matchable routes point at a team that has actually been flipped live). The Readiness page under On-Call renders it.

Prefer it to the SQL below, which is here so you can see the mechanism and check a database directly. It cannot tell you about precondition 1.

1. A role must actually hold monitors:*​

New permissions ship as code; granting them is an operator action. Verified 2026-08-21: no role in production holds any monitors: permission — all 19 lack it. Until a role is granted them, nobody but a super-admin can even see the Uptime section, let alone flip a paging switch.

-- Which roles can manage monitors at all? (roles.permissions is a TEXT[])
SELECT name, permissions FROM roles
WHERE EXISTS (SELECT 1 FROM unnest(permissions) p WHERE p LIKE 'monitors:%')
ORDER BY name;

An empty result is the shipped state, not a bug. Grant the permissions through the role editor in Admin → Roles, not with an UPDATE.

2. The client must have an escalation route that matches​

Check that routes exist, then flip. Never flip and see.

Verified 2026-08-21: only 2 of 40+ clients have any escalation route at all. For every other client, turning paging_enabled on produces silence, not pages — and it will look exactly like a broken feature, discovered during the first outage it was supposed to catch.

-- Does this client have any route that could PAGE?
-- The predicate is armEscalation's own, and it is two filters: a disabled NON-default route
-- is skipped while a default catch-all is considered whether or not it is enabled (filtering
-- on `enabled` alone would HIDE a live route -- a disabled default still matches everything),
-- and a route must name an escalation policy. A notify-only route (escalation_policy_id NULL)
-- matches for notify and arms no chain, so counting it here would say "covered" about a client
-- that pages nobody.
SELECT id, position, severity, environment_id, host_id, is_default, enabled, match_labels
FROM escalation_routes
WHERE client_id = '<client id>' AND (enabled OR is_default)
AND escalation_policy_id IS NOT NULL
ORDER BY position;

Zero rows means the flip pages nobody. Fix the routing first.

The unrouted case is not silent once it happens — Console counts it as proxima_escalation_unrouted_total, and the Alert Groups page carries an unrouted-client banner. Neither helps you here: the counter only ticks once an alert has already gone nowhere, and the banner is on a different page from the one you flip the switch on. A check that takes ten seconds beforehand is the whole of your warning.

3. Entity-scoped routes match — the monitor's own environment is carried​

A probe alert arrives already scoped to its monitor's environment_id, and to its host_id when the monitor has one. Both are read from the monitor row and carried in the ingest payload, so escalation.routeMatches is handed real entity ids and an environment- or host-scoped route matches normally. No alert_label_mapping and no catch-all route are needed for this.

The distinction is worth understanding, because it is the reason this works without configuration. Most alert sources can only describe an entity: an Alertmanager rule emits env="prod" and the client's own alert_label_mapping turns that string into an environment id. Uptime monitors are a first-party source — monitors.environment_id is NOT NULL and a foreign key — so the emitter holds the authoritative id before it ever builds a label, and carries it rather than describing it.

-- Which routes a probe alert on this monitor could match. Entity-scoped routes are no
-- longer excluded: the alert carries the monitor's environment, so they are candidates.
-- escalation_policy_id is selected rather than filtered here: a NULL one is a notify-only
-- route, which matches for notify and pages nobody -- worth seeing, never worth counting as
-- paging coverage.
SELECT r.id, r.position, r.severity, r.environment_id, r.host_id, r.is_default, r.enabled,
r.escalation_policy_id, r.notify_telegram, r.notify_feed
FROM escalation_routes r
WHERE r.client_id = '<client id>' AND (r.enabled OR r.is_default)
ORDER BY r.position;

-- The monitor's own scope, which is what an entity-scoped route is matched against.
SELECT id, name, environment_id, host_id FROM monitors WHERE id = '<monitor id>';

A route scoped to a different environment than the monitor's still will not match, which is correct rather than a gap — that is the route saying it does not cover this monitor.

If you add a default catch-all, put it LAST

Not required any more, but if a client has one: MatchRoute sorts by ascending position and returns the first default it meets, without consulting any of its match fields. Nothing on the server forces a default to sort last — position is whatever the request body says. So a catch-all at a low position silently swallows every route after it, and every alert in that client, of every severity and every source, is routed to the catch-all's policy. Give it the highest position on that client, and re-read the ordered route list afterwards.

Before 2026-08-23 this section was a cut-over blocker

Probe alerts used to carry only monitor_* labels, entity resolution ran solely through per-client alert_label_mappings, and the matcher rejects any route whose environment_id or host_id is non-nil when the alert's is nil. A client whose routes were all environment-scoped therefore got silence from uptime monitors — with paging_enabled on, a route present, and nothing obviously wrong on any surface. The documented workaround was a default catch-all per tenant. If you are reading an older runbook that says to author one for this reason, you no longer need to.

4. The client's oncall_enabled feature must be on​

This one has nothing to do with monitors and is the easiest to miss, because nothing on any monitor page mentions it.

Escalation is armed from publishL1Notification, whose paging branch runs only if the oncallEnabled predicate passes, and that predicate resolves the client's oncall_enabled feature — default false for every client (LaunchDefaults, pinned by a test). With it off, the alert group is created, the alert is stored, the timeline is correct, and no escalation is ever armed. Not "no route matched" — the matcher is never reached.

This used to be called l1_agent, and the AI switch is not a precondition

oncall_enabled is one half of the retired l1_agent, which gated paging and AI investigation with a single boolean. Its other half, ai_triage, is not on this list and is not a paging precondition: a client with ai_triage off pages for a down monitor exactly as one with it on. A client whose stored features still hold only l1_agent resolves oncall_enabled from it through a legacy alias, so nothing about this precondition changed for an existing client — only its name did. See Paging Control.

It is not silent after the fact: the gate records a no-page event against the group whose reason is l1_disabled — that exact string is what lands in the group's no-page log and what labels proxima_page_suppressed_total, so it is what you grep for. The reason token deliberately keeps the retired name: historical alert_group_log rows carry it and a partial unique index is built on the literal, so renaming it would orphan both. The predicate itself writes no log line on this branch; the oncall_enabled disabled; side-effects suppressed INFO you may also see comes from the escalation timer's own dispatch-time re-check, which is a second emitter of the same l1_disabled reason. Either way it only tells you afterwards.

Read it from the coverage endpoint's oncall_enabled field (the response also still carries a deprecated l1_agent key holding the identical value, for frontends that have not migrated). Do not try to read it from /auth/me: for a super-admin that returns a single wildcard entry with every feature true, and the frontend's own helper short-circuits to true as well — both would report paging as on for a client where it is off.

5. The page-owner team must be flipped live, not in shadow​

Also nothing to do with monitors. The escalation timer suppresses every mutating side effect — page, declare-incident and resolve alike — while the frozen page-owner team is in shadow. team_cutover_state.escalation_live defaults to false, and a team with no row at all reads as shadow (fail-safe), so this is off until somebody deliberately turns it on per team.

A plan whose frozen team is nil fails safe the same way, recorded as policy_invalid rather than as shadow.

-- Which teams page for real? A team absent from this table is in shadow.
SELECT t.id, t.name, COALESCE(c.escalation_live, false) AS live, c.flipped_at
FROM teams t LEFT JOIN team_cutover_state c ON c.team_id = t.id
ORDER BY live DESC, t.name;

The coverage endpoint folds this into live_route_count: a client with matchable_route_count > 0 and live_route_count = 0 has routing that looks healthy and pages nobody.

Is that all of them?​

No, and this page will not pretend otherwise. Five is what has been traced from the emit path to the phone and verified in source; it is not a proof that nothing else can intervene. The chain continues past this page into notification chains, per-user notification policies, a verified phone number, and the voice trunk. Treat the coverage endpoint and a real end-to-end test page as the authority, and this list as the things that will bite first.

Then, in order​

  1. Confirm a role holds monitors:read/:write and the people who need it have that role.
  2. Read GET /api/v1/oncall/readiness/coverage for the client and confirm matchable_route_count > 0, oncall_enabled: true and live_route_count > 0.
  3. If the matching route is not a default, confirm it is not entity-scoped.
  4. Let the shadow window run. proxima_probe_would_page_total is accumulating against real traffic with paging_enabled off; compare it against what actually happened.
  5. Tune tier 2's threshold from that evidence — it shipped as a guess and the shadow window is the only thing that can correct it.
  6. Flip one monitor on the proxima client — one of the two with a route, and our own on-call at the far end. Then take that monitor's target down on purpose and confirm a phone actually rang. Nothing short of that closes the list in the previous section. Dogfood before any tenant.

A group's flag is its own, and the five preconditions still apply​

monitor_groups.paging_enabled ships false exactly as a monitor's does, and it is a separate flag: enabling paging on every member of a group leaves the group itself silent, and enabling it on the group does not require any member to have it.

A group page is an ordinary alert from the same first-party source, so all five preconditions above apply to it unchanged — the monitors:* grant, a matchable route, the route actually matching, oncall_enabled, and the page-owner team being live. Only one of the five reads differently: precondition 3 matches on the group's environment_id, which is NOT NULL on the group row, and a group carries no host_id at all, so a host-scoped route will never match a group page. A client whose only route is host-scoped has routing that works for member monitors and pages nobody for their group.

GET /api/v1/oncall/readiness/coverage answers the same four for a group as it does for a monitor; it knows nothing about the group flag itself.

The flags are exclusive, so turning one on can require turning others off

While a group holds paging authority, none of its members may — and vice versa. The exclusivity is enforced by refusal at authoring time, never by suppression at runtime, so enabling the group's flag while a member still holds its own is rejected with a 400 naming the other side rather than silently accepted and quietly de-duplicated later. That is deliberate: grouping must add no new silence mechanism that can fail. If a group's page is lost downstream — no matching route, oncall_enabled off, team in shadow — that is the existing, already-instrumented loss above, not a fresh one hidden behind grouping.

The shadow-window step is the same, with one extra reading: proxima_monitor_group_would_page_total against proxima_probe_would_page_total is what tells you how many member pages one group page would actually have replaced. Flip the group only after that ratio looks like what you expected.

One gate that used to be open is now closed. The API's own group evaluator was wired with a permanent nil fleet guard when group derivation first shipped — the tier-2 verdict is in-process state built when the backend connects to NATS, after the router has constructed every handler — so a group transition caused by an API write during a fleet-wide prober failure would have paged anyway. It now reaches the API through a write-once handle set the moment the guard exists, and both composition roots name it as an argument. An unset handle still reports "nothing suppressed", which is fail-open on purpose: a wiring mistake that silenced every page in the fleet would be worse than the bug tier 2 exists for, and it would be invisible.

One gate is now closed; one thing is still open

A group deleted while down used to publish no resolve, and that is fixed. The delete cascades the group's state row and timeline and detaches its members, so once the row is gone there is nothing left to derive from and no later evaluation could ever find its way back to it — the grouping key (proxima_probe_group:<group_id>) is derived from a UUID that no longer belongs to anything. DeleteMonitorGroup now calls uptime.MonitorGroupEvaluator.ResolveOnDelete before the row goes, publishing the resolve from the group's last persisted state while the key can still be minted. This is deliberately not a re-evaluation — no membership read runs and no row is written, since the row is about to be deleted regardless of what a fresh read would say — and it is a silent no-op on a group that was never down, one that has already recovered, or one already gone in a race with a concurrent delete. A resolve failure blocks the delete rather than letting the row go with nothing to close it, so a group whose resolve could not be published still exists for a retry to find. This was one of the preconditions blocking paging_enabled from being flipped true for any group; it no longer does.

And nothing here has been validated against a live server. Every behaviour on this page is pinned by unit and integration tests, and the group path has driven a real alert through the real alert worker and a real Postgres — but no group has been authored, filled, failed and resolved against a running deployment, and no group page has reached a phone. The dogfooding step in the list above is not optional for groups either, and it is the step that has not been taken.

When a page is lost, it is lost permanently​

emit is best-effort by necessity, and a failed publish is a page nobody gets

The state change and its monitor_events row are committed before the publish. The transition is durable, so a redelivered result re-enters evaluation, finds the monitor already in the state it judged, and returns before reaching the emit path — a NAK could therefore never retry the lost page, only replay the ingest work. Retrying was not rejected as too expensive; it is unreachable.

So a failed publish is an ERROR log (probe alert publish failed — this transition paged nobody) plus proxima_nats_publish_total{status="error"}, and nothing else. Detect it with the identity the three emission counters are built to satisfy:

# Should be flat at zero. Anything else is transitions that reached NO outcome at all.
sum(increase(proxima_probe_would_page_total[1h]))
- sum(increase(proxima_probe_paged_total[1h]))
- sum(increase(proxima_probe_page_suppressed_total[1h]))

would_page = paged + Σ suppressed + failures, so a persistent non-zero residual is the failure count. Alert on it once any monitor has paging on.

An outage and the alert it raised​

Every transition row (monitor_events, monitor_group_events) is inserted with an id the backend mints before the write. A firing page carries the same id as an annotation: proxima_monitor_event_id for a monitor, proxima_monitor_group_event_id for a service. It is an annotation and not a label on purpose. Labels become the alert group's common_labels, which route matching and the card read, and a value that changes every outage would make those differ between episodes.

The write happens as the LAST step of processIngest, right after the correlate publish, bounded by a 5s timeout (outageLinkTimeout) so a slow or stuck lock costs this consumer seconds, never its whole backlog. After the ingest recount (the first moment the alert group's firing_epoch is final; a reopen bumps it there), the AlertWorker writes alert_group_id and firing_epoch onto that row. The write:

  • is its own statement, outside the recount's transaction, which holds a row lock on the paging hot path. A descriptive write must never be able to stall or roll back a page;
  • runs only for a proxima_probe source (read off the source row, never the payload) and a firing delivery. A resolve names nothing;
  • also matches the UPDATE against the monitor or group id the alert's own label names, not the event id alone — a mismatch there writes nothing and is unreachable in correct operation (see the metrics row below);
  • is written once. A redelivery, or a later episode, finds the row already linked and leaves it alone;
  • is best-effort. It never NAKs the ingest. A link it could not write is a WARN and proxima_uptime_outage_link_failed_total{reason}, and stays NULL for good, so the counter is its only record.
On the rowMeaning
alert_group_id setThis outage paged. (alert_group_id, firing_epoch) is the outage on the alert side, the same key remediation, triage claims and cost use
alert_group_id NULLThis outage never produced an alert. That is the normal case while paging_enabled is off
NULL, and the counter tickedThe page happened but the link was not written: a broken link, diagnosable from the WARN

A down that merges into an alert group that is still open is linked to that group's current epoch. firing_epoch is meaningful only while alert_group_id is set: deleting an alert group clears the id (ON DELETE SET NULL) and leaves the epoch.

Incidents for monitors without a host​

A monitor with a host has always joined that host's incident. A paged monitor without one, and every monitor-group page, used to have no incident at all. Each now opens a one-member incident scoped to the monitor (scope_kind = monitor) or the service (scope_kind = monitor_group).

  • Paging is unchanged. A new incident's leader fans out exactly as the singleton did, and route matching reads the alert group, never the incident.
  • An unknown tier opens a P5 incident. incidents.severity admits only P1–P5. The alert group keeps unknown, which is what routing reads.
  • Flapping: another outage of the same monitor while its incident is still open joins it, and once the incident resolves, the next outage opens a new one. A one-member incident resolves with its outage, so in practice each outage is its own incident unless the resolve was held.
  • A probe page is never suppressed when it joins an open monitor incident as a member. A monitor-scoped incident has no storm to control — suppressing the monitor's own recurring page would mean it stops paging while its incident is open, which is not what "one incident per outage" is supposed to do. A leader is never a member (see the next point), so this is specifically the flap case: a second outage of the same monitor arriving while its incident is still open. Membership is still recorded, and in active mode the join is recorded as its own decision, decision=probe_member_unsuppressed, distinct from an ordinary suppressed member; in shadow mode it is recorded as an ordinary decision=member, since shadow mode suppresses nothing to begin with.
  • A re-correlated incident leader (any scope kind, including a monitor's own incident) never joins as its own member. A redelivery after a transient DB error, or a reopen while the incident is still open, gets the full fan-out again instead of being folded in as a member — recorded as decision=leader_recorrelated. Before this, an active-mode leader re-correlated this way was suppressed as its own member and paged nobody; it now pages once per reopen. docs/runbooks/incident-grouping-shadow-to-active.md covers what changes for clients already flipped to active.
  • Blast radius now runs for these incidents and, with one member, shows nothing.

scope_kind = monitor_investigation is reserved for investigating an outage that never paged (a later release). Those rows are not incidents that happened, so they MUST be excluded by any incident-count aggregate — via grouping.IsIncidentScopeKind, not a hand-rolled kind check, since that is the one place the classification is made and tested against the live incidents_scope_kind_check constraint. No such aggregate exists yet; this is a gate for whoever writes the first one, and in particular for the later Investigate feature, whose whole premise is a monitor_investigation row that must never inflate an incident count, MTTR or precision denominator.

Monitor labels​

A monitor can carry the same free-form key=value labels a host can — edited from the monitor detail page, and searchable from the monitors list via ?label_key=/?label_value=. Omit label_value to match any value for that key; passing label_value without label_key is a 400, the same reasoning the kind/state filters already use — a filter that looks applied but silently matches nothing is the worst answer to give during an incident.

Labels are stored in the same entity_tags table hosts already use — entity_type = 'monitor' rather than 'host' — so the same key format rules apply: lowercase alphanumeric with dots, hyphens or underscores, starting with a letter or digit, up to 100 characters; values up to 500 characters.

Monitor tags gate on monitors:read/monitors:write directly, not on the host system's separate tags:write. That is a deliberate simplification rather than an oversight: unlike hosts, which predate monitors and already had tags:write wired through several other tag consumers, monitor tags are monitor-owned metadata with no cross-entity-type permission story to preserve — a fourth permission would only add a grant nobody needs to hold. GET/PUT /api/v1/monitors/{monitorID}/tags and DELETE /api/v1/monitors/{monitorID}/tags/{key} follow the monitor's own three-permission model (see Permissions); GET /api/v1/monitors/tags/batch?monitor_ids=… mirrors the host batch read, silently dropping monitors the caller cannot access rather than refusing the whole batch.

That three-permission model is not the only door onto this data, though: the cross-entity autocomplete endpoints — GET /api/v1/tags/keys and GET /api/v1/tags/values, which back the key/value suggestions in the label editor and the monitors-list label filter — are gated on tags:read instead, not monitors:read, and since this plan they also union in monitor label keys/values (see AutocompleteKeysScopedTenant/AutocompleteValuesScopedTenant). So a caller holding tags:read without monitors:read can enumerate a tenant's monitor label vocabulary through autocomplete even though they cannot list the monitors themselves. This is not a tenant-isolation leak — the tenant scoping still applies — just a second, pre-existing permission surface on the same underlying data that monitor labels now sit behind too.

Labels are also the mechanism notify routing matches on — a monitor's labels reaching an escalation_routes match_labels decision. That is built and live, but only for notify, never for paging: paging's own route matching still reads a probe alert's environment_id and host_id directly, carried from the monitor row (see Entity-scoped routes match), and never touches entity_tags. notify is the one path anywhere in this feature that reads a monitor's labels at all.

Notify routing​

Every escalation route can independently do two things on a match: page — the on-call apparatus the rest of this page assumes: an escalation policy, an on-call schedule, a per-user notification chain, ending in something that must be acknowledged and re-escalates on silence — and notify — a fire-and-forget action that posts once, to the in-app feed, and expects nothing back. Nothing escalates on a notify's silence, because nothing is waiting for a reply, and no on-call schedule needs to exist for a route to notify at all.

notify no longer reaches Telegram. escalation_routes.notify_telegram is ignored for a monitor's or a group's transition — Telegram delivery is the client's explicit Telegram targets now, sent independently of route matching altogether. The column, and the database CHECK (escalation_routes_page_or_notify_chk, migration 000205) that still accepts it as satisfying "at least one of a policy, notify_telegram or notify_feed", both remain — a route saved before this change keeps its stored value — but the route editor no longer offers a way to set a new one.

The two coexist on the same route rather than substituting for each other — a route can page, notify, or both. escalation_routes.escalation_policy_id is nullable specifically so a route can notify with no policy at all.

Why notify cannot ride the paging pipeline

A monitor's -> down transition only reaches escalation_routes matching — only creates an alert_groups row at all — when paging_enabled is on: pageRefusedBecause holds a firing publish before the ingest payload is even built. paging_enabled ships false on every monitor, so the population notify exists to serve — a client who wants a lightweight down/up notice without opting into the full paging apparatus — would never reach a route match if notify rode that same gated pipeline. So it does not: notify is evaluated by a second, independent call, made unconditionally alongside the paging gate rather than through it, that never touches alert_groups and is never gated by paging_enabled or by ProbeSanityGuard. The one thing it does inherit from the paging pipeline is which transitions count — the same -> down / leaving-down pair that reaches the pager; -> pending fires neither.

Route matching: labels reach a monitor through entity_tags​

match_labels is not new — it has been a route field, a JSONB column, subset-equality matched against an alert's labels (escalation.routeMatches: every key the route names must equal the alert's value for it), since before this feature. What is new is that a monitor's own labels are now one of the things that can populate the alert side of that match — previously only hosts had labels to match against at all.

The read happens at evaluation time, outside the alert-ingest pipeline entirely: notify's dispatcher reads the monitor's (or, for a member of a group — see below — the group's) entity_tags rows fresh, on every transition, and hands them to the same escalation.MatchRoute function paging already uses. Notify is unfiltered where paging is not. Paging first narrows to routes naming a policy (escalation.PagingRoutes) and then runs first-match-wins over that narrowed set; notify runs first-match-wins over every enabled route, policy-bearing or notify-only alike — so the route that matches for notify on a given transition is not guaranteed to be the same route that matches for page on it. A notify-only route sitting earlier in position never shadows a paging route further down (paging never sees notify-only routes at all), but the reverse is not excluded: the first route by position that matches at all is notify's answer, whatever else that route does.

A route's own severity, environment_id and host_id filters apply to notify exactly as they do to page — a route scoped to P1 or to one environment only ever notifies for the transitions it would have paged for, had it named a policy.

Groups inherit the same exclusivity notify already respects for paging — not a new rule

A monitor inside a group never independently reaches notify — not conditioned on whether the group actually holds paging authority, but unconditionally, the moment the monitor's group_id is set. The group's own aggregate transition reaches notify instead, carrying the group's entity_tags (entity_type = 'monitor_group') rather than any member's, and it does so on every group transition regardless of the group's own paging_enabled — because a group's derived state re-evaluates on every member transition regardless of who holds paging, so the group's own call is guaranteed to fire for the same underlying event a suppressed member call would otherwise have duplicated. A group carries no host_id into the match, same as for paging — a group spans hosts by construction.

The two targets​

A matched route fires notify_feed only — a monitor.state_changed event on the in-app notification feed, read under monitors:read (see Shipped kinds). The message text is a past-tense fact — "Checkout is down" or "Checkout recovered" — never a live claim about now, per the feed's own events, not state rule.

Routes no longer decide Telegram. escalation_routes.notify_telegram is ignored everywhere, not just for uptime — nothing anywhere reads it to decide whether or where a Telegram message is sent — and it has no toggle in the route editor (the column stays until a later migration drops it; existing rows and pc apply'd files keep their stored value and keep applying). Telegram delivery is the client's explicit Telegram targets, sent on every confirmed down/up whether or not a route matches.

The per-monitor override​

A monitor may pin its own Telegram chat (telegram_chat_id_override). It replaces the client target's chat for that one monitor. It must be one of the monitor's own client's audience = 'client' chats — refused with a 400 at authoring and re-checked at send (a mismatch is skipped and counted, never delivered). It never adds a second send, never decides whether a client message goes out (no client target → no client message, override or not), and never touches the internal target. client_notify = false on the pinned chat still mutes it.

What this does not do​

  • Does not touch paging_enabled, pageRefusedBecause, or anything downstream of a page. Every existing route, policy, schedule and chain behaves identically; notify runs alongside the gate, never through it.
  • Does not rate-limit a flapping target. A monitor confirming down and up repeatedly notifies on every confirmed transition, exactly as it would page on every one if paging were on. A per-route or per-monitor minimum notify interval is the likely shape of a fix, and is deliberately not built ahead of real usage showing it is needed.
  • Does not give the feed a severity. monitor.state_changed carries none, for the same reason no kind ever will.
  • Does not implement or unblock StepNotifyChannel. That escalation-step type — a dormant notification-chain step, unreachable today regardless — stays exactly as unreachable as it was. notify is a property of the route, evaluated outside the escalation-chain machinery entirely, never a step a policy executes.
notify now works for push monitors too — closed 2026-09-15

Everything above is true for a probe monitor's own transition (ProbeResultWorker.emit), for a monitor group's aggregate transition (MonitorGroupEmitter.EmitGroup), and now for a push monitor's transition (MonitorEmitter.Emit) — all three call the notify evaluation described above, unconditionally, alongside the paging gate chain.

Both places a push monitor's state actually changes — the public heartbeat check-in (GET /monitors/heartbeat/{token} / .../fail) and the silence-sweep auditor that confirms down when nothing arrives inside the grace window — go through MonitorEmitter.Emit, which now carries the same monitorNotifyTimeout-bounded sibling call the other two paths already had:

func (e *MonitorEmitter) Emit(ctx context.Context, m *domain.Monitor, t *domain.MonitorTransition) {
emitMonitorTransition(ctx, m, t, e.alerts, e.sources, e.fleetSuppressed)

notifyCtx, cancel := context.WithTimeout(ctx, monitorNotifyTimeout)
defer cancel()
notifyMonitorTransition(notifyCtx, m, t, e.notify, e.fleetSuppressed)
}

This gap was found and disclosed here during the feature's own testing pass, went live on real push-heartbeat traffic before it was closed, and stayed disclosed until it was — a monitor kind this section didn't cover was worth naming plainly rather than assuming nobody would hit it.

Telegram targets​

Every client has at most two uptime Telegram targets (client_uptime_notify), set on the client's Settings → Uptime notifications card, PUT /api/v1/clients/{id}/uptime-notify, or the client_uptime_notify kind:

TargetMay beGets
Internala team-bound chat, or an audience = 'internal' group chat — never a chat with an audience = 'client' binding, even if it is also team-boundan HTML card: status emoji + name, target (<code>), cause (escaped, truncated to 300 characters), confirming locations, down-since or duration, and an Open in Console link button — all secret-sanitized
Clientone of this client's audience = 'client' group chatsa status line only — no buttons, no links, no internal detail

Unset is a choice; broken is loud. No target → no message of that kind, no metric, no warning; the card says so ("No internal chat — uptime messages for this client are not sent to the team"). A target that was set and no longer resolves — its binding deleted, re-bound to another client or team, or now failing the audience rule — is skipped, counted in proxima_uptime_notify_invalid_total{target}, logged at WARN with the client id, and flagged Unrouted on the card. It is never silently re-pointed.

Threads. An explicit thread wins; otherwise the chat's uptime topic, else its alerts topic, else the binding's own thread, else the root. The card shows which one a target will use.

Who may change what. Reading needs monitors:read on the client. The client target needs telegram_ops:write held client-wide — an environment-scoped grant is not enough. The internal target needs super-admin or the global telegram_ops:write — callers without it see only whether an internal chat is set, never which one. Every change writes one audit_log row per target (client_uptime_notify.internal_changed / .client_changed, old → new chat and thread).

Services (monitor groups) use their client's targets. Paged uptime alerts do not post a status card to client chats — the client gets one message per outage, this one.

The close is a delivery receipt, not a state re-check (uptime_client_notice, migration 000238). A client "down" that was actually delivered opens a notice keyed by the subject (monitor + the monitor's id, or monitor_group + the group's id) recording the chat, thread and client it went to. The matching "back up" (or "monitoring stopped", or "no longer down") is sent only if that notice is found — to the exact chat and thread the down went to, regardless of later target changes (re-pointed, cleared, or the whole client_uptime_notify row deleted), as long as that chat is still this client's audience = 'client' binding with client_notify on. A down that was never delivered (no target, held by the fleet guard, opted out, or the send itself failed) opens no notice, so it is never followed by a phantom "back up".

Known gap: a monitor joining a group mid-outage strands its notice

A standalone monitor's notice is keyed by the monitor's id. If that monitor is added to a group while its "down" notice is still open, every later transition for it is dispatched under the group's subject instead (see Derived state) — nothing ever closes under the monitor's own id again. The client keeps the stale "down" it already received with no matching close; the group's own transitions behave correctly from the moment it starts owning notifications. Not yet fixed; there is no migration path for an open notice across a subject change.

Permissions​

Monitors use three per-client permissions, checked on every request against the monitor's own client — not one the caller named:

PermissionGrants
monitors:readList monitors, read one monitor and its state, read its event timeline and its uptime history.
monitors:writeCreate, edit, pause and resume.
monitors:deleteDelete. Deliberately separate from :write: deleting a monitor takes its state and its whole incident timeline with it, which is not the same authority as editing a threshold.

Monitor groups ride the same three, and there is deliberately no fourth permission. monitors:read lists and reads a group and its timeline, monitors:write creates, edits and takes paging authority, and monitors:delete deletes one — for the same reason it is separate on a monitor: deleting a group takes its incident timeline with it. A group is monitor configuration rather than authority an operator holds separately from the monitors it folds; a second permission would let a role edit services it cannot see inside. Membership is authored through the monitor routes, so putting a monitor in a service needs monitors:write on that monitor, not on the group.

No role holds any of them in production

Verified 2026-08-21: all 19 roles lack every monitors: permission. New permissions ship as code; granting them to a role is an operator action that no deploy step performs, so the feature is invisible to everyone except super-admins until somebody does it. Treat the grant as part of enabling uptime monitoring for a client — see Turning paging on.

Tenant scoping is enforced by the store, not by the caller:

  • Listing another client's monitors is a 403, not a filtered list. Naming a client you cannot read is refused rather than silently narrowed.
  • Omitting client_id lists every client you can read — never every client. The scope is ClientIDsWithPermission(monitors:read), so it is exactly your own grants: a client where you hold only monitors:write is not in it, and a caller holding monitors:read nowhere gets an empty list rather than everything. See The cross-client list.
  • A monitor id belonging to another client is a 404, indistinguishable from one that does not exist, on every per-monitor route. It does not leak whether the id is real.
  • A group id outside your scope is a 404 on every per-group route, on the same principle, and a group_id naming another tenant's or another environment's group is the same 400 a non-existent one gets — so neither route can be used to probe which groups exist elsewhere.
  • An environment_id, host_id, asset_id or cluster_id belonging to another client is a 400 with the same message a non-existent id gets, so the API cannot be used to probe which ids exist elsewhere.

The Uptime section appears in the sidebar only for users holding monitors:read, and the create/edit and delete controls are gated on monitors:write and monitors:delete.

API​

Every route is mounted twice — at the top level with the client in the query string, and nested under a client — and both enforce the same checks. Responses use the standard {"data": …} envelope; DELETE returns 204 with no body. The two heartbeat routes below are the exception on both counts, and deliberately: they are public and session-less, mounted once, carry no client anywhere, and answer a bare {"status":"ok"} — see Checking in.

MethodPathReturns
GET/api/v1/monitors?client_id=…The client's monitors, each with current_state.
—rate_limit on the list, the detail and the status rowsPresent only while at least one location is backing off from a 429: { "locations": [...], "checking_every_seconds": 240, "retry_after_seconds": 600 | null }. Derived from rows whose last result was rate-limited and that are fresh under their own widened window. checking_every_seconds is the largest announced next check (or the interval when none is backing off yet); retry_after_seconds is what the site asked for, before the 240 s cap, and is display only. The current_state stays up.
GET/api/v1/monitorsWithout client_id: every client the caller holds monitors:read on — never every client. Grouped by client, with per-client totals in meta. See The cross-client list.
POST/api/v1/monitors201 and the created monitor (plus push_token for a push monitor).
GET/api/v1/monitors/heartbeat/{token}Public, token-authenticated. 200 {"status":"ok"} on any accepted call, 404 on an unknown token. See Checking in.
GET/api/v1/monitors/heartbeat/{token}/failSame as above, but confirms down immediately.
GET/api/v1/monitors/{monitorID}One monitor plus its nested current state.
PATCH/api/v1/monitors/{monitorID}The updated monitor. Every field optional; an omitted field is left alone.
DELETE/api/v1/monitors/{monitorID}204. Cascades to state, events and notice history.
POST/api/v1/monitors/{monitorID}/pauseThe monitor with enabled: false.
POST/api/v1/monitors/{monitorID}/resumeThe monitor with enabled: true.
GET/api/v1/monitors/{monitorID}/eventsThe transition timeline, newest first, limit capped at 200. Each row carries quorum_degraded, computed server-side. meta.total is the monitor's TRUE transition count — larger than meta.per_page (== len(data)) exactly when the timeline was truncated to fit limit.
GET/api/v1/monitors/{monitorID}/incidentsThe SAME rows, folded into incidents, newest first, limit capped at 100. Cursor-paged by ?before= (an RFC3339 instant) rather than total-counted — see Incident timeline.
GET/api/v1/monitors/{monitorID}/uptime?window=&buckets=Availability over a window, including history — the per-location bar at the requested (or default) segment count; each bucket carries maintenance_checks (checks left out of total_checks because a window held the monitor). 503 when no metrics backend is configured.
GET/api/v1/monitors/{monitorID}/maintenanceThe monitor's shortcut window, every other window covering it ("also covered by"), and those holding it now. See Maintenance windows.
GET/api/v1/monitors/uptime?client_id=…&monitor_ids=…&window=Many monitors at once, as an object keyed by monitor id: {has_data, uptime_percent} each. client_id is required; ids outside it are dropped rather than refused. See The batch uptime read.
GET/api/v1/monitor-groups?client_id=…The client's services with their derived state and counts. client_id is required here — this endpoint never lists across tenants — and the list is not paginated.
POST/api/v1/monitor-groups201 and the created group, which starts empty and unknown. There is no members field.
GET/api/v1/monitor-groups/{groupID}The group, its state, and its members with each member's own state and enabled flag.
PATCH/api/v1/monitor-groups/{groupID}Partial: environment_id, name, severity, paging_enabled. No client_id — a group cannot be re-homed to another tenant.
DELETE/api/v1/monitor-groups/{groupID}204. Cascades to the group's state and timeline, and detaches its members rather than deleting them.
GET/api/v1/monitor-groups/{groupID}/eventsThe service's incident timeline, newest first, limit capped at 200.

The six /monitor-groups routes are mounted the same two ways — top level with ?client_id=, and nested under /api/v1/clients/{id}/monitor-groups — and their path parameter is {groupID} and never {id}, because under the nested mount id is already the client. Authoring a group covers them in full, and Joining and leaving a group covers membership, which moves through the monitor routes rather than these.

Behaviours worth knowing before you write a client:

  • current_state on a list row is nullable, and the null is not unknown. unknown means "never checked"; null means the monitor has no state row at all, which the backend makes impossible at create time — so it is a storage anomaly worth chasing, and clients must render it as its own thing.
  • state on the detail response is nullable for the same reason. The endpoint answers 200 with "state": null rather than contradicting the read that just proved the monitor exists.
  • The kind and state filters are validated, not forwarded. An unrecognised value is a 400 naming the valid set. Answering 200 {"data":[]} to a mistyped filter reads as "you have no monitors" during an incident, which is the worst possible thing to conclude.
  • Paging: list limit defaults to 100 and is clamped to 500. The events limit defaults to 50 and is capped at 200. The incidents limit defaults to 20 and is capped at 100, and pages by ?before= rather than an offset — see Incident timeline.
  • A duplicate name is a 409, on create and on rename alike.
  • Re-homing is impossible. There is no client_id in the update body: moving a monitor to another tenant is not an edit.
  • Entity links are three-state on update. Send a uuid to set host_id/asset_id/cluster_id, an explicit null to clear it, or omit the key to leave it alone.
  • group_id is the fourth three-state field, and the one it is dangerous to get wrong. It joins, moves or leaves a monitor group; an omitted key leaves membership alone, and an explicit null is the only way out. A group outside the monitor's client or environment is a 400, moving a still-grouped monitor's environment_id is a 400 telling you to detach it first, and a paging_enabled that conflicts with the group's is a 400 naming the other side. On responses group_id is nullable — the ungrouped majority is every monitor that has not been put in a service.

The cross-client list​

GET /api/v1/monitors without client_id answers with every client the caller can read, rather than the 400 it used to. It is one endpoint in two modes, and the difference is worth stating precisely because the loosening is the whole risk of it.

The scope is the caller's grants. It comes from ClientIDsWithPermission(monitors:read), never AllowedClientIDs(), so holding monitors:read on one client does not expose a client where you hold only monitors:write or only hosts:read. A super-admin sees every tenant; a caller with monitors:read on three of their eight clients sees three; a caller holding it nowhere gets {"data": []} and not everything. That last case is the one worth naming: an empty scope means no access, never no filter, and the store enforces it with clientIDs != nil rather than len(clientIDs) > 0 — the two values a length check collapses into one, on the one monitor query where the collapse hands over every tenant.

The route is still gated on monitors:read at the middleware, which with no client_id asks "do you hold it on any client?". So a caller with none is refused with a 403 before the handler runs; the handler's own empty-scope path is the second line of the same defence rather than the visible one.

A purely environment-scoped user is refused here but served with client_id

That middleware branch iterates client-wide grants only (ClientAccess). A user whose monitors:read comes solely from an environment-scoped grant — no client-wide row at all, which is constructible: POST /users/{id}/clients with an environment_id writes only user_client_environments, and nothing requires a matching client-wide grant — therefore gets a 403 without client_id and a 200 with it, for the same monitors. It fails closed, so it is not a leak; it is an inconsistency. It was inert before this endpoint accepted a request without client_id (every such request was a 400), so the cross-client list is what turns that branch into the gate on the feature. The fix belongs in the middleware, not here.

Rows arrive grouped by client (ORDER BY client_id, name), so each client's monitors are contiguous and a section can be rendered by walking the page once. The order is by client id, not name — this table has no client name to sort by — so a UI that wants alphabetical sections reorders the sections it was given.

meta is the standard pagination envelope with the section headers added. It is the {total, page, per_page} the repo standard reserves (docs/standards/api.md) — the same shape the frontend mirrors as PaginationMeta — plus client_totals:

{
"data": [ /* … monitors, grouped by client … */ ],
"meta": {
"total": 512,
"page": 1,
"per_page": 100,
"client_totals": [
{ "client_id": "…", "total": 340, "shown": 100 }
]
}
}

client_totals is what makes a truncated section honest. The list pages at 1–500 (default 100), so a page boundary can fall in the middle of a client's monitors. Each entry's total is counted by the database over the filtered set before the page was cut (a COUNT(*) OVER (PARTITION BY client_id) on the same query the rows come from, so the two can never disagree); shown is counted from the rows this response actually carries. A section header must render both. "340 monitors" above 100 rows is a lie the UI tells silently and nothing downstream can detect — "showing 100 of 340" is the only honest rendering of a cut section. There is one entry per client with a row on this page; a client whose rows all fall outside the page window has no entry, because there is no section for it to sit above.

meta.total is the only thing that can say a whole client was cut. client_totals describes the page; a scope of seven clients whose first owns 101 monitors renders one section, and neither the rows nor the section headers mention the other six. total counts the whole filtered scope, so "showing 100 of 512" is renderable above the sections. It survives a page past the end too — the per-row window count cannot, so an empty page has it counted separately rather than reported as zero.

meta.per_page is the page size that was actually served, not the one requested. A ?limit=501 is clamped to 100 and says so; echoing the request would advertise a page that never existed, which is the same silent lie as a header describing rows it is not showing.

meta.page is meaningful only when offset is a multiple of the effective limit. This list pages by limit/offset; page/per_page is the envelope's vocabulary, not this endpoint's input, so the two can only agree on an aligned window. page is offset / per_page + 1, which for an unaligned offset reports the page the window starts in — ?limit=10&offset=5 answers page: 1 — so a client that round-trips (page - 1) * per_page back into offset asks for a different window than it was served. Page by offset, and read page as a label rather than a cursor.

Passing client_id is unchanged in every respect, response shape included: that mode answers the plain {"data": […]} envelope with no meta key at all. Its absence is the signal that the response describes the one client you asked about.

The batch uptime read​

GET /api/v1/monitors/uptime answers many monitors' availability in one request, and it exists for exactly one caller: the monitors list's Uptime column, where reading /{monitorID}/uptime per row would be up to 500 round trips on a page load.

{
"data": {
"0f9c…": { "has_data": true, "uptime_percent": 99.95 },
"3a71…": { "has_data": false, "uptime_percent": null }
}
}

An object keyed by monitor id, not an array, because the caller is a table looking rows up rather than iterating — and an array would additionally have to define what its order meant, which the request never asked about.

One query pair for the whole page. Both instant queries are grouped by (monitor_id, location), so a page of 500 monitors costs the same two round trips to VictoriaMetrics as one monitor does. Grouping by location as well as by monitor is not incidental: the result is then folded by the same code path the per-monitor endpoint uses, keeping the per-location skip, the per-location clamp and the upside_down inversion identical. sum by (monitor_id) alone would have been a second, subtly different definition of "uptime" — and a row disagreeing with the monitor page it links to is the kind of bug nobody reports and everybody stops trusting.

client_id is required. It is the tenant the whole batch is authorized against, so there is no default that would be safe; omitting it is a 400, not a cross-client mode.

A monitor id outside that client is dropped, not refused. The rows are loaded through a tenant-scoped read, so a well-formed id belonging to another client — or to no monitor at all — is simply absent from the response. Both are the same nothing, which is what keeps the endpoint from being an existence oracle for other tenants' monitors; a stale list page holds such ids routinely, so refusing the batch would break the column and confirm the id is real. Read the answer from what came back, never from what you asked for.

A malformed id is still a 400, and the asymmetry is deliberate: a foreign id is a legitimate request you may not be answered about, while an unparseable one is a caller bug. Answering {"data":{}} to it would render as "nothing on this page has any availability", which during an incident is indistinguishable from a fleet-wide outage. monitor_ids is capped at 500 — the list's own page cap, read from the same constant, so the cap can never be the reason a full page has a column it cannot fill.

has_data carries the per-monitor endpoint's contract unchanged. A monitor that has never been probed is present with has_data: false and uptime_percent: null; 0 means every check in the window failed. Never coalesce the null — uptime_percent ?? 0 renders an unmonitored service as a failing one.

Push monitors need no special case. They are checked in to rather than probed, so no probe_success series exists for them and the ordinary no-data path already answers has_data: false. Nothing in this endpoint inspects kind; the console simply does not request them.

window is the same parameter, parsed by the same function, as the per-monitor endpoint: Go duration syntax plus a whole-number d/w suffix, bounded to [5m, 30d], defaulting to 24h, with an out-of-range or unparseable value a 400 rather than a silent clamp. A 503 means no metrics backend is configured — the column could not be read, which is not the same answer as "these monitors have no history".

In the console​

Uptime is a top-level sidebar entry rather than a row under Observability, because monitors are an alert source rather than a dashboard. A monitor group is called a service in the UI: "group" is what the table is, "service" is what it means.

New service and New monitor sit in the page header, both gated on monitors:write. Both need a client chosen first: the list defaults to the cross-client mode, which has no single client to author into, so with none selected the header says to choose one rather than offering a button that cannot work.

The monitors list (/uptime)​

The list is client → service → monitor, and it reads the cross-client mode by default — every client the caller holds monitors:read on, never every client. Its columns are Monitor, Kind, Environment, Status, Uptime (24h), Severity and Paging.

Eight things about it are honesty obligations rather than layout choices, and each is pinned by a test:

  • No header states a count it is not showing. A client section whose monitors were cut by the page boundary reads "Showing 100 of 340 monitors", from meta.client_totals; a complete one reads "340 monitors". Above the sections, meta.total gives the same treatment to the whole scope — "Showing 1–100 of 512 monitors" — because that is the only figure that can say a whole client fell off the page. client_totals has an entry only for a client with a row on this page, so seven clients whose first owns 101 monitors render one section and mention the other six nowhere else.
  • "This window is not the whole scope" and "there is more after this window" are different questions, and the scope line answers the first. The last page of a cut list is still a page: offset=100 over a scope of 102 reads "Showing 101–102 of 102 monitors. This is the last page…", not "102 monitors." above two rows. The count of clients ("across 2 clients") is printed only when the window is genuinely the whole scope, because client_totals.length counts clients with a row on this page and would otherwise be a page-scoped number presented as the scope's.
  • The list pages by offset, never by meta.page, which is a label rather than a cursor. Paging past the end is undoable by clicking: the pager renders whenever offset > 0, including on a window with no rows, and that window says "No monitors on this page — you have paged past the end" rather than "No monitors yet", which would be a claim about the whole client.
  • A service is never counted as a monitor. Not in a section header, not in a total, not anywhere. Modelling a group as a monitor is what makes Uptime Kuma's dashboard report "3 down" for two monitors and their group (louislam/uptime-kuma#4597); a service here is a heading over monitors, never one of them.
  • A service header says how much of itself is on screen — "Showing 2 of 5 members" — because member_count is the database's count of every member while the rows are a page, and it carries the service's derived verdict, its down_count, its severity and whether it pages.
  • A verdict folded from a partial sample looks different from a complete one. A non-zero excluded_count renders a dashed "Partial — 2 of 5 not measured" badge beside the verdict. A current_state of null is rendered as "State unavailable" and suppresses every count: the zeroes beside a missing state row are the absence of a row, not a measurement.
  • The Paging cell tells the truth about who pages. A member of a paging service reads "Paged by Checkout" — not a blank, and not a claim that the member pages, because it does not: the service pages once for the whole thing. A monitor that holds its own flag and sits in a paging service reads "Pages on-call — and so does Checkout", which is the double-paging pair that nothing at runtime can see.
  • The Uptime column distinguishes four kinds of nothing. It is filled from GET /monitors/uptime — one request per client section, not one per row — and every state that cannot show a figure says which nothing it is, in the aria-label as well as the tooltip, because for a screen reader the reason is the whole cell. A monitor that was answered and has never been probed reads "never probed in this window" and never 0%, which is a verdict rather than the absence of one. A read that failed reads "could not be read", which is a statement about us and says nothing about the monitor. A push monitor reads "checked in to, not probed" and is never requested at all: it has no probe_success series by construction, so the row already knows the answer from its own kind. The percentage itself is printed exactly as the server sent it, so a row and the monitor page it links to cannot disagree about the same window.

A page holding a single client renders no client tier at all — a lone accordion wrapping everything is friction for the common case. The tier is decided from the monitor response's own scope, never from the client list the filter is drawn from: that list is the clients you can see, while this page's scope is the clients you hold monitors:read on, and deciding from the wider one gave a caller who reads monitors on one of two clients the very accordion the tier exists to avoid — and left rows unattributed entirely whenever that list failed or merely arrived after the monitors. Client names still come from it, and a section whose name is unavailable falls back to the raw client id rather than to no heading. A client that has never made a service gets a flat table with no service headings. While a client's services cannot be read, nothing is filed under "Ungrouped": a monitor whose service is unknown is not a monitor with no service, and the heading says so.

A monitor in maintenance reads "Maintenance · until 03:00 · still checking (failing)" (or up). The time is the covering occurrence's end (maintenance_until on the status row, from the in-memory index; prefixed with the weekday when it is not today, or the date when it is a week or more away). The failing / up part is last_result_ok — plain "still checking" when that is unknown — and it is a display-only approximation: the latest sample per location folded by location quorum over the fresh locations, without the evaluator's prober-level detail (several probers at one location: the latest writer wins), its healthy-and-capable denominator, suppression, or 429-widened freshness. It can occasionally differ from what the evaluator will decide when the window ends, and it pages nothing.

A monitor whose locations are backing off from a 429 keeps its green Up and gains an amber "Rate limited · checking every 4 min" chip. Its tooltip names the limited locations and the Retry-After the site sent, worded so it never implies Console waits longer than its 4-minute cap.

The service page (/uptime/groups/{id})​

Four cards: Current state (the verdict, the partial badge when the sample was incomplete, and the four counts — members, measured, down, excluded — or an explicit "unreadable here, not zero" when there is no state row), Members (each with its own state and whether it counts toward the verdict; a paused member reads Paused rather than its stale state row), State transitions (newest first, naming the members that were down and excluded at the time, from the names stored on the event), and Configuration.

Deleting a service says what deleting one does: the members are kept and detached, and if the service was the thing holding paging, the confirm says plainly that nothing will page until somebody authors it per monitor or makes another service.

The monitor detail page​

Four cards: Current state (the chip, the confirming count, the degraded badge when it applies, and the timings), Availability and latency (the measured figure with the check counts behind it, the mean response time, the per-location availability bar, the latency chart and the per-location breakdown — or an honest "no probe data yet" — over a window chosen from four presets that the URL remembers), Configuration, and State transitions — the incident timeline, day-grouped, newest first, naming the locations that confirmed each incident's start and end.

The header carries the same amber rate-limit chip beside the severity while any location is backing off from a 429.

The dialogs​

  • The monitor create/edit dialog drives its options from kind, mirrors the backend's numeric floors client-side (as a convenience — the API remains authoritative), and reveals a minted push token, with the two check-in URLs built from it, in a panel that must be acknowledged. For a push monitor it also shows the grace window its interval and retry fields imply, recomputed as you type, so the silence budget is visible while it is being set rather than inferred afterwards. Its Service select is where membership is authored — there is no group-side join — and it offers only the services in the monitor's own environment, because monitors_group_scope_fkey refuses the rest. Until the service list has actually arrived the select is disabled and says it is keeping the current service: a request still in flight and a request that failed read differently, and neither is allowed to look like "this monitor is in no service". Turning a monitor's paging on while its chosen service already pages says "this will be refused" before the save is attempted.
  • The service create/edit dialog creates a service empty, and says so: it starts unknown, never up. Its paging toggle names the members already holding their own paging so the exclusivity refusal is explained before it is hit — and it withholds the negative claim ("no member is in the way") when the monitor list it checked was capped, because reassurance on incomplete evidence is the same class of error as a header describing a page. A refusal is rendered verbatim, with error.details.members listed beneath it.

Backend metrics​

All bounded, and not one carries a monitor_id or a group_id — /metrics is a pull exporter with no series expiry, so a per-entity label there is unbounded heap growth ending in an OOM and a dark metrics plane. It bites harder for groups than for monitors, because groups are cheap to create. Per-monitor data is in VictoriaMetrics, where series do expire; which group did or did not page is in the log line and on the alert.

MetricLabelsNotes
proxima_probe_results_totalkind, location, resultresult is success or the error class (timeout, dns, connection_refused, tls, status, keyword, internal) — a bare success/failure split would throw away the one thing a failure counter is read for.
proxima_probe_state_transitions_totalkind, to_stateCounted only where the transition was actually applied, so a verdict another replica wrote first is one transition, not N.
proxima_probe_partial_failure_totallocationMinority-failure observations. Names the pop it indicts.
proxima_probe_rate_limited_totallocationAccepted results a prober recorded as rate-limited (429, counted as up). A rising rate at one location usually means a target is throttling that pop's egress IP.
proxima_probe_quorum_degradedkindPer-replica gauge — see the warning above.
proxima_probe_results_rejected_totalreasonmalformed, invalid_subject, unknown_agent, location_disabled, unassigned, no_longer_assigned, monitor_gone, stale_generation, lookup_error, persist_error.
proxima_probe_results_corrected_totalfieldWhat the backend had to correct in what a pop reported. next_check_out_of_range and retry_after_out_of_range are the 429 hint's; see Freshness.
proxima_probe_assignments_skipped_totalreasonMonitors the resolver could not place.
proxima_probe_no_eligible_locations_totalkindEvaluations of a monitor with no location left able to vote — probed every interval, structurally unable to confirm anything.
proxima_probe_would_page_totalkindConfirmed down-transitions that would have paged, counted before the gate chain. The shadow window's whole output.
proxima_probe_paged_totalkindProbe alerts actually published to the ingest path.
proxima_probe_page_suppressed_totalkind, reasonPages the chain refused: paging_disabled, in_maintenance, fleet_suppressed.
proxima_probe_location_suppressedlocationTier 1 suppression, 1/0. Per-replica gauge — alert on max().
proxima_probe_fleet_suppressed—Tier 2 halt, 1/0. Per-replica gauge — alert on max().
proxima_probe_guard_inerttierA guard tier that cannot fire at all. location, fleet. Nothing alerts on this yet — add the rule.
proxima_probe_sanity_sweeps_totalresultGuard proof of life, clean|error.
proxima_probe_sanity_last_success_timestamp_seconds—Last sweep that read the failure rates. Alert on min() — every replica runs its own guard.
proxima_probe_sanity_withheld_totallocation, reasonLocations that met the failure bar and were not suppressed: single_tenant_failure, single_tenant_fleet.
proxima_probe_prober_stale_totallocationA prober machine silent past the audit window.
proxima_probe_location_dark_totallocationEvery enabled prober at a location silent — the vantage point has left every denominator.
proxima_probe_location_disagreement_totallocationRows sharing a code disagree on kind/client_id.
proxima_probe_location_auditor_sweeps_totalresultAuditor proof of life, clean|error.
proxima_probe_location_auditor_last_success_timestamp_seconds—Last sweep that read the fleet. Alert on max() — the auditor is a singleton, so only the lock winner stamps it.
proxima_push_heartbeat_totalresultHeartbeat calls by outcome: ok|fail|unknown_token. Counted at the handler, before any state-machine work, so a flood of bad tokens is visible even where it never reaches a monitor.
proxima_push_auditor_sweeps_totalresultPush-auditor proof of life, clean|error. Only the tick's lock winner counts clean; a replica that loses the lock counts nothing rather than claiming a sweep it did not run.
proxima_push_auditor_last_success_timestamp_seconds—Last sweep that read the push monitors. Alert on max() — the auditor is a singleton, so only the lock winner stamps it.
proxima_monitor_group_would_page_total—Group down-transitions that would have paged, counted before the chain.
proxima_monitor_group_paged_total—Group alerts actually published.
proxima_monitor_group_page_suppressed_totalreasonThe chain refused a group page: paging_disabled, in_maintenance, fleet_suppressed.
proxima_monitor_group_transitions_totalto_stateApplied group state changes, including the edges that page nobody.
proxima_monitor_group_drift_totalcorrectionGroup states the drift auditor corrected: state (a missed page or resolve), counts (stale evidence).
proxima_monitor_group_authority_conflicts_total—Groups holding paging_enabled beside a member that holds it too — a pair that double-pages, and never self-healing.
proxima_monitor_group_auditor_sweeps_totalresultDrift-auditor proof of life, clean|partial|error.
proxima_monitor_group_auditor_last_success_timestamp_seconds—Last sweep that ran. Alert on max() — the auditor is a singleton, so only the lock winner stamps it.
proxima_uptime_notify_invalid_totaltargetUptime Telegram sends skipped because a configured target (internal/client) no longer resolves — binding deleted, re-bound, or failing the audience rule. Unset targets are never counted.
proxima_uptime_outage_link_failed_totalreasonA firing uptime page whose outage ↔ alert link was not written: write_error, no_row, bad_id, subject_mismatch (unreachable in correct operation). The link stays NULL for good, so this is its only record. Should be flat at zero.

A pop is trusted with what it saw and nothing else — not with metric-label cardinality (an unrecognised error class is folded to internal, because an unbounded label forks the series and every existing rate() flatlines to zero, which on a failure counter reads as "healthy"), and not with its own clock (a checked_at in the future is clamped to ingest time; one in the past is kept, because that is what a buffer replay looks like).

What has not shipped​

ICMP/gRPC/SMTP check kinds, private locations that check databases, certificate-expiry notices, per-location scaling of the fleet-audit window, and the pop-side self-monitoring channel that would let a prober report its own health rather than being inferred from silence.

Paging itself has shipped, and is off on every monitor and on every group. See Turning paging on.

Monitor groups have shipped as well — derivation, emission, the exclusivity refusal, six REST endpoints, membership on the monitor bodies, and the console — so they are no longer on this list. Three group features are, and each is deferred rather than merely missing:

  • Group pause. A group cannot be paused as a unit. Pausing every member one at a time leaves the group unknown, never paused — monitor_group_state.state can hold paused and nothing can reach it. It is the most-requested group operation, which is why it is named here rather than quietly absent, and it is deferred because it needs a cascade semantics decision of its own (does pausing a service pause its monitors, or only stop the service judging them?) plus the tests that pin whichever answer wins. Better Stack cascades it to members; Kuma's is the known-broken one (#7341), where the children of a paused group keep checking and keep alerting.
  • all_children_down. Not a knob nobody got to — a refused one. Under the group's exclusive paging authority it is subtractive rather than additive: one member going down would page nobody at all. Adding it needs its own withheld-page counter and an unmistakable warning in the form.
  • Nested groups. One level only, and the schema says so: monitors.group_id points at a group, and a group points at nothing. A service cannot contain a service.

One smaller gap is worth knowing before you use groups in anger: the monitor list has no ?group_id= filter, so a service's members are found by opening the service rather than by narrowing the list.