Skip to main content

Maintenance Windows

A maintenance window tells Console that work is planned: during it, covered alerts are recorded but nobody is paged, covered uptime monitors keep checking and show Maintenance instead of Down, and maintenance checks are left out of availability. Anything still broken when the window ends pages at once, through the ordinary routing.

Quiet, but never invisible: every held alert is still in the alert list, marked as held, and the window's page lists everything it held.

Find it under On-Call → Maintenance. The badge next to it counts windows in progress.

Scheduling a window​

Schedule maintenance opens a form in three sections, plus Notify (see Service desk ticket and client chat):

  • Basic information — What's going on? (a short summary) and a description (markdown). The on-call team sees both in Console and in Telegram. Show on status page puts the window on every status page of the client whose components it covers. It is saved through its own switch after the form saves. Turning it on needs status_pages:write on the whole client, and a monitor's own shortcut window can never be shown.
  • Maintenance window — From / To and a time zone (the form shows the current time in it). Turn on Recurring maintenance to repeat weekly on ticked days: each occurrence runs from From's time to To's time, and may cross midnight.
  • What's affected — a pick list: the whole client, environments, hosts, monitors, or services (monitor groups). Label rules (key = value, key != value) sit under Label rules (advanced, optional); a rule's matchers are ANDed and a window's rules are ORed.

The Covers right now panel updates as you edit: hosts, monitors and open alerts the draft would hold, and a rule that matches nothing (with the nearest real value, e.g. did you mean payments?).

What each target covers​

TargetCovers
Clientevery alert and monitor of the client
Environmentalerts, monitors and services in that environment
Hostalerts and monitors on that host in the host's own environment — a monitor in another environment that merely links to the host is not covered (give it its own target); an alert that names the host but carries no environment is covered too (treated as the host's environment); one naming a different environment is not
Monitorthat monitor and its uptime alert
Servicethe service's own page and its member monitors
Label rulealerts by their labels, monitors by their tags. A != matcher never matches a subject whose labels could not be read — a negative rule must not cover everything during a lookup failure
Webhook alerts and host windows

An alert's environment comes from its source: Console's own alerts (uptime, agent) carry it, but a webhook alert (Alertmanager, Grafana, …) gets one only from an environment_slug label mapping. A host window does not need that mapping: an alert resolved to the host with no environment is held as if it were in the host's environment. The mapping is still needed for an environment target to hold a webhook alert, and for environment-scoped escalation routing. An alert whose mapping names a different environment than the host's is not held by the host target.

Visibility: such an alert has no environment, so a reader whose access is limited to some environments does not see it in Alerts. A reader with maintenance:read on the host's environment does see its title, severity and host in the window's held list and in the Covers right now preview — it names a host that reader is authorized on.

A window only ever covers its own client. Label rules are evaluated on the paging path and in the preview; a service's page is held by its pick-list entry, client or environment, not by a label rule.

Limits​

One-off windowat most 7 days
Recurring occurrenceat most 24 hours
Extendat most 24 hours past the occurrence's scheduled end
Scopeat least one pick-list entry or rule (at most 200 entries, 20 rules of up to 20 matchers)
Name / description / update200 / 10 000 / 2 000 characters

Time zones and DST​

A recurring window runs at a wall-clock time in its zone. On a DST change, a start that falls in the skipped hour starts at the end of the gap; a start in the repeated hour starts at the first of the two.

What a window does​

Alerts​

A covered alert fires and is recorded as usual. What does not happen:

  • no escalation is armed and nobody is paged;
  • no Telegram alert card is posted;
  • the L1 agent does not triage it.

The alert shows In maintenance · not paged in the list and on its page, linking to the window, and its timeline carries Not paged: in maintenance naming the window. If a window opens mid-escalation, the next escalation step is held until the window ends.

Coverage is read from the database at every paging decision, so a window created seconds before an alert arrives holds it, on any replica.

Incident grouping. With active grouping an incident escalates once, through its first alert. When a window holds that first alert before its escalation started, a later alert in the same incident that the window does not cover escalates on its own rather than waiting for the window to end; a later alert the window covers is held like any other. If the first alert's escalation had already started and the window only holds its next step, later alerts stay part of that escalation as usual.

That can mean two escalations for one incident: the uncovered alert's while the window runs, and the held first alert's when the window ends if it is still firing. Both are stated — the held alert's timeline reads Maintenance ended: paging resumed — another alert in this incident has already escalated, and the Maintenance ended notice adds (the incident has already escalated for …) on that alert's line. "Escalated" means the escalation started; it does not say whether anyone was reached.

Readiness-drill pages are never held. The readiness drill pages when on-call itself is unmanned; a window — even a whole-client one — must not mute the page that says nobody would be reached.

When the window ends​

Within 30 seconds of the end (early or on time), Console releases what it held. It acts first and clears second: the hold is cleared only after the page went out, so a crash in between re-runs the release rather than losing the page.

  • an alert still firing is paged at once, exactly as if it had just fired — routing, oncall_enabled and the team's live flag all apply, so an alert with no matching route is still not_routed;
  • an alert that was already escalating resumes at the step it was held on;
  • an alert that recovered during the window is simply left alone;
  • an alert that is still covered by another window (overlapping, or extended past this one) stays held and moves to that window — it pages when that window ends, and its timeline names it;
  • an alert silenced during the window before it ever paged stays held until the silence ends, then pages as above (a silence's own expiry never pages an alert that was never armed);
  • if the release runs more than 24 hours after the window ended (Console was down) — or, for an alert held through a silence, after the silence ended — held alerts are resolved as maintenance_stale rather than paging about something a day old;
  • if the database cannot be read when an occurrence is released, nothing is paged that pass and the release is retried on the next 30-second pass — a 30-second delay is better than a burst of individual pages on a database blip.

Burst rule: when more than 10 held alerts are still firing when one occurrence ends, Console pages once — the most severe alert among those that would actually page — and lists the rest on that page as and N more still firing. The others record Not paged: folded into the maintenance release page. A folded alert is paged on its own if it is still firing and unacknowledged 30 minutes after the paging alert was acknowledged, resolved or silenced — or 30 minutes after the release if the paging alert never actually paged (for example it was not_routed), so a fold never waits on a page nobody will answer. The decision is taken once per occurrence and stored, so replicas and retries follow the same one.

Ending a window early (End now), shortening it or deleting it releases held escalations immediately.

Uptime monitors​

A covered monitor moves to Maintenance whatever its locations say, and stays there while covered. It keeps checking: results are recorded and latency charts are unaffected, but checks during maintenance are written to probe_success_maintenance instead of probe_success, so uptime_percent, total_checks and the availability grid leave them out. The grid draws a bucket whose every check was a maintenance check in the maintenance colour, and the availability table gains a Maintenance column (maintenance time is neither downtime nor measured time). The list row reads Maintenance · until 03:00 · still checking (failing); the failing / up part is a display-only approximation of the latest checks and pages nothing.

When coverage ends, the monitor leaves Maintenance through the normal evaluation of its current votes: up if the locations agree it is up, pending if they are failing, unknown if nothing is decided. Failures during the window count toward the monitor's retries, so a target that is still broken usually confirms down — and pages — on the next result after the window, not after a fresh round of retries. A paused monitor stays paused.

A monitor that was already down when the window started has its alert resolved on entering Maintenance (its events read down → maintenance, never down → up); if it is still broken after the window, it pages afresh.

A push monitor enters and leaves Maintenance at its next evaluation — when a heartbeat arrives or the missed-heartbeat check finds it overdue. Its pages are held for the whole window either way.

During a window​

The window page shows the description, the schedule, Post an update, the updates timeline, and Held during this window (each held alert and whether it recovered). Updates can be posted only while an occurrence is in progress and are also sent to the on-call Telegram chat. Extend 1h moves the current occurrence's end; End now ends it.

Telegram notices​

To the client's internal alert chats (where alert cards go), once per occurrence:

  • 🔧 Maintenance started: name, until 03:00 Asia/Tashkent · covers 2 hosts, 4 monitors.
  • 🔧 Maintenance update · name — each posted update, with its author and time.
  • ✅ Maintenance ended: name · 2 recovered · 3 still firing (1 paging now).

The ended notice waits until the occurrence's release has run, so it can say truthfully what is paging. It lists each still-firing alert (up to 20, then and N more linking to the window page) as paging, in the page for the burst's paging alert, or not paged (for example no matching route); an early end ("ended early") or an extension ("extended to 04:00") is said as such. A client with no internal alert chat gets no notices. These notices are for the team; what the client's chats are told is opt-in, below.

Service desk ticket and client chat​

The schedule form's Notify section has two independent options. Both are off by default, both can be changed when editing, and monitor shortcut windows have neither.

Create service desk ticket

  • Opens one SUP Task for the window, raised on behalf of the person who scheduled it.
    • Summary: "Плановые работы: title".
    • Maintenance Window = Yes.
    • Start date: the first occurrence (for a recurring window, the running or next one).
    • Due date: the end of a one-off window; none for a recurring one.
    • Description: the period with its time zone and the window's description (up to 500 characters) — the same as the client chat message. The client sees the summary and description on their portal, so keep host names out of the description.
    • The Console link to the window goes in an internal comment, which the client does not see.
  • Later changes become internal comments:
    • A schedule change also updates the dates ("Rescheduled: … Dates updated.").
    • "Extended to 03:30.", "Ended early at 02:41 from Console (End now)." and "Cancelled: the window was deleted in Console." when the window is deleted.
  • Turning the option off stops the comments and leaves the ticket as it is.
  • Console never changes the ticket's status.
  • Needs a service desk organization mapped to the client, and servicedesk:write on the client to turn it on.
  • The window page's Ticket line shows the ticket as a chip linking to the portal ("SUP-2418 ↗"), "Creating…", or "Ticket failed" with the reason and Retry (for people with servicedesk:write). A failed ticket that JSM did create still shows and links its key.
  • "A previous attempt may already have created this ticket" means Console asked JSM to create the ticket but never heard back (a timeout, a restart), so the ticket may exist without Console knowing it. Look in SUP. If it is there, add the label the message names (console-src-…) to it, then press Retry: Console adopts it. If it is not, just press Retry. Pressing Retry without adding the label to a ticket that exists creates a second one. After adding the label in SUP, wait about a minute before pressing Retry, so Jira's search can see it.
  • If ticket creation is not set up for this deployment (PROXIMA_TICKETING_REQUEST_TYPE_ID), the option is disabled with "Ticket creation is not configured".

Notify client chat

  • Messages go to the client's bound Telegram chats with audience client and client notices on, in each chat's alerts topic, in the chat's language: Russian for a chat set to Русский, English for every other chat. The language is set on Admin → Telegram Groups.
  • What is sent:
    • an announcement, once per window, when the option is on and the window has an occurrence running or still to come. Turning the option off and on again does not announce it again;
    • a start and an end message for every occurrence (the end message goes out only if that occurrence's start message did, and does not wait for alerts);
    • "rescheduled", once per schedule change made after the announcement;
    • "cancelled" when a window that was announced and still had an occurrence to come is deleted.
  • Extend sends nothing to the client.
  • The messages contain only the title, the period, the time zone and the description (up to 500 characters). They never name hosts, monitors, alerts or what the window covers, so keep host names out of the description.
  • The window page's Client chat line shows the last message sent and what is next, or why a message failed. A failed message is retried automatically (every 30 seconds).

Saving a window never waits for JSM or Telegram. A ticket or message that cannot go out is shown on the window page and leaves the window saved.

Per-monitor daily window​

A monitor's edit form has Advanced → Maintenance: "Maintenance window between [HH:MM] and [HH:MM]", a time zone and Mon–Sun. It creates one recurring window that covers only that monitor; it is listed on the Maintenance page like any other and editable from either place. The section also lists every other window that covers the monitor (through its host, environment, client or a label rule). Without maintenance:write it is read-only.

Permissions​

PermissionSeeded toAllows
maintenance:readevery role with alerts:read (custom roles included)seeing windows, their updates and what they held
maintenance:writeSuper Admin, Admin, Team Leadcreating, editing, extending, ending, deleting, posting updates

Both are checked on every environment a window names or reaches through a host, monitor or service. A whole-client target or any label rule can reach any environment, so it needs the permission client-wide. A window you cannot read in full is not shown at all — it is a 404, never a 403, so its existence does not leak.

Turning Create service desk ticket on (when creating, or later when editing) also needs servicedesk:write on the client (403 servicedesk_write_required). Keeping an existing ticket while editing does not, and turning the option off never does. Notify client chat needs nothing beyond maintenance:write. Retrying a failed ticket needs servicedesk:write and read access to the window.

API​

MethodPath
GET / POST/api/v1/clients/{id}/maintenance-windowslist (?from&to adds occurrences, ≤ 62 days), create
GET / PATCH / DELETE/api/v1/clients/{id}/maintenance-windows/{windowID}read, replace, soft-delete
POST…/{windowID}/end, …/{windowID}/extend {minutes}end now, extend (409 not_in_occurrence outside one)
GET / POST…/{windowID}/updatesprogress updates
POST/api/v1/clients/{id}/maintenance-windows/previewwhat a draft covers now
GET/api/v1/maintenance-windows?active=truewindows in progress across your clients
GET/api/v1/monitors/{monitorID}/maintenancea monitor's shortcut window and "also covered by"
GET/api/v1/clients/{id}/maintenance-windows/client-chatshow many chats would get client notices (maintenance:write)
GET/api/v1/clients/{id}/ticketing/availabilitycan Console open tickets for this client (servicedesk:read, or maintenance:write with ?source_kind=maintenance_window)
POST/api/v1/ticketing/maintenance_window/{windowID}/retryretry a failed ticket (servicedesk:write)

A window's GET carries open_ticket, notify_client, ticket (null when none) and client_notice (null before the first message).

Troubleshooting​

  • Rolling back migration 000251 — first end every active window and wait one release sweep (30 seconds), then check SELECT count(*) FROM alert_groups WHERE maintenance_held_at IS NOT NULL is 0. The down migration drops the held columns, so a group still held at that moment is never paged. An alert silenced before it ever paged stays held after its window ends, until its silence ends — so before rolling back, unsilence those alerts (or wait for their silences to end) and let one more sweep run.

  • "Why didn't I get paged?" — open the alert: Not paged: in maintenance names the window.

  • "The window ended and nothing paged" — the alert recovered during the window, or the release found no matching route (not_routed on the timeline), or paging is off for the client or team, exactly as for any other alert. If the release ran more than 24 hours late, the alert was resolved maintenance_stale.

  • "I was paged during a window" — check that the alert really is covered (a host target does not hold an alert filed under a different environment, see above), then proxima_maintenance_gate_errors_total{at}: when a gate cannot read coverage or cannot record the hold it fails open and the page goes out, by design. at="release" counts a release decision that could not be taken; those alerts stay held and are retried next pass.

  • Metrics: proxima_page_suppressed_total{reason="in_maintenance"} and {reason="folded_into_maintenance_release"}; proxima_probe_page_suppressed_total{reason="in_maintenance"} and proxima_monitor_group_page_suppressed_total{reason="in_maintenance"} for monitors and services.