Service Upgrades
A service upgrade upgrades one apt package, and restarts its systemd unit, on every host of an environment that has the package installed. Console dry-runs the change on each host first, shows you every package the change would touch, and then upgrades one host (the canary) before moving on in waves. Each host is checked with readings taken only after its upgrade. A host that fails is rolled back automatically, and the upgrade pauses until a person looks at it.
This page describes the first slice (M1), the production approvals added to it (M2a), judging by monitors and alerts (M2b-1), and restarting the other services that use the upgraded files (M2b-2). It is deliberately narrow: it only runs where you have said it may, and it refuses anything it cannot undo.
Two independent conditions have to hold before Console changes anything on a host:
- The host opts in. Its agent must run with
PROXIMA_AGENT_UPGRADES_ENABLED=true. Every other host refuses and is shown as blocked. - A production environment needs a second person. Every environment is production until someone with
clients:writeon the whole project marks it otherwise. A production upgrade starts only after a second person approves it, or when an approved standing rule covers it. See Production: approval and standing rules.
What it does and does not do
| It does | It does not (yet) |
|---|---|
| Upgrade one apt package to an exact version on Debian-family hosts | dnf/yum hosts: the agent refuses with package_manager_unsupported |
| Dry-run every host and store the transaction apt would run | Upgrades that remove a package, or involve another architecture's packages (libfoo:i386): refused at the dry run |
| Canary first, then waves of a configurable size | Maintenance windows: a standing rule does not restrict when an upgrade starts (M2d) |
| Production upgrades after a second person's approval, or under a standing rule (M2a) | An AI review of the plan (M2c), or an approver signing a pre-flight override (M2b) |
| Verify each host after the change (unit, version, your HTTP/TCP checks, the host's uptime monitors) and soak it, watching the host's alerts (M2b) | Silence the alerts the upgrade's own restart may raise: that needs maintenance windows (M2d) |
| Roll back a failed host automatically and pause | Paging on rollback_failed: M1 records it on the timeline and pauses, but pages nobody |
| Predict which other running services use the files the upgrade replaces, restart the ones you pick, and report which services on the host still use replaced files (M2b-2) | Restart anything you did not pick, or anything on the never-restart list |
| Record every step on an append-only timeline and in the audit log | Snapshots, docker images, GitOps targets, live logs |
Opting a host in
Every host agent and k8s-node-monitor agent answers upgrade commands (a probe agent and the collector never register the handler; see below), but it acts only when both of these hold:
PROXIMA_AGENT_UPGRADES_ENABLEDis exactlytrueor1. Any other value, includingTRUEoryes, leaves it off. It is an environment variable only (the agent reads no YAML for it) and is read at startup, so it takes an agent restart.- The agent is a host agent (
PROXIMA_AGENT_TYPE=host). Other agents run in a container or against other systems, where dpkg would upgrade the wrong machine, so they never upgrade anything, whatever the variable says. Ak8s-node-monitoragent answers every upgrade command withupgrades_disabled. Aprobeagent and thecollector(cluster agent) do not handle upgrade commands at all, so they never answer: their preflight times out and the host is blocked withno_preflight_response.
The agent logs service upgrades at startup with enabled, opted_in and agent_type, so you can confirm the setting took. A host that has not opted in answers every upgrade command with upgrades_disabled and is shown as blocked with that reason; nothing on it changes.
An agent too old to know the upgrade command never answers at all. Its preflight times out and it is blocked with no_preflight_response.
Production and non-production environments
Every environment carries a production marker, is_production. It defaults to true, including for every environment that existed before this feature: an environment is treated as non-production only because a person said so, never by accident. A production upgrade is drafted and dry-run like any other, but it needs a second person's approval or a standing rule to start (see Production: approval and standing rules). The marker is read again at start and at resume, not only when the upgrade was drafted, so marking an environment production again means an upgrade drafted before needs approval to start, and a second person to resume. It does not stop an upgrade that is already running: the worker keeps driving it. Pause that upgrade yourself. Marking an environment non-production leaves an approved upgrade unable to start (the approval is for production only): run its dry run again and start it from ready.
To mark an environment non-production:
- Open the project, edit the environment, and untick Production environment.
- Confirm the dialog.
The change needs clients:write held on the whole project. A grant scoped to one environment is not enough, even if its role holds clients:write: otherwise an administrator of a staging environment could mark the project's production environment non-production. The checkbox is disabled for users with no clients:write on the project. A user whose clients:write is scoped to one environment still sees it enabled, but saving the change is refused with 403. Every change of the marker is written to the audit log (environment.production_marker_changed, with from and to). The entry is written after the change is saved: if writing it fails, the change stands and the failure is logged as an error. (A backend with no audit log configured at all refuses marker changes.)
Through the API it is the is_production field of PUT /api/v1/environments/{envID}. Omitting the field leaves the marker unchanged.
Permissions
| Permission | Allows |
|---|---|
upgrades:read | List upgrades and view one, with its hosts and timeline |
upgrades:write | Draft, dry-run, submit for approval, start, pause, resume and cancel |
upgrades:approve | Approve or reject someone else's production upgrade, resume a production upgrade someone else drafted, and create, approve and revoke standing rules |
Every upgrade route also needs upgrades:read, which gates the whole /api/v1/upgrades group, so the permissions stack: an approver needs upgrades:read and upgrades:approve, and a production resumer needs upgrades:write and upgrades:approve.
upgrades:read and upgrades:write are granted to the Super Admin, Admin, Engineer and Team Lead system roles. upgrades:approve is granted to Super Admin, Admin and Team Lead, not Engineer (migration 000247 adds it to those three roles). All three are checked against the upgrade's own environment, so an environment-scoped grant covers upgrades in that environment only.
The lifecycle
An upgrade moves through draft → dry_running → ready → running → completed, with paused and cancelled on the side. A production upgrade has three more statuses between ready and running: pending_approval, approved and rejected (see Production: approval and standing rules). Each host (a target) has its own status.
1. Draft
Pick an environment, a package, the exact version to install and the systemd unit to restart. From version currency, Plan upgrade prefills the package, the project and the environment: it is the button on a product's page, and the Plan upgrade… item in a row's ⋯ menu on the list (offered on rows that are behind, run as an OS package on at least one host, and where you hold upgrades:write). Neither prefills the target version — the catalog's latest release (1.27.3) is not the apt version a host installs (1.26.2-1ubuntu1). Check the package too: a product's name there is not always the apt package name (postgresql vs postgresql-16). You can also set:
- Checks (up to 10): HTTP checks (a URL, an expected status, optionally a string the body must contain) and TCP checks (host and port). Each check needs its own target: two checks on the same URL, or on the same host and port, are refused (combine them into one check), because they would share one reading. See Checks.
- Soak: how long a verified host is watched before it counts as healthy, 60 to 3600 seconds (default 600).
- Wave size: hosts per wave after the canary, 1 to 50 (default 2).
The targets are the hosts of the environment whose inventory lists the package and that have an agent.
2. Dry run
The dry run asks every host's agent to preflight the change without making it. The agent checks that the package is cleanly installed and the unit is active, runs your operator checks, simulates the install, and downloads and verifies the currently installed .deb of every package the change touches, so that a rollback will have them. The transaction it would run is stored on the target. A host that cannot be upgraded safely is blocked with a reason; the others go back to pending with their transaction. A host where one of your checks already fails is blocked with checks_failing_before_upgrade: verify would fail on that check after the upgrade and roll back an upgrade that caused nothing, so fix the check or the service first. A failing check is run once more, about 2 seconds later, before it counts, so one blip does not block a host. Pre-flight runs your checks one after another and retries each failing one once, so with many slow checks it can take a few minutes: the worst case, 10 HTTP checks that all time out, is about 220 seconds. Keep PROXIMA_UPGRADE_STEP_TIMEOUT comfortably above that, or a slow pre-flight is blocked with no_preflight_response. A check that tests something only the new version has (a body that names the new version, an endpoint the new release adds) cannot pass before the upgrade: mark it Only expected to pass after the upgrade (after_upgrade_only: true in the API). Preflight still runs it but does not block on it, and verify still requires it. (An agent too old to run checks at preflight returns no check readings; the host is then not blocked for them, and a failing check is found only at verify.)
When every host has answered (or timed out), the upgrade is ready. Running the dry run again re-checks every pending and blocked host, so a host that has opted in since is picked up.
3. Start
On a production environment, see Production: approval and standing rules first: Start runs only on an approved upgrade, or on a ready one that an active standing rule covers.
Review the transactions. If any host's transaction changes packages other than the one you are upgrading, starting requires you to acknowledge them (the API answers 422 extra_packages_not_acknowledged with the list until you send acknowledge_extra_packages: true). An upgrade with every host blocked cannot start.
4. Canary and waves
The first host is the canary, alone in wave 0. Each later wave has wave size hosts. A wave starts only when every host of the previous wave has finished, and all hosts of a wave start together.
Until some host of the upgrade has become healthy, Console starts only one host at a time. If the canary is blocked (at the dry run or when it is preflighted again at start), the next pending host becomes the canary, alone, rather than its whole wave starting without a proven host. Once any host is healthy, waves start whole.
For each host:
- Preflight again. If the transaction apt would run now differs from the one you reviewed, the host is blocked with
transaction_changed, and if one of your checks now fails, withchecks_failing_before_upgrade. Either way nothing changes. - Apply. The agent checks the transaction a last time, keeps and re-verifies the old
.debfiles, backs up the configuration files of every package in the transaction, installs, and restarts the unit. - Verify. Readings are taken after the apply until the host passes, fails or times out.
- Soak. A verified host is watched for the soak period and verified once more at the end, with readings sampled after the soak ended. Then it is healthy. If that final verify has not passed within one step timeout after the soak ended, the host becomes unknown and the upgrade pauses, as for any other timeout.
When every host is finished and nothing failed, the upgrade is completed.
What "verified" means
A host passes only when every one of these is true of readings sampled after the apply finished. A reading from before the apply can never make a host pass.
- The unit is active, and has been active for at least 30 seconds (stable). An active but not yet stable unit makes Console wait, not fail.
- The package's installed version is the target version.
- Every operator check you added passes, including those marked only after the upgrade. Each one needs a fresh reading, so a check that has not been read yet makes Console wait. When an HTTP check gets the expected status but the body lacks your text, its reading says so (
200, body missing "…"). - Every uptime monitor linked to the host is
up, from a probe that ran after the apply, unless you excluded it. See Monitors and alerts during an upgrade. - Every service you picked to restart that was restarted on this host is active, or has stopped cleanly on its own since (a
unit_active:<unit>reading). Unlike the package's own unit, it is not required to have been active for 30 seconds. See Affected services.
While the host is verified and while it soaks, Console also watches the alerts firing on it: a new alert about the upgraded service rolls the host back, and any other new alert pauses the upgrade (see Alerts).
A wrong version, an inactive unit, a failing check or a down monitor means the host fails and is rolled back. How long Console waits is bounded by the step timeout (below).
Operator checks run from the host itself, by its agent, so http://127.0.0.1:8080/health checks the service on that host. Link-local targets (169.254.0.0/16, fe80::/10) are refused, both when you save the check and, for names that resolve to one, by the agent when it connects. Loopback is allowed.
Automatic rollback
A host that fails verification, or whose install or restart fails, is rolled back:
- The old packages are reinstalled from the
.debfiles the agent kept (and verified) before the apply, then any package the upgrade newly installed is removed. A rollback that would need apt to remove anything else fails rather than removing it. - The configuration files backed up before the apply are restored, with their mode and ownership.
- The unit is restarted, then every service picked to restart is restarted again so it loads the old files (all of them, whether or not the apply got to them; see Rollback and affected services), and the host is verified again, against the old version.
The rollback is judged only on what it controls: the unit is active and stable, and the installed version is the old one again (readings sampled after the rollback finished). Uptime monitors and alerts are not part of this verdict either. Your operator checks are still read but are not part of this verdict: preflight checks them when the agent can, so a check that still fails after a correct rollback was most likely not broken by the rollback. If the built-ins pass, the host is rolled_back; a check still failing is written to its last error and the timeline ("rolled back to X; check Y still failing after rollback: …") for a person to look at. A check marked only after the upgrade is expected to fail on the old version and is not noted. The host is rollback_failed, and needs a person, only when the rollback itself failed, the old version was not restored ("rollback did not restore X: …"), or the unit is not active ("rollback left the unit inactive: …"). Either way the upgrade pauses; hosts that are already healthy stay upgraded.
Monitors and alerts during an upgrade
An upgrade is judged by the host's own uptime monitors and alerts as well as by its operator checks, with no setup on the upgrade. This section says which monitors count, when they block or fail a host, and what a new alert does. For what a monitor's state means, see Uptime monitors.
Which monitors count
Every enabled uptime monitor that is linked to the host (the monitor's host, monitors.host_id) in the upgrade's environment counts. A monitor linked to no host, or to another host, does not. A disabled monitor does not.
The dry run records the monitors it found on each host, with their state then. The run page shows them in the Monitors panel, and the timeline has a monitors_snapshot row for each host whose monitors it read. A host the agent's own pre-flight already refused (for example upgrades_disabled or unhealthy_before_upgrade) is blocked before its monitors are read, so it has no snapshot. So the approver of a production upgrade sees which monitors will judge it.
The dry run's list is what you review. It does not decide what is judged. Verify judges the monitors that are enabled and linked to the host at verify time, minus your exclusions, so the plan can only get stricter:
- a monitor added to the host after the dry run is judged too;
- a monitor disabled or unlinked after the dry run is dropped.
Either change is written to the timeline (monitor_set_changed), normally once per host and attempt. Each backend replica remembers what it wrote only in memory, so a restart or a second replica may write the row once more.
Before the upgrade: the dry run and the pre-flight
At the dry run, and again at the pre-flight just before a host's apply, a host whose agent passed its own pre-flight (see Blocked reasons for what the agent refuses) is blocked by the first of these that applies, in this order. Nothing is changed on a blocked host.
checks_failing_before_upgrade: one of your operator checks already fails (see Dry run).monitors_unreadable: Console could not read the host's monitors or the upgrade's exclusions. It never proceeds without them. Run the dry run again.monitor_cannot_verify_in_time: a monitor could not give a fresh verdict within the step timeout, so verify would be certain to time out. That is a push (heartbeat) monitor, which reports only when its job runs, or a monitor whose check interval plus 30 seconds reaches the step timeout. With the default 10-minute step timeout, that is an interval of 570 seconds or more. Exclude the monitor with a reason, shorten its interval, or raisePROXIMA_UPGRADE_STEP_TIMEOUT.monitors_failing_before_upgrade: a monitor is notup, or isupbut has not been probed recently. Not up includesdown, and alsopending,unknown,pausedandmaintenance. Not probed recently means it has no result at all, or its last result is older than its interval plus 30 seconds: a monitor whose probers stopped keeps showingupfrom its last result, and that says nothing about now. Verify would fail on it, or wait on it, and roll back or time out an upgrade that caused nothing. Fix the service or the monitor, or exclude the monitor with a reason, then dry-run again. The host's last error names the monitors and their states, and for one not probed recently, how long ago its last result was.
An excluded monitor is skipped by rules 3 and 4.
At the pre-flight, once these gates and the later ones pass (no_dry_run, transaction_changed, host_moved_environment), and just before the host is changed, Console reads the host's firing alerts to record the alert baseline. If it cannot read them, the host is blocked with monitors_unreadable too. So a host with a monitor that is not up and unreadable alerts is blocked monitors_failing_before_upgrade.
Verify and soak
Each monitor that counts is judged by its current state. That state is the one the monitor already confirmed across its probe locations (its quorum), so a single probe's blip does not decide it.
| Monitor state | At verify and soak |
|---|---|
up | Passes. |
down | Fails: the host rolls back. Its last error names the monitor (monitor <name>). |
pending, unknown, paused, maintenance | No verdict yet: Console waits. |
A verdict counts only if it comes from a probe that ran after the apply finished, and at the end of the soak only from one that ran after the soak ended. A monitor that has not been probed since then makes Console wait, as a check that has not been read yet does.
The wait is bounded by the step timeout (PROXIMA_UPGRADE_STEP_TIMEOUT, default 10 minutes). A monitor that is still waiting when it runs out leaves the host unknown, and the upgrade pauses (see Unknown). A monitor that never settles is never taken as healthy and never rolls the host back on its own.
A down verdict at any time during the soak fails the host, as a failing check does.
If Console cannot read the host's monitors or alerts on a verify round, that round makes no decision and Console tries again on the next one. The step timeout bounds that too.
Alerts
The baseline. When a host moves to applying, Console records the alert groups firing on it at that moment (timeline: alert_baseline). Those groups never fail or pause this upgrade, even if they keep firing, or fire again into the same group. An alert group counts as new only if it is not in the baseline and first fired at or after the moment the host moved to applying. Only alerts on the upgraded host count: an alert on another host does nothing to this host's upgrade.
On every verify round, while the host is verified and while it soaks, Console looks at the new alert groups on the host. There are two kinds.
A service alert rolls the host back. A new alert group is about the upgraded service when either of these holds:
- it is a probe alert from an uptime monitor linked to the host (its
monitor_idlabel names one), including a monitor you excluded: an excluded monitor is not judged, but its alert is still about this service; - one of its
service,job,unitorsystemd_unitlabels equals the upgrade's unit name or its package name, or the name of a service this upgrade restarted on this host for you (see Affected services). The comparison ignores case, surrounding spaces, and a.servicesuffix on the label or on the unit name.
The host fails and is rolled back. Its last error, and a service_alert timeline row, say service alert fired: <alertname>.
For an upgrade of package nginx-core with unit nginx.service:
| Alert labels | Result |
|---|---|
service="nginx" | Service alert: rolls back |
unit="Nginx.service" | Service alert: rolls back |
job="NGINX" | Service alert: rolls back |
systemd_unit="nginx-core" | Service alert (the package name): rolls back |
monitor_id="<a monitor linked to this host>" | Service alert: rolls back |
job="node", or no such label at all | Any other alert: pauses |
With package nginx and the same unit, service="nginx-core" is not a service alert, and it pauses.
Any other new alert pauses the whole upgrade. Console cannot tell whether the upgrade caused it, so it stops starting new hosts and asks a person:
- The upgrade moves to
pausedwith the reasonnew_alert_on_host: <alertname> on <host>, where<alertname>is the group'salertnamelabel (its id when it has none). The run page's banner says a new alert fired and that the host kept soaking. The timeline has anew_alert_pausedrow. - The host is not rolled back. It keeps being verified and soaking, and it can still end healthy, or fail and roll back, on its own evidence.
- The alert joins that host's baseline, so it never pauses this upgrade again, not on the next round and not after a resume. A new alert while the upgrade is already paused, by a person or by an earlier alert, does not pause it a second time. It joins the baseline too, and the host's timeline gets a
new_alert_while_pausedrow naming the alert: "New alert<alertname>fired while the upgrade was paused; check it before resuming." The pause banner does not change, so it still names only the first alert (or none, for a pause by hand). Read the hosts' timelines, and look at the hosts' alerts, before you resume. - Resume clears the reason. On a production environment, resuming needs a second person, as always (see Resuming a production upgrade).
A host with no baseline. A host that reached verify without a baseline (for example, a backend restart at the wrong moment) gets one on its first verify round, holding everything firing then, and no alert is judged on that round. The timeline row says the baseline was recorded late. Console never judges alerts against a missing baseline, under which every firing alert would look new.
Excluding a monitor
A monitor that cannot judge this upgrade fairly can be taken out of it, with a reason. For example, a monitor that checks an endpoint the new release removes, or a slow heartbeat. On the run page, use Exclude… on the monitor in the Monitors panel. Include again takes it back.
- Only while the upgrade is
draftorready. Once approval is requested, or the upgrade starts, the exclusions are fixed (409illegal_transition), so an approver always sees the exclusions they approve. A new dry run brings the upgrade back toready, where they can be changed again; on production that also clears the approval, as any new dry run does. - The reason is 1 to 500 characters on one line. Console records who excluded the monitor and when, and writes every change to the audit log (
upgrade.monitor_exclusions_set). The approver sees each excluded monitor and its reason. - A monitor can be newly excluded only if the last dry run saw it on one of the upgrade's hosts (400
unknown_monitorotherwise). An existing exclusion whose monitor is no longer linked may stay, or be removed; it excludes nothing. The panel marks it "no longer linked to these hosts". - An excluded monitor is not gated before the upgrade and not judged after it. Its alerts still count as service alerts (see Alerts).
- A standing rule never covers an upgrade with an exclusion. Excluding a monitor loosens verification, so a person must approve it. Start answers 422
approval_required; indetails.reasons, a rule that lists the upgrade's environment and package says the plan excludes monitors. If a monitor is excluded in the moment between the rule check and the start, the start is refused with 409monitor_exclusions_changed. Reload the page and submit the upgrade for approval.
A restart may page on-call
An upgrade silences nothing. Restarting the unit can make a monitor go down, or fire an alert, for as long as the restart takes. That alert pages on-call exactly as it would at any other time, through the usual escalation routes. Silencing an upgrade's expected alerts needs maintenance windows (M2d). Console's silences today apply only to an alert that has already fired, so there is nothing to pre-arm. Until M2d ships, tell whoever is on call before you start a production upgrade, or run it when a short page is acceptable.
A page and a rollback are separate. The same service alert that pages also rolls the host back, and a paused upgrade does not acknowledge or resolve anything.
Clock skew
A monitor's last result time is on the probe machine's clock, and the apply's finish time is on the host agent's clock. Console compares them as they are, and expects both machines to be NTP-synced. A skew of a few seconds can let a probe that ran just before the apply finished count as after it, or make Console wait one more check interval. It is tolerated and not corrected.
Affected services
A running process keeps the old version of a shared library (or of its own executable) mapped in memory until it restarts. Upgrading libssl3 changes the file on disk, but postfix and nginx keep using the old copy until they restart. The package's own unit is restarted as always. For every other service, Console tells you which ones use the files the upgrade replaces. It restarts the ones you pick, inside the upgrade, and reports afterwards which services on the host still use replaced files.
What is predicted, and how
At the dry run, each host's agent works out which running services use files of the packages in the transaction:
- It lists the files of every package the transaction touches (
dpkg -L). It keeps the shared libraries (.so,.so.1.2, …) and the executables. - It reads every process's memory map (
/proc/<pid>/maps) and keeps the processes that map one of those files. - It names each process's systemd service from its cgroup (
/proc/<pid>/cgroup). A process that is not in a.service(a session scope, for example) is not listed. A process of a user's own service manager (user@<uid>.service) is not listed either: a system restart cannot reach it.
The prediction always comes from /proc, on every host, even where needrestart is installed: needrestart only sees files that have already been replaced, so it cannot predict anything before the install.
The dry run of each host stores what it found:
affected_detection:proc_maps(the scan worked), orunavailable.affected_services: for each service, itsunit, up to 20 of itspids, up to 5 of the matchedfiles, and whether it isrestartable. At most 50 services per host are reported, restartable ones first, each group sorted by name. The panel says when a list has reached that cap.
unavailable means "couldn't tell", never "nothing is affected". The agent reports it when it cannot read /proc, when no process of any .service other than the agent itself is visible (an agent in a container without the host's PID namespace, or a /proc mounted with hidepid), or when dpkg could list none of the packages. Unreadable single processes are skipped. A scan problem never fails the dry run.
A dry run with no affected_detection at all came from an agent that predates this feature. See Rollout order.
Picking what to restart
The run page has an Affected services panel, under the Monitors panel. After the dry run it lists every service you can pick, with the number of hosts on which it can be restarted (postfix.service — 3 hosts), and, per host, everything the dry run found. With more than 3 hosts the per-host detail is collapsed behind a summary line.
- Nothing is ticked by default. A service you do not pick is not restarted. It keeps using the old files and will pick up the change at its next restart. The approver sees both lists, what will restart and what will not.
- A service can be picked if at least one host that will run (not blocked, not cancelled) predicted it as restartable. A service the dry run lists as not restartable is greyed out with "Console never restarts this service".
- At most 20 services. They are restarted in the order you picked them.
- Only while the upgrade is
draftorready, withupgrades:write. Once approval is requested, or the upgrade starts, the picks are fixed (409illegal_transition), so an approver always sees the restarts they approve. Every save is written to the audit log (upgrade.restart_units_set) and the timeline (restart_units_set). - A new dry run clears the approval, as always (see A new dry run clears the approval), and keeps the picks. A pick that the new dry run no longer predicts as restartable on any host is marked No longer predicted: untick it and save. Saving a list that still contains it is refused with 400
unknown_unit(orunit_not_restartable, when the new dry run lists it only as not restartable). - Each host restarts only what its own dry run predicted. A pick that a host's stored dry run (the one you reviewed, not a newer preflight) did not list as restartable is skipped on that host only, and the host's timeline gets a
restart_units_skippedrow naming it. - A standing rule never covers an upgrade with restart picks. Restarting other services goes beyond the package's own unit, so a person must approve it. Start answers 422
approval_required, and the rule's reason says "the plan restarts other services (…); a person must approve that". If services are picked in the moment between the rule check and the start, the start is refused with 409restart_units_changed. An upgrade with no picks can still be covered: the report alone changes nothing.
Saving the picks (PUT /api/v1/upgrades/{id}/restart-units) refuses:
| Status | Code | When |
|---|---|---|
| 400 | bad_request | The body is not valid JSON, or not the expected shape (for example, units is not a list of strings). |
| 400 | invalid_input | The units field is missing, there are more than 20, one is listed twice, or one is not a plain unit name. A systemd-escaped name (openvpn-client@my\x2dvpn.service) is shown in the dry run but can never be picked. |
| 400 | unit_not_restartable | The service is on the never-restart list, or the dry runs reported it only as not restartable. |
| 400 | unknown_unit | No dry run of a host that will run predicted the service. |
| 409 | illegal_transition, status_changed | The upgrade is not draft or ready, or it moved on while you were saving. |
What is never restarted
Whatever you pick, Console never restarts:
- the agent's own unit, and any
proxima-agent*unit; systemd-*anddbus*units, andinit.scope;user@*units (a user's whole service manager);- anything that is not a
.service; - a systemd-escaped name, which is shown but never pickable.
The backend refuses these when you save the picks (400 unit_not_restartable, or invalid_input for an escaped name). The agent refuses them again, before it changes anything: its apply answers unit_not_restartable and the host is blocked, with nothing installed and no rollback.
During the upgrade
For each host, the apply:
- records the picks beside the transaction, before installing anything (so a rollback still knows them after an agent restart);
- installs the package and restarts its own unit, as before;
- restarts each picked service, in order. The first failure stops the apply with
affected_restart_failed, and the host is rolled back; - checks which services on the host still use replaced files (see The still-affected report).
How long it may take. The agent bounds every systemctl restart at 5 minutes: each picked service, and the package's own unit too, in the apply and in the rollback. A restart that takes longer counts as failed. The agent also ends the whole apply (and the whole rollback) after 30 minutes, its limit for one operation. On the backend, a host in applying or rolling back has the step timeout, plus 5 minutes for each picked service sent to that host, plus 2 minutes for the still-affected check, before it times out to unknown; at most 30 minutes plus the step timeout. With the default 10-minute step timeout, a host restarting 3 picked services has 27 minutes, and one with no picks 12. A pick skipped on a host adds nothing there.
Behaviour change: the package's own unit is bounded too. Before this release, its restart had no bound of its own, only the 30-minute operation limit. Now a unit that takes more than 5 minutes to restart (a database doing crash recovery, say) fails the apply with restart_failed, and the host is rolled back; in a rollback, the same makes the host rollback_failed.
Verify, and the final verify at the end of the soak, then require every restarted service to be up: active, or stopped cleanly on its own. A restarted service that later stops cleanly on its own, such as an on-demand service exiting when idle (fwupd, packagekit), is not a failure; one that fails is. Precisely, the unit_active:<unit> reading passes when systemd reports ActiveState=active, or ActiveState=inactive with Result=success; it fails on failed, on any other Result, on any other state (activating, reloading), and on a unit systemd no longer knows. Its value shows what systemd reported (ActiveState=inactive, Result=success), so an idle exit and a crash read differently. A failing reading fails the host, and it is rolled back. The package's own unit, picked or not, must still be active.
An apply that skipped the restarts is rolled back. If the agent reports success but did not restart a service it was sent (for example, an agent downgraded between the dry run and the apply), the host is rolled back and its last error starts with restart_units_ignored:, naming the services.
Alerts on a restarted service
A service this upgrade restarted on a host gets the new library, so a new alert about it is treated like one about the package's own unit: it rolls that host back instead of pausing the upgrade. Alerts on services that were not restarted still pause, as before (see Alerts).
The match uses the same labels as for the package's own unit: service, job, unit and systemd_unit. That includes job. If you restart postfix.service and a scrape job is also named postfix, a new alert from that job on the host rolls the host back.
Rollback and affected services
A rollback reinstalls the old packages, restores the configuration files and restarts the package's own unit, as before. Then it restarts every picked service again, so they load the old files: every pick the apply recorded, including the ones the apply never reached (those after a failed restart, or all of them when the install or the package's own restart failed). Restarting a service that was never restarted only reloads the old files it already used. The agent reads the picks it recorded at the apply, not the command: the backend does not send them again. The picks are restarted even when the agent itself restarted in between.
- A rollback restarts every picked service it can, even after one fails. If any restart fails, or the recorded picks cannot be read, the rollback answers
restart_failedand the host is rollback_failed: the old packages are back, but some services may still use the new files, so a person needs to look. The run page says so in those words, distinct from the package's own unit failing to restart. The rollback's timeline row carries the services that did restart (restarted_units). - If the package's own unit fails to restart, no picked service is restarted.
- The rollback verdict does not look at the picked services, only at the package's own unit and version (see Automatic rollback).
The still-affected report
After an apply that succeeded, the agent looks again at which services on the host still use replaced files. The check covers the whole host, not just this upgrade's files: a service can be listed because it still uses a library an earlier upgrade replaced.
- with needrestart, where it is installed (
needrestart -b -r l); - otherwise from
/proc: processes that still map a deleted shared library (.so*, shown by the kernel as(deleted)); - or
unavailablewhen neither could tell.
Console stores the report on the host (affected_report: detection, restarted_units, still_affected) and writes an affected_services timeline row, for example "restarted postfix.service, sshd.service; 1 service on this host still uses replaced files: app-api.service". restarted_units keeps only the services this host was sent (at most 20); anything else the agent names is left out. On the run page it is the host's After apply group: what was restarted, and the services on this host still using replaced files, which will pick up the change at their next restart. The summary line above the per-host detail, which stays visible when the detail is collapsed, also counts the hosts still using replaced files and the restart failures. A host with a rollback after the apply (rolling back, rolled back, rollback failed, or unknown after a rollback) keeps its report, marked as what it reported right after the apply. Only an apply that succeeded stores a report; after affected_restart_failed, the services that did restart are in the timeline row.
needrestart's results carry no PIDs or files. The panel says "PIDs not reported" and "files not reported", not "none".
Detection limits
- The
/procfallback only flags deleted shared libraries (.so*). It skips transient mappings that no upgrade explains: files under/tmp,/var/tmp,/runand/dev/shm(and/dev), and the same scratch-file patterns needrestart ignores. - A service whose replaced file is its own executable is caught after the apply only when needrestart is installed. The
/procfallback does not look at executables. (The prediction does list services that run one of the package's executables.) - needrestart results carry no PIDs or files.
- Output is bounded. The agent reads at most 4 MiB from needrestart and from each
dpkg -L. Past that, needrestart counts as failed (the/procscan answers), and that package counts as not listed. - Statically linked binaries never appear. They map no shared library.
- Containers. A container's own libraries are not covered: they are not the host package's files. Matching is by path, so a containerised process that maps a file at the same path as a host package's file (its image's own
libssl.so.3) can be counted as using the host package's file. - An agent running in a container gets an incomplete view. dpkg and needrestart run where the agent runs, and without the host's PID namespace it sees few or no host processes. With none visible, it reports
unavailable.
Rollout order
The backend ships first; the agent follows.
- Deploy the backend. It accepts results without the new fields.
- Release an agent with detection and restarts, and roll it out to the fleet (
pc rollout create).
Until a host's agent is updated, its dry run carries no affected_detection. The panel says "This host's agent predates affected-service detection; update it to use restarts". An upgrade with no picks runs on that host exactly as before. An upgrade with picks blocks that host with agent_cannot_restart_affected before anything is sent, rather than upgrading it without the restarts. The run page asks you to cancel the upgrade, then update the agent and create a new upgrade, or create one without restart picks. A running upgrade never returns to draft or ready, so a new upgrade is the way forward. Nothing breaks in the meantime.
Production: approval and standing rules
On a production environment, an upgrade needs a second person before it starts: either that person approves this upgrade, or they approved, in advance, a standing rule that covers it.
The flow
-
Dry run, exactly as elsewhere. The upgrade reaches
ready. -
Submit for approval (
upgrades:write). The upgrade moves topending_approval. Submitting needs at least one host that is not blocked. -
A second person approves or rejects it, with a note of 1 to 1000 characters. The note is required and is kept on the upgrade and in the audit log.
- Approved: the upgrade moves to
approved. - Rejected: the upgrade moves to
rejected, which is final. To try again, draft a new upgrade.
The decision is bound to the submission the approver reviewed. Approve and reject carry the upgrade's
approval_requested_atas the page showed it (the API answers 400invalid_inputwithout it). If the drafter dry-ran the upgrade again and re-submitted it after the approver opened the page, the plan may have changed, so the decision is refused with 409approval_staleand nothing is recorded. The run page then reloads the new plan and clears the note: review it and decide again. - Approved: the upgrade moves to
-
Anyone with
upgrades:writestarts it, the drafter included. From here it runs like any other upgrade.
The run page has an Approval panel for a production upgrade. It says what comes next; once approved, it says that a second person approved it and shows their note (it does not name the approver); once rejected, it shows the rejection note. The person who can approve sees a note field with Approve and Reject. The drafter sees why they can't. On the Upgrades list, the Waiting for approval (this page) filter shows the pending_approval upgrades among the rows already loaded. It filters the current page only, not every upgrade.
Start on a production upgrade that is ready and has not been approved runs only if an active standing rule covers it (below). Otherwise it answers 422 approval_required. Its details.reasons lists why each active standing rule of the environment did not cover the upgrade. When the environment has no active rule, details carries no reasons.
Who can approve
- The approver needs
upgrades:approveon the upgrade's environment. By default the Super Admin, Admin and Team Lead roles hold it; Engineer does not. - The approver is never the drafter, and super-admins are bound by this too (403
self_approval). The same applies to rejecting: the drafter can't reject their own plan either. - The approver is always a person. A service account can neither approve nor reject (403
service_account_cannot_approve). - A service account can draft, and a draft by a service account has no drafter on record. The drafter rule compares the approver with the person who drafted the upgrade, so it cannot see through a service account. A person who holds
upgrades:approveand also controls a service-account key withupgrades:writecan draft through that key and then approve the upgrade themselves. Keepupgrades:writeservice-account keys away from people who approve production upgrades, or treat an approval of a service-account draft by the key's owner as a self-approval when you review the audit log. - This is enforced in the handler, again in the database update that records the decision, and by
CHECKconstraints on the table. A race between two requests ends in a 409, never in a self-approval. - A project with only one holder of
upgrades:approvecan't approve its own production upgrades, or create a standing rule that anyone can approve. M2a leaves that open. Grant a second person the permission.
A new dry run clears the approval
A dry run is allowed from pending_approval and approved too. It clears the approval, and the upgrade goes back through dry_running to ready, because the new dry run may show a different plan from the one that was approved. Submit it again. Starting an upgrade whose approval a dry run cleared a moment before is a 409, never a start.
Resuming a production upgrade
A production upgrade pauses after a host fails, like any other. Resume then starts new hosts after that failure, so it needs the same kind of second person as an approval. The person who resumes must:
- hold
upgrades:approveon the environment, as well asupgrades:write(resume is anupgrades:writeroute); - not be the drafter;
- not be a service account.
This holds whichever way the upgrade started, by a peer's approval or under a standing rule. Anyone else gets 403 peer_resume_required, and the run page says who can resume.
Hosts that moved environment
An approval, or the lack of a need for one, is given for the upgrade's environment. The hosts of an upgrade are fixed when it is drafted, but a host can move to another environment later, for example when its agent re-registers elsewhere. Before Console preflights a host, in the dry run and again when the host's turn comes, it checks that the host is still in the upgrade's environment. It checks again when that preflight's result arrives, before the host moves to applying and the apply is sent. A host that has moved, or that no longer exists, is blocked with host_moved_environment, and nothing more is sent to it. To upgrade it, draft an upgrade in its current environment. Only a move during the seconds of the apply itself is not checked. The agent's own check that the transaction is unchanged does not cover a move: moving to another environment does not change what apt would install.
Standing rules
A standing rule approves, in advance, routine production upgrades of named packages in named environments. It is proposed by one person and approved by another, once. While it is active, a ready production upgrade that it covers starts on Start with no approval for that upgrade. The upgrade records which rule it started under, and the run page links to the rule. Standing rules are on the Standing rules page (/upgrades/standing-rules, linked from the Upgrades list). A link to /upgrades/standing-rules#<rule-id> scrolls to and highlights that rule.
A rule lists:
- a name;
- 1 to 20 environments, all of one project;
- 1 to 20 exact apt package names;
- the kinds of change it allows,
patch,minoror both; - an expiry.
The expiry is at most 180 days after the rule is created. The API requires it to be at least 1 minute ahead and at most 180 days less 1 minute away. The form keeps a 5-minute margin at both ends: at least 5 minutes ahead and at most 180 days less 5 minutes.
Who can do what:
- Create a rule: a person (never a service account) holding
upgrades:approveon every environment of the rule. - Approve a rule: a different person, holding
upgrades:approveon every environment of the rule, with a note. The creator never approves their own rule, super-admins included (403self_approval). A service account never approves one. - Revoke a rule: any person holding
upgrades:approveon every environment of the rule, with a note. The creator may revoke their own rule, because revoking only ever narrows what is allowed. A revoked rule never covers another start. Revoking does not stop an upgrade that already started under the rule: the worker keeps driving it, wave after wave. To stop one, pause or cancel that upgrade yourself; resuming it then needs a second person, as for any production upgrade. - Revoke a rule whose environment was deleted: someone holding
upgrades:approveon the whole project, or a super-admin. For anyone else, the deleted environment answers 404.
A rule's state is one of:
revoked;expired, once the expiry has passed, including a rule that expired before anyone approved it;pending, not approved yet;active.
The states are checked in that order. Only an active rule covers anything.
What a rule covers. The kind of change is judged on the upstream part of the Debian version, [epoch:]upstream[-revision], for every host that will run:
- patch: the upstream changes only after its second number (
1.24.0→1.24.2); - minor: the second number of the upstream changes (
1.24.2→1.26.0).
A rebuild that changes only the Debian revision counts as a patch. That needs a byte-identical upstream, a revision on both sides, and a new revision that is higher under dpkg's own ordering (2.10-2 → 2.10-3).
A rule never covers any of the following, whatever it says:
- a major change (the first number of the upstream changes);
- an epoch change, a downgrade, or a change between two different upstreams that are not both plain dot-separated numbers (
1.2.3~rc1→1.2.3). Console can't classify these, so it treats them as needing a person; - an upgrade whose dry run, on any host, changes packages other than the one being upgraded;
- a host that needs a snapshot before the upgrade (snapshots arrive in a later milestone; today no host is marked as needing one);
- an operator check other than HTTP or TCP (today drafting accepts only those two kinds, so this guards future check kinds);
- an upgrade that excludes a monitor from verification (see Excluding a monitor);
- an upgrade that restarts other services (any restart picks; see Picking what to restart). Report-only plans, with no picks, can still be covered;
- a package or an environment the rule does not list.
When no rule covers an upgrade, it goes through the approval flow above.
Until maintenance windows ship (M2d), a standing rule does not restrict when an upgrade starts. A covered upgrade can start at any time of day.
Every create, approve and revoke is written to the audit log (upgrade_standing_rule.created, .approved, .revoked), and a start under a rule is also recorded as upgrade.started_under_rule.
Pause, resume and cancel
Pause means no new host starts. Hosts that are already mid-change keep being driven to a safe end: a host being verified keeps being verified, a soaking host keeps soaking, and a host that fails after the pause is still rolled back. An upgrade pauses itself when a host fails, and when a new alert that is not about the upgraded service fires on one of its hosts (see Alerts); you can also pause a running upgrade by hand. Resume clears the reason for such an alert pause.
Resume acknowledges the failed hosts and carries on with the rest. Look at a rollback_failed or unknown host before resuming: resume does not retry it. On a production environment, resuming needs a second person; see Resuming a production upgrade.
Cancel is accepted for a draft, ready, pending_approval, approved or paused upgrade, and cancels its pending hosts. It is refused (409 targets_in_flight) while any host is mid-change (preflight, applying, verifying, soaking or rolling back): a cancelled upgrade is never driven again, so those hosts would be abandoned halfway. Pause it, wait for the in-flight hosts to finish, then cancel.
Unknown
A host is unknown when its agent stopped answering while it was being changed: applying, verifying, soaking (past the soak's end plus the step timeout) or rolling back went past its deadline (the step timeout; for applying and rolling back, plus the time its picked restarts may take, see How long it may take), or the agent reports that an apply was interrupted (for example, the agent restarted mid-install). A host whose uptime monitor never gives a fresh verdict before the step timeout also ends unknown. When the monitors were all it was still waiting on, its last error names them: no fresh verdict from monitor <name> within the step timeout. Otherwise, or when the backend replica that timed it out had not seen its last verify round (after a restart, say), the last error is no response within step timeout. Console does not know what state the host is in, so it does not retry, roll back or guess. A person checks the host (dpkg -l <package>, the unit's status) and fixes it. A late answer from the agent is recorded on the timeline but changes nothing.
The step timeout is PROXIMA_UPGRADE_STEP_TIMEOUT on the backend (default 10m). A preflight that times out changes nothing on the host, so that host is blocked (no_preflight_response) rather than unknown.
A result can be lost on the way back (for example, while every backend replica is restarting). The agent keeps the result of every preflight, apply and rollback, and while a host waits in one of those steps Console asks the agent for it every 15 seconds, so a lost result is recovered and the host moves on. Only when the agent has no result for the step by the deadline does the host become blocked or unknown.
Blocked reasons
A blocked host was not changed. The run page shows the text of the Meaning column below; text in italics is context this page adds and the run page does not show. A code the page does not know (the last row) is shown as the raw code.
| Reason | Meaning |
|---|---|
upgrades_disabled | This host hasn't opted in to upgrades (PROXIMA_AGENT_UPGRADES_ENABLED). A k8s-node-monitor agent always answers this. |
package_manager_unsupported | Only apt hosts are supported so far. |
package_not_installed_or_broken | The package isn't installed, or dpkg reports it in a broken state. Repair it on the host first. |
unhealthy_before_upgrade | The service wasn't healthy before the upgrade, so it wasn't touched. |
checks_failing_before_upgrade | One of your checks already fails before the upgrade, so the upgrade would roll back. Fix the check or the service first. The host's last error names the check and what it saw. |
monitors_failing_before_upgrade | An uptime monitor on this host wasn't up before the upgrade, or had not been probed recently. Fix it, or exclude it with a reason, and dry-run again. Not up includes pending, unknown, paused and maintenance, and an up monitor with no probe result in the last interval plus 30 seconds. The host's last error names the monitors. |
monitor_cannot_verify_in_time | An uptime monitor on this host checks too rarely (or is a push heartbeat) to confirm the upgrade in time. Exclude it with a reason, or shorten its interval, then dry-run again. Too rarely: its interval plus 30 seconds reaches PROXIMA_UPGRADE_STEP_TIMEOUT. |
monitors_unreadable | Console couldn't read this host's monitors, so it didn't touch the host. Try the dry run again. Also used when the upgrade's exclusions or, at the pre-flight, the host's firing alerts could not be read. |
transaction_changed | The packages to install changed since the dry run. Run the dry run again. |
transaction_removes_packages | Installing this would remove other packages. That isn't supported yet. |
unsupported_transaction | The change involves packages for another architecture. That isn't supported yet. |
rollback_artifact_missing | The old package files couldn't be kept, so a safe rollback isn't guaranteed. |
rollback_artifact_failed | Keeping the old package files failed, so a safe rollback isn't guaranteed. |
conffile_backup_failed | Backing up the package's config files failed, so nothing was changed. |
dry_run_failed | The package manager's dry run failed on this host. |
dpkg_query_failed | Reading the installed package failed on this host. |
state_error | The agent couldn't save its upgrade state, so nothing was changed. |
invalid_input | The plan's package, version or unit name was rejected. |
no_preflight_response | The agent didn't answer. It may be offline or too old for upgrades. Probe and collector agents never answer upgrade commands, so they end up here. |
no_dry_run | No dry run is stored for this host. Run the dry run again. |
target_older_than_installed | The host already runs a newer version than the target. M1 never downgrades. |
already_at_target | The host already runs the target version. |
host_moved_environment | This host moved to another environment, or no longer exists, since the upgrade was drafted, so it was left alone. If it moved, draft a new upgrade for its current environment. Its last error says which. Checked before the preflight and again before the apply; nothing more was sent to it. |
agent_cannot_restart_affected | This host's agent is too old to restart the services picked for this upgrade, so the host was left alone rather than upgraded without them. Cancel this upgrade, then update the agent and create a new one, or create one without services to restart. Decided by Console from the host's dry run, which carries no affected_detection; nothing is sent to the host. |
restart_units_unreadable | Console couldn't read which services this upgrade restarts, so it didn't touch the host. |
unit_not_restartable | The agent refused to restart one of the picked services (it never restarts itself, systemd, D-Bus or anything that isn't a service), so nothing was changed. The agent refuses at the start of the apply, before anything is installed, so nothing is rolled back. |
rollback_not_ready, preflight_failed | Shown as raw codes: the preflight answered without an error code but did not confirm that a rollback is ready, or was not a clean pass. Run the dry run again; if it repeats, check the agent's log. |
Known limits
- Removals and foreign-architecture packages are refused. A transaction that removes a package, or names a
pkg:archpackage, blocks the host at the dry run. Rollback also refuses to remove anything apt wants to remove, and fails instead. - apt only. dnf/yum hosts answer
package_manager_unsupported. - Restoring configuration files replaces them by rename. That keeps mode and ownership but breaks hard links and drops extended attributes and ACLs on the restored files.
- No paging on
rollback_failed. It pauses the upgrade and is on the timeline and the run page, but nobody is paged. - An upgrade silences nothing, so a restart may page on-call. See A restart may page on-call. Silencing waits for maintenance windows (M2d).
- Only monitors linked to the host count. A monitor of a service that depends on the upgraded one, on another host, is not judged, and neither is an alert on another host.
- Operator checks run from the host and refuse link-local targets (IPv4
169.254.0.0/16, IPv6fe80::/10). A check cannot reach something only reachable from elsewhere. Only those ranges are refused: a metadata endpoint a provider exposes at another address is not. - A draft for an environment you cannot see is answered 404, exactly as for one that does not exist; 403 means you can see the environment but lack
upgrades:writeon it. - The agent's upgrade files are never pruned, and free space is not checked first. Everything under the agent's data directory in
upgrades/(the kept.debfiles, the configuration-file backups, the stored results and the in-flight markers) stays there after the upgrade ends, and nothing checks free disk space before an apply downloads or backs anything up. Watch the disk of hosts you upgrade often, and remove an old upgrade'supgrades/<upgrade-id>_<target-id>directory and itsupgrades/results/<upgrade-id>_<target-id>_*files yourself once you no longer need to roll it back. - Deleting a project, environment or host deletes its upgrade history. An upgrade belongs to its project and environment, and a target to its host, with
ON DELETE CASCADE: deleting one removes the upgrades or targets under it, with their timeline rows. The audit log keeps its own entries. - Affected-service detection has limits. Statically linked binaries never appear, a container's own libraries are not covered, and the report after the apply is only as good as needrestart or
/proccan make it. See Detection limits. - One change per host at a time. A host that another upgrade is changing right now is not started: its target stays pending, and is picked up on a later tick once that change has finished.
Configuration
| Where | Variable | Default | Meaning |
|---|---|---|---|
| Agent | PROXIMA_AGENT_UPGRADES_ENABLED | off | Opts the host in. Only true or 1; host agents only. |
| Backend | PROXIMA_UPGRADE_STEP_TIMEOUT | 10m | How long a host may stay in preflight, applying, verifying or rolling back, and how long after the soak's end its final verify may take. Applying and rolling back get 2 minutes more, plus 5 minutes per picked service sent to the host, at most 30 minutes more (During the upgrade). A monitor whose interval plus 30 seconds reaches it blocks the host (monitor_cannot_verify_in_time). |
| Backend | PROXIMA_UPGRADE_WORKER_INTERVAL | 5s | How often the upgrade worker ticks. |
API
All under /api/v1/upgrades:
upgrades:read:GET /,GET /{id}andGET /standing-rules.upgrades:write:POST /, andPOST /{id}/dry-run,/request-approval,/start,/pause,/resumeand/cancel, andPUT /{id}/monitor-exclusionsandPUT /{id}/restart-units.upgrades:approve:POST /{id}/approveand/reject, andPOST /standing-rules,/standing-rules/{id}/approveand/standing-rules/{id}/revoke.
Approve and reject take {"note": "...", "approval_requested_at": "..."}, and revoke takes {"note": "..."}. GET /{id} carries is_production, whether the upgrade's environment is production now, and the upgrade's monitor_exclusions and pause_reason; each target carries its monitor_snapshot.
PUT /{id}/monitor-exclusions takes the full list, {"exclusions": [{"monitor_id": "...", "reason": "..."}]}, at most 50 items; an empty list clears it. The server adds by and at to each.
PUT /{id}/restart-units takes the full list, {"units": ["postfix.service", ...]}, at most 20, in the order they are restarted; an empty list clears it, and the field is required. It is audited as upgrade.restart_units_set. GET /{id} carries the upgrade's restart_units (always a list, never null); each target's dry_run carries affected_services and affected_detection, and each target carries affected_report after an apply that succeeded. See Affected services.
A caller without access to an upgrade or a rule gets 404, as if it did not exist.
The main error codes:
| Status | Codes |
|---|---|
| 400 | invalid_note, invalid_input, unknown_monitor, unknown_unit, unit_not_restartable |
| 403 | self_approval, service_account_cannot_approve, peer_resume_required |
| 409 | illegal_transition; targets_in_flight; standing_rule_changed (the rule was revoked or expired as the upgrade started); monitor_exclusions_changed (a monitor was excluded as the upgrade started under a standing rule); restart_units_changed (services to restart were picked as the upgrade started under a standing rule); approval_stale; status_changed (someone else changed the upgrade or rule first; reload and retry) |
| 422 | no_targets, no_runnable_targets, extra_packages_not_acknowledged, approval_required, approval_not_needed (approve, reject or submit on a non-production environment) |
The full schemas are in the API reference.