Skip to main content

Config as Code

Config as code lets an external tool — today, the pc CLI, run by hand or from CI — own the definition of a monitor, an escalation route, an escalation policy or a Telegram group binding as a YAML file instead of a UI form. You edit the file, run pc apply, and Console's state matches the file. This is milestone 1 of the config-as-code/GitOps track: a fixed set of resource kinds, one generic HTTP API, and a one-shot CLI — no continuous reconciler, no CRD operator, and no Terraform provider yet (those are later roadmap phases).

It solves a narrower problem than "manage all of Console from Git": it exists so a client's escalation routes and monitors can be defined once, reviewed in a merge request, and applied idempotently — and so that a pc apply run can never silently clobber something a person configured by hand in the UI, or something another tool already owns.

The resource envelope​

Every config-as-code resource is a resource envelope — the same shape pc get -o yaml prints and pc apply -f reads back in:

kind: monitor
version: v1alpha1
metadata:
name: web-checkout-https
client: acme-corp
spec:
target: https://checkout.acme.example.com/health
interval_seconds: 60
timeout_seconds: 10
kind: escalation_route
version: v1alpha1
metadata:
name: primary
client: acme-corp
spec:
escalation_policy_id: 11111111-1111-1111-1111-111111111111
notify_feed: true
enabled: true
These spec blocks are trimmed for readability

A real pc get monitor web-checkout-https --client acme-corp -o yaml prints a spec that is the full stored entity, not a minimal create/update subset — see the spec bullet below.

A file may hold one envelope or several, separated by a line containing only --- (the standard YAML document separator). The two envelopes above can therefore live in one acme-corp.yaml instead of two files — which is what you want when a client's whole on-call configuration is a single unit of review. Kinds and even clients may be mixed freely within one file; each document is resolved and applied on its own.

  • kind — monitor, escalation_route, escalation_policy, telegram_group_binding or client_uptime_notify. The kind vocabulary is expected to grow in later phases; treat it as a stable, versioned contract, not an implementation detail.

  • version — currently always v1alpha1, and required: pc apply refuses any file declaring anything else (including an omitted version) with unsupported version "..." (expected v1alpha1), before it makes any network call. The check is client-side only — pc apply still doesn't send the version on the wire, and the backend doesn't read one — so it pins which schema pc will accept, not how the backend behaves.

  • metadata.name — the resource's identity, unique per client and kind. This is what pc apply upserts by, and what --prune diffs against.

    telegram_group_binding is the exception, and it is not a small one: its metadata.name is the Telegram chat id, stringified — name: "-1001234567890", not a friendly slug. That table has no name column at all; its primary key is the chat, so the (client, name) address degenerates to (client, chat_id). Quote it — unquoted YAML reads it as a number and the file is refused on the envelope, before any network call. A name that is not a whole number is refused with a 400 naming the rule rather than reported as "not found", because the fix is in the file.

  • metadata.client — the client's name, not its UUID — a portable, human-readable label for a file that will live in Git. pc resolves it to a client_id for you (the same case-insensitive name/slug match pc ls --client already uses); the resource API itself only ever speaks client_id.

  • spec — on write, kind-specific: for monitor it's the same JSON shape the monitor create/update API accepts — see Uptime Monitors for the full field set, including the required kind (http/tcp/dns/tls/push), severity (P1-P5), and the numeric floors (interval_seconds >= 20, retry_interval_seconds >= 10, timeout_seconds >= 1 and less than interval_seconds); for escalation_route, the shape the escalation route API accepts — see Escalation Chains. A route needs at least one of escalation_policy_id, notify_telegram, or notify_feed — never none of the three — and a named escalation_policy_id must reference a known, team-owned policy. The short examples above show only the fields you're likely to hand-author.

    Some spec fields accept a name instead of a UUID — see References by name below.

    On read, spec is not filtered down to that shape. pc get -o yaml prints a raw json.Marshal of the entire stored domain.Monitor/domain.EscalationRoute row — every column, including read-only ones: id, client_id, environment_id, origin, created_at, updated_at (and for routes, also position, source, is_default), plus name/client_id themselves (metadata.name/metadata.client are not subtracted out of spec the way the original design intended — they're simply duplicated). This is harmless for the pc get -o yaml > file.yaml → edit → pc apply -f file.yaml round trip: Upsert force- overrides id/client_id/name/origin from the URL path and its own policy decision regardless of what the spec says, and creating a new object never reads an incoming id. But don't treat a pc get -o yaml dump as the minimal shape to hand-author from scratch — extra fields in it are read-only noise, not required input.

References by name​

metadata.client has always taken a client name rather than its UUID. The same now applies to some fields inside spec: where a field holds a reference to another Console object, you may write that object's name or slug and pc apply resolves it to the UUID before sending. Raw UUIDs are not something a human can source correctly for a file that lives in Git.

These fields are covered:

KindFieldCandidates come fromMatched on
monitor, escalation_routespec.environment_idthe file's own client's environmentsname or slug
escalation_routespec.escalation_policy_idevery escalation policy you can seename
escalation_policyspec.owner_idteamsname
escalation_policyspec.steps[].schedule_idthe policy's own team's on-call schedulesname
escalation_policyspec.steps[].user_ids[]usersemail

escalation_policy's spec.steps[].type is covered too, but it is a different kind of reference — a name for a number rather than for an object. See Step types below.

kind: monitor
version: v1alpha1
metadata:
name: web-checkout-https
client: acme-corp
spec:
environment_id: production # instead of 33333333-3333-3333-3333-333333333333
kind: http
target: https://checkout.acme.example.com/health
severity: P2
interval_seconds: 60
timeout_seconds: 10

The rules:

  • Matching is case-insensitive, on the field(s) the table names — for environments that is name or slug, the same match metadata.client and pc ls --client already use, so one spelling convention covers the whole file. Users match on email only, deliberately: an email address is unique in Console, a display name is not, and "who gets paged at 3am" must not depend on which of two colleagues sharing a name the lookup happened to see first.
  • The lookup is scoped as narrowly as the object's own ownership allows. Environments come from GET /api/v1/clients/{client_id}/environments, resolved from metadata.client, so environment_id: production can only ever mean this client's production — never another tenant's environment that happens to share the name. A step's schedule_id is restricted to the policy's own owning team, enforced against each returned row rather than trusted to a query parameter (see Steps reference their own team's schedule). You need the ordinary read permission on whatever is being listed — environments:read, teams:read, users:read, or an on-call permission for schedules and policies; a 403 is reported as the backend's own message.
  • A name that matches two objects is refused, not guessed. escalation_policy_id: critical where two teams each own a chain called critical fails with escalation_policy_id: "critical" is ambiguous … it names at least two objects (<id> and <id>); rename one, or write the id you mean. Silently taking the first row would mean an apply rewriting — or paging — something the file never named. This is the same answer the backend gives when a policy name is ambiguous under a client.
  • A value that is already a UUID is passed straight through, with no lookup and no request. That is what every pc get -o yaml → edit → pc apply round trip does, so the round trip costs nothing extra and nothing about it changes.
  • A name that matches nothing is a hard error, before anything is written: environment_id: no match for "staging" at /api/v1/clients/<id>/environments. pc never forwards an unresolved name for the backend to reject with a confusing "must be a UUID".
  • Absent, null and empty values are left exactly as written, so the backend's own validation still produces the error (a monitor without an environment_id is refused with environment_id is required; an escalation route's environment_id is legitimately optional).
  • --dry-run diffs the resolved spec, not the names — a dry run shows exactly the bytes a real apply would send.

spec.host_id still requires a UUID. It is the next addition to the same mechanism, not a different one.

escalation_policy step types​

An escalation step's type is stored — and sent on the wire — as an integer. pc apply accepts the name instead, so a chain reads as what it does rather than as a row of magic numbers:

NameWire valueWhat the step does
wait0delay before the next step (the only delay step)
notify_users1notify every user in user_ids at once
notify_users_queue2notify the next single user, round-robin
notify_schedule3resolve who is on call now and notify them
repeat9loop back to the start, capped by repeat_count

These five are not a subset — they are everything Console accepts. The domain declares twelve step types, but the escalation worker only has handlers for these; every write path, the UI's included, refuses the rest. So a name outside this table is a hard client-side error naming the step and the vocabulary (steps[1].type: unknown step type "declare_incident" — supported: notify_schedule, notify_users, notify_users_queue, repeat, wait) rather than a file that applies and then fails server-side.

A type that is already a number is left alone — that is what pc get -o yaml prints, so both spellings apply cleanly and a round trip is unaffected. The two spellings may be mixed freely within one steps: block, which is what editing a single step of a dumped policy produces.

A step's type also decides which of its other fields are resolved. schedule_id is resolved only on a notify_schedule step, and user_ids only on notify_users / notify_users_queue — exactly the fields Console itself reads for each type. So a step left carrying a key that no longer applies (changed from notify_schedule to wait without deleting its schedule_id, say) is passed through as written rather than failing the apply: Console ignores that key, and pc will not refuse a file the API would accept. The flip side is that a name in such a field is not resolved and will be stored as-is — harmless, but tidy it up rather than relying on it.

Steps reference their own team's schedule​

GET /api/v1/oncall/schedules is scoped by owner, not by client: asked without a filter it returns every team's schedules that you can see. Schedule names repeat across teams — several teams each calling their rotation primary is normal — so resolving schedule_id: primary against that unfiltered list could attach another team's rotation to this policy.

pc apply therefore resolves a step's schedule_id only against the schedules owned by the policy's own team, and it does that by resolving spec.owner_id first. It enforces that in two places, because one of them is not enough: it asks the server to narrow the list (?owner_type=team&owner_id=…), and it re-checks the owner_id every returned row carries before considering that row a candidate. The server-side filter is honoured only for super-admins — every other caller gets the schedules for all of their visible clients whatever the query string says — so the client-side check is what actually holds for the CI service account or client-scoped operator running this in practice.

Two consequences worth knowing:

  • A spec that names a schedule but carries no owner_id is refused, rather than falling back to an unscoped lookup: steps[0].schedule_id: cannot resolve a schedule name without the policy's owner_id — add owner_id (the owning team's name is enough) to the spec, or write the schedule's UUID. Omitting owner_id is otherwise legitimate on an update (the stored owner is kept), so this only ever bites a file that needs the scope it did not supply.
  • A step carrying a schedule UUID needs no lookup at all, so it applies with or without owner_id.

The backend enforces the same boundary independently — a notify_schedule step pointing at another team's rotation is refused with schedule_id belongs to a different owner. Scoping the lookup in pc is what keeps you from hitting that message for a name that is correct under your own team.

A worked escalation_policy​

kind: escalation_policy
version: v1alpha1
metadata:
name: critical
client: acme-corp
spec:
owner_type: team
owner_id: platform # the owning team's name
repeat_count: 2
ack_timeout_minutes: 15
steps:
- type: notify_schedule # instead of: type: 3
schedule_id: primary # platform's rotation, not another team's
- type: wait
wait_seconds: 300
- type: notify_users
user_ids:
- [email protected] # instead of a user UUID
- sre-[email protected]
- type: repeat

One pc apply resolves all of it: owner_id → the team's UUID, then schedule_id → that team's primary schedule, user_ids → the two users' UUIDs, and each type → its integer. pc apply -f policy.yaml --dry-run prints the resolved spec, so you can see exactly what would be sent before sending it.

owner_type is team in everything this milestone authors; a spec declaring any other owner is refused when it also names a schedule by name, since the owner axis of the schedule lookup is the team.

Origin: who owns what​

Every monitor and escalation route carries an origin column: ui (hand-authored, the default for anything created in the web app), pc, file, k8s, or terraform. It records which tool last claimed a resource, and it's the mechanism that keeps pc apply from touching anything it doesn't own:

  • Applying to a new name always succeeds and stamps origin as the caller's own (pc, when you run pc apply).

  • Applying to an existing resource owned by a different origin is refused with 409 origin_conflict — "web-checkout-https is owned by ui — pass adopt=true to claim it." Adopting reassigns origin to the caller's.

  • Deleting (including via pc apply --prune) never has an adopt escape hatch: a resource owned by a different origin is always refused with 409 origin_conflict, never removed. Origin is therefore the boundary that keeps --prune from touching anything another tool created: it can only ever prune what its own origin (or another run under the same origin) created.

    Origin is not the whole safety story for --prune, though, and the gap is worth knowing before you automate it. Origin answers "which tool made this"; it says nothing about "which client does this belong to", and --prune needs both. For monitor, escalation_route and telegram_group_binding the second question has an answer in the schema — each carries its own client_id — so a --prune run scoped to one client sees exactly that client's objects and the unattended case is genuinely safe. escalation_policy is the one exception: a policy has no client at all (it belongs to a team), so --prune cannot tell which client's file set was supposed to describe it, and it skips the kind entirely rather than guess. See --prune below for what that looks like and why.

pc apply exposes both overrides as flags:

pc apply -f route.yaml --adopt   # claim ownership of an object owned by a different origin
pc apply -f route.yaml --force # overwrite an object with an active break-glass edit

Both are per-invocation, not per-file: whichever you pass is sent on every upsert pc apply does in that run (--dry-run never sends either — it never writes). --adopt answers the origin_conflict case above; --force answers the break_glass_active case in the next section. They're independent — an object that's both differently-owned and break-glass-active needs both flags, since the break-glass check runs first.

Break-glass: when a hand edit is in the way​

Two independent locks share the same audit columns (break_glass_reason, break_glass_by, break_glass_at), and they protect against different things:

Editing a k8s-owned object by hand, in the UI. Only k8s-origin objects lock the UI — a continuously-reconciling CRD operator would otherwise clobber your edit on its next pass. pc- and file-origin objects are fire-and-forget (nothing comes back on its own to reassert them), so they never lock UI edits. If you PATCH/DELETE a locked monitor or route without a break_glass_reason, you get 409 origin_locked: "managed by k8s — edit via k8s, or set break_glass_reason to override" (DELETE's wording is "…delete via k8s, or pass ?break_glass_reason= to override"). Supply a reason (in the request body for PATCH, as ?break_glass_reason= for DELETE, since it has no body) and the edit proceeds — origin is left unchanged, and the reason, your user ID, and the timestamp are stamped as the audit trail.

Applying over a hand-edited object. If a monitor or route was broken-glass-edited this way and never resolved — today that only happens to k8s-origin objects, since that's the only origin the UI locks — the next pc apply targeting it refuses with 409 break_glass_active: "web-checkout-https was hand-edited ("bumping timeout for a known-slow deploy") — re-apply with force to overwrite." pc apply -f file.yaml --force clears the break-glass columns and reasserts the file's spec — the tool is taking ownership back. Because the object is k8s-owned and you're applying as pc, this case also needs --adopt in the same run (see the note at the end of Origin) — --force alone clears the break-glass lock but still hits 409 origin_conflict on the very next check.

pc get​

pc get monitor --client acme-corp                          # list, table output
pc get monitor web-checkout-https --client acme-corp # one, table output
pc get monitor web-checkout-https --client acme-corp -o yaml # one, YAML envelope
pc get escalation_route --client acme-corp -o yaml # list, YAML (multi-doc stream)

--client is required and resolved the same way pc ls --client resolves it. Default output is a table:

NAME                  ORIGIN    BREAK_GLASS
web-checkout-https pc
primary ui bumping timeout for a known-slow deploy

-o yaml prints the canonical envelope(s) shown above — this is the round trip: pc get monitor web-checkout-https --client acme-corp -o yaml > web-checkout-https.yaml, edit the file, then pc apply -f web-checkout-https.yaml. The list form round-trips too: its ----separated stream is a multi-document file, and pc apply -f applies every document in it.

pc apply​

pc apply -f route.yaml                  # upsert every envelope in the file
pc apply -f dir/ # upsert every *.yaml/*.yml directly under the
# directory (not its subdirectories — see -R below)
pc apply -f dir/ -R # ...and now including every subdirectory too
pc apply -f route.yaml --dry-run # print a diff, make no changes
pc apply -f dir/ --prune # also delete pc-owned objects absent from this run's
# applied set

-f/--filename is required and takes a file or a directory. A successful apply prints one line per resource: escalation_route/primary applied. On a directory, files are applied in sorted filename order, stopping at the first failure (the failing file's path is named in the error) — partial application on a multi-file failure is possible, by design (v1 has no transaction across files). -R/--recursive extends the file collection into every subdirectory (still sorted, depth-first) — off by default, so an existing pc apply -f dir/ keeps applying only dir/'s own files even if it happens to contain unrelated subdirectories.

A single file may carry several ----separated envelopes, so one client's escalation route and the monitor that pages through it can be one reviewable file:

kind: escalation_route
version: v1alpha1
metadata:
name: primary
client: acme-corp
spec:
escalation_policy_id: 11111111-1111-1111-1111-111111111111
enabled: true
---
kind: monitor
version: v1alpha1
metadata:
name: web-checkout-https
client: acme-corp
spec:
kind: http
target: https://checkout.acme.example.com/health
severity: P2
interval_seconds: 60
timeout_seconds: 10

Documents are applied top to bottom, and — exactly like the file-level loop — applying stops at the first document that fails, so a partly-applied file is possible by design. The error names the file and, when the file holds more than one document, the failing document's position: acme-corp.yaml: document 2: unsupported version "v1alpha2" (expected v1alpha1). A chunk that describes no resource — a leading ---, a trailing one, two in a row, and anything that is only blank lines and comments, such as a file header or a commented section break — is ignored rather than treated as an empty resource, and does not consume a position in that numbering either. --prune counts every document in the file as applied, so a multi-document file never prunes a resource it just applied.

A separator is a line holding only --- (trailing spaces, tabs and a CR are fine, so CRLF files work). A line carrying anything after the dashes — --- # routes below, a tag, or content on the same line — is legal YAML but is not recognized here, and a file using that form is applied as a single document: only its first resource lands. Write the separator bare, the way pc get -o yaml emits it.

Every request pc apply sends carries X-Proxima-Origin: pc, so everything it creates or touches is pc-owned.

Push monitors: the heartbeat token prints exactly once​

Applying a kind: push monitor mints its heartbeat credential, and pc apply prints the plaintext on the line under the applied line:

monitor/nightly-backup applied
push token (save now — it will not be shown again): <token>

Only the token's SHA-256 hash is stored, so this apply's output is the only place the plaintext ever exists. pc get, pc get -o yaml and pc apply --dry-run are all reads, and no read path returns it — lose the line and the only way to get a working token is to delete and recreate the monitor. It is deliberately not written into the -o yaml envelope either: a one-time credential must never end up in a file you commit. The line appears only on the apply that actually mints a token — a create, an edit turning an existing monitor into kind: push, or a re-apply of a push monitor that has no stored token hash yet, which is how a push monitor created before this behaviour existed repairs itself into a working one. Re-applying a push monitor whose token already exists prints the usual single applied line and mints nothing: the token you already saved keeps working.

--dry-run​

Replaces the write with a read: for each file, pc fetches the resource's current state and prints a diff instead of applying.

$ pc apply -f route.yaml --dry-run
escalation_route/primary diff:
{
"escalation_policy_id": "11111111-1111-1111-1111-111111111111",
- "enabled": false,
+ "enabled": true,
}

Unchanged lines print too, with a two-space prefix — this is a line-oriented diff over each side's pretty-printed JSON spec ( unchanged, - removed, + added), not a unified diff with hunk headers.

A resource that doesn't exist yet prints the whole spec as a creation, not a diff:

$ pc apply -f new-route.yaml --dry-run
escalation_route/staging-secondary would be created:
+ {
+ "escalation_policy_id": "22222222-2222-2222-2222-222222222222",
+ "notify_feed": true
+ }

No PUT is sent either way — --dry-run never calls a server-side dry-run mode (none exists in this milestone); it's entirely a client-side GET + diff.

--prune​

Deletes pc-owned resources that are absent from the file(s) you applied. pc groups whatever it successfully applied by (kind, client), asks the backend what pc already owns for each group (?origin=pc), and deletes any name in that group's current state that isn't in the applied set:

$ pc apply -f dir/ --prune
monitor/web-checkout-https applied
escalation_route/primary applied
escalation_route/stale-route pruned

A single pc apply -f dir/ --prune can cover several kinds and clients in one invocation — prune groups by whatever kinds and clients actually appeared in the applied files, not by a fixed per-command scope. It only ever considers pc-origin resources: a ui- or k8s-owned object of the same name is never touched, no matter how the file set changes.

--prune does not operate on escalation_policy​

--prune handles monitor, escalation_route and telegram_group_binding. It skips escalation_policy, and prints a line saying so:

$ pc apply -f dir/ -R --prune
escalation_policy/team-critical applied
escalation_route/critical applied
escalation_policy: not pruned — an escalation_policy is team-owned and not scoped to one client, so --prune cannot safely operate on it
escalation_route/stale-route pruned

The rest of the run is unaffected: every other kind in the same tree still prunes exactly as it otherwise would, and the notice prints under --dry-run --prune too, so a rehearsal and a real run agree about what happens to your policies.

Why. --prune's whole method is to ask the backend what pc owns for a (kind, client) pair and delete whatever the files did not name. That is only sound when the backend's answer covers exactly one client. A policy has no client_id — its identity is (owner_type, owner_id, name), scoped to the owning team — so listing policies "for a client" means listing every policy owned by every team assigned to that client. One team serving two clients (the normal arrangement for a shared DevOps or SRE team) is enough for that list to contain chains the applied files were never meant to describe:

  • devops-team-1 serves both acme-retail and initech-logistics.
  • Its chains are authored with metadata.client: acme-retail; devops-team-2's are authored with metadata.client: initech-logistics.
  • Applying that tree builds an (escalation_policy, initech-logistics) group naming only devops-team-2's chains — but the backend's list for that client also returns devops-team-1's, because that team serves it too.
  • Without the skip, --prune would read those as "absent from the files" and delete them, while acme-retail's routes still pointed at them.

The skip is deliberately blunt because the alternative is not a smarter diff: policies would need a per-client anchor recording which client's file set authored them, and that does not exist yet. Until it does, delete an obsolete escalation policy explicitly — in the web app, or with a direct DELETE /api/v1/resources/escalation_policy/{name}?client_id=... (pc has no delete command in this milestone) — rather than expecting --prune to notice it is gone from Git. That delete still goes through the origin check, so a policy pc does not own is refused there too.

--dry-run --prune together​

Combining the two flags never issues a DELETE — a deliberate safety choice, since --dry-run's whole contract is "make no changes," and a prune that deletes anyway would silently break that for exactly the destructive half of this command. Instead of pruning, it reports what would be pruned:

$ pc apply -f dir/ --dry-run --prune
escalation_route/primary diff:
{
- "enabled": false,
+ "enabled": true,
}
escalation_route/stale-route would be pruned

--dry-run alone (no --prune) behaves as described above and is unaffected by this.

A worked example: a whole on-call configuration​

docs-site/docs/config-as-code/examples/ in this repository holds a complete on-call configuration — four teams' escalation chains and three clients' routing, seventeen resources across eleven applied files — in the shape a real repository would have. It is modelled on the Grafana OnCall Terraform setup described under Migrating from another on-call tool: the same four teams, the same eight chains, the same step sequence, the same critical/default route pair per client. The names, emails and chat ids are fictional; the structure is not. Every file in it has been applied against a real backend.

examples/                          # <- NOT the apply target; `oncall/` is
├── oncall/ # the applied tree — pc apply -f oncall/ -R
│ ├── 10-policies/
│ │ ├── devops-team-1-critical.yaml
│ │ ├── devops-team-1-default.yaml
│ │ ├── devops-team-2-critical.yaml
│ │ ├── devops-team-2-default.yaml
│ │ ├── devops-team-3-critical.yaml
│ │ ├── devops-team-3-default.yaml
│ │ ├── internal-infra-critical.yaml
│ │ └── internal-infra-default.yaml
│ └── 20-clients/
│ ├── acme-retail.yaml # one multi-document file: binding + 2 routes
│ ├── globex-fintech.yaml
│ └── internal-infra.yaml
└── oncall-prerequisites/ # NOT applied — see rule 2 below
├── devops-team-1.yaml
├── devops-team-2.yaml
├── devops-team-3.yaml
└── internal-infra.yaml

Apply oncall/, never examples/. The two directories are siblings, so -R on the examples/ root would walk oncall-prerequisites/ as well — and because oncall/ sorts first and applies perfectly, the refusal lands after all eleven real files have already been written. A failure at the end of a mostly-successful run is a far worse way to learn this than the one-word difference in the command.

Three layout rules the tree encodes​

1. -R applies in lexical order, so a tree with references has to be named in dependency order. pc apply -f dir/ -R walks the tree depth-first in lexical order and stops at the first failure — it does not sort by kind, and it has no dependency graph. A client's route names its team's chain, so every chain has to be applied before any route that points at one. With the obvious names (policies/, clients/) clients sorts first and the very first client file fails with escalation_policy_id: no match for "devops-team-1-critical". The numeric prefixes are therefore load-bearing, not cosmetic: 10- before 20- is how the directory names carry the ordering constraint that the tool does not know about.

2. Everything under the applied tree must be an applyable resource, so the prerequisites live outside oncall/. -R picks up every *.yaml beneath whatever path you give it, and a file that is not a resource envelope — including one that is nothing but comments — is refused with unsupported version "" (expected v1alpha1). The teams, on-call schedules and users a policy file references are not config-as-code resource kinds (see What's not here yet); they have to exist already. oncall-prerequisites/ is the written record of what that is — one file per team, listing the team, its client assignments, its schedule and the people its chains name — kept beside oncall/ rather than inside it, which is what keeps pc apply -f oncall/ -R clean. It is what a reviewer checks a policy file's by-name references against without opening the UI. (Being a sibling is only safe for the command that names oncall/; see the note above the rules.)

3. One file per client. A client's Telegram binding and both of its routes are one unit of review, so they are one multi-document file. That is what keeps the tree's file count at 8 + <clients> instead of 8 + 3 × <clients>.

The chains​

Each of the eight policy files is one team's critical or default chain. The two variants share a ladder — on-call engineer, team lead, CTO, CEO, then repeat — and differ in three ways, all three inherited from the Grafana chains they reproduce: the critical chain waits 5 minutes between steps where the default waits 30/60/60, every critical step sets important: true (which selects each responder's important notification chain), and the default chain has one extra wait after the CEO step that the critical one does not.

kind: escalation_policy
version: v1alpha1
metadata:
name: devops-team-1-critical
client: acme-retail
spec:
owner_type: team
owner_id: devops-team-1 # the owning TEAM's name
role: null
repeat_count: 2
ack_timeout_minutes: 15
steps:
- type: notify_schedule
schedule_id: "TEAM 1" # resolved against THIS TEAM's schedules only
important: true
- type: wait
wait_seconds: 300
- type: notify_users
user_ids:
- lead-team-[email protected]
important: true
# …CTO, CEO, same shape…
- type: repeat

Three things in that file are easy to get wrong:

  • metadata.client is not the policy's owner. escalation_policies has no client_id at all — owner_id is what owns the chain. metadata.client is the tenancy anchor this API addresses it under, and it has to name a client the owning team is assigned to: the resource API resolves a policy name by listing the policies visible under that client, which is exactly the ones owned by a team serving it. Naming a client the team does not serve is refused with 403 access denied rather than filed somewhere no later apply could find it again. Pick one of the team's own clients and keep it — (client, name) is the address a re-apply resolves by.
  • Put the team in the chain's name. escalation_policy_id is resolved against every policy the caller can see, and a policy name is unique only per owning team — so two teams each owning a bare critical is refused as ambiguous rather than guessed at. devops-team-1-critical can never collide.
  • Write null for anything you want cleared. An upsert load-copies the stored row and unmarshals the spec onto it, so an omitted key preserves whatever is there. role: null and ack_timeout_minutes: null are how a file keeps saying "deliberately unset" on every re-apply instead of silently inheriting an earlier version of itself.

The client files​

kind: telegram_group_binding
version: v1alpha1
metadata:
name: "-1001110001111" # the chat id, QUOTED
client: acme-retail
spec:
audience: client
notify_alerts: true
notify_jira: false
language: en
---
kind: escalation_route
version: v1alpha1
metadata:
name: critical
client: acme-retail
spec:
severity: P1
escalation_policy_id: devops-team-1-critical
notify_feed: true
enabled: true
---
kind: escalation_route
version: v1alpha1
metadata:
name: default
client: acme-retail
spec:
severity: null # any severity — the catch-all
escalation_policy_id: devops-team-1-default
notify_feed: true
enabled: true
  • Quote the chat id. A telegram_group_binding's metadata.name is the Telegram chat id (see metadata.name), and unquoted YAML reads -1001110001111 as a number, which fails on the envelope before anything is sent.
  • Route order is position order, and position is set on create. Routes match first-match-wins by position, and a new route takes one past the client's current maximum — so applying this file to a client with no routes gives critical position 0 and default position 1, which is the order they must be in, since default matches every severity including P1. pc apply never moves a route afterwards (that is PUT /api/v1/escalation/routes/order's job), so re-ordering the documents later does nothing to a client that is already applied.
  • severity is one tier or null, not a list. Grafana OnCall's payload.groupLabels.severity in ['P2','P3',…] becomes an absence here: Console stores a route's severity as a single canonical tier (P1…P5, unknown) or nothing at all.
  • notify_telegram still satisfies the "at least one of" rule, but sends nothing. It is accepted on write for compatibility — an existing file that sets it keeps applying — but no code path reads it to decide whether or where a Telegram message goes; see Notify routing for where Telegram delivery actually lives now (the client's client_uptime_notify targets). notify_feed is the field that still does something on a match.

Applying it​

pc apply -f docs-site/docs/config-as-code/examples/oncall/ -R
escalation_policy/devops-team-1-critical applied
escalation_policy/devops-team-1-default applied
…
telegram_group_binding/-1001110001111 applied
escalation_route/critical applied
escalation_route/default applied

One pc apply resolves every name in the tree: four team names to team ids, four schedule names to schedule ids (each scoped to its own policy's team), six email addresses to user ids, and every step type to its wire integer. Re-running it is a no-op that reports applied again — pc owns what it created, so nothing is refused as an origin conflict.

--dry-run cannot pre-flight a tree that creates its own references

A dry run replaces the write with a read, so nothing the run would have created exists while it is running. On a tree like this one the eight policies print as would be created, and then the first client file fails: escalation_policy_id: no match for "devops-team-1-critical". That is correct behaviour, not a defect — but it means a dry run is a review tool for changes to an already-applied tree, not a pre-flight check for a new one. To pre-flight a first apply, apply the policies for real and dry-run the clients.

Also note that a dry run of an already-applied file is not a clean, empty diff: the current state it diffs against is the full stored entity (id, created_at, origin, …) while the file is the hand-authored subset, so the two sides differ in shape as well as in content. See the spec bullet under The resource envelope.

Migrating from another on-call tool​

proxima/proxima_oncall is a Terraform repository that drives on-call routing for 21 client modules today through a self-hosted Grafana OnCall instance — per client, one Telegram-integrated alert receiver plus a critical and a default route, each pointing at one of four teams' escalation chains, applied through GitLab-native Terraform CI. The worked example above is a faithful translation of its shape into Console's own resource kinds, and the point of translating rather than writing a one-off import script is to keep the property that setup already has: every change to who gets paged is a reviewable diff in Git.

The actual cutover is not part of this feature and is not described here. Writing the real client files, obtaining each client's Alertmanager alert-source token out of band, pointing each client's Alertmanager at Console, firing a test alert through the new path, and finally deleting each module block from proxima_oncall and tearing its Grafana OnCall integration down — that is operational rollout work, done client by client and ordered by risk, lowest first. The runbook for it lives in the repository's design spec, docs/superpowers/specs/2026-09-21-oncall-config-as-code-migration-design.md, which also records what was deliberately left unbuilt.

Four structural differences are worth knowing before anyone starts, because they are properties of Console's model rather than gaps in this tooling:

  • A Telegram chat belongs to exactly one client. telegram_group_binding's primary key is chat_id, so a single team-wide chat receiving every one of that team's clients' alerts — which is how the Grafana setup is wired, one telegram_id per team shared across its clients — cannot be reproduced. Each Console client needs its own chat.
  • There is no per-client alert message template. Grafana OnCall lets a client override the shared Jinja2 Telegram template; Console's delivery has no equivalent field, and adding one is out of scope.
  • There is no cross-client global fallback chain. Grafana OnCall's Default chain fires when no client route matches. Console's escalation_route is always client-scoped, so there is nothing for it to attach to. In practice a client's critical and default routes already cover every severity between them, so this is a real but rarely-reached structural difference — it is flagged, not designed.
  • An alert-source token is a secret and is never in a file. It is minted once, per client, through POST /api/v1/clients/{client_id}/alert-sources, and handled the way the Terraform setup already handles its own secrets — out of band, not in the repository.

What's not here yet​

  • No continuous reconciler: pc apply is one-shot, like terraform apply. Nothing re-applies a file on its own.
  • No CRD operator (k8s origin already exists in the schema and is UI-locked, but nothing writes it yet) and no Terraform provider (terraform origin exists in the schema for the same reason).
  • No resource kinds beyond monitor, escalation_route, escalation_policy, telegram_group_binding and client_uptime_notify — teams, schedules, client features, and notification channels are later-phase work. In particular an on-call schedule cannot be authored here, only referenced, so a policy's rotation still has to exist before the file that points at it. Nor can a team, a team's client assignment, or a user — which is why a worked tree needs a prerequisites file alongside it (see Three layout rules).
  • By-name references are nearly there. Everything an on-call configuration references — environments, escalation policies, owning teams, schedules, users — resolves from a name (see References by name); spec.host_id is the one reference field that still carries a Console UUID, and telegram_group_binding has no by-name fields at all, so a binding that sets environment_id or team_id has to write the UUID. --prune also performs no referenced-object check.
  • No dependency ordering. pc apply applies files in lexical order and stops at the first failure; it does not sort by kind and it builds no reference graph. A tree whose files reference each other has to encode the order in its own directory names — see Three layout rules.

client_uptime_notify​

A client's two uptime Telegram targets — the internal target (a team/staff chat) and the client target (one of the client's own bound chats) that a monitor's down/up transition notifies, independent of Notify routing's escalation-route notify_telegram. There is exactly one per client, so metadata.name must be default.

kind: client_uptime_notify
version: v1alpha1
metadata:
name: default
client: acme-retail
spec:
internal_chat_id: -1001234567890 # a team chat or internal-audience chat; super-admin or global telegram_ops:write
internal_thread_id: 17118 # optional; omit for the uptime → alerts topic ladder
client_chat_id: -1009876543210 # one of this client's client-audience chats
client_thread_id: null

Reading needs monitors:read. Writing the client target needs telegram_ops:write held client-wide — an environment-scoped grant is not enough, the same rule PUT /clients/{id}/uptime-notify enforces. Writing the internal target needs super-admin or the global telegram_ops:write; a caller without one of those is refused for internal_chat_id/internal_thread_id the moment either key is mentioned at all — matching, changed, or explicit null — not only when it would actually change something, so a guessed value can never be told apart from a wrong one by whether the apply succeeds.

A key you omit keeps its stored value; null clears it — except that changing a *_chat_id to a genuinely different chat while omitting its sibling *_thread_id resets the thread to null rather than carrying the OLD chat's thread onto the new one; give both keys together to pin a specific thread on the new chat. Deleting the object clears both targets outright, and clearing a configured internal target through delete needs the same super-admin/global bar as clearing it through an apply — deleting a row that has no internal target configured needs only the client-wide bar. pc get for a caller who cannot set the internal target omits internal_chat_id/internal_thread_id from the spec entirely — there is no separate "configured" flag the way the HTTP settings API's redacted field has one, so the absence of the keys is the only signal, and it means exactly what it says for pc apply: nothing here to preserve or to change. --prune does not delete this kind.