Config as Code
Config as code lets an external tool — today, the pc CLI, run by hand or from CI — own the
definition of a monitor, an escalation route, an escalation policy or a Telegram
group binding as a YAML file instead of a UI form. You edit the file, run pc apply, and
Console's state matches the file. This is milestone 1 of the config-as-code/GitOps track: a
fixed set of resource kinds, one generic HTTP API, and a one-shot CLI — no continuous
reconciler, no CRD operator, and no Terraform provider yet (those are later roadmap phases).
It solves a narrower problem than "manage all of Console from Git": it exists so a client's
escalation routes and monitors can be defined once, reviewed in a merge request, and applied
idempotently — and so that a pc apply run can never silently clobber something a person
configured by hand in the UI, or something another tool already owns.
The resource envelope
Every config-as-code resource is a resource envelope — the same shape pc get -o yaml prints
and pc apply -f reads back in:
kind: monitor
version: v1alpha1
metadata:
name: web-checkout-https
client: acme-corp
spec:
target: https://checkout.acme.example.com/health
interval_seconds: 60
timeout_seconds: 10
kind: escalation_route
version: v1alpha1
metadata:
name: primary
client: acme-corp
spec:
escalation_policy_id: 11111111-1111-1111-1111-111111111111
notify_feed: true
enabled: true
spec blocks are trimmed for readabilityA real pc get monitor web-checkout-https --client acme-corp -o yaml prints a spec that is the
full stored entity, not a minimal create/update subset — see the spec bullet below.
A file may hold one envelope or several, separated by a line containing only --- (the
standard YAML document separator). The two envelopes above can therefore live in one
acme-corp.yaml instead of two files — which is what you want when a client's whole on-call
configuration is a single unit of review. Kinds and even clients may be mixed freely within one
file; each document is resolved and applied on its own.
-
kind—monitor,escalation_route,escalation_policy,telegram_group_bindingorclient_uptime_notify. The kind vocabulary is expected to grow in later phases; treat it as a stable, versioned contract, not an implementation detail. -
version— currently alwaysv1alpha1, and required:pc applyrefuses any file declaring anything else (including an omittedversion) withunsupported version "..." (expected v1alpha1), before it makes any network call. The check is client-side only —pc applystill doesn't send the version on the wire, and the backend doesn't read one — so it pins which schemapcwill accept, not how the backend behaves. -
metadata.name— the resource's identity, unique per client and kind. This is whatpc applyupserts by, and what--prunediffs against.telegram_group_bindingis the exception, and it is not a small one: itsmetadata.nameis the Telegram chat id, stringified —name: "-1001234567890", not a friendly slug. That table has no name column at all; its primary key is the chat, so the(client, name)address degenerates to(client, chat_id). Quote it — unquoted YAML reads it as a number and the file is refused on the envelope, before any network call. A name that is not a whole number is refused with a400naming the rule rather than reported as "not found", because the fix is in the file. -
metadata.client— the client's name, not its UUID — a portable, human-readable label for a file that will live in Git.pcresolves it to aclient_idfor you (the same case-insensitive name/slug matchpc ls --clientalready uses); the resource API itself only ever speaksclient_id. -
spec— on write, kind-specific: formonitorit's the same JSON shape the monitor create/update API accepts — see Uptime Monitors for the full field set, including the requiredkind(http/tcp/dns/tls/push),severity(P1-P5), and the numeric floors (interval_seconds >= 20,retry_interval_seconds >= 10,timeout_seconds >= 1and less thaninterval_seconds); forescalation_route, the shape the escalation route API accepts — see Escalation Chains. A route needs at least one ofescalation_policy_id,notify_telegram, ornotify_feed— never none of the three — and a namedescalation_policy_idmust reference a known, team-owned policy. The short examples above show only the fields you're likely to hand-author.Some
specfields accept a name instead of a UUID — see References by name below.On read,
specis not filtered down to that shape.pc get -o yamlprints a rawjson.Marshalof the entire storeddomain.Monitor/domain.EscalationRouterow — every column, including read-only ones:id,client_id,environment_id,origin,created_at,updated_at(and for routes, alsoposition,source,is_default), plusname/client_idthemselves (metadata.name/metadata.clientare not subtracted out ofspecthe way the original design intended — they're simply duplicated). This is harmless for thepc get -o yaml > file.yaml→ edit →pc apply -f file.yamlround trip:Upsertforce- overridesid/client_id/name/originfrom the URL path and its own policy decision regardless of what the spec says, and creating a new object never reads an incomingid. But don't treat apc get -o yamldump as the minimal shape to hand-author from scratch — extra fields in it are read-only noise, not required input.
References by name
metadata.client has always taken a client name rather than its UUID. The same now applies
to some fields inside spec: where a field holds a reference to another Console object, you
may write that object's name or slug and pc apply resolves it to the UUID before sending.
Raw UUIDs are not something a human can source correctly for a file that lives in Git.
These fields are covered:
| Kind | Field | Candidates come from | Matched on |
|---|---|---|---|
monitor, escalation_route | spec.environment_id | the file's own client's environments | name or slug |
escalation_route | spec.escalation_policy_id | every escalation policy you can see | name |
escalation_policy | spec.owner_id | teams | name |
escalation_policy | spec.steps[].schedule_id | the policy's own team's on-call schedules | name |
escalation_policy | spec.steps[].user_ids[] | users | email |
escalation_policy's spec.steps[].type is covered too, but it is a different kind of
reference — a name for a number rather than for an object. See
Step types below.
kind: monitor
version: v1alpha1
metadata:
name: web-checkout-https
client: acme-corp
spec:
environment_id: production # instead of 33333333-3333-3333-3333-333333333333
kind: http
target: https://checkout.acme.example.com/health
severity: P2
interval_seconds: 60
timeout_seconds: 10
The rules:
- Matching is case-insensitive, on the field(s) the table names — for environments that is
nameorslug, the same matchmetadata.clientandpc ls --clientalready use, so one spelling convention covers the whole file. Users match onemailonly, deliberately: an email address is unique in Console, a display name is not, and "who gets paged at 3am" must not depend on which of two colleagues sharing a name the lookup happened to see first. - The lookup is scoped as narrowly as the object's own ownership allows. Environments come
from
GET /api/v1/clients/{client_id}/environments, resolved frommetadata.client, soenvironment_id: productioncan only ever mean this client's production — never another tenant's environment that happens to share the name. A step'sschedule_idis restricted to the policy's own owning team, enforced against each returned row rather than trusted to a query parameter (see Steps reference their own team's schedule). You need the ordinary read permission on whatever is being listed —environments:read,teams:read,users:read, or an on-call permission for schedules and policies; a403is reported as the backend's own message. - A name that matches two objects is refused, not guessed.
escalation_policy_id: criticalwhere two teams each own a chain calledcriticalfails withescalation_policy_id: "critical" is ambiguous … it names at least two objects (<id> and <id>); rename one, or write the id you mean. Silently taking the first row would mean an apply rewriting — or paging — something the file never named. This is the same answer the backend gives when a policy name is ambiguous under a client. - A value that is already a UUID is passed straight through, with no lookup and no request.
That is what every
pc get -o yaml→ edit →pc applyround trip does, so the round trip costs nothing extra and nothing about it changes. - A name that matches nothing is a hard error, before anything is written:
environment_id: no match for "staging" at /api/v1/clients/<id>/environments.pcnever forwards an unresolved name for the backend to reject with a confusing "must be a UUID". - Absent,
nulland empty values are left exactly as written, so the backend's own validation still produces the error (a monitor without anenvironment_idis refused withenvironment_id is required; an escalation route'senvironment_idis legitimately optional). --dry-rundiffs the resolved spec, not the names — a dry run shows exactly the bytes a real apply would send.
spec.host_id still requires a UUID. It is the next addition to the same mechanism, not a
different one.
escalation_policy step types
An escalation step's type is stored — and sent on the wire — as an integer. pc apply
accepts the name instead, so a chain reads as what it does rather than as a row of magic
numbers:
| Name | Wire value | What the step does |
|---|---|---|
wait | 0 | delay before the next step (the only delay step) |
notify_users | 1 | notify every user in user_ids at once |
notify_users_queue | 2 | notify the next single user, round-robin |
notify_schedule | 3 | resolve who is on call now and notify them |
repeat | 9 | loop back to the start, capped by repeat_count |
These five are not a subset — they are everything Console accepts. The domain declares
twelve step types, but the escalation worker only has handlers for these; every write path,
the UI's included, refuses the rest. So a name outside this table is a hard client-side error
naming the step and the vocabulary (steps[1].type: unknown step type "declare_incident" — supported: notify_schedule, notify_users, notify_users_queue, repeat, wait) rather than a file
that applies and then fails server-side.
A type that is already a number is left alone — that is what pc get -o yaml prints, so
both spellings apply cleanly and a round trip is unaffected. The two spellings may be mixed
freely within one steps: block, which is what editing a single step of a dumped policy
produces.
A step's type also decides which of its other fields are resolved. schedule_id is
resolved only on a notify_schedule step, and user_ids only on notify_users /
notify_users_queue — exactly the fields Console itself reads for each type. So a step left
carrying a key that no longer applies (changed from notify_schedule to wait without
deleting its schedule_id, say) is passed through as written rather than failing the apply:
Console ignores that key, and pc will not refuse a file the API would accept. The flip side
is that a name in such a field is not resolved and will be stored as-is — harmless, but tidy
it up rather than relying on it.
Steps reference their own team's schedule
GET /api/v1/oncall/schedules is scoped by owner, not by client: asked without a filter it
returns every team's schedules that you can see. Schedule names repeat across teams — several
teams each calling their rotation primary is normal — so resolving schedule_id: primary
against that unfiltered list could attach another team's rotation to this policy.
pc apply therefore resolves a step's schedule_id only against the schedules owned by the
policy's own team, and it does that by resolving spec.owner_id first. It enforces that in
two places, because one of them is not enough: it asks the server to narrow the list
(?owner_type=team&owner_id=…), and it re-checks the owner_id every returned row carries
before considering that row a candidate. The server-side filter is honoured only for
super-admins — every other caller gets the schedules for all of their visible clients whatever
the query string says — so the client-side check is what actually holds for the CI service
account or client-scoped operator running this in practice.
Two consequences worth knowing:
- A spec that names a schedule but carries no
owner_idis refused, rather than falling back to an unscoped lookup:steps[0].schedule_id: cannot resolve a schedule name without the policy's owner_id — add owner_id (the owning team's name is enough) to the spec, or write the schedule's UUID. Omittingowner_idis otherwise legitimate on an update (the stored owner is kept), so this only ever bites a file that needs the scope it did not supply. - A step carrying a schedule UUID needs no lookup at all, so it applies with or without
owner_id.
The backend enforces the same boundary independently — a notify_schedule step pointing at
another team's rotation is refused with schedule_id belongs to a different owner. Scoping the
lookup in pc is what keeps you from hitting that message for a name that is correct under
your own team.
A worked escalation_policy
kind: escalation_policy
version: v1alpha1
metadata:
name: critical
client: acme-corp
spec:
owner_type: team
owner_id: platform # the owning team's name
repeat_count: 2
ack_timeout_minutes: 15
steps:
- type: notify_schedule # instead of: type: 3
schedule_id: primary # platform's rotation, not another team's
- type: wait
wait_seconds: 300
- type: notify_users
user_ids:
- [email protected] # instead of a user UUID
- sre-[email protected]
- type: repeat
One pc apply resolves all of it: owner_id → the team's UUID, then schedule_id → that
team's primary schedule, user_ids → the two users' UUIDs, and each type → its integer.
pc apply -f policy.yaml --dry-run prints the resolved spec, so you can see exactly what would
be sent before sending it.
owner_type is team in everything this milestone authors; a spec declaring any other owner
is refused when it also names a schedule by name, since the owner axis of the schedule lookup
is the team.
Origin: who owns what
Every monitor and escalation route carries an origin column: ui (hand-authored, the default
for anything created in the web app), pc, file, k8s, or terraform. It records which tool
last claimed a resource, and it's the mechanism that keeps pc apply from touching anything it
doesn't own:
-
Applying to a new name always succeeds and stamps
originas the caller's own (pc, when you runpc apply). -
Applying to an existing resource owned by a different origin is refused with
409 origin_conflict— "web-checkout-httpsis owned byui— passadopt=trueto claim it." Adopting reassignsoriginto the caller's. -
Deleting (including via
pc apply --prune) never has an adopt escape hatch: a resource owned by a different origin is always refused with409 origin_conflict, never removed. Origin is therefore the boundary that keeps--prunefrom touching anything another tool created: it can only ever prune what its own origin (or another run under the same origin) created.Origin is not the whole safety story for
--prune, though, and the gap is worth knowing before you automate it. Origin answers "which tool made this"; it says nothing about "which client does this belong to", and--pruneneeds both. Formonitor,escalation_routeandtelegram_group_bindingthe second question has an answer in the schema — each carries its ownclient_id— so a--prunerun scoped to one client sees exactly that client's objects and the unattended case is genuinely safe.escalation_policyis the one exception: a policy has no client at all (it belongs to a team), so--prunecannot tell which client's file set was supposed to describe it, and it skips the kind entirely rather than guess. See--prunebelow for what that looks like and why.
pc apply exposes both overrides as flags:
pc apply -f route.yaml --adopt # claim ownership of an object owned by a different origin
pc apply -f route.yaml --force # overwrite an object with an active break-glass edit
Both are per-invocation, not per-file: whichever you pass is sent on every upsert pc apply does
in that run (--dry-run never sends either — it never writes). --adopt answers the
origin_conflict case above; --force answers the break_glass_active case in the next
section. They're independent — an object that's both differently-owned and break-glass-active
needs both flags, since the break-glass check runs first.
Break-glass: when a hand edit is in the way
Two independent locks share the same audit columns (break_glass_reason, break_glass_by,
break_glass_at), and they protect against different things:
Editing a k8s-owned object by hand, in the UI. Only k8s-origin objects lock the UI — a
continuously-reconciling CRD operator would otherwise clobber your edit on its next pass. pc-
and file-origin objects are fire-and-forget (nothing comes back on its own to reassert them), so
they never lock UI edits. If you PATCH/DELETE a locked monitor or route without a
break_glass_reason, you get 409 origin_locked: "managed by k8s — edit via k8s, or set
break_glass_reason to override" (DELETE's wording is "…delete via k8s, or pass
?break_glass_reason= to override"). Supply a reason (in the request body for PATCH, as
?break_glass_reason= for DELETE, since it has no body) and the edit proceeds — origin is
left unchanged, and the reason, your user ID, and the timestamp are stamped as the audit trail.
Applying over a hand-edited object. If a monitor or route was broken-glass-edited this way and
never resolved — today that only happens to k8s-origin objects, since that's the only origin the
UI locks — the next pc apply targeting it refuses with 409 break_glass_active:
"web-checkout-https was hand-edited ("bumping timeout for a known-slow deploy") — re-apply
with force to overwrite." pc apply -f file.yaml --force clears the break-glass columns and
reasserts the file's spec — the tool is taking ownership back. Because the object is k8s-owned
and you're applying as pc, this case also needs --adopt in the same run (see the note at the
end of Origin) — --force alone clears the break-glass lock but still
hits 409 origin_conflict on the very next check.
pc get
pc get monitor --client acme-corp # list, table output
pc get monitor web-checkout-https --client acme-corp # one, table output
pc get monitor web-checkout-https --client acme-corp -o yaml # one, YAML envelope
pc get escalation_route --client acme-corp -o yaml # list, YAML (multi-doc stream)
--client is required and resolved the same way pc ls --client resolves it. Default output is
a table:
NAME ORIGIN BREAK_GLASS
web-checkout-https pc
primary ui bumping timeout for a known-slow deploy
-o yaml prints the canonical envelope(s) shown above — this is the round trip:
pc get monitor web-checkout-https --client acme-corp -o yaml > web-checkout-https.yaml, edit
the file, then pc apply -f web-checkout-https.yaml. The list form round-trips too: its
----separated stream is a multi-document file, and pc apply -f applies every document in it.
pc apply
pc apply -f route.yaml # upsert every envelope in the file
pc apply -f dir/ # upsert every *.yaml/*.yml directly under the
# directory (not its subdirectories — see -R below)
pc apply -f dir/ -R # ...and now including every subdirectory too
pc apply -f route.yaml --dry-run # print a diff, make no changes
pc apply -f dir/ --prune # also delete pc-owned objects absent from this run's
# applied set
-f/--filename is required and takes a file or a directory. A successful apply prints one line
per resource: escalation_route/primary applied. On a directory, files are applied in sorted
filename order, stopping at the first failure (the failing file's path is named in the error) —
partial application on a multi-file failure is possible, by design (v1 has no transaction across
files). -R/--recursive extends the file collection into every subdirectory (still sorted,
depth-first) — off by default, so an existing pc apply -f dir/ keeps applying only dir/'s own
files even if it happens to contain unrelated subdirectories.
A single file may carry several ----separated envelopes, so one client's escalation route and
the monitor that pages through it can be one reviewable file:
kind: escalation_route
version: v1alpha1
metadata:
name: primary
client: acme-corp
spec:
escalation_policy_id: 11111111-1111-1111-1111-111111111111
enabled: true
---
kind: monitor
version: v1alpha1
metadata:
name: web-checkout-https
client: acme-corp
spec:
kind: http
target: https://checkout.acme.example.com/health
severity: P2
interval_seconds: 60
timeout_seconds: 10
Documents are applied top to bottom, and — exactly like the file-level loop — applying stops at
the first document that fails, so a partly-applied file is possible by design. The error names
the file and, when the file holds more than one document, the failing document's position:
acme-corp.yaml: document 2: unsupported version "v1alpha2" (expected v1alpha1). A chunk that
describes no resource — a leading ---, a trailing one, two in a row, and anything that is only
blank lines and comments, such as a file header or a commented section break — is ignored rather
than treated as an empty resource, and does not consume a position in that numbering either.
--prune counts every document in the file as applied, so a multi-document file never prunes
a resource it just applied.
A separator is a line holding only --- (trailing spaces, tabs and a CR are fine, so CRLF files
work). A line carrying anything after the dashes — --- # routes below, a tag, or content on
the same line — is legal YAML but is not recognized here, and a file using that form is applied as
a single document: only its first resource lands. Write the separator bare, the way
pc get -o yaml emits it.
Every request pc apply sends carries X-Proxima-Origin: pc, so everything it creates or
touches is pc-owned.
Push monitors: the heartbeat token prints exactly once
Applying a kind: push monitor mints its heartbeat credential, and pc apply prints the
plaintext on the line under the applied line:
monitor/nightly-backup applied
push token (save now — it will not be shown again): <token>
Only the token's SHA-256 hash is stored, so this apply's output is the only place the
plaintext ever exists. pc get, pc get -o yaml and pc apply --dry-run are all reads, and no
read path returns it — lose the line and the only way to get a working token is to delete and
recreate the monitor. It is deliberately not written into the -o yaml envelope either: a
one-time credential must never end up in a file you commit. The line appears only on the apply
that actually mints a token — a create, an edit turning an existing monitor into kind: push, or
a re-apply of a push monitor that has no stored token hash yet, which is how a push monitor
created before this behaviour existed repairs itself into a working one. Re-applying a push
monitor whose token already exists prints the usual single applied line and mints nothing: the
token you already saved keeps working.
--dry-run
Replaces the write with a read: for each file, pc fetches the resource's current state and
prints a diff instead of applying.
$ pc apply -f route.yaml --dry-run
escalation_route/primary diff:
{
"escalation_policy_id": "11111111-1111-1111-1111-111111111111",
- "enabled": false,
+ "enabled": true,
}
Unchanged lines print too, with a two-space prefix — this is a line-oriented diff over each
side's pretty-printed JSON spec ( unchanged, - removed, + added), not a unified diff
with hunk headers.
A resource that doesn't exist yet prints the whole spec as a creation, not a diff:
$ pc apply -f new-route.yaml --dry-run
escalation_route/staging-secondary would be created:
+ {
+ "escalation_policy_id": "22222222-2222-2222-2222-222222222222",
+ "notify_feed": true
+ }
No PUT is sent either way — --dry-run never calls a server-side dry-run mode (none exists in
this milestone); it's entirely a client-side GET + diff.
--prune
Deletes pc-owned resources that are absent from the file(s) you applied. pc groups whatever
it successfully applied by (kind, client), asks the backend what pc already owns for each
group (?origin=pc), and deletes any name in that group's current state that isn't in the
applied set:
$ pc apply -f dir/ --prune
monitor/web-checkout-https applied
escalation_route/primary applied
escalation_route/stale-route pruned
A single pc apply -f dir/ --prune can cover several kinds and clients in one invocation — prune
groups by whatever kinds and clients actually appeared in the applied files, not by a fixed
per-command scope. It only ever considers pc-origin resources: a ui- or k8s-owned object of
the same name is never touched, no matter how the file set changes.
--prune does not operate on escalation_policy
--prune handles monitor, escalation_route and telegram_group_binding. It skips
escalation_policy, and prints a line saying so:
$ pc apply -f dir/ -R --prune
escalation_policy/team-critical applied
escalation_route/critical applied
escalation_policy: not pruned — an escalation_policy is team-owned and not scoped to one client, so --prune cannot safely operate on it
escalation_route/stale-route pruned
The rest of the run is unaffected: every other kind in the same tree still prunes exactly as it
otherwise would, and the notice prints under --dry-run --prune too, so a rehearsal and a real
run agree about what happens to your policies.
Why. --prune's whole method is to ask the backend what pc owns for a (kind, client)
pair and delete whatever the files did not name. That is only sound when the backend's answer
covers exactly one client. A policy has no client_id — its identity is
(owner_type, owner_id, name), scoped to the owning team — so listing policies "for a
client" means listing every policy owned by every team assigned to that client. One team serving
two clients (the normal arrangement for a shared DevOps or SRE team) is enough for that list to
contain chains the applied files were never meant to describe:
devops-team-1serves bothacme-retailandinitech-logistics.- Its chains are authored with
metadata.client: acme-retail;devops-team-2's are authored withmetadata.client: initech-logistics. - Applying that tree builds an
(escalation_policy, initech-logistics)group naming onlydevops-team-2's chains — but the backend's list for that client also returnsdevops-team-1's, because that team serves it too. - Without the skip,
--prunewould read those as "absent from the files" and delete them, whileacme-retail's routes still pointed at them.
The skip is deliberately blunt because the alternative is not a smarter diff: policies would need
a per-client anchor recording which client's file set authored them, and that does not exist yet.
Until it does, delete an obsolete escalation policy explicitly — in the web app, or with a
direct DELETE /api/v1/resources/escalation_policy/{name}?client_id=... (pc has no delete
command in this milestone) — rather than expecting --prune to notice it is gone from Git. That
delete still goes through the origin check, so a policy pc does not own is refused there too.
--dry-run --prune together
Combining the two flags never issues a DELETE — a deliberate safety choice, since
--dry-run's whole contract is "make no changes," and a prune that deletes anyway would silently
break that for exactly the destructive half of this command. Instead of pruning, it reports what
would be pruned:
$ pc apply -f dir/ --dry-run --prune
escalation_route/primary diff:
{
- "enabled": false,
+ "enabled": true,
}
escalation_route/stale-route would be pruned
--dry-run alone (no --prune) behaves as described above and is unaffected by this.
A worked example: a whole on-call configuration
docs-site/docs/config-as-code/examples/
in this repository holds a complete on-call configuration — four teams' escalation chains and
three clients' routing, seventeen resources across eleven applied files — in the shape a real
repository would have. It is modelled on the Grafana OnCall Terraform setup described under
Migrating from another on-call tool: the same four
teams, the same eight chains, the same step sequence, the same critical/default route pair per
client. The names, emails and chat ids are fictional; the structure is not. Every file in it
has been applied against a real backend.
examples/ # <- NOT the apply target; `oncall/` is
├── oncall/ # the applied tree — pc apply -f oncall/ -R
│ ├── 10-policies/
│ │ ├── devops-team-1-critical.yaml
│ │ ├── devops-team-1-default.yaml
│ │ ├── devops-team-2-critical.yaml
│ │ ├── devops-team-2-default.yaml
│ │ ├── devops-team-3-critical.yaml
│ │ ├── devops-team-3-default.yaml
│ │ ├── internal-infra-critical.yaml
│ │ └── internal-infra-default.yaml
│ └── 20-clients/
│ ├── acme-retail.yaml # one multi-document file: binding + 2 routes
│ ├── globex-fintech.yaml
│ └── internal-infra.yaml
└── oncall-prerequisites/ # NOT applied — see rule 2 below
├── devops-team-1.yaml
├── devops-team-2.yaml
├── devops-team-3.yaml
└── internal-infra.yaml
Apply oncall/, never examples/. The two directories are siblings, so -R on the
examples/ root would walk oncall-prerequisites/ as well — and because oncall/ sorts first
and applies perfectly, the refusal lands after all eleven real files have already been written.
A failure at the end of a mostly-successful run is a far worse way to learn this than the
one-word difference in the command.
Three layout rules the tree encodes
1. -R applies in lexical order, so a tree with references has to be named in dependency
order. pc apply -f dir/ -R walks the tree depth-first in lexical order and stops at the
first failure — it does not sort by kind, and it has no dependency graph. A client's route
names its team's chain, so every chain has to be applied before any route that points at one.
With the obvious names (policies/, clients/) clients sorts first and the very first
client file fails with escalation_policy_id: no match for "devops-team-1-critical". The
numeric prefixes are therefore load-bearing, not cosmetic: 10- before 20- is how the
directory names carry the ordering constraint that the tool does not know about.
2. Everything under the applied tree must be an applyable resource, so the prerequisites live
outside oncall/. -R picks up every *.yaml beneath whatever path you give it, and a file
that is not a resource envelope — including one that is nothing but comments — is refused with
unsupported version "" (expected v1alpha1). The teams, on-call schedules and users a policy
file references are not config-as-code resource kinds (see
What's not here yet); they have to exist already. oncall-prerequisites/
is the written record of what that is — one file per team, listing the team, its client
assignments, its schedule and the people its chains name — kept beside oncall/ rather than
inside it, which is what keeps pc apply -f oncall/ -R clean. It is what a reviewer checks a
policy file's by-name references against without opening the UI. (Being a sibling is only safe
for the command that names oncall/; see the note above the rules.)
3. One file per client. A client's Telegram binding and both of its routes are one unit of
review, so they are one multi-document file. That is what keeps the tree's file count at
8 + <clients> instead of 8 + 3 × <clients>.
The chains
Each of the eight policy files is one team's critical or default chain. The two variants share
a ladder — on-call engineer, team lead, CTO, CEO, then repeat — and differ in three ways, all
three inherited from the Grafana chains they reproduce: the critical chain waits 5 minutes
between steps where the default waits 30/60/60, every critical step sets important: true
(which selects each responder's important notification chain), and the default chain has one
extra wait after the CEO step that the critical one does not.
kind: escalation_policy
version: v1alpha1
metadata:
name: devops-team-1-critical
client: acme-retail
spec:
owner_type: team
owner_id: devops-team-1 # the owning TEAM's name
role: null
repeat_count: 2
ack_timeout_minutes: 15
steps:
- type: notify_schedule
schedule_id: "TEAM 1" # resolved against THIS TEAM's schedules only
important: true
- type: wait
wait_seconds: 300
- type: notify_users
user_ids:
- lead-team-[email protected]
important: true
# …CTO, CEO, same shape…
- type: repeat
Three things in that file are easy to get wrong:
metadata.clientis not the policy's owner.escalation_policieshas noclient_idat all —owner_idis what owns the chain.metadata.clientis the tenancy anchor this API addresses it under, and it has to name a client the owning team is assigned to: the resource API resolves a policy name by listing the policies visible under that client, which is exactly the ones owned by a team serving it. Naming a client the team does not serve is refused with403 access deniedrather than filed somewhere no later apply could find it again. Pick one of the team's own clients and keep it —(client, name)is the address a re-apply resolves by.- Put the team in the chain's name.
escalation_policy_idis resolved against every policy the caller can see, and a policy name is unique only per owning team — so two teams each owning a barecriticalis refused as ambiguous rather than guessed at.devops-team-1-criticalcan never collide. - Write
nullfor anything you want cleared. An upsert load-copies the stored row and unmarshals the spec onto it, so an omitted key preserves whatever is there.role: nullandack_timeout_minutes: nullare how a file keeps saying "deliberately unset" on every re-apply instead of silently inheriting an earlier version of itself.
The client files
kind: telegram_group_binding
version: v1alpha1
metadata:
name: "-1001110001111" # the chat id, QUOTED
client: acme-retail
spec:
audience: client
notify_alerts: true
notify_jira: false
language: en
---
kind: escalation_route
version: v1alpha1
metadata:
name: critical
client: acme-retail
spec:
severity: P1
escalation_policy_id: devops-team-1-critical
notify_feed: true
enabled: true
---
kind: escalation_route
version: v1alpha1
metadata:
name: default
client: acme-retail
spec:
severity: null # any severity — the catch-all
escalation_policy_id: devops-team-1-default
notify_feed: true
enabled: true
- Quote the chat id. A
telegram_group_binding'smetadata.nameis the Telegram chat id (seemetadata.name), and unquoted YAML reads-1001110001111as a number, which fails on the envelope before anything is sent. - Route order is position order, and position is set on create. Routes match first-match-wins
by position, and a new route takes one past the client's current maximum — so applying this
file to a client with no routes gives
criticalposition 0 anddefaultposition 1, which is the order they must be in, sincedefaultmatches every severity including P1.pc applynever moves a route afterwards (that isPUT /api/v1/escalation/routes/order's job), so re-ordering the documents later does nothing to a client that is already applied. severityis one tier ornull, not a list. Grafana OnCall'spayload.groupLabels.severity in ['P2','P3',…]becomes an absence here: Console stores a route's severity as a single canonical tier (P1…P5,unknown) or nothing at all.notify_telegramstill satisfies the "at least one of" rule, but sends nothing. It is accepted on write for compatibility — an existing file that sets it keeps applying — but no code path reads it to decide whether or where a Telegram message goes; see Notify routing for where Telegram delivery actually lives now (the client'sclient_uptime_notifytargets).notify_feedis the field that still does something on a match.
Applying it
pc apply -f docs-site/docs/config-as-code/examples/oncall/ -R
escalation_policy/devops-team-1-critical applied
escalation_policy/devops-team-1-default applied
…
telegram_group_binding/-1001110001111 applied
escalation_route/critical applied
escalation_route/default applied
One pc apply resolves every name in the tree: four team names to team ids, four schedule
names to schedule ids (each scoped to its own policy's team), six email addresses to user ids,
and every step type to its wire integer. Re-running it is a no-op that reports applied
again — pc owns what it created, so nothing is refused as an origin conflict.
--dry-run cannot pre-flight a tree that creates its own referencesA dry run replaces the write with a read, so nothing the run would have created exists
while it is running. On a tree like this one the eight policies print as would be created,
and then the first client file fails: escalation_policy_id: no match for "devops-team-1-critical". That is correct behaviour, not a defect — but it means a dry run is
a review tool for changes to an already-applied tree, not a pre-flight check for a new one.
To pre-flight a first apply, apply the policies for real and dry-run the clients.
Also note that a dry run of an already-applied file is not a clean, empty diff: the current
state it diffs against is the full stored entity (id, created_at, origin, …) while the
file is the hand-authored subset, so the two sides differ in shape as well as in content. See
the spec bullet under The resource envelope.
Migrating from another on-call tool
proxima/proxima_oncall is a Terraform repository that drives on-call routing for 21 client
modules today through a self-hosted Grafana OnCall instance — per client, one
Telegram-integrated alert receiver plus a critical and a default route, each pointing at one of
four teams' escalation chains, applied through GitLab-native Terraform CI. The worked example
above is a faithful translation of its shape into Console's own resource kinds, and the point of
translating rather than writing a one-off import script is to keep the property that setup
already has: every change to who gets paged is a reviewable diff in Git.
The actual cutover is not part of this feature and is not described here. Writing the real
client files, obtaining each client's Alertmanager alert-source token out of band, pointing each
client's Alertmanager at Console, firing a test alert through the new path, and finally deleting
each module block from proxima_oncall and tearing its Grafana OnCall integration down — that
is operational rollout work, done client by client and ordered by risk, lowest first. The runbook
for it lives in the repository's design spec,
docs/superpowers/specs/2026-09-21-oncall-config-as-code-migration-design.md,
which also records what was deliberately left unbuilt.
Four structural differences are worth knowing before anyone starts, because they are properties of Console's model rather than gaps in this tooling:
- A Telegram chat belongs to exactly one client.
telegram_group_binding's primary key ischat_id, so a single team-wide chat receiving every one of that team's clients' alerts — which is how the Grafana setup is wired, onetelegram_idper team shared across its clients — cannot be reproduced. Each Console client needs its own chat. - There is no per-client alert message template. Grafana OnCall lets a client override the shared Jinja2 Telegram template; Console's delivery has no equivalent field, and adding one is out of scope.
- There is no cross-client global fallback chain. Grafana OnCall's
Defaultchain fires when no client route matches. Console'sescalation_routeis always client-scoped, so there is nothing for it to attach to. In practice a client's critical and default routes already cover every severity between them, so this is a real but rarely-reached structural difference — it is flagged, not designed. - An alert-source token is a secret and is never in a file. It is minted once, per client,
through
POST /api/v1/clients/{client_id}/alert-sources, and handled the way the Terraform setup already handles its own secrets — out of band, not in the repository.
What's not here yet
- No continuous reconciler:
pc applyis one-shot, liketerraform apply. Nothing re-applies a file on its own. - No CRD operator (
k8sorigin already exists in the schema and is UI-locked, but nothing writes it yet) and no Terraform provider (terraformorigin exists in the schema for the same reason). - No resource kinds beyond
monitor,escalation_route,escalation_policy,telegram_group_bindingandclient_uptime_notify— teams, schedules, client features, and notification channels are later-phase work. In particular an on-call schedule cannot be authored here, only referenced, so a policy's rotation still has to exist before the file that points at it. Nor can a team, a team's client assignment, or a user — which is why a worked tree needs a prerequisites file alongside it (see Three layout rules). - By-name references are nearly there. Everything an on-call configuration references —
environments, escalation policies, owning teams, schedules, users — resolves from a name (see
References by name);
spec.host_idis the one reference field that still carries a Console UUID, andtelegram_group_bindinghas no by-name fields at all, so a binding that setsenvironment_idorteam_idhas to write the UUID.--prunealso performs no referenced-object check. - No dependency ordering.
pc applyapplies files in lexical order and stops at the first failure; it does not sort by kind and it builds no reference graph. A tree whose files reference each other has to encode the order in its own directory names — see Three layout rules.
client_uptime_notify
A client's two uptime Telegram targets — the internal target (a team/staff chat) and the client target (one of the client's own bound chats) that a monitor's down/up transition notifies, independent of Notify routing's escalation-route notify_telegram. There is exactly one per client, so metadata.name must be default.
kind: client_uptime_notify
version: v1alpha1
metadata:
name: default
client: acme-retail
spec:
internal_chat_id: -1001234567890 # a team chat or internal-audience chat; super-admin or global telegram_ops:write
internal_thread_id: 17118 # optional; omit for the uptime → alerts topic ladder
client_chat_id: -1009876543210 # one of this client's client-audience chats
client_thread_id: null
Reading needs monitors:read. Writing the client target needs telegram_ops:write held client-wide — an environment-scoped grant is not enough, the same rule PUT /clients/{id}/uptime-notify enforces. Writing the internal target needs super-admin or the global telegram_ops:write; a caller without one of those is refused for internal_chat_id/internal_thread_id the moment either key is mentioned at all — matching, changed, or explicit null — not only when it would actually change something, so a guessed value can never be told apart from a wrong one by whether the apply succeeds.
A key you omit keeps its stored value; null clears it — except that changing a *_chat_id to a genuinely different chat while omitting its sibling *_thread_id resets the thread to null rather than carrying the OLD chat's thread onto the new one; give both keys together to pin a specific thread on the new chat. Deleting the object clears both targets outright, and clearing a configured internal target through delete needs the same super-admin/global bar as clearing it through an apply — deleting a row that has no internal target configured needs only the client-wide bar. pc get for a caller who cannot set the internal target omits internal_chat_id/internal_thread_id from the spec entirely — there is no separate "configured" flag the way the HTTP settings API's redacted field has one, so the absence of the keys is the only signal, and it means exactly what it says for pc apply: nothing here to preserve or to change. --prune does not delete this kind.