L1 Incident Agent
The L1 incident agent is the layer between an alert fired and a human reads a root cause. It assembles the evidence deterministically, investigates it with an LLM under hard budgets, verifies its own conclusion against that evidence, and remembers what actually fixed the problem so the next occurrence starts from the answer instead of from zero.
It replaced the incident-response layer that previously sat behind Grafana OnCall. Paging itself is a separate concern — see Paging Control and Escalation.
The switch: ai_triage
Everything on this page is behind one per-client feature flag, ai_triage (Client Settings →
Features → "AI investigation", super-admin, default off). With it off, alerts still ingest,
correlate, group and page exactly as they did — nothing on this page runs: no auto-triage, no
context pack assembled for a model, no memory draft, no remediation proposal, and no reactive
/engage reply in a bound Telegram chat.
In a Telegram chat bound to a team rather than to one client, the flag is a filter on what
that chat may ask about instead of an on/off switch for the chat: the room is the team's clients
with ai_triage on, and a team whose clients all have it off has an empty room and answers
nothing at all. That is the flag working, not a fault — see
Reactive AI Investigator for the rest of that
chat's gate chain, none of which ai_triage can substitute for.
It is not a paging switch, and paging is not gated by it. That is worth stating plainly
because it was not always true. ai_triage and the manual-mode switch under it replace a single
boolean named l1_agent, which was labelled "L1 Agent" and read as an AI switch while quietly
deciding whether anyone got paged. Turning the AI off turned the pager off, and the only evidence
was a no-page reason named l1_disabled. Paging now has its own flag, oncall_enabled, and the
two are independent in both directions:
| You want | oncall_enabled | ai_triage |
|---|---|---|
| Paged, with LLM spend | on | on |
| Paged, no LLM spend — this page does nothing | on | off |
| Investigated but never paged (a staging or read-only client) | off | on |
| Neither | off | off |
The middle two rows were unreachable before the split. See Paging Control for the paging side, and Per-Client Feature Flags for the legacy alias that kept every existing client on its previous behaviour.
Two flags sit under ai_triage and are meaningless without it:
ai_triage_manual— investigation mode. Off (the default) = automatic; on = the correlation worker skips auto-triage and an operator triggers each investigation from the alert. Never affects paging. See Auto-triage vs. manual triage.remediation— the proposal path below. The worker re-checksai_triageandremediationfail-closed, soremediationalone proposes nothing.
The pipeline
Each stage is separately bounded and separately observable. Nothing in this chain pages anyone —
paging is decided by escalation routes behind oncall_enabled, independently of whether triage
ran, succeeded, or is enabled at all.
The context pack
Before any model runs, Console assembles a deterministic, LLM-free context pack for the alert group. The same pack is what the web UI shows you, what the triage worker sends to the model, and what the evaluation harness replays — one shape, three consumers, so what you see is what the model saw.
| Section | Contents |
|---|---|
| Anchor | Severity, host, environment, grouping key, and fired_at as an explicit "now" so the model can judge the recency of every piece of evidence |
| Metrics | A metric snapshot plus the window it was sampled over (e.g. ±30m around fire) |
| Recent changes | Host-scoped changes in the pre-fire window, each stamped with its offset from the fire time and kept in newest-first order. Empty when the alert group has no resolved host. See Change Detection |
| Deploy changes | The git / CI / deploy subset, client-scoped and proximity-ranked, weighted highest as causes |
| Cloud changes | Client-scoped, host-less control-plane changes (Hetzner, Cloudflare) that host-scoped correlation cannot reach |
| Upstream causality | Observed dependency edges whose upstream host is itself firing |
| Prior resolutions | Similar past incidents retrieved from memory |
GET /api/v1/alerts/:alertGroupID/context returns the pack. Reading it is the fastest way to
answer "why did L1 conclude that?" — if the evidence was not in the pack, the model did not have
it.
Investigation, then verification
The investigation loop lets the model call tools — metrics, logs, changes, compliance, Kubernetes events — until it can state a root cause. It is bounded three ways at once, and whichever bound trips first ends the loop:
| Bound | Default | Variable |
|---|---|---|
| Wall clock | 300s | PROXIMA_TRIAGE_DEADLINE |
| Tokens | 120,000 | PROXIMA_TRIAGE_TOKEN_BUDGET |
| Tool loops | 20 | PROXIMA_TRIAGE_MAX_LOOPS |
The verify pass then re-reads the model's cited claims against the evidence and decides whether they hold. It gets its own deadline and token budget so a long investigation cannot starve verification into a fail-open pass:
| Bound | Default | Variable |
|---|---|---|
| Wall clock | 60s | PROXIMA_TRIAGE_VERIFY_DEADLINE |
| Tokens | 8,000 | PROXIMA_TRIAGE_VERIFY_TOKEN_BUDGET |
| Retries on timeout | 1 | PROXIMA_TRIAGE_VERIFY_MAX_RETRIES |
Trust levels
Every RCA carries a trust level, and the engine — never the model — stamps it:
| Trust level | Meaning |
|---|---|
verified | The verifier ran, supported the claims, and calibrated confidence cleared the threshold |
unconfirmed | Verification was inconclusive, or it failed while the investigator had cited at least PROXIMA_TRIAGE_TRUST_MIN_EVIDENCE (default 2) pieces of evidence |
unverified | No usable verification — the claim stands on the model's word alone |
A granular verify_status (verified, inconclusive, refuted, timed_out, errored,
not_run, no_analysis) records why. The RCA also keeps unsupported_claims — the specific
cited claims the evidence did not back — which is the first thing to read when the trust rate
drops.
confidence is the post-verification value. The model's own self-report is kept separately as
self_confidence, for telemetry only. Do not treat a high self-report as a verified conclusion.
The short-circuit
The short-circuit is behind a kill-switch. PROXIMA_TRIAGE_MEMORY_SHORTCIRCUIT_DISTANCE defaults to
0, and 0 disables it entirely — Pack.ShortCircuitCandidate returns no candidate and the
worker never even evaluates the branch. With stock configuration every triage takes the full
investigation path. Production pods set it to 0.15.
Once armed, the variable is a cosine-distance ceiling: when memory returns a prior resolution at or below that distance — and that resolution is held, has not recurred, and was not marked failed — L1 skips the full investigation and runs a cheap confirm pass instead (2 loops, 15,000 tokens, 60s by default). Those budgets are a net saving only because they sit far below the full run's; that relationship is the point, so keep it when tuning. A confirm that does not adopt falls through to the full investigation unchanged.
The memory loop
A triage that produced a real answer is only valuable if the next occurrence can use it.
- Draft. After resolution, L1 drafts a resolution stub: the symptom, the root cause, the fix, and optionally a runbook slug.
- Human confirmation. A person approves it —
POST /api/v1/alerts/:alertGroupID/resolution/approve. The corpus is human-confirmed by construction; every stub records who confirmed it and when. - Retrieval. The symptom is embedded (1536-dim) and stored in
resolution_stubwith a pgvector cosine ANN index. On the next similar alert, the pack retrieves it. - Verdict.
POST /api/v1/alerts/:alertGroupID/fix-workedrecords whether the known fix actually worked, closing the loop.
PROXIMA_MEMORY_IVFFLAT_PROBESThe corpus is small, and PostgreSQL's default ivfflat.probes of 1 usually probes an empty cell —
silently dropping the true nearest neighbour and making memory look empty when it is not. The
default here is 10 for that reason. Lower it only with measurements in hand.
Recurrence and regression
| Signal | Default | Variable |
|---|---|---|
| Chronic lookback | 168h (7d) | PROXIMA_RECURRENCE_CHRONIC_WINDOW |
| Occurrences that mark an alert family chronic | 3 | PROXIMA_RECURRENCE_CHRONIC_THRESHOLD |
| Window in which a fresh occurrence counts as the fix regressing | 48h | PROXIMA_RECURRENCE_FIX_REGRESSION_WINDOW |
Remediation
Where a fix is proposable, L1 produces a remediation proposal against the incident. It is never executed automatically:
GET /api/v1/incidents/:incidentID/remediation
POST /api/v1/incidents/:incidentID/remediation/approve
POST /api/v1/incidents/:incidentID/remediation/reject
POST /api/v1/incidents/:incidentID/remediation/fix-worked
Approval executes through the runbook engine and requires
runbooks:execute. It is a human-only path — a service-account caller gets 403. Every step is
re-validated against the non-destructive catalog immediately before execution, so a proposal cannot
widen its own blast radius between drafting and approval.
Approve returns 409 when the proposal is already approved, executed, or rejected. It
deliberately does not block on failed: a remediation that failed for a transient reason — the
agent was offline, the run timed out — can simply be approved again. The fix-worked verdict feeds
the same learning loop as the memory path, and requires the proposal to be in the executed state
(409 otherwise).
Incident lifecycle
Alert groups are grouped into incidents so that ten symptoms of one failure read as one event.
An open incident with no firing member activity within PROXIMA_INCIDENT_STALE_TTL
(default 6h) is auto-resolved by the reaper, which sweeps every
PROXIMA_INCIDENT_REAP_INTERVAL (default 15m). A chronic alert that keeps re-firing keeps its
members fresh and is never reaped.
Reliability
Triage runs off the ALERTS JetStream stream, and the failure classes are handled differently on
purpose:
- Terminal errors — a content / bad-request or unknown-model gateway error (a 4xx carrying no
operational marker). A retry cannot help, so the message is
Termed immediately and dead-letters with reasonterminal. - Transient errors — rate limits, overload, 5xx, timeouts, and generic quota errors. Up to 35
deliveries with an escalating NAK backoff (15s, 30s, … capped at 5m), summing to roughly two
hours, then dead-lettering as
transient_exhausted. That window is deliberately long enough for capacity to recover. - Operator-action failures — a provider credit balance, "plans & billing", or "budget has
been exceeded" error. These are not retried: they
Termon the first delivery and dead-letter asoperator_action_required, because every redelivery re-bills the tool calls that still succeed. Triage resumes only once a human tops up or raises the cap, and the alert is re-triaged. - Other operational failures use a flat
PROXIMA_TRIAGE_OPERATIONAL_NAK_BACKOFF(default 5m) floor, because capacity pressure will not clear in seconds. - A bad triage model alias falls back once to the code-default expert model rather than dead-lettering the alert.
A malformed NATS payload is neither of the above: it is ACKed and dropped to avoid a redelivery loop, with no dead-letter metric recorded.
Observability
| Metric | Tells you |
|---|---|
proxima_triage_trust_total | Trust-level distribution — the headline quality signal |
proxima_triage_verify_total, proxima_triage_verify_duration_seconds, proxima_triage_verify_retries_total | Verify-pass health |
proxima_triage_confidence | Calibrated confidence distribution |
proxima_triage_investigation_loops | Loop depth per triage |
proxima_triage_token_budget_exceeded_total, proxima_triage_timeout_total | Runs hitting their bounds |
proxima_triage_dead_letter_total, proxima_triage_error_class_total, proxima_triage_llm_error_total | Failure classes |
proxima_triage_shortcircuit_total, proxima_triage_resolution_used_total | Memory-loop payoff |
proxima_triage_change_cause_used_total, proxima_triage_upstream_cause_used_total | Which causal inputs are actually landing |
proxima_triage_pack_enrichment_degraded_total | The pack was assembled with a section missing |
proxima_incident_grouping_total, proxima_incident_reaped_total, proxima_incident_retriage_total | Incident lifecycle |
proxima_triage_pack_enrichment_degraded_total is the one to alert on first: a degraded pack
means the model reasoned without evidence it should have had, and the resulting RCA can look
confident and be wrong.
API
| Method | Path | Description |
|---|---|---|
| GET | /api/v1/alerts/:alertGroupID/context | The deterministic context pack |
| POST | /api/v1/alerts/:alertGroupID/triage | Trigger an on-demand investigation |
| POST | /api/v1/alerts/:alertGroupID/resolution/approve | Approve a resolution into memory |
| POST | /api/v1/alerts/:alertGroupID/fix-worked | Record whether the known fix worked |
| GET | /api/v1/incidents/:incidentID | Incident detail |
| GET | /api/v1/incidents/:incidentID/session | The investigation session |
| GET | /api/v1/incidents/:incidentID/session/stream | Stream the session (SSE) |
| GET | /api/v1/incidents/:incidentID/dependencies | Host dependency graph |
| POST | /api/v1/incidents/:incidentID/feedback | Submit RCA feedback |
| GET | /api/v1/l1/models | List allowlisted triage models (admin) |
| PUT | /api/v1/l1/settings/triage-model | Set the global triage model (admin) |
Full list in the API Endpoints reference.
See also
- Alerting & Correlation — ingestion, label mapping, alert lifecycle
- Change Detection — where cited causes come from
- Escalation — who gets paged, independently of triage
- Paging Control —
oncall_enabled, routes, and the paging control surface - Per-Client Feature Flags — every flag and its default
- Environment Variables — every knob above, with defaults
- Ops Bot — the Telegram surface for acting on incidents