Skip to main content

L1 Incident Agent

The L1 incident agent is the layer between an alert fired and a human reads a root cause. It assembles the evidence deterministically, investigates it with an LLM under hard budgets, verifies its own conclusion against that evidence, and remembers what actually fixed the problem so the next occurrence starts from the answer instead of from zero.

It replaced the incident-response layer that previously sat behind Grafana OnCall. Paging itself is a separate concern — see Paging Control and Escalation.

The switch: ai_triage​

Everything on this page is behind one per-client feature flag, ai_triage (Client Settings → Features → "AI investigation", super-admin, default off). With it off, alerts still ingest, correlate, group and page exactly as they did — nothing on this page runs: no auto-triage, no context pack assembled for a model, no memory draft, no remediation proposal, and no reactive /engage reply in a bound Telegram chat.

In a Telegram chat bound to a team rather than to one client, the flag is a filter on what that chat may ask about instead of an on/off switch for the chat: the room is the team's clients with ai_triage on, and a team whose clients all have it off has an empty room and answers nothing at all. That is the flag working, not a fault — see Reactive AI Investigator for the rest of that chat's gate chain, none of which ai_triage can substitute for.

It is not a paging switch, and paging is not gated by it. That is worth stating plainly because it was not always true. ai_triage and the manual-mode switch under it replace a single boolean named l1_agent, which was labelled "L1 Agent" and read as an AI switch while quietly deciding whether anyone got paged. Turning the AI off turned the pager off, and the only evidence was a no-page reason named l1_disabled. Paging now has its own flag, oncall_enabled, and the two are independent in both directions:

You wantoncall_enabledai_triage
Paged, with LLM spendonon
Paged, no LLM spend — this page does nothingonoff
Investigated but never paged (a staging or read-only client)offon
Neitheroffoff

The middle two rows were unreachable before the split. See Paging Control for the paging side, and Per-Client Feature Flags for the legacy alias that kept every existing client on its previous behaviour.

Two flags sit under ai_triage and are meaningless without it:

  • ai_triage_manual — investigation mode. Off (the default) = automatic; on = the correlation worker skips auto-triage and an operator triggers each investigation from the alert. Never affects paging. See Auto-triage vs. manual triage.
  • remediation — the proposal path below. The worker re-checks ai_triage and remediation fail-closed, so remediation alone proposes nothing.

The pipeline​

Each stage is separately bounded and separately observable. Nothing in this chain pages anyone — paging is decided by escalation routes behind oncall_enabled, independently of whether triage ran, succeeded, or is enabled at all.

The context pack​

Before any model runs, Console assembles a deterministic, LLM-free context pack for the alert group. The same pack is what the web UI shows you, what the triage worker sends to the model, and what the evaluation harness replays — one shape, three consumers, so what you see is what the model saw.

SectionContents
AnchorSeverity, host, environment, grouping key, and fired_at as an explicit "now" so the model can judge the recency of every piece of evidence
MetricsA metric snapshot plus the window it was sampled over (e.g. ±30m around fire)
Recent changesHost-scoped changes in the pre-fire window, each stamped with its offset from the fire time and kept in newest-first order. Empty when the alert group has no resolved host. See Change Detection
Deploy changesThe git / CI / deploy subset, client-scoped and proximity-ranked, weighted highest as causes
Cloud changesClient-scoped, host-less control-plane changes (Hetzner, Cloudflare) that host-scoped correlation cannot reach
Upstream causalityObserved dependency edges whose upstream host is itself firing
Prior resolutionsSimilar past incidents retrieved from memory

GET /api/v1/alerts/:alertGroupID/context returns the pack. Reading it is the fastest way to answer "why did L1 conclude that?" — if the evidence was not in the pack, the model did not have it.

Investigation, then verification​

The investigation loop lets the model call tools — metrics, logs, changes, compliance, Kubernetes events — until it can state a root cause. It is bounded three ways at once, and whichever bound trips first ends the loop:

BoundDefaultVariable
Wall clock300sPROXIMA_TRIAGE_DEADLINE
Tokens120,000PROXIMA_TRIAGE_TOKEN_BUDGET
Tool loops20PROXIMA_TRIAGE_MAX_LOOPS

The verify pass then re-reads the model's cited claims against the evidence and decides whether they hold. It gets its own deadline and token budget so a long investigation cannot starve verification into a fail-open pass:

BoundDefaultVariable
Wall clock60sPROXIMA_TRIAGE_VERIFY_DEADLINE
Tokens8,000PROXIMA_TRIAGE_VERIFY_TOKEN_BUDGET
Retries on timeout1PROXIMA_TRIAGE_VERIFY_MAX_RETRIES

Trust levels​

Every RCA carries a trust level, and the engine — never the model — stamps it:

Trust levelMeaning
verifiedThe verifier ran, supported the claims, and calibrated confidence cleared the threshold
unconfirmedVerification was inconclusive, or it failed while the investigator had cited at least PROXIMA_TRIAGE_TRUST_MIN_EVIDENCE (default 2) pieces of evidence
unverifiedNo usable verification — the claim stands on the model's word alone

A granular verify_status (verified, inconclusive, refuted, timed_out, errored, not_run, no_analysis) records why. The RCA also keeps unsupported_claims — the specific cited claims the evidence did not back — which is the first thing to read when the trust rate drops.

Confidence is calibrated, not self-reported

confidence is the post-verification value. The model's own self-report is kept separately as self_confidence, for telemetry only. Do not treat a high self-report as a verified conclusion.

The short-circuit​

Ships disabled

The short-circuit is behind a kill-switch. PROXIMA_TRIAGE_MEMORY_SHORTCIRCUIT_DISTANCE defaults to 0, and 0 disables it entirely — Pack.ShortCircuitCandidate returns no candidate and the worker never even evaluates the branch. With stock configuration every triage takes the full investigation path. Production pods set it to 0.15.

Once armed, the variable is a cosine-distance ceiling: when memory returns a prior resolution at or below that distance — and that resolution is held, has not recurred, and was not marked failed — L1 skips the full investigation and runs a cheap confirm pass instead (2 loops, 15,000 tokens, 60s by default). Those budgets are a net saving only because they sit far below the full run's; that relationship is the point, so keep it when tuning. A confirm that does not adopt falls through to the full investigation unchanged.

The memory loop​

A triage that produced a real answer is only valuable if the next occurrence can use it.

  1. Draft. After resolution, L1 drafts a resolution stub: the symptom, the root cause, the fix, and optionally a runbook slug.
  2. Human confirmation. A person approves it — POST /api/v1/alerts/:alertGroupID/resolution/approve. The corpus is human-confirmed by construction; every stub records who confirmed it and when.
  3. Retrieval. The symptom is embedded (1536-dim) and stored in resolution_stub with a pgvector cosine ANN index. On the next similar alert, the pack retrieves it.
  4. Verdict. POST /api/v1/alerts/:alertGroupID/fix-worked records whether the known fix actually worked, closing the loop.
Recall depends on PROXIMA_MEMORY_IVFFLAT_PROBES

The corpus is small, and PostgreSQL's default ivfflat.probes of 1 usually probes an empty cell — silently dropping the true nearest neighbour and making memory look empty when it is not. The default here is 10 for that reason. Lower it only with measurements in hand.

Recurrence and regression​

SignalDefaultVariable
Chronic lookback168h (7d)PROXIMA_RECURRENCE_CHRONIC_WINDOW
Occurrences that mark an alert family chronic3PROXIMA_RECURRENCE_CHRONIC_THRESHOLD
Window in which a fresh occurrence counts as the fix regressing48hPROXIMA_RECURRENCE_FIX_REGRESSION_WINDOW

Remediation​

Where a fix is proposable, L1 produces a remediation proposal against the incident. It is never executed automatically:

GET  /api/v1/incidents/:incidentID/remediation
POST /api/v1/incidents/:incidentID/remediation/approve
POST /api/v1/incidents/:incidentID/remediation/reject
POST /api/v1/incidents/:incidentID/remediation/fix-worked

Approval executes through the runbook engine and requires runbooks:execute. It is a human-only path — a service-account caller gets 403. Every step is re-validated against the non-destructive catalog immediately before execution, so a proposal cannot widen its own blast radius between drafting and approval.

Approve returns 409 when the proposal is already approved, executed, or rejected. It deliberately does not block on failed: a remediation that failed for a transient reason — the agent was offline, the run timed out — can simply be approved again. The fix-worked verdict feeds the same learning loop as the memory path, and requires the proposal to be in the executed state (409 otherwise).

Incident lifecycle​

Alert groups are grouped into incidents so that ten symptoms of one failure read as one event. An open incident with no firing member activity within PROXIMA_INCIDENT_STALE_TTL (default 6h) is auto-resolved by the reaper, which sweeps every PROXIMA_INCIDENT_REAP_INTERVAL (default 15m). A chronic alert that keeps re-firing keeps its members fresh and is never reaped.

Reliability​

Triage runs off the ALERTS JetStream stream, and the failure classes are handled differently on purpose:

  • Terminal errors — a content / bad-request or unknown-model gateway error (a 4xx carrying no operational marker). A retry cannot help, so the message is Termed immediately and dead-letters with reason terminal.
  • Transient errors — rate limits, overload, 5xx, timeouts, and generic quota errors. Up to 35 deliveries with an escalating NAK backoff (15s, 30s, … capped at 5m), summing to roughly two hours, then dead-lettering as transient_exhausted. That window is deliberately long enough for capacity to recover.
  • Operator-action failures — a provider credit balance, "plans & billing", or "budget has been exceeded" error. These are not retried: they Term on the first delivery and dead-letter as operator_action_required, because every redelivery re-bills the tool calls that still succeed. Triage resumes only once a human tops up or raises the cap, and the alert is re-triaged.
  • Other operational failures use a flat PROXIMA_TRIAGE_OPERATIONAL_NAK_BACKOFF (default 5m) floor, because capacity pressure will not clear in seconds.
  • A bad triage model alias falls back once to the code-default expert model rather than dead-lettering the alert.

A malformed NATS payload is neither of the above: it is ACKed and dropped to avoid a redelivery loop, with no dead-letter metric recorded.

Observability​

MetricTells you
proxima_triage_trust_totalTrust-level distribution — the headline quality signal
proxima_triage_verify_total, proxima_triage_verify_duration_seconds, proxima_triage_verify_retries_totalVerify-pass health
proxima_triage_confidenceCalibrated confidence distribution
proxima_triage_investigation_loopsLoop depth per triage
proxima_triage_token_budget_exceeded_total, proxima_triage_timeout_totalRuns hitting their bounds
proxima_triage_dead_letter_total, proxima_triage_error_class_total, proxima_triage_llm_error_totalFailure classes
proxima_triage_shortcircuit_total, proxima_triage_resolution_used_totalMemory-loop payoff
proxima_triage_change_cause_used_total, proxima_triage_upstream_cause_used_totalWhich causal inputs are actually landing
proxima_triage_pack_enrichment_degraded_totalThe pack was assembled with a section missing
proxima_incident_grouping_total, proxima_incident_reaped_total, proxima_incident_retriage_totalIncident lifecycle

proxima_triage_pack_enrichment_degraded_total is the one to alert on first: a degraded pack means the model reasoned without evidence it should have had, and the resulting RCA can look confident and be wrong.

API​

MethodPathDescription
GET/api/v1/alerts/:alertGroupID/contextThe deterministic context pack
POST/api/v1/alerts/:alertGroupID/triageTrigger an on-demand investigation
POST/api/v1/alerts/:alertGroupID/resolution/approveApprove a resolution into memory
POST/api/v1/alerts/:alertGroupID/fix-workedRecord whether the known fix worked
GET/api/v1/incidents/:incidentIDIncident detail
GET/api/v1/incidents/:incidentID/sessionThe investigation session
GET/api/v1/incidents/:incidentID/session/streamStream the session (SSE)
GET/api/v1/incidents/:incidentID/dependenciesHost dependency graph
POST/api/v1/incidents/:incidentID/feedbackSubmit RCA feedback
GET/api/v1/l1/modelsList allowlisted triage models (admin)
PUT/api/v1/l1/settings/triage-modelSet the global triage model (admin)

Full list in the API Endpoints reference.

See also​