AI Chat Observability
Prometheus Metrics
| Metric | Type | Labels | Description |
|---|---|---|---|
proxima_chat_messages_total | counter | role | Total chat messages (user + assistant) |
proxima_chat_tokens_total | counter | direction (input/output), model | Token consumption |
proxima_chat_cost_usd_total | counter | model | Estimated AI cost from LiteLLM |
proxima_chat_tool_calls_total | counter | tool, status | Tool call count |
proxima_chat_tool_duration_seconds | histogram | tool | Per-tool execution latency |
proxima_chat_tool_dedup_hits_total | counter | tool | Repeated identical tool calls served from the dedup cache |
proxima_chat_response_duration_seconds | histogram | model | Full response latency (user-perceived) |
proxima_chat_model_routing_total | counter | tier, reason | Routing decisions (llm/keyword_fallback/user_override) |
proxima_chat_conversations_active | gauge | — | Active conversations |
proxima_chat_errors_total | counter | error_type | Error breakdown. Only two values are ever emitted: llm_stream and rate_limited. |
proxima_chat_credential_source_total | counter | source | Credential source — currently always platform (BYOK is planned, not yet implemented) |
Cost tracking: proxima_chat_cost_usd_total reads the cost from LiteLLM's
x-litellm-response-cost-original response header, falling back to the legacy
x-litellm-model-response-cost — no manual price table maintenance needed.
client_id labelPer-client breakdowns of chat volume, cost or tool calls are not available from these metrics.
client_id is set as a span attribute on ChatHandler.Chat, not as a metric attribute, so a
PromQL by (client_id) returns a single empty-label series.
OpenTelemetry Traces
ChatHandler.Chat (root span; client_id attribute)
├── chat.context_prefetch
├── chat.router (model routing)
├── chat.llm_call (model, tokens_in, tokens_out, cost)
├── chat.investigator
├── chat.encryption.dek_resolve (cache hit/miss attribute)
├── chat.encryption.encrypt / .decrypt
└── chat.encryption.vault_encrypt / .vault_decrypt
ChatHandler.Confirm (tool-confirmation path)
chat.nl_translate (natural language → PQL)
There is no per-tool span. Tool latency and outcome are visible only through the
proxima_chat_tool_duration_seconds and proxima_chat_tool_calls_total metrics.
Structured Logging
All chat operations log with component=api.chat or component=chat.*:
| Level | Event | When |
|---|---|---|
| INFO | AI chat enabled | Backend startup with LiteLLM configured |
| INFO | model routing via LLM | Every classification with tier + message preview |
| INFO | DEK created, encrypted, and cached | First chat for a new client |
| WARN | LLM classification failed, falling back to keywords | LiteLLM unreachable |
| WARN | context pre-fetch failed | Store error during pre-fetch (graceful) |
| WARN | Valkey unavailable, auto-rejecting confirmation | Valkey down during write tool |
| ERROR | LLM API error | LiteLLM/provider failure with status code |
| ERROR | failed to encrypt/decrypt DEK | Vault Transit failure |
Grafana Dashboard
The backend observability dashboard includes an "AI Chat" row with the following panels:
- Token usage over time — legend
{{direction}} - {{model}}(there is no per-client series) - AI cost — legend
{{model}} - Chat messages
- Tool call popularity (top tools by call count)
- Response latency P50/P95/P99
- Credential source (currently always
platform— BYOK is planned) - Error rate by
error_type - Routing tier distribution
- Active conversations (a stat panel, not a gauge)
LLM Evaluation
Promptfoo evaluation suite in tests/llm-eval/:
# Model routing accuracy (23 test cases)
DEEPSEEK_API_KEY=sk-... npx promptfoo eval -c tests/llm-eval/promptfooconfig.yaml
# Prompt injection defense (10 adversarial cases)
DEEPSEEK_API_KEY=sk-... npx promptfoo eval -c tests/llm-eval/promptfoo-redteam.yaml
# View HTML dashboard
npx promptfoo view