Skip to main content

AI Chat Observability

Prometheus Metrics​

MetricTypeLabelsDescription
proxima_chat_messages_totalcounterroleTotal chat messages (user + assistant)
proxima_chat_tokens_totalcounterdirection (input/output), modelToken consumption
proxima_chat_cost_usd_totalcountermodelEstimated AI cost from LiteLLM
proxima_chat_tool_calls_totalcountertool, statusTool call count
proxima_chat_tool_duration_secondshistogramtoolPer-tool execution latency
proxima_chat_tool_dedup_hits_totalcountertoolRepeated identical tool calls served from the dedup cache
proxima_chat_response_duration_secondshistogrammodelFull response latency (user-perceived)
proxima_chat_model_routing_totalcountertier, reasonRouting decisions (llm/keyword_fallback/user_override)
proxima_chat_conversations_activegauge—Active conversations
proxima_chat_errors_totalcountererror_typeError breakdown. Only two values are ever emitted: llm_stream and rate_limited.
proxima_chat_credential_source_totalcountersourceCredential source — currently always platform (BYOK is planned, not yet implemented)

Cost tracking: proxima_chat_cost_usd_total reads the cost from LiteLLM's x-litellm-response-cost-original response header, falling back to the legacy x-litellm-model-response-cost — no manual price table maintenance needed.

No chat metric carries a client_id label

Per-client breakdowns of chat volume, cost or tool calls are not available from these metrics. client_id is set as a span attribute on ChatHandler.Chat, not as a metric attribute, so a PromQL by (client_id) returns a single empty-label series.

OpenTelemetry Traces​

ChatHandler.Chat (root span; client_id attribute)
├── chat.context_prefetch
├── chat.router (model routing)
├── chat.llm_call (model, tokens_in, tokens_out, cost)
├── chat.investigator
├── chat.encryption.dek_resolve (cache hit/miss attribute)
├── chat.encryption.encrypt / .decrypt
└── chat.encryption.vault_encrypt / .vault_decrypt

ChatHandler.Confirm (tool-confirmation path)
chat.nl_translate (natural language → PQL)
Tool execution is not traced

There is no per-tool span. Tool latency and outcome are visible only through the proxima_chat_tool_duration_seconds and proxima_chat_tool_calls_total metrics.

Structured Logging​

All chat operations log with component=api.chat or component=chat.*:

LevelEventWhen
INFOAI chat enabledBackend startup with LiteLLM configured
INFOmodel routing via LLMEvery classification with tier + message preview
INFODEK created, encrypted, and cachedFirst chat for a new client
WARNLLM classification failed, falling back to keywordsLiteLLM unreachable
WARNcontext pre-fetch failedStore error during pre-fetch (graceful)
WARNValkey unavailable, auto-rejecting confirmationValkey down during write tool
ERRORLLM API errorLiteLLM/provider failure with status code
ERRORfailed to encrypt/decrypt DEKVault Transit failure

Grafana Dashboard​

The backend observability dashboard includes an "AI Chat" row with the following panels:

  • Token usage over time — legend {{direction}} - {{model}} (there is no per-client series)
  • AI cost — legend {{model}}
  • Chat messages
  • Tool call popularity (top tools by call count)
  • Response latency P50/P95/P99
  • Credential source (currently always platform — BYOK is planned)
  • Error rate by error_type
  • Routing tier distribution
  • Active conversations (a stat panel, not a gauge)

LLM Evaluation​

Promptfoo evaluation suite in tests/llm-eval/:

# Model routing accuracy (23 test cases)
DEEPSEEK_API_KEY=sk-... npx promptfoo eval -c tests/llm-eval/promptfooconfig.yaml

# Prompt injection defense (10 adversarial cases)
DEEPSEEK_API_KEY=sk-... npx promptfoo eval -c tests/llm-eval/promptfoo-redteam.yaml

# View HTML dashboard
npx promptfoo view