Metrics Collected
The agent collects system metrics via gopsutil v4 and Linux /proc files, plus plugin metrics from configured integrations. All metrics follow the OpenMetrics naming convention.
Raw values: The agent emits raw counters and gauges. The derived ratio metrics the UI shows (cpu_usage_ratio, memory_used_ratio, disk_used_ratio, …) are computed at query time in the backend from these raw values, scoped to the caller's tenant — see Derived metrics.
System Metrics
Collected automatically on every managed host. No configuration required.
CPU
| Metric | Type | Unit | Description |
|---|---|---|---|
cpu_user_seconds_total | Counter | seconds | Time spent in user mode |
cpu_system_seconds_total | Counter | seconds | Time spent in kernel mode |
cpu_nice_seconds_total | Counter | seconds | Time spent in low-priority user mode |
cpu_idle_seconds_total | Counter | seconds | Time spent idle |
cpu_iowait_seconds_total | Counter | seconds | Time spent waiting for I/O |
cpu_irq_seconds_total | Counter | seconds | Time spent servicing hardware interrupts |
cpu_softirq_seconds_total | Counter | seconds | Time spent servicing software interrupts |
cpu_steal_seconds_total | Counter | seconds | Time stolen by hypervisor |
cpu_load1 | Gauge | — | 1-minute load average |
cpu_load5 | Gauge | — | 5-minute load average |
cpu_load15 | Gauge | — | 15-minute load average |
cpu_usage_ratio and cpu_iowait_ratio are computed at query time in the backend from the per-mode counters above. They produce the same 0.0–1.0 ratios the frontend displays.
Memory
| Metric | Type | Unit | Description |
|---|---|---|---|
memory_total_bytes | Gauge | bytes | Total physical memory |
memory_used_bytes | Gauge | bytes | Memory in use (excluding buffers/cache) |
memory_available_bytes | Gauge | bytes | Memory available without swapping |
memory_buffers_bytes | Gauge | bytes | Kernel buffer cache |
memory_cached_bytes | Gauge | bytes | Page cache (reclaimable) |
memory_swap_total_bytes | Gauge | bytes | Total swap space |
memory_swap_free_bytes | Gauge | bytes | Free swap space |
memory_used_ratio and memory_swap_used_ratio are computed at query time in the backend from raw bytes.
Disk
Per-partition and per-device metrics. Pseudo-filesystems (tmpfs, sysfs, proc, overlay) are excluded.
Usage (per partition):
| Metric | Type | Unit | Labels | Description |
|---|---|---|---|---|
disk_total_bytes | Gauge | bytes | mount, device | Total partition size |
disk_used_bytes | Gauge | bytes | mount, device | Used space |
disk_free_bytes | Gauge | bytes | mount, device | Free space |
disk_inodes_total | Gauge | — | mount, device | Total inodes |
disk_inodes_free | Gauge | — | mount, device | Free inodes |
I/O (per device):
| Metric | Type | Unit | Labels | Description |
|---|---|---|---|---|
disk_read_bytes_total | Counter | bytes | device | Total bytes read |
disk_written_bytes_total | Counter | bytes | device | Total bytes written |
disk_reads_total | Counter | — | device | Total read operations |
disk_writes_total | Counter | — | device | Total write operations |
disk_io_time_ms_total | Counter | ms | device | Time spent doing I/O |
disk_io_weighted_ms_total | Counter | ms | device | Weighted I/O time (saturation signal) |
disk_io_in_progress | Gauge | — | device | I/O operations currently in flight |
disk_read_time_ms_total | Counter | ms | device | Total read latency |
disk_write_time_ms_total | Counter | ms | device | Total write latency |
disk_used_ratio and disk_inodes_used_ratio are computed at query time in the backend from raw values.
disk_io_in_progress > 0 consistently indicates disk saturation. disk_io_weighted_ms_total is the USE method's saturation signal for storage.
Network
Per-interface metrics. The loopback interface (lo) is excluded.
| Metric | Type | Unit | Labels | Description |
|---|---|---|---|---|
network_receive_bytes_total | Counter | bytes | interface | Total bytes received |
network_transmit_bytes_total | Counter | bytes | interface | Total bytes transmitted |
network_receive_packets_total | Counter | — | interface | Total packets received |
network_transmit_packets_total | Counter | — | interface | Total packets transmitted |
network_receive_errors_total | Counter | — | interface | Total receive errors |
network_transmit_errors_total | Counter | — | interface | Total transmit errors |
network_receive_drops_total | Counter | — | interface | Total receive drops (buffer exhaustion) |
network_transmit_drops_total | Counter | — | interface | Total transmit drops |
Pressure Stall Information (PSI)
Available on Linux kernel 4.20+. Skipped silently on older kernels.
| Metric | Type | Description |
|---|---|---|
psi_cpu_some_avg10 | Gauge | % of time tasks stalled on CPU (10s avg) |
psi_cpu_some_avg60 | Gauge | % of time tasks stalled on CPU (60s avg) |
psi_cpu_some_total_microseconds | Counter | Total CPU stall time |
psi_memory_some_avg10 | Gauge | % of time some tasks stalled on memory (10s avg) |
psi_memory_some_avg60 | Gauge | % of time some tasks stalled on memory (60s avg) |
psi_memory_full_avg10 | Gauge | % of time all tasks stalled on memory (10s avg) |
psi_memory_full_avg60 | Gauge | % of time all tasks stalled on memory (60s avg) |
psi_io_some_avg10 | Gauge | % of time some tasks stalled on I/O (10s avg) |
psi_io_some_avg60 | Gauge | % of time some tasks stalled on I/O (60s avg) |
psi_io_full_avg10 | Gauge | % of time all tasks stalled on I/O (10s avg) |
psi_io_full_avg60 | Gauge | % of time all tasks stalled on I/O (60s avg) |
PSI is the single most valuable metric for automated root cause analysis. Unlike CPU usage (which can be 100% and healthy), PSI measures actual stall time — when tasks are blocked waiting for resources.
Kernel & System State
| Metric | Type | Description |
|---|---|---|
vmstat_oom_kill_total | Counter | OOM kill events (from /proc/vmstat) |
vmstat_pgmajfault_total | Counter | Major page faults (memory pressure → disk I/O) |
vmstat_pswpin_total | Counter | Pages swapped in |
vmstat_pswpout_total | Counter | Pages swapped out |
context_switches_total | Counter | System-wide context switches |
processes_created_total | Counter | Total processes created (fork rate) |
processes_running | Gauge | Processes currently running |
processes_blocked | Gauge | Processes blocked on I/O |
entropy_available | Gauge | Available entropy (always 256 on kernel 5.4+) |
TCP & Resource Limits
| Metric | Type | Description |
|---|---|---|
sockstat_tcp_inuse | Gauge | TCP sockets in use |
sockstat_tcp_orphan | Gauge | Orphaned TCP sockets (no process) |
sockstat_tcp_tw | Gauge | TCP sockets in TIME_WAIT state |
sockstat_sockets_used | Gauge | Total sockets in use |
filefd_allocated | Gauge | System-wide allocated file descriptors |
filefd_max | Gauge | System-wide file descriptor limit |
conntrack_entries | Gauge | Conntrack table entries (skipped if netfilter not loaded) |
Process Metrics
Top-N process metrics are exported to VictoriaMetrics alongside system metrics. Only the top 20 processes by CPU usage are exported per collection cycle to control cardinality.
| Metric | Type | Unit | Labels | Description |
|---|---|---|---|---|
process_cpu_usage_ratio | Gauge | ratio | pid, process_name, username | Per-process CPU usage (0.0–1.0) |
process_memory_usage_ratio | Gauge | ratio | pid, process_name, username | Per-process memory usage (0.0–1.0) |
process_memory_rss_bytes | Gauge | bytes | pid, process_name, username | Resident Set Size |
process_threads | Gauge | — | pid, process_name, username | Thread count |
The full process list is stored in PostgreSQL and available via GET /api/v1/hosts/{id}/processes. Only the top 20 by CPU are exported as time-series metrics to VictoriaMetrics.
Drill-down endpoint: GET /api/v1/hosts/{id}/processes/{pid}/metrics?process_name=X returns all 4 process metrics as time-series for a specific process. The process_name parameter is required to disambiguate PID reuse across collection cycles.
Integration Metrics
Plugin collectors extend monitoring beyond system metrics. Each plugin injects a collector=<name> label into all its metrics for differentiation.
PostgreSQL
See the PostgreSQL Integration page for setup and configuration.
| Metric | Type | Unit | Labels | Description |
|---|---|---|---|---|
pg_connections_active | Gauge | — | — | Active connections |
pg_connections_idle | Gauge | — | — | Idle connections |
pg_connections_max | Gauge | — | — | Maximum allowed connections |
pg_blocks_hit_total | Counter | — | — | Buffer cache block hits |
pg_blocks_read_total | Counter | — | — | Disk block reads |
pg_replication_lag_seconds | Gauge | seconds | — | Replication lag |
pg_database_size_bytes | Gauge | bytes | database | Database size |
pg_deadlocks_total | Counter | — | — | Total deadlocks |
pg_xact_commit_total | Counter | — | — | Total committed transactions |
pg_xact_rollback_total | Counter | — | — | Total rolled-back transactions |
pg_dead_tuples | Gauge | — | table | Dead rows (top 50 tables) |
pg_stat_statements_count | Gauge | — | — | Tracked query statements |
pg_stat_statements_calls_total | Counter | — | — | Total statement executions |
pg_stat_statements_time_seconds_total | Counter | seconds | — | Total statement execution time |
pg_stat_statements_slow_queries | Gauge | — | — | Statements with mean > 100ms |
pg_cache_hit_ratio is computed at query time in the backend from pg_blocks_hit_total and pg_blocks_read_total.
pg_stat_statements metrics require the PostgreSQL extension to be enabled. If the extension is not installed, these metrics are silently skipped.
Nginx
See the Nginx Integration page for setup and configuration.
| Metric | Type | Unit | Labels | Description |
|---|---|---|---|---|
nginx_connections_active | Gauge | — | — | Current active connections |
nginx_connections_accepted_total | Counter | — | — | Total accepted connections |
nginx_connections_handled_total | Counter | — | — | Total handled connections |
nginx_requests_total | Counter | — | — | Total HTTP requests |
nginx_connections_reading | Gauge | — | — | Connections reading request header |
nginx_connections_writing | Gauge | — | — | Connections writing response |
nginx_connections_waiting | Gauge | — | — | Idle keep-alive connections |
Docker
See the Docker Integration page for setup and configuration.
| Metric | Type | Unit | Labels | Description |
|---|---|---|---|---|
docker_containers | Gauge | — | state | Container count by state (running, exited, created, etc.) |
docker_container_cpu_seconds_total | Counter | seconds | container_name, container_id | Cumulative CPU time consumed |
docker_container_memory_used_bytes | Gauge | bytes | container_name, container_id | Memory usage minus cache |
docker_container_memory_limit_bytes | Gauge | bytes | container_name, container_id | Memory limit (omitted if unlimited) |
docker_container_network_receive_bytes_total | Counter | bytes | container_name, container_id | Total bytes received (all interfaces) |
docker_container_network_transmit_bytes_total | Counter | bytes | container_name, container_id | Total bytes transmitted (all interfaces) |
docker_container_block_read_bytes_total | Counter | bytes | container_name, container_id | Total block device bytes read |
docker_container_block_write_bytes_total | Counter | bytes | container_name, container_id | Total block device bytes written |
docker_container_cpu_usage_ratio and docker_container_memory_usage_ratio are computed at query time in the backend.
Per-container metrics are only collected for running containers. Memory limit metrics are omitted when the container has no memory limit set.
Agent Self-Monitoring
The agent collects metrics about its own resource usage. These are always enabled and require no configuration. All metrics carry agent_version and hostname labels, enabling cross-version regression detection and multi-host comparison.
A pre-built Grafana dashboard is available at infra/grafana/dashboards/proxima/agent-resources.json.
Process
| Metric | Type | Unit | Description |
|---|---|---|---|
agent_build_info | Gauge | — | Constant 1; build identity carried in labels (agent_version, hostname, go_version). Use count by (agent_version) (agent_build_info) to see the fleet's version distribution and observe a rollout landing. |
agent_uptime_seconds | Gauge | seconds | Agent process uptime |
agent_cpu_seconds_total | Counter | seconds | Cumulative CPU time (user + system) |
agent_memory_rss_bytes | Gauge | bytes | Resident set size |
agent_memory_vms_bytes | Gauge | bytes | Virtual memory size |
agent_open_fds | Gauge | — | Open file descriptors |
agent_threads | Gauge | — | OS thread count |
agent_ctx_switches_voluntary_total | Counter | — | Voluntary context switches |
agent_ctx_switches_involuntary_total | Counter | — | Involuntary context switches |
agent_io_read_bytes_total | Counter | bytes | Process disk read bytes |
agent_io_write_bytes_total | Counter | bytes | Process disk write bytes |
agent_cpu_ratio is computed at query time in the backend from agent_cpu_seconds_total.
Go Runtime
| Metric | Type | Unit | Description |
|---|---|---|---|
agent_go_heap_alloc_bytes | Gauge | bytes | Go heap allocation |
agent_go_heap_sys_bytes | Gauge | bytes | Go heap obtained from OS |
agent_go_heap_objects | Gauge | — | Live heap object count |
agent_go_stack_bytes | Gauge | bytes | Go stack memory in use |
agent_go_sys_bytes | Gauge | bytes | Total Go memory from OS |
agent_go_memlimit_bytes | Gauge | bytes | Soft memory limit (only present when PROXIMA_AGENT_MEMORY_LIMIT_MB is set) |
agent_goroutines | Gauge | — | Goroutine count |
agent_gc_cycles_total | Counter | — | GC cycle count |
agent_gc_pause_seconds_total | Counter | seconds | Total GC pause time |
NATS Transport
| Metric | Type | Unit | Description |
|---|---|---|---|
agent_nats_messages_out_total | Counter | — | Messages published to NATS |
agent_nats_bytes_out_total | Counter | bytes | Bytes published to NATS |
agent_nats_messages_in_total | Counter | — | Messages received from NATS (request-reply) |
agent_nats_bytes_in_total | Counter | bytes | Bytes received from NATS |
agent_nats_reconnects_total | Counter | — | NATS reconnection count |
agent_nats_outbox_length | Gauge | — | Outbox queue depth (indicates backpressure when > 0) |
agent_nats_rtt_seconds | Gauge | seconds | Round-trip time to NATS server |
Offline Buffer
These metrics are present only when the transport exposes disk-buffer stats. While NATS is unreachable, the agent buffers outgoing messages to disk and replays them on reconnect.
| Metric | Type | Unit | Description |
|---|---|---|---|
agent_buffer_messages | Gauge | — | Messages currently buffered to disk |
agent_buffer_bytes | Gauge | bytes | Bytes currently buffered to disk |
agent_buffer_dropped_total | Counter | — | Messages evicted because the disk buffer was over its byte bound |
agent_buffer_expired_total | Counter | — | Buffered messages discarded because they aged out (PROXIMA_AGENT_BUFFER_MAX_AGE, default 1h) before NATS came back. Non-zero means an outage outlasted the buffer and that data is gone. The cluster collector exports the same count as proxima_agent_buffer_expired_total on its own metrics endpoint |
agent_buffer_append_failed_total | Counter | — | Messages the agent could neither send nor write to the buffer (for example, a full disk). Non-zero means the host lost data |
agent_buffer_replayed_total | Counter | — | Buffered messages successfully replayed after reconnect |
Collector Health
One pair per collector the agent runs, labelled collector.
| Metric | Type | Unit | Description |
|---|---|---|---|
agent_collector_errors_total | Counter | — | Collection errors |
agent_collector_last_success_timestamp | Gauge | seconds | Unix timestamp of the collector's last successful collection (0 until the first success). Stored under exactly this name — no _seconds suffix is added to agent metrics |
Stability
| Metric | Type | Unit | Description |
|---|---|---|---|
agent_goroutine_panics_total | Counter | — | Panics the agent recovered from instead of crashing. Non-zero means it survived something it should not have hit — worth a bug report |
Healthcheck
These metrics are present only when the healthcheck subsystem is active (i.e., healthcheck rules are configured for the host).
| Metric | Type | Unit | Description |
|---|---|---|---|
agent_healthcheck_scans_total | Counter | — | Total healthcheck scans completed |
agent_healthcheck_last_duration_seconds | Gauge | seconds | Duration of last healthcheck scan |
agent_healthcheck_last_passed | Gauge | — | Checks passed in last scan |
agent_healthcheck_last_failed | Gauge | — | Checks failed in last scan |
agent_healthcheck_last_total | Gauge | — | Total checks in last scan |
agent_healthcheck_last_findings | Gauge | — | Detailed scanner findings (e.g. CVEs, OpenSCAP rule results) recorded in last scan |
Config Cache
These metrics are present only when the agent config sync subsystem is active.
| Metric | Type | Unit | Description |
|---|---|---|---|
agent_configcache_configs_cached | Gauge | — | Number of cached config types |
agent_configcache_version | Gauge | — | Max version across cached configs |
Resource Limits
The agent supports configurable resource limits:
| Environment Variable | Default | Description |
|---|---|---|
PROXIMA_AGENT_MEMORY_LIMIT_MB | — (disabled) | Go soft memory limit via debug.SetMemoryLimit |
PROXIMA_AGENT_COLLECTOR_TIMEOUT | 30s | Per-collector context deadline; timed-out collectors degrade gracefully |
Hetzner Cloud objects (backend-written)
Not collected by an agent: the backend reads these from the Hetzner Cloud metrics API for each pull source's load balancers and agentless servers, every PROXIMA_HETZNER_METRICS_INTERVAL (default 5 minutes), one point a minute. Labels: provider, resource_type, resource_id, provider_id, pull_source_id, and environment_id when the source is pinned to an environment. The registry, with the Hetzner series each comes from, is in docs/standards/metrics.md ("Backend-written series: hcloud_*").
| Metric | Unit | Object |
|---|---|---|
hcloud_lb_requests_per_second | /s | load balancer |
hcloud_lb_connections_per_second | /s | load balancer |
hcloud_lb_open_connections | count | load balancer |
hcloud_lb_bandwidth_in_bytes_per_second, hcloud_lb_bandwidth_out_bytes_per_second | bytes/s | load balancer |
hcloud_server_cpu_percent | % (Hetzner's figure) | server |
hcloud_server_disk_bandwidth_read_bytes_per_second, hcloud_server_disk_bandwidth_write_bytes_per_second | bytes/s | server |
hcloud_server_disk_iops_read, hcloud_server_disk_iops_write | ops/s | server |
hcloud_server_network_in_bytes_per_second, hcloud_server_network_out_bytes_per_second | bytes/s | server |
hcloud_server_network_in_packets_per_second, hcloud_server_network_out_packets_per_second | packets/s | server |
Derived metrics (computed at query time)
The ratio metrics the UI and pc CLI show are not stored — the backend computes them from the raw values above at query time, scoped to the caller's VictoriaMetrics tenant. A request for a derived name is rewritten into the equivalent MetricsQL expression, so callers keep querying the same names:
| Derived Metric | Source Expression |
|---|---|
cpu_usage_ratio | Sum of non-idle CPU modes / total CPU time |
cpu_iowait_ratio | iowait / total CPU time |
memory_used_ratio | memory_used_bytes / memory_total_bytes |
memory_swap_used_ratio | 1 - (swap_free / swap_total) |
disk_used_ratio | disk_used_bytes / disk_total_bytes |
disk_inodes_used_ratio | 1 - (inodes_free / inodes_total) |
docker_container_cpu_usage_ratio | rate(cpu_seconds_total) |
docker_container_memory_usage_ratio | memory_used / memory_limit |
pg_cache_hit_ratio | rate(blocks_hit) / (rate(blocks_hit) + rate(blocks_read)) |
agent_cpu_ratio | rate(agent_cpu_seconds_total) |
Why query-time, not recording rules
These were originally produced by vmalert recording rules, which wrote the computed series back into VictoriaMetrics. That does not work with the platform's per-client VM tenancy: host metrics live in one tenant per client, but OSS vmalert is single-tenant per instance and cannot write derived series into each client tenant (per-group tenant / -clusterMode is a VictoriaMetrics Enterprise feature, and this cluster's multitenant query endpoint won't enumerate the client tenants). Computing at query time needs no recording rules, no per-tenant infrastructure, and covers any new client automatically — so vmalert and its rules were removed entirely. Background: docs/superpowers/specs/2026-06-05-vmalert-multitenant-recording-rules-design.md.
Adding a derived metric
Add an entry to the computedMetrics registry in backend/internal/metrics/computed.go — a build function that renders the MetricsQL expression for a label scope (set container: true for per-container metrics, seriesSource for the raw metric whose labels_key enumerates its series). The host / environment / client / container query builders and the metric-name lists pick it up automatically.
Backend Self-Monitoring
The backend exposes Prometheus metrics at /metrics via the OpenTelemetry SDK. These are scraped by vmagent and stored in VictoriaMetrics under account 0 (separate from tenant-scoped agent metrics). Subsystems include HTTP, Auth, PQL Search, NATS Workers, VictoriaMetrics Client, Tenant Cache, Change Detection, Log Collection, Healthcheck, Compliance, Alerting, Credentials, Config Sync, Terminal, Fleet, and Database connection pools.
For the full metric list with types, labels, and descriptions, see Prometheus Metrics Reference.
Backend metric names on this page are the names you query. The OpenTelemetry exporter adds _total to a counter and _seconds to a seconds-unit instrument when the code's name lacks them, so the code may say proxima_push_auditor_last_success_timestamp while /metrics — and every query — says proxima_push_auditor_last_success_timestamp_seconds. A query on the code's name returns nothing rather than an error.
Fleet (agent self-update)
| Metric | Type | Labels | Description |
|---|---|---|---|
proxima_fleet_rollouts_total | counter | status (created/paused/completed/aborted) | Rollout lifecycle transitions |
proxima_fleet_rollouts_active | gauge | — | Currently active rollouts |
proxima_fleet_agent_updates_total | counter | status (sent/healthy/failed/timed_out/skipped_offline) | Per-agent update transitions |
proxima_fleet_update_duration_seconds | histogram | — | Rollout duration from start to terminal status |
proxima_fleet_auth_denied_total | counter | — | Tenant/scope denials on rollout requests |
Version currency (EOL sync)
| Metric | Type | Labels | Description |
|---|---|---|---|
proxima_version_sync_duration_seconds | histogram | — | How long an end-of-life catalogue sync took |
proxima_version_sync_products | gauge | status (matched/unmatched) | Products seen on the last sync, by whether they matched the catalogue |
proxima_version_sync_errors_total | counter | step | Sync errors, by the step that failed |
proxima_version_export_total | counter | format | Version-currency report exports |
Tenant usage
Sampled every 15 minutes; tenant is the client slug.
| Metric | Type | Labels | Description |
|---|---|---|---|
proxima_tenant_usage_series | gauge | tenant | Active VictoriaMetrics series in the client's account |
proxima_tenant_usage_log_rows | gauge | tenant | Log rows ingested over the last 24 hours |
proxima_tenant_usage_log_rows_total | gauge | tenant | Log rows stored over the full retention window — a gauge despite its name; do not rate() it |
proxima_tenant_usage_db_rows | gauge | tenant | PostgreSQL rows owned by the client in the tracked tables |
proxima_tenant_usage_db_bytes_est | gauge | tenant | Estimated PostgreSQL bytes for the client |
Audit retention
| Metric | Type | Labels | Description |
|---|---|---|---|
proxima_audit_retention_runs_total | counter | result (ok/error/skipped) | Daily retention passes |
proxima_audit_log_partitions_pruned_total | counter | — | Months of audit log archived to R2 and removed from PostgreSQL |
proxima_audit_log_partitions_overdue | gauge | — | Months past the retention period still in PostgreSQL — nonzero while pruning is disabled |
proxima_audit_log_default_partition_rows | gauge | — | Audit rows that had no monthly partition to land in; should be zero |
Audit integrity verifier
| Metric | Type | Labels | Description |
|---|---|---|---|
proxima_audit_verify_problems | gauge | check | Problems found by the hourly independent verification of the audit log |
proxima_audit_verify_run_error | gauge | — | 1 when the last verification could not run |
proxima_audit_verify_last_run_timestamp_seconds | gauge | — | When the last verification finished |
proxima_audit_verify_rows_verified | gauge | — | Audit entries recomputed by the last verification |
proxima_audit_verify_anchors_checked | gauge | — | Chain-head anchors compared by the last verification |
Synthetic uptime probes
Two different things, deliberately kept apart.
Per-monitor series, written into the monitor's own client's VictoriaMetrics account by the ingest worker (not scraped from anything):
| Metric | Type | Unit | Labels | Description |
|---|---|---|---|---|
probe_success | gauge | — | monitor_id, location, kind | 1/0 — whether the check itself succeeded. upside_down is applied when reading, so this metric always means what its name says |
probe_duration_seconds | gauge | seconds | monitor_id, location, kind | How long the check took |
probe_ssl_cert_expiry_seconds | gauge | seconds | monitor_id | Leaf certificate expiry. No location label — every pop sees the same certificate |
Backend counters on /metrics, none of which carries a monitor_id (an unbounded label on a pull exporter with no series expiry is an OOM waiting to happen): proxima_probe_results_total{kind,location,result}, proxima_probe_state_transitions_total{kind,to_state}, proxima_probe_partial_failure_total{location}, proxima_probe_quorum_degraded{kind}, proxima_probe_results_rejected_total{reason}, proxima_probe_results_corrected_total{field} and proxima_probe_assignments_skipped_total{reason}.
proxima_probe_quorum_degraded is a per-replica gauge — alert on sum(proxima_probe_quorum_degraded) > 0, never on the number. See Uptime Monitors.
Emission counters, the paging gate chain's three signals. proxima_probe_would_page_total{kind} is ticked before the chain on every confirmed down-transition, so it measures what would have woken somebody whether or not anything did — with paging_enabled shipping false on every monitor, it is the shadow window's entire output and the evidence tier 2's guessed threshold gets tuned against. proxima_probe_paged_total{kind} is the publish the broker accepted, and proxima_probe_page_suppressed_total{kind,reason} is what the chain refused, under paging_disabled or fleet_suppressed. The three satisfy an identity — would_page = paged + Σ suppressed + failures — and that residual is the only way to see a lost page: emission is best-effort by necessity (the transition is durable and a redelivery short-circuits before the emit path, so a NAK could never retry it), so a failed publish is an ERROR log and nothing more. Only the firing half is counted at all; a resolve is never gated and never counted here. Resolves published during the shadow window land on the ingest worker's find-only path and tick proxima_alert_resolve_no_open_group_total{source_type="proxima_probe"} — expected, not a fault, for as long as paging is off.
Fleet-audit counters, emitted by the probe location auditor — the fleet's dead-man's switch. A location is a vantage point that several prober machines may sit behind, so it can lose machines, or all of them, while every monitor still looks green and is simply confirmed by fewer locations than its operator asked for: proxima_probe_prober_stale_total{location} (a prober that reported nothing inside the audit window), proxima_probe_location_dark_total{location} (no prober at a location reported inside it), proxima_probe_location_disagreement_total{location} (probers sharing one code disagree on kind or client_id, which makes that code ineligible for every tenant's monitors), plus the auditor's own pulse proxima_probe_location_auditor_sweeps_total{result} and proxima_probe_location_auditor_last_success_timestamp_seconds. The sweep is an advisory-lock singleton, so a finding is counted once per fleet rather than once per replica. The audit window is derived from the same source as the quorum's own seat window, so the auditor is never quieter than the denominator it reports on: it is exact for monitors at intervals up to 270s and reports early above that, since one fleet-wide window cannot scale the way a per-monitor one does. Per-location scaling is a follow-up.
Independence-guard signals, emitted by the probe sanity guard. Quorum treats several locations agreeing as evidence about the target, which holds only while their failures are independent — a pop that has lost its own transit fails nearly everything it probes at once, which looks exactly like hundreds of simultaneous customer outages. proxima_probe_location_suppressed{location} is 1 while such a location's votes are being dropped and 0 once they count again; it is per-replica, so alert on max() rather than min() — the guard is deliberately not a singleton, because its output is in-process state the result worker in that same replica reads on every result. proxima_probe_no_eligible_locations_total{kind} counts evaluations of a monitor left with no location able to vote at all: probed every interval, structurally unable to confirm anything, and — before the gauge was widened to cover it — reporting itself as not degraded. The guard carries its own proof of life, proxima_probe_sanity_sweeps_total{result} and proxima_probe_sanity_last_success_timestamp_seconds (alert on min() across replicas, since every replica runs its own guard): if it stops, the suppression set freezes while the suppressed gauge keeps reporting stale values, and with nothing suppressed it would otherwise die silently. proxima_probe_sanity_withheld_total{location,reason} records the verdicts it declined — single_tenant_failure (the failures are confined to one client, so this is a customer outage and not the pop's transit) and single_tenant_fleet (every monitor at that pop belongs to one client, so it can never be judged at all — which is not the same as being healthy).
Tier 2 — a broken fleet. Tier 1 cannot see the failure that breaks the independence premise hardest: a bug shipped to every prober fails all locations together, so each looks individually broken, every monitor confirms down on what reads as unanimous agreement between independent vantage points, and no surviving location is left to disagree. proxima_probe_fleet_suppressed is 1 while the guard is holding every synthetic page for that reason, and 0 otherwise — per-replica, so alert on max(). It suppresses emission and never state: monitors still transition and events are still written, so the history read during the incident review stays true; only the wake-someone-up action is held. It counts a location as failing using tier 1's own verdict rather than re-deriving it, so a pop tier 1 judged to be one customer's outage (single_tenant_failure) is not counted toward halting every tenant's pages — otherwise a single tenant who dominates most pops would take the whole product's paging down when their own datacentre went dark. A pop that is undecidable (single_tenant_fleet) is still counted, so a private-pop fleet stays trippable. proxima_probe_guard_inert{tier} answers the question the first gauge cannot: a tier can be structurally unable to fire, and there are two routes into that — the sample is too small to be evidence (a location below MinSample, a fleet below FleetMinLocations), or the tier's threshold has been set above 1.0, which is the documented way to switch it off. Both are correct and both mean a safety guard can be entirely inert while looking configured and healthy — on a three-pop pilot fleet that is the likely state, not the edge case. An operator reading proxima_probe_fleet_suppressed 0 is entitled to know whether that means "healthy" or "this check has never been able to run". Two shapes to expect concretely: tier 1 is two gates deep (MinSample, and failures spanning at least two tenants), so on a thin or single-tenant fleet it may never fire and the reason is split between this gauge and proxima_probe_sanity_withheld_total; and tier 2 at the default FleetThreshold of 0.75 needs 3 of 3 on the current pilot fleet, since 2/3 = 0.667 — one trippable state, unanimity, and inert entirely at two pops. Nothing in the repository alerts on this gauge yet — max by (tier) (proxima_probe_guard_inert) == 1 is the rule that makes it useful, and without it an inert guard is indistinguishable from a healthy one. Every threshold behind both tiers is an explicit guess awaiting the shadow window and is tunable from the environment: PROXIMA_PROBE_SANITY_INTERVAL, PROXIMA_PROBE_SANITY_WINDOW, PROXIMA_PROBE_LOCATION_THRESHOLD, PROXIMA_PROBE_LOCATION_MIN_SAMPLE, PROXIMA_PROBE_FLEET_THRESHOLD, PROXIMA_PROBE_FLEET_MIN_LOCATIONS. A non-positive value is refused and logged rather than silently replaced by the default — to disable a tier, set its threshold above 1.0.
Push-monitor counters: proxima_push_heartbeat_total{result} (ok, fail, unknown_token) counts heartbeat calls that got past the rate limiter; proxima_push_auditor_sweeps_total{result} and proxima_push_auditor_last_success_timestamp_seconds are the push auditor's pulse. The auditor is what turns a missed heartbeat into a down verdict, so a stale pulse means push monitors cannot go down — alert on max() (only the lock winner stamps it).
Monitor-group counters, emitted where a service — a set of one client's monitors judged and paged as a single thing — is derived and paged. None of them carries a group_id, for the reason none of the probe counters carries a monitor_id, and it bites harder because groups are cheap to create; which group did or did not page is in the log line and on the alert. proxima_monitor_group_would_page_total, proxima_monitor_group_paged_total and proxima_monitor_group_page_suppressed_total{reason} are the same gate chain the monitor path reports, with the same two refusal reasons (paging_disabled, fleet_suppressed) and the same identity and resolve asymmetry — monitor_groups.paging_enabled also ships false, so would_page against proxima_probe_would_page_total is what says how many member pages one group page replaced. proxima_monitor_group_transitions_total{to_state} counts applied group state changes including the edges that page nobody. The remaining four belong to the drift auditor, the dead-man's switch behind edge-triggered evaluation: proxima_monitor_group_drift_total{correction} counts what a sweep had to correct (state — a page or a resolve was missed and has just been sent, late; counts — the state was right and the evidence beside it was stale), and a corrected drift that is not counted looks exactly like a system with no drift, so a sustained rate is a bug at a mutation path rather than a threshold to widen. proxima_monitor_group_authority_conflicts_total counts groups holding paging_enabled beside a member that holds it too — a pair that double-pages every outage of that service, that the API's authoring-time refusal structurally cannot catch (two concurrent PATCHes, or hand-run SQL), and that never self-heals, because the auditor deliberately does not choose which side loses its page. proxima_monitor_group_auditor_sweeps_total{result} carries three values, not two — clean, partial (it ran but could not finish some unit of work, and still stamps the pulse) and error (it read nothing and withholds the pulse) — and proxima_monitor_group_auditor_last_success_timestamp_seconds is the pulse: the auditor is an advisory-lock singleton, so alert on max() across replicas, the opposite of the sanity guard's rule. Full behaviour in Uptime Monitors.
Public status pages
Backend counters on /metrics, registered in backend/internal/observability/metrics_status_page.go. No label carries a page ID or slug: pages are client-authored and unbounded, and /metrics never expires a series. Every label comes from a closed set, so an unknown value folds into a known one (result → error, tab → status) and a crafted path can never mint a series.
| Metric | Type | Labels | Description |
|---|---|---|---|
proxima_status_page_requests_total | counter | result (ok, not_found, rate_limited, error), tab (status, maintenance, incidents, asset, css) | Requests to <slug>.{PROXIMA_STATUS_PAGE_DOMAIN}. not_found deliberately covers an unknown slug, a draft, a Console-only page, a wrong or missing key, a refused IP and a malformed host alike. rate_limited is over 120 a minute per trusted client IP. error is any 503: a failed page lookup, snapshot build or asset read, or a recovered panic. asset covers logos, the favicon, /static/* and robots.txt; css is /page.css |
proxima_status_page_render_seconds | histogram | — | Snapshot (cached or built) plus template render, for one HTML tab |
proxima_status_page_snapshot_cache_total | counter | result (hit, miss, error) | Snapshot lookups. miss builds the snapshot and stores it in Valkey for 30 s. error means Valkey was unreachable, or skipped for the 10 s after a Valkey error, and the snapshot came from the 10 s in-process cache, so a sustained error rate means every replica is building its own |
See Status Pages.
Label Schema
Every metric stored in VictoriaMetrics includes these system labels:
| Label | Description |
|---|---|
__name__ | Metric name (e.g., cpu_user_seconds_total) |
host_id | Host UUID |
environment_id | Environment UUID |
labels_key | Deterministic key from sorted label pairs (e.g., device=sda,mount=/) |
Plus the original agent labels (e.g., device, mount, interface, collector, database, table).
The labels_key enables series discovery — the API endpoint GET /hosts/:id/metrics/:name/series returns all distinct labels_key values, allowing the frontend to enumerate and query each series individually.