Skip to main content

Metrics Collected

The agent collects system metrics via gopsutil v4 and Linux /proc files, plus plugin metrics from configured integrations. All metrics follow the OpenMetrics naming convention.

Raw values: The agent emits raw counters and gauges. The derived ratio metrics the UI shows (cpu_usage_ratio, memory_used_ratio, disk_used_ratio, …) are computed at query time in the backend from these raw values, scoped to the caller's tenant — see Derived metrics.

System Metrics​

Collected automatically on every managed host. No configuration required.

CPU​

MetricTypeUnitDescription
cpu_user_seconds_totalCountersecondsTime spent in user mode
cpu_system_seconds_totalCountersecondsTime spent in kernel mode
cpu_nice_seconds_totalCountersecondsTime spent in low-priority user mode
cpu_idle_seconds_totalCountersecondsTime spent idle
cpu_iowait_seconds_totalCountersecondsTime spent waiting for I/O
cpu_irq_seconds_totalCountersecondsTime spent servicing hardware interrupts
cpu_softirq_seconds_totalCountersecondsTime spent servicing software interrupts
cpu_steal_seconds_totalCountersecondsTime stolen by hypervisor
cpu_load1Gauge—1-minute load average
cpu_load5Gauge—5-minute load average
cpu_load15Gauge—15-minute load average
Derived (computed at query time)

cpu_usage_ratio and cpu_iowait_ratio are computed at query time in the backend from the per-mode counters above. They produce the same 0.0–1.0 ratios the frontend displays.

Memory​

MetricTypeUnitDescription
memory_total_bytesGaugebytesTotal physical memory
memory_used_bytesGaugebytesMemory in use (excluding buffers/cache)
memory_available_bytesGaugebytesMemory available without swapping
memory_buffers_bytesGaugebytesKernel buffer cache
memory_cached_bytesGaugebytesPage cache (reclaimable)
memory_swap_total_bytesGaugebytesTotal swap space
memory_swap_free_bytesGaugebytesFree swap space
Derived (computed at query time)

memory_used_ratio and memory_swap_used_ratio are computed at query time in the backend from raw bytes.

Disk​

Per-partition and per-device metrics. Pseudo-filesystems (tmpfs, sysfs, proc, overlay) are excluded.

Usage (per partition):

MetricTypeUnitLabelsDescription
disk_total_bytesGaugebytesmount, deviceTotal partition size
disk_used_bytesGaugebytesmount, deviceUsed space
disk_free_bytesGaugebytesmount, deviceFree space
disk_inodes_totalGauge—mount, deviceTotal inodes
disk_inodes_freeGauge—mount, deviceFree inodes

I/O (per device):

MetricTypeUnitLabelsDescription
disk_read_bytes_totalCounterbytesdeviceTotal bytes read
disk_written_bytes_totalCounterbytesdeviceTotal bytes written
disk_reads_totalCounter—deviceTotal read operations
disk_writes_totalCounter—deviceTotal write operations
disk_io_time_ms_totalCountermsdeviceTime spent doing I/O
disk_io_weighted_ms_totalCountermsdeviceWeighted I/O time (saturation signal)
disk_io_in_progressGauge—deviceI/O operations currently in flight
disk_read_time_ms_totalCountermsdeviceTotal read latency
disk_write_time_ms_totalCountermsdeviceTotal write latency
Derived (computed at query time)

disk_used_ratio and disk_inodes_used_ratio are computed at query time in the backend from raw values.

Disk saturation

disk_io_in_progress > 0 consistently indicates disk saturation. disk_io_weighted_ms_total is the USE method's saturation signal for storage.

Network​

Per-interface metrics. The loopback interface (lo) is excluded.

MetricTypeUnitLabelsDescription
network_receive_bytes_totalCounterbytesinterfaceTotal bytes received
network_transmit_bytes_totalCounterbytesinterfaceTotal bytes transmitted
network_receive_packets_totalCounter—interfaceTotal packets received
network_transmit_packets_totalCounter—interfaceTotal packets transmitted
network_receive_errors_totalCounter—interfaceTotal receive errors
network_transmit_errors_totalCounter—interfaceTotal transmit errors
network_receive_drops_totalCounter—interfaceTotal receive drops (buffer exhaustion)
network_transmit_drops_totalCounter—interfaceTotal transmit drops

Pressure Stall Information (PSI)​

Available on Linux kernel 4.20+. Skipped silently on older kernels.

MetricTypeDescription
psi_cpu_some_avg10Gauge% of time tasks stalled on CPU (10s avg)
psi_cpu_some_avg60Gauge% of time tasks stalled on CPU (60s avg)
psi_cpu_some_total_microsecondsCounterTotal CPU stall time
psi_memory_some_avg10Gauge% of time some tasks stalled on memory (10s avg)
psi_memory_some_avg60Gauge% of time some tasks stalled on memory (60s avg)
psi_memory_full_avg10Gauge% of time all tasks stalled on memory (10s avg)
psi_memory_full_avg60Gauge% of time all tasks stalled on memory (60s avg)
psi_io_some_avg10Gauge% of time some tasks stalled on I/O (10s avg)
psi_io_some_avg60Gauge% of time some tasks stalled on I/O (60s avg)
psi_io_full_avg10Gauge% of time all tasks stalled on I/O (10s avg)
psi_io_full_avg60Gauge% of time all tasks stalled on I/O (60s avg)
Root cause analysis

PSI is the single most valuable metric for automated root cause analysis. Unlike CPU usage (which can be 100% and healthy), PSI measures actual stall time — when tasks are blocked waiting for resources.

Kernel & System State​

MetricTypeDescription
vmstat_oom_kill_totalCounterOOM kill events (from /proc/vmstat)
vmstat_pgmajfault_totalCounterMajor page faults (memory pressure → disk I/O)
vmstat_pswpin_totalCounterPages swapped in
vmstat_pswpout_totalCounterPages swapped out
context_switches_totalCounterSystem-wide context switches
processes_created_totalCounterTotal processes created (fork rate)
processes_runningGaugeProcesses currently running
processes_blockedGaugeProcesses blocked on I/O
entropy_availableGaugeAvailable entropy (always 256 on kernel 5.4+)

TCP & Resource Limits​

MetricTypeDescription
sockstat_tcp_inuseGaugeTCP sockets in use
sockstat_tcp_orphanGaugeOrphaned TCP sockets (no process)
sockstat_tcp_twGaugeTCP sockets in TIME_WAIT state
sockstat_sockets_usedGaugeTotal sockets in use
filefd_allocatedGaugeSystem-wide allocated file descriptors
filefd_maxGaugeSystem-wide file descriptor limit
conntrack_entriesGaugeConntrack table entries (skipped if netfilter not loaded)

Process Metrics​

Top-N process metrics are exported to VictoriaMetrics alongside system metrics. Only the top 20 processes by CPU usage are exported per collection cycle to control cardinality.

MetricTypeUnitLabelsDescription
process_cpu_usage_ratioGaugeratiopid, process_name, usernamePer-process CPU usage (0.0–1.0)
process_memory_usage_ratioGaugeratiopid, process_name, usernamePer-process memory usage (0.0–1.0)
process_memory_rss_bytesGaugebytespid, process_name, usernameResident Set Size
process_threadsGauge—pid, process_name, usernameThread count
note

The full process list is stored in PostgreSQL and available via GET /api/v1/hosts/{id}/processes. Only the top 20 by CPU are exported as time-series metrics to VictoriaMetrics.

Drill-down endpoint: GET /api/v1/hosts/{id}/processes/{pid}/metrics?process_name=X returns all 4 process metrics as time-series for a specific process. The process_name parameter is required to disambiguate PID reuse across collection cycles.

Integration Metrics​

Plugin collectors extend monitoring beyond system metrics. Each plugin injects a collector=<name> label into all its metrics for differentiation.

PostgreSQL​

See the PostgreSQL Integration page for setup and configuration.

MetricTypeUnitLabelsDescription
pg_connections_activeGauge——Active connections
pg_connections_idleGauge——Idle connections
pg_connections_maxGauge——Maximum allowed connections
pg_blocks_hit_totalCounter——Buffer cache block hits
pg_blocks_read_totalCounter——Disk block reads
pg_replication_lag_secondsGaugeseconds—Replication lag
pg_database_size_bytesGaugebytesdatabaseDatabase size
pg_deadlocks_totalCounter——Total deadlocks
pg_xact_commit_totalCounter——Total committed transactions
pg_xact_rollback_totalCounter——Total rolled-back transactions
pg_dead_tuplesGauge—tableDead rows (top 50 tables)
pg_stat_statements_countGauge——Tracked query statements
pg_stat_statements_calls_totalCounter——Total statement executions
pg_stat_statements_time_seconds_totalCounterseconds—Total statement execution time
pg_stat_statements_slow_queriesGauge——Statements with mean > 100ms
Derived (computed at query time)

pg_cache_hit_ratio is computed at query time in the backend from pg_blocks_hit_total and pg_blocks_read_total.

note

pg_stat_statements metrics require the PostgreSQL extension to be enabled. If the extension is not installed, these metrics are silently skipped.

Nginx​

See the Nginx Integration page for setup and configuration.

MetricTypeUnitLabelsDescription
nginx_connections_activeGauge——Current active connections
nginx_connections_accepted_totalCounter——Total accepted connections
nginx_connections_handled_totalCounter——Total handled connections
nginx_requests_totalCounter——Total HTTP requests
nginx_connections_readingGauge——Connections reading request header
nginx_connections_writingGauge——Connections writing response
nginx_connections_waitingGauge——Idle keep-alive connections

Docker​

See the Docker Integration page for setup and configuration.

MetricTypeUnitLabelsDescription
docker_containersGauge—stateContainer count by state (running, exited, created, etc.)
docker_container_cpu_seconds_totalCountersecondscontainer_name, container_idCumulative CPU time consumed
docker_container_memory_used_bytesGaugebytescontainer_name, container_idMemory usage minus cache
docker_container_memory_limit_bytesGaugebytescontainer_name, container_idMemory limit (omitted if unlimited)
docker_container_network_receive_bytes_totalCounterbytescontainer_name, container_idTotal bytes received (all interfaces)
docker_container_network_transmit_bytes_totalCounterbytescontainer_name, container_idTotal bytes transmitted (all interfaces)
docker_container_block_read_bytes_totalCounterbytescontainer_name, container_idTotal block device bytes read
docker_container_block_write_bytes_totalCounterbytescontainer_name, container_idTotal block device bytes written
Derived (computed at query time)

docker_container_cpu_usage_ratio and docker_container_memory_usage_ratio are computed at query time in the backend.

note

Per-container metrics are only collected for running containers. Memory limit metrics are omitted when the container has no memory limit set.

Agent Self-Monitoring​

The agent collects metrics about its own resource usage. These are always enabled and require no configuration. All metrics carry agent_version and hostname labels, enabling cross-version regression detection and multi-host comparison.

A pre-built Grafana dashboard is available at infra/grafana/dashboards/proxima/agent-resources.json.

Process​

MetricTypeUnitDescription
agent_build_infoGauge—Constant 1; build identity carried in labels (agent_version, hostname, go_version). Use count by (agent_version) (agent_build_info) to see the fleet's version distribution and observe a rollout landing.
agent_uptime_secondsGaugesecondsAgent process uptime
agent_cpu_seconds_totalCountersecondsCumulative CPU time (user + system)
agent_memory_rss_bytesGaugebytesResident set size
agent_memory_vms_bytesGaugebytesVirtual memory size
agent_open_fdsGauge—Open file descriptors
agent_threadsGauge—OS thread count
agent_ctx_switches_voluntary_totalCounter—Voluntary context switches
agent_ctx_switches_involuntary_totalCounter—Involuntary context switches
agent_io_read_bytes_totalCounterbytesProcess disk read bytes
agent_io_write_bytes_totalCounterbytesProcess disk write bytes
Derived (computed at query time)

agent_cpu_ratio is computed at query time in the backend from agent_cpu_seconds_total.

Go Runtime​

MetricTypeUnitDescription
agent_go_heap_alloc_bytesGaugebytesGo heap allocation
agent_go_heap_sys_bytesGaugebytesGo heap obtained from OS
agent_go_heap_objectsGauge—Live heap object count
agent_go_stack_bytesGaugebytesGo stack memory in use
agent_go_sys_bytesGaugebytesTotal Go memory from OS
agent_go_memlimit_bytesGaugebytesSoft memory limit (only present when PROXIMA_AGENT_MEMORY_LIMIT_MB is set)
agent_goroutinesGauge—Goroutine count
agent_gc_cycles_totalCounter—GC cycle count
agent_gc_pause_seconds_totalCountersecondsTotal GC pause time

NATS Transport​

MetricTypeUnitDescription
agent_nats_messages_out_totalCounter—Messages published to NATS
agent_nats_bytes_out_totalCounterbytesBytes published to NATS
agent_nats_messages_in_totalCounter—Messages received from NATS (request-reply)
agent_nats_bytes_in_totalCounterbytesBytes received from NATS
agent_nats_reconnects_totalCounter—NATS reconnection count
agent_nats_outbox_lengthGauge—Outbox queue depth (indicates backpressure when > 0)
agent_nats_rtt_secondsGaugesecondsRound-trip time to NATS server

Offline Buffer​

These metrics are present only when the transport exposes disk-buffer stats. While NATS is unreachable, the agent buffers outgoing messages to disk and replays them on reconnect.

MetricTypeUnitDescription
agent_buffer_messagesGauge—Messages currently buffered to disk
agent_buffer_bytesGaugebytesBytes currently buffered to disk
agent_buffer_dropped_totalCounter—Messages evicted because the disk buffer was over its byte bound
agent_buffer_expired_totalCounter—Buffered messages discarded because they aged out (PROXIMA_AGENT_BUFFER_MAX_AGE, default 1h) before NATS came back. Non-zero means an outage outlasted the buffer and that data is gone. The cluster collector exports the same count as proxima_agent_buffer_expired_total on its own metrics endpoint
agent_buffer_append_failed_totalCounter—Messages the agent could neither send nor write to the buffer (for example, a full disk). Non-zero means the host lost data
agent_buffer_replayed_totalCounter—Buffered messages successfully replayed after reconnect

Collector Health​

One pair per collector the agent runs, labelled collector.

MetricTypeUnitDescription
agent_collector_errors_totalCounter—Collection errors
agent_collector_last_success_timestampGaugesecondsUnix timestamp of the collector's last successful collection (0 until the first success). Stored under exactly this name — no _seconds suffix is added to agent metrics

Stability​

MetricTypeUnitDescription
agent_goroutine_panics_totalCounter—Panics the agent recovered from instead of crashing. Non-zero means it survived something it should not have hit — worth a bug report

Healthcheck​

These metrics are present only when the healthcheck subsystem is active (i.e., healthcheck rules are configured for the host).

MetricTypeUnitDescription
agent_healthcheck_scans_totalCounter—Total healthcheck scans completed
agent_healthcheck_last_duration_secondsGaugesecondsDuration of last healthcheck scan
agent_healthcheck_last_passedGauge—Checks passed in last scan
agent_healthcheck_last_failedGauge—Checks failed in last scan
agent_healthcheck_last_totalGauge—Total checks in last scan
agent_healthcheck_last_findingsGauge—Detailed scanner findings (e.g. CVEs, OpenSCAP rule results) recorded in last scan

Config Cache​

These metrics are present only when the agent config sync subsystem is active.

MetricTypeUnitDescription
agent_configcache_configs_cachedGauge—Number of cached config types
agent_configcache_versionGauge—Max version across cached configs

Resource Limits​

The agent supports configurable resource limits:

Environment VariableDefaultDescription
PROXIMA_AGENT_MEMORY_LIMIT_MB— (disabled)Go soft memory limit via debug.SetMemoryLimit
PROXIMA_AGENT_COLLECTOR_TIMEOUT30sPer-collector context deadline; timed-out collectors degrade gracefully

Hetzner Cloud objects (backend-written)​

Not collected by an agent: the backend reads these from the Hetzner Cloud metrics API for each pull source's load balancers and agentless servers, every PROXIMA_HETZNER_METRICS_INTERVAL (default 5 minutes), one point a minute. Labels: provider, resource_type, resource_id, provider_id, pull_source_id, and environment_id when the source is pinned to an environment. The registry, with the Hetzner series each comes from, is in docs/standards/metrics.md ("Backend-written series: hcloud_*").

MetricUnitObject
hcloud_lb_requests_per_second/sload balancer
hcloud_lb_connections_per_second/sload balancer
hcloud_lb_open_connectionscountload balancer
hcloud_lb_bandwidth_in_bytes_per_second, hcloud_lb_bandwidth_out_bytes_per_secondbytes/sload balancer
hcloud_server_cpu_percent% (Hetzner's figure)server
hcloud_server_disk_bandwidth_read_bytes_per_second, hcloud_server_disk_bandwidth_write_bytes_per_secondbytes/sserver
hcloud_server_disk_iops_read, hcloud_server_disk_iops_writeops/sserver
hcloud_server_network_in_bytes_per_second, hcloud_server_network_out_bytes_per_secondbytes/sserver
hcloud_server_network_in_packets_per_second, hcloud_server_network_out_packets_per_secondpackets/sserver

Derived metrics (computed at query time)​

The ratio metrics the UI and pc CLI show are not stored — the backend computes them from the raw values above at query time, scoped to the caller's VictoriaMetrics tenant. A request for a derived name is rewritten into the equivalent MetricsQL expression, so callers keep querying the same names:

Derived MetricSource Expression
cpu_usage_ratioSum of non-idle CPU modes / total CPU time
cpu_iowait_ratioiowait / total CPU time
memory_used_ratiomemory_used_bytes / memory_total_bytes
memory_swap_used_ratio1 - (swap_free / swap_total)
disk_used_ratiodisk_used_bytes / disk_total_bytes
disk_inodes_used_ratio1 - (inodes_free / inodes_total)
docker_container_cpu_usage_ratiorate(cpu_seconds_total)
docker_container_memory_usage_ratiomemory_used / memory_limit
pg_cache_hit_ratiorate(blocks_hit) / (rate(blocks_hit) + rate(blocks_read))
agent_cpu_ratiorate(agent_cpu_seconds_total)

Why query-time, not recording rules​

These were originally produced by vmalert recording rules, which wrote the computed series back into VictoriaMetrics. That does not work with the platform's per-client VM tenancy: host metrics live in one tenant per client, but OSS vmalert is single-tenant per instance and cannot write derived series into each client tenant (per-group tenant / -clusterMode is a VictoriaMetrics Enterprise feature, and this cluster's multitenant query endpoint won't enumerate the client tenants). Computing at query time needs no recording rules, no per-tenant infrastructure, and covers any new client automatically — so vmalert and its rules were removed entirely. Background: docs/superpowers/specs/2026-06-05-vmalert-multitenant-recording-rules-design.md.

Adding a derived metric​

Add an entry to the computedMetrics registry in backend/internal/metrics/computed.go — a build function that renders the MetricsQL expression for a label scope (set container: true for per-container metrics, seriesSource for the raw metric whose labels_key enumerates its series). The host / environment / client / container query builders and the metric-name lists pick it up automatically.

Backend Self-Monitoring​

The backend exposes Prometheus metrics at /metrics via the OpenTelemetry SDK. These are scraped by vmagent and stored in VictoriaMetrics under account 0 (separate from tenant-scoped agent metrics). Subsystems include HTTP, Auth, PQL Search, NATS Workers, VictoriaMetrics Client, Tenant Cache, Change Detection, Log Collection, Healthcheck, Compliance, Alerting, Credentials, Config Sync, Terminal, Fleet, and Database connection pools.

For the full metric list with types, labels, and descriptions, see Prometheus Metrics Reference.

Query the exposed name

Backend metric names on this page are the names you query. The OpenTelemetry exporter adds _total to a counter and _seconds to a seconds-unit instrument when the code's name lacks them, so the code may say proxima_push_auditor_last_success_timestamp while /metrics — and every query — says proxima_push_auditor_last_success_timestamp_seconds. A query on the code's name returns nothing rather than an error.

Fleet (agent self-update)​

MetricTypeLabelsDescription
proxima_fleet_rollouts_totalcounterstatus (created/paused/completed/aborted)Rollout lifecycle transitions
proxima_fleet_rollouts_activegauge—Currently active rollouts
proxima_fleet_agent_updates_totalcounterstatus (sent/healthy/failed/timed_out/skipped_offline)Per-agent update transitions
proxima_fleet_update_duration_secondshistogram—Rollout duration from start to terminal status
proxima_fleet_auth_denied_totalcounter—Tenant/scope denials on rollout requests

Version currency (EOL sync)​

MetricTypeLabelsDescription
proxima_version_sync_duration_secondshistogram—How long an end-of-life catalogue sync took
proxima_version_sync_productsgaugestatus (matched/unmatched)Products seen on the last sync, by whether they matched the catalogue
proxima_version_sync_errors_totalcounterstepSync errors, by the step that failed
proxima_version_export_totalcounterformatVersion-currency report exports

Tenant usage​

Sampled every 15 minutes; tenant is the client slug.

MetricTypeLabelsDescription
proxima_tenant_usage_seriesgaugetenantActive VictoriaMetrics series in the client's account
proxima_tenant_usage_log_rowsgaugetenantLog rows ingested over the last 24 hours
proxima_tenant_usage_log_rows_totalgaugetenantLog rows stored over the full retention window — a gauge despite its name; do not rate() it
proxima_tenant_usage_db_rowsgaugetenantPostgreSQL rows owned by the client in the tracked tables
proxima_tenant_usage_db_bytes_estgaugetenantEstimated PostgreSQL bytes for the client

Audit retention​

MetricTypeLabelsDescription
proxima_audit_retention_runs_totalcounterresult (ok/error/skipped)Daily retention passes
proxima_audit_log_partitions_pruned_totalcounter—Months of audit log archived to R2 and removed from PostgreSQL
proxima_audit_log_partitions_overduegauge—Months past the retention period still in PostgreSQL — nonzero while pruning is disabled
proxima_audit_log_default_partition_rowsgauge—Audit rows that had no monthly partition to land in; should be zero

Audit integrity verifier​

MetricTypeLabelsDescription
proxima_audit_verify_problemsgaugecheckProblems found by the hourly independent verification of the audit log
proxima_audit_verify_run_errorgauge—1 when the last verification could not run
proxima_audit_verify_last_run_timestamp_secondsgauge—When the last verification finished
proxima_audit_verify_rows_verifiedgauge—Audit entries recomputed by the last verification
proxima_audit_verify_anchors_checkedgauge—Chain-head anchors compared by the last verification

Synthetic uptime probes​

Two different things, deliberately kept apart.

Per-monitor series, written into the monitor's own client's VictoriaMetrics account by the ingest worker (not scraped from anything):

MetricTypeUnitLabelsDescription
probe_successgauge—monitor_id, location, kind1/0 — whether the check itself succeeded. upside_down is applied when reading, so this metric always means what its name says
probe_duration_secondsgaugesecondsmonitor_id, location, kindHow long the check took
probe_ssl_cert_expiry_secondsgaugesecondsmonitor_idLeaf certificate expiry. No location label — every pop sees the same certificate

Backend counters on /metrics, none of which carries a monitor_id (an unbounded label on a pull exporter with no series expiry is an OOM waiting to happen): proxima_probe_results_total{kind,location,result}, proxima_probe_state_transitions_total{kind,to_state}, proxima_probe_partial_failure_total{location}, proxima_probe_quorum_degraded{kind}, proxima_probe_results_rejected_total{reason}, proxima_probe_results_corrected_total{field} and proxima_probe_assignments_skipped_total{reason}.

proxima_probe_quorum_degraded is a per-replica gauge — alert on sum(proxima_probe_quorum_degraded) > 0, never on the number. See Uptime Monitors.

Emission counters, the paging gate chain's three signals. proxima_probe_would_page_total{kind} is ticked before the chain on every confirmed down-transition, so it measures what would have woken somebody whether or not anything did — with paging_enabled shipping false on every monitor, it is the shadow window's entire output and the evidence tier 2's guessed threshold gets tuned against. proxima_probe_paged_total{kind} is the publish the broker accepted, and proxima_probe_page_suppressed_total{kind,reason} is what the chain refused, under paging_disabled or fleet_suppressed. The three satisfy an identity — would_page = paged + Σ suppressed + failures — and that residual is the only way to see a lost page: emission is best-effort by necessity (the transition is durable and a redelivery short-circuits before the emit path, so a NAK could never retry it), so a failed publish is an ERROR log and nothing more. Only the firing half is counted at all; a resolve is never gated and never counted here. Resolves published during the shadow window land on the ingest worker's find-only path and tick proxima_alert_resolve_no_open_group_total{source_type="proxima_probe"} — expected, not a fault, for as long as paging is off.

Fleet-audit counters, emitted by the probe location auditor — the fleet's dead-man's switch. A location is a vantage point that several prober machines may sit behind, so it can lose machines, or all of them, while every monitor still looks green and is simply confirmed by fewer locations than its operator asked for: proxima_probe_prober_stale_total{location} (a prober that reported nothing inside the audit window), proxima_probe_location_dark_total{location} (no prober at a location reported inside it), proxima_probe_location_disagreement_total{location} (probers sharing one code disagree on kind or client_id, which makes that code ineligible for every tenant's monitors), plus the auditor's own pulse proxima_probe_location_auditor_sweeps_total{result} and proxima_probe_location_auditor_last_success_timestamp_seconds. The sweep is an advisory-lock singleton, so a finding is counted once per fleet rather than once per replica. The audit window is derived from the same source as the quorum's own seat window, so the auditor is never quieter than the denominator it reports on: it is exact for monitors at intervals up to 270s and reports early above that, since one fleet-wide window cannot scale the way a per-monitor one does. Per-location scaling is a follow-up.

Independence-guard signals, emitted by the probe sanity guard. Quorum treats several locations agreeing as evidence about the target, which holds only while their failures are independent — a pop that has lost its own transit fails nearly everything it probes at once, which looks exactly like hundreds of simultaneous customer outages. proxima_probe_location_suppressed{location} is 1 while such a location's votes are being dropped and 0 once they count again; it is per-replica, so alert on max() rather than min() — the guard is deliberately not a singleton, because its output is in-process state the result worker in that same replica reads on every result. proxima_probe_no_eligible_locations_total{kind} counts evaluations of a monitor left with no location able to vote at all: probed every interval, structurally unable to confirm anything, and — before the gauge was widened to cover it — reporting itself as not degraded. The guard carries its own proof of life, proxima_probe_sanity_sweeps_total{result} and proxima_probe_sanity_last_success_timestamp_seconds (alert on min() across replicas, since every replica runs its own guard): if it stops, the suppression set freezes while the suppressed gauge keeps reporting stale values, and with nothing suppressed it would otherwise die silently. proxima_probe_sanity_withheld_total{location,reason} records the verdicts it declined — single_tenant_failure (the failures are confined to one client, so this is a customer outage and not the pop's transit) and single_tenant_fleet (every monitor at that pop belongs to one client, so it can never be judged at all — which is not the same as being healthy).

Tier 2 — a broken fleet. Tier 1 cannot see the failure that breaks the independence premise hardest: a bug shipped to every prober fails all locations together, so each looks individually broken, every monitor confirms down on what reads as unanimous agreement between independent vantage points, and no surviving location is left to disagree. proxima_probe_fleet_suppressed is 1 while the guard is holding every synthetic page for that reason, and 0 otherwise — per-replica, so alert on max(). It suppresses emission and never state: monitors still transition and events are still written, so the history read during the incident review stays true; only the wake-someone-up action is held. It counts a location as failing using tier 1's own verdict rather than re-deriving it, so a pop tier 1 judged to be one customer's outage (single_tenant_failure) is not counted toward halting every tenant's pages — otherwise a single tenant who dominates most pops would take the whole product's paging down when their own datacentre went dark. A pop that is undecidable (single_tenant_fleet) is still counted, so a private-pop fleet stays trippable. proxima_probe_guard_inert{tier} answers the question the first gauge cannot: a tier can be structurally unable to fire, and there are two routes into that — the sample is too small to be evidence (a location below MinSample, a fleet below FleetMinLocations), or the tier's threshold has been set above 1.0, which is the documented way to switch it off. Both are correct and both mean a safety guard can be entirely inert while looking configured and healthy — on a three-pop pilot fleet that is the likely state, not the edge case. An operator reading proxima_probe_fleet_suppressed 0 is entitled to know whether that means "healthy" or "this check has never been able to run". Two shapes to expect concretely: tier 1 is two gates deep (MinSample, and failures spanning at least two tenants), so on a thin or single-tenant fleet it may never fire and the reason is split between this gauge and proxima_probe_sanity_withheld_total; and tier 2 at the default FleetThreshold of 0.75 needs 3 of 3 on the current pilot fleet, since 2/3 = 0.667 — one trippable state, unanimity, and inert entirely at two pops. Nothing in the repository alerts on this gauge yet — max by (tier) (proxima_probe_guard_inert) == 1 is the rule that makes it useful, and without it an inert guard is indistinguishable from a healthy one. Every threshold behind both tiers is an explicit guess awaiting the shadow window and is tunable from the environment: PROXIMA_PROBE_SANITY_INTERVAL, PROXIMA_PROBE_SANITY_WINDOW, PROXIMA_PROBE_LOCATION_THRESHOLD, PROXIMA_PROBE_LOCATION_MIN_SAMPLE, PROXIMA_PROBE_FLEET_THRESHOLD, PROXIMA_PROBE_FLEET_MIN_LOCATIONS. A non-positive value is refused and logged rather than silently replaced by the default — to disable a tier, set its threshold above 1.0.

Push-monitor counters: proxima_push_heartbeat_total{result} (ok, fail, unknown_token) counts heartbeat calls that got past the rate limiter; proxima_push_auditor_sweeps_total{result} and proxima_push_auditor_last_success_timestamp_seconds are the push auditor's pulse. The auditor is what turns a missed heartbeat into a down verdict, so a stale pulse means push monitors cannot go down — alert on max() (only the lock winner stamps it).

Monitor-group counters, emitted where a service — a set of one client's monitors judged and paged as a single thing — is derived and paged. None of them carries a group_id, for the reason none of the probe counters carries a monitor_id, and it bites harder because groups are cheap to create; which group did or did not page is in the log line and on the alert. proxima_monitor_group_would_page_total, proxima_monitor_group_paged_total and proxima_monitor_group_page_suppressed_total{reason} are the same gate chain the monitor path reports, with the same two refusal reasons (paging_disabled, fleet_suppressed) and the same identity and resolve asymmetry — monitor_groups.paging_enabled also ships false, so would_page against proxima_probe_would_page_total is what says how many member pages one group page replaced. proxima_monitor_group_transitions_total{to_state} counts applied group state changes including the edges that page nobody. The remaining four belong to the drift auditor, the dead-man's switch behind edge-triggered evaluation: proxima_monitor_group_drift_total{correction} counts what a sweep had to correct (state — a page or a resolve was missed and has just been sent, late; counts — the state was right and the evidence beside it was stale), and a corrected drift that is not counted looks exactly like a system with no drift, so a sustained rate is a bug at a mutation path rather than a threshold to widen. proxima_monitor_group_authority_conflicts_total counts groups holding paging_enabled beside a member that holds it too — a pair that double-pages every outage of that service, that the API's authoring-time refusal structurally cannot catch (two concurrent PATCHes, or hand-run SQL), and that never self-heals, because the auditor deliberately does not choose which side loses its page. proxima_monitor_group_auditor_sweeps_total{result} carries three values, not two — clean, partial (it ran but could not finish some unit of work, and still stamps the pulse) and error (it read nothing and withholds the pulse) — and proxima_monitor_group_auditor_last_success_timestamp_seconds is the pulse: the auditor is an advisory-lock singleton, so alert on max() across replicas, the opposite of the sanity guard's rule. Full behaviour in Uptime Monitors.

Public status pages​

Backend counters on /metrics, registered in backend/internal/observability/metrics_status_page.go. No label carries a page ID or slug: pages are client-authored and unbounded, and /metrics never expires a series. Every label comes from a closed set, so an unknown value folds into a known one (result → error, tab → status) and a crafted path can never mint a series.

MetricTypeLabelsDescription
proxima_status_page_requests_totalcounterresult (ok, not_found, rate_limited, error), tab (status, maintenance, incidents, asset, css)Requests to <slug>.{PROXIMA_STATUS_PAGE_DOMAIN}. not_found deliberately covers an unknown slug, a draft, a Console-only page, a wrong or missing key, a refused IP and a malformed host alike. rate_limited is over 120 a minute per trusted client IP. error is any 503: a failed page lookup, snapshot build or asset read, or a recovered panic. asset covers logos, the favicon, /static/* and robots.txt; css is /page.css
proxima_status_page_render_secondshistogram—Snapshot (cached or built) plus template render, for one HTML tab
proxima_status_page_snapshot_cache_totalcounterresult (hit, miss, error)Snapshot lookups. miss builds the snapshot and stores it in Valkey for 30 s. error means Valkey was unreachable, or skipped for the 10 s after a Valkey error, and the snapshot came from the 10 s in-process cache, so a sustained error rate means every replica is building its own

See Status Pages.

Label Schema​

Every metric stored in VictoriaMetrics includes these system labels:

LabelDescription
__name__Metric name (e.g., cpu_user_seconds_total)
host_idHost UUID
environment_idEnvironment UUID
labels_keyDeterministic key from sorted label pairs (e.g., device=sda,mount=/)

Plus the original agent labels (e.g., device, mount, interface, collector, database, table).

The labels_key enables series discovery — the API endpoint GET /hosts/:id/metrics/:name/series returns all distinct labels_key values, allowing the frontend to enumerate and query each series individually.