Skip to main content

Kubernetes Inventory

The Kubernetes inventory system collects and stores cluster infrastructure data — clusters, nodes, namespaces, workloads, pods, services, ingresses, PVCs, events, jobs, HPAs, network policies, resource quotas, and endpoint slices — and exposes it through REST API endpoints with RBAC client scoping. Clusters are also registered as CMDB assets, enabling unified inventory views alongside hosts.

Phase

CMDB-2a is complete. The system collects 13 Kubernetes resource types via the cluster agent, stores them in PostgreSQL, and exposes read-only API endpoints. The cluster agent binary runs as a single-replica Deployment inside each K8s cluster.

Overview​

Kubernetes inventory data flows from the cluster agent (deployed inside each K8s cluster) through NATS into the backend, where the KubernetesWorker processes it into PostgreSQL tables. Each cluster also gets an entry in the assets table for unified CMDB views.

Collected Resource Types​

The cluster agent gathers 13 resource types from the K8s API:

#ResourceK8s API GroupTableDescription
1NodesCoreV1k8s_nodesStatus, capacity, allocatable, usage, addresses
2NamespacesCoreV1k8s_namespacesStatus, labels, workload/pod/service counts
3WorkloadsAppsV1 (+ dynamic client for argoproj.io/v1alpha1 rollouts)k8s_workloadsDeployments, StatefulSets, DaemonSets, Argo Rollouts (kind Rollout)
4PodsCoreV1k8s_podsPhase, containers, restarts, resources, owner
5ServicesCoreV1k8s_servicesType, ports, selector, cluster IP
6IngressesNetworkingV1k8s_ingressesRules, TLS, ingress class, load balancer
7PVCsCoreV1k8s_pvcsPhase, storage class, capacity, access modes
8EventsCoreV1k8s_eventsType, reason, involved object, count, timestamps
9Jobs/CronJobsBatchV1k8s_jobsStatus, completions, schedule, containers
10HPAsAutoscalingV2k8s_hpasTarget, min/max/current replicas, metrics
11Network PoliciesNetworkingV1k8s_network_policiesPod selector, ingress/egress rules, policy types
12Resource QuotasCoreV1k8s_resource_quotasHard limits, current usage per namespace
13Endpoint SlicesDiscoveryV1k8s_endpoint_slicesAddresses, ports, service association

Data Flow​

NATS Message Path​

Kubernetes inventory messages use the standard telemetry subject pattern with kubernetes as the message type:

proxima.{client_slug}.{env_slug}.{host_id}.kubernetes

The message is routed through the EVENTS JetStream stream to the EventsWorker, which dispatches kubernetes message types to the KubernetesWorker.

Upsert-Then-Prune Sync Pattern​

The KubernetesWorker uses an upsert-then-prune pattern rather than delete-reinsert. This is important for data consistency and referential integrity:

  1. Upsert phase — Each resource is inserted or updated using ON CONFLICT ... DO UPDATE. The worker tracks all seen entity IDs during the sync.
  2. Prune phase — After all upserts, the worker deletes any rows for the cluster that were not seen in this sync cycle. This removes resources that no longer exist in the cluster.

The prune follows FK dependency order (children before parents):

k8s_endpoint_slices → k8s_resource_quotas → k8s_network_policies →
k8s_hpas → k8s_jobs → k8s_events → k8s_pvcs → k8s_ingresses →
k8s_services → k8s_pods → k8s_workloads → k8s_namespaces → k8s_nodes

All upserts and prunes run inside a single database transaction with a 60-second timeout.

Processing Steps​

  1. Asset upsert — Creates or updates an assets record with asset_type = 'cluster' and maps the cluster status to an asset status.
  2. Cluster upsert — Inserts or updates the clusters record with identity, K8s metadata, cloud context, and status fields.
  3. Node sync — Upserts all nodes for the cluster. For each node, attempts to link it to an existing host via hostname or IP address matching.
  4. Namespace sync — Upserts all namespaces with workload, pod, and service counts.
  5. Workload sync — Upserts all workloads (Deployments, StatefulSets, DaemonSets, Argo Rollouts) with replica status and Helm metadata.
  6. Pod sync — Upserts all pods with phase, container specs, resources, restarts, and owner references.
  7. Service sync — Upserts all services with type, ports, selectors, and IPs.
  8. Ingress sync — Upserts all ingresses with rules, TLS, and class metadata.
  9. PVC sync — Upserts all PVCs with phase, storage class, and capacity.
  10. Event sync — Upserts all events with type, reason, and involved object.
  11. Job sync — Upserts all jobs and cron jobs with schedule, completions, and status.
  12. HPA sync — Upserts all HPAs with target reference, scaling bounds, and current metrics.
  13. Network policy sync — Upserts all network policies with selectors and rules.
  14. Resource quota sync — Upserts all quotas with hard limits and current usage.
  15. Endpoint slice sync — Upserts all endpoint slices with addresses and ports.
  16. Prune — Removes entities not seen in this sync cycle.
  17. Count update — Updates denormalized counts on the cluster record.

Collection Limits and Filtering​

Namespace exclusion: The collector skips namespaces listed in PROXIMA_K8S_EXCLUDE_NAMESPACES (comma-separated). Excluded namespaces and their resources are not collected or stored. Specifically:

  • Every namespaced kind is filtered, not just the namespace row: pods, workloads, services, ingresses, PVCs, events, jobs and CronJobs, HPAs, network policies, resource quotas and endpoint slices.
  • Exclusion is applied before the pod and event caps below, so an excluded namespace cannot consume the budget that kept namespaces need.
  • An Event also records the namespace of the object it concerns. An event about a resource in an excluded namespace is dropped even when the event object itself lives elsewhere.
  • Nodes and cluster-level facts are cluster-scoped and therefore unaffected.
  • Adding a namespace to the list removes what was already stored for it: the sync prunes rows the payload no longer contains, so the next successful sync after the change deletes that namespace's existing resources. No manual cleanup is needed.

Pod limit: The collector enforces a configurable pod limit (PROXIMA_K8S_POD_LIMIT, default: 5000). When the limit is exceeded, pods are sorted with non-Running pods first (then by restart count descending), and the list is truncated. This ensures problematic pods are always included.

Event limit: Events are capped at 1000 per sync (hardcoded maxEvents). Events are sorted by their last-seen time descending, keeping only the most recent. From cluster agent v0.8.0 the last-seen time of an event without lastTimestamp (one written through events.k8s.io, such as kube-scheduler's FailedScheduling) falls back to series.lastObservedTime, then eventTime, then firstTimestamp, so those events are no longer the first dropped.

Annotation stripping: The annotation kubectl.kubernetes.io/last-applied-configuration is stripped from all resources before transmission. This annotation often contains the full resource manifest and can be very large. The stripping is applied to workloads, services, ingresses, PVCs, jobs, and network policies.

Workload Spec Facts​

Each workload (Deployment, StatefulSet, DaemonSet, Argo Rollout) carries a spec_facts block: the facts of its spec that a "what changed?" diff compares between two syncs. The schema is api/proto/kube/v1/workload_spec.proto; both ends exchange it as JSON through wire.WorkloadSpecFacts.

FieldContents
replicasspec.replicas. Absent for a DaemonSet, which has none. A Rollout with no replicas set reports 1, as Argo defaults it. 0 (scaled to zero) is kept as 0.
strategytype (RollingUpdate / Recreate / OnDelete; canary / blueGreen for a Rollout), plus max_surge and max_unavailable as written ("25%", "1") when the spec sets them.
containers[]Main containers sorted by name, then init containers (init: true) in spec order. Init containers run one after another, so reordering them is a real change and shows up as one.
containers[].name, imageAs in the pod template.
containers[].requests, limitscpu and memory as Kubernetes quantity strings in canonical form ("500m", "1Gi"). The agent canonicalizes them. For Deployments, StatefulSets and DaemonSets that is the form the apiserver already returns, so it matches kubectl get -o yaml. A Rollout is a CRD, whose quantities the apiserver keeps as written, so a Rollout's cpu: 0.5 is reported as "500m". Absent when not set.
containers[].env_namesEnv var names, literal and valueFrom alike, sorted and de-duplicated.
containers[].env_fromenvFrom sources as configmap:<name> / secret:<name>, sorted. When the source sets a variable-name prefix, it is appended after #: configmap:app-config#APP_. The prefix renames every injected variable, and it is a name, not a value. The object name cannot contain #, but under relaxed env var validation the prefix may, so split at the first #.
containers[].volume_refsWhat backs each volume the container mounts: configmap:<name>, secret:<name>, pvc:<claimName>, and pvc-template:<name> for a StatefulSet volumeClaimTemplate (the PVC itself is named per replica). Projected volumes contribute one entry per ConfigMap/Secret source. A CSI volume contributes secret:<name> for its nodePublishSecretRef when it has one; its driver and volumeAttributes are never read. emptyDir, hostPath, downwardAPI and similar volumes are not listed. Sorted.
facts_versionThe version of the agent's normalization rules (currently 1). Any agent change to how facts are normalized (quantity spelling, list ordering, which refs are listed) bumps it. The backend diffs only facts of equal versions and treats a mismatch as a new baseline. Absent (0) from an agent that predates the field.

Values are never sent. An env entry contributes only its name, whether it is a literal or a valueFrom. The following are not read:

  • the key selector of a secretKeyRef/configMapKeyRef;
  • a fieldRef path;
  • the items (keys and paths) of Secret, ConfigMap and projected volumes;
  • CSI volumeAttributes;
  • container args and command;
  • annotations.

ConfigMap and Secret objects are never read at all: a reference is the name written in the workload's own spec. The agent's tests enforce this in three ways:

  • a frozen allow-list of every field of the facts type, so a new field cannot be added without a test edit;
  • a fixture that writes a unique marker into every string of a Deployment (pod template and its metadata included) and fails if any marker from outside the allowed source paths (names, images, references, the envFrom prefix, strategy) reaches the output;
  • a scan of the whole serialized payload for planted values.

Every list is sorted (init containers keep their spec order), so the same spec always produces the same bytes and a diff between syncs is never noise.

An oversized inventory falls back to no facts. The whole inventory travels as one NATS message. If it exceeds the server's max_payload, the agent drops spec_facts from every workload and publishes once more without them. That sync then records no workload changes, but the rest of the inventory arrives. The agent logs a WARN and counts the event on proxima_collector_inventory_too_large_total{outcome="spec_facts_dropped"}. If the inventory is still too large without the facts, nothing is published, as before, and outcome="inventory_dropped" counts it. proxima_collector_inventory_payload_bytes shows the encoded size of the last full inventory, so the headroom is visible before the limit is reached.

Absent means unknown. An agent older than this field sends no spec_facts, and a Rollout that uses spec.workloadRef (no inline pod template) reports none. The backend reads both as "unknown", never as an empty set, so an agent upgrade does not look like every fact changing. The backend turns differences between two syncs into change records (next section). The facts add about 1 KB per workload: about 200 KB for 200 workloads with three containers and around 17 env names each.

Workload change records​

Each sync, the Kubernetes worker compares every workload's incoming spec_facts with the last known set, stored in k8s_workloads.spec_facts together with the time of the snapshot it came from (spec_facts_at). It writes at most one change_events row per changed workload per sync. The rows appear on the Changes page, in the change feed the API serves, and in the cluster's "What changed?" timeline (which the workload page's History tab renders).

ColumnValue
sourcek8s-config
event_typeworkload_scaled for a manual replicas-only change, otherwise workload_changed
client_id, environment_idThe cluster's own, so the usual changes:read scoping applies
host_idNULL: a workload belongs to a cluster, not a host
scanned_atThe snapshot's time (the inventory envelope's timestamp), not the time the backend processed it
severityinfo
summaryDeployment shop/checkout-api: image checkout-api 2.40.3 → 2.41.0; memory limit 1Gi → 768Mi; +2 more. It names at most three changes, then +N more, and is cut at 300 characters. A manual scale reads Deployment shop/checkout-api scaled 3 → 0.
details{cluster_id, cluster, kind, namespace, name, change_count, truncated, changes: [{field, before, after}]}

changes holds at most 50 entries. change_count is the real total, and truncated is true when entries were cut. before or after is null when that side had no value: the fact was added, removed, or unset. Quantities appear exactly as the agent sent them (1Gi, 500m).

field is a dotted path. Container names are DNS-1123 labels and never contain a dot, so the path is unambiguous:

fieldMeaning
replicasspec.replicas (never present for a workload an HPA targets, see below)
strategy.type, strategy.max_surge, strategy.max_unavailableThe update strategy
containers.<c>Container added (after = image) or removed (before = image)
containers.<c>.imageImage reference
containers.<c>.requests.cpu / .requests.memory / .limits.cpu / .limits.memoryResources
containers.<c>.envOne env var name added (after) or removed (before)
containers.<c>.env_fromOne envFrom ref added or removed. A prefix change on the same ref is one entry: configmap:app#OLD_ → configmap:app#NEW_.
containers.<c>.volumesOne volume ref added or removed
init_containers.<c>…The same, for init containers
init_containers.orderInit containers reordered. They run in order, so this is a real change. Reordering main containers is not a change.

HPA scaling is not recorded. When a HorizontalPodAutoscaler in the same snapshot targets the workload (same namespace, scaleTargetRef kind and name), its replica count belongs to the HPA. A replicas-only change writes no row, and a mixed change (an image bump during a scale) records everything except replicas. The replica history is already the VictoriaMetrics series k8s_workload_replicas, and the scaling events are in the Kubernetes event stream. As change rows, HPA moves would arrive every few minutes during a load incident and push real causes out of triage's change fetch. The match uses only the HPAs in the current snapshot, so a scale in the same sync that deletes its HPA is recorded as manual. Scaling without an HPA is an operator action and is recorded as workload_scaled.

Rules the diff keeps:

  • First sight is a baseline. A workload with no stored facts gets its facts stored and no record. This covers the first sync after this backend is deployed, the first sync after a cluster agent is upgraded to one that sends facts, and a newly created workload. Without this rule, rolling out the backend would write one "changed" row for every workload in every cluster.
  • A facts format change is a baseline. The agent stamps its facts with facts_version, the version of its normalization rules (how it spells quantities, how it orders lists). When the stored and incoming versions differ, the incoming facts are stored and nothing is recorded, so an agent upgrade that changes normalization does not report every spelling difference as a change.
  • Absent means unknown, and the stored set is kept. When a workload arrives with no spec_facts (old agent, a workloadRef Rollout, facts dropped for size), nothing is diffed and the stored set is not overwritten. If the facts come back unchanged later, the comparison is against the last known set, so nothing is recorded.
  • Only a newer snapshot counts. Facts are diffed, and overwritten, only when the incoming snapshot's time is strictly later than spec_facts_at. A snapshot that was redelivered after a later one had already been processed is ignored for change purposes: it does not record a rollback that never happened, and it does not rewind the stored facts. An equal timestamp is a replay, not news. Skipped workloads are counted on proxima_k8s_spec_facts_not_newer_total.
  • A clock ahead is clamped. The snapshot time is the agent's wall clock. A stamp more than one minute ahead of the backend's clock is clamped to backend time plus one minute, and the backend logs a WARN (kubernetes snapshot stamped in the future, clamped). Without the clamp, one snapshot from a node whose clock jumped a day ahead would make every later honest snapshot "not newer" and stop recording for that workload for a day.
  • No values. The facts carry names and refs only, and the records carry only what the facts carry.
  • Atomic and serialized per cluster. The workload sync locks the cluster row (FOR NO KEY UPDATE, which serializes syncs without blocking inserts that reference the cluster), then reads the stored facts, overwrites them and writes the change rows in one transaction. A crash before commit loses nothing and writes nothing, so the redelivered snapshot produces the record once. Two snapshots of the same cluster processed at the same time (for example a backlog spread over several backend replicas) wait for each other instead of diffing against the same stored facts. The prune step takes the same lock first, so every transaction locks the cluster before its child rows.
  • Deletions and creations are not recorded. A workload missing from one snapshot may be a partial collection (for example, a Rollouts read that was refused) as much as a deletion. Its row is pruned as before, but no "deleted" record is written. A new workload is a first sight and gets a baseline only.

Weights for L1. workload_changed has weight class Config (0.9), like a NetworkPolicy edit. workload_scaled is Neutral (0.5). A manual scale-down or scale-to-zero shortly before an alert is a prime cause, and at 0.5 it is citable only while recent: its score falls below the ranker's minimum about 28 minutes before the alert. The triage change fetch (ListDeployChanges, newest 100 rows taken before weighting) excludes manual scales, so a burst of them cannot push an older deploy out. Manual scales reach the ranker through a separate fetch capped at 15 rows. Scale-ups and scale-downs share the weight; telling them apart would need a second event type, and a recent scale-up an operator made in response to the incident is easy for the verification pass to dismiss.

Volume. Per cluster, expect one workload_changed row per workload per rollout (an image bump in a Deployment is one row, whatever the replica count), plus one row per manual scale or config edit. A cluster deploying 30 services twice a day writes about 60 rows a day. HPA activity adds nothing. Change retention applies as for every other source.

Security​

The cluster agent follows a strict security posture:

  • No secrets or configmaps collected — These resource types are not in the ClusterRole permissions.
  • Read-only RBAC — The ClusterRole only grants get, list, watch verbs.
  • Annotation stripping — kubectl.kubernetes.io/last-applied-configuration is removed to avoid leaking full resource manifests, and every surviving annotation goes through the agent's key-aware secret detector (the same pass as the live manifest read; see Kubernetes access → Security).
  • Workload spec facts are names, never values — env var names and ConfigMap/Secret/PVC names only (see Workload Spec Facts).
  • ArgoCD Applications without their values — no Helm values, parameters or plugin env, and repository URLs with credentials stripped, failing closed (see ArgoCD Drift).
  • Container logs are never stored — the live logs route proxies one read per request and keeps nothing (see Live Reads).

Table Schemas​

clusters​

The primary cluster identity table. Linked to assets via asset_id.

ColumnTypeDescription
idUUIDPrimary key
asset_idUUIDFK to assets.id (ON DELETE CASCADE)
client_idUUIDFK to clients.id
environment_idUUIDFK to environments.id
nameTEXTCluster name (e.g., k8s-prod-01)
display_nameTEXTHuman-readable display name
slugTEXTURL-safe identifier (unique per environment)
versionTEXTKubernetes version (e.g., v1.29.3)
platformTEXTDistribution platform (rke2, eks, gke, aks, k3s, kubeadm)
distributionTEXTDistribution detail
api_server_urlTEXTAPI server endpoint
cluster_cidrTEXTPod CIDR range
service_cidrTEXTService CIDR range
dns_domainTEXTCluster DNS domain (default: cluster.local)
cloud_providerTEXTCloud provider (aws, gcp, azure, hetzner)
cloud_regionTEXTCloud region
cloud_account_idTEXTCloud account identifier
statusTEXTCluster health status (healthy, degraded, unreachable, unknown)
agent_statusTEXTAgent connection status (connected, disconnected)
agent_versionTEXTReporting agent version
node_countINTEGERDenormalized node count
namespace_countINTEGERDenormalized namespace count
workload_countINTEGERDenormalized workload count
last_seen_atTIMESTAMPTZLast inventory message timestamp
created_atTIMESTAMPTZRecord creation time
updated_atTIMESTAMPTZLast modification time

Indexes: client_id, environment_id, asset_id, status, agent_status, last_seen_at, unique on (environment_id, slug).

k8s_nodes​

Kubernetes nodes with capacity, allocatable resources, and usage data.

ColumnTypeDescription
idUUIDPrimary key
cluster_idUUIDFK to clusters.id (ON DELETE CASCADE)
host_idUUIDFK to hosts.id (nullable, set by host linkage)
nameTEXTNode name
uidTEXTKubernetes UID
rolesTEXT[]Node roles (e.g., control-plane, worker)
labelsJSONBKubernetes labels
taintsJSONBNode taints
statusTEXTNode condition status (Ready, NotReady, Unknown)
conditionsJSONBFull conditions array
unschedulableBOOLEANWhether scheduling is disabled
kubelet_versionTEXTKubelet version
container_runtimeTEXTContainer runtime (e.g., containerd://1.7.2)
os_imageTEXTOS image (e.g., Ubuntu 22.04.3 LTS)
kernel_versionTEXTKernel version
capacity_cpu_millicoresINTEGERTotal CPU capacity
capacity_memory_bytesBIGINTTotal memory capacity
capacity_podsINTEGERMax pod count
allocatable_cpu_millicoresINTEGERAllocatable CPU
allocatable_memory_bytesBIGINTAllocatable memory
allocatable_podsINTEGERAllocatable pod count
usage_cpu_millicoresINTEGERCurrent CPU usage
usage_memory_bytesBIGINTCurrent memory usage
pod_countINTEGERCurrent pod count
provider_idTEXTCloud provider instance ID
instance_typeTEXTCloud instance type
zoneTEXTAvailability zone

Indexes: cluster_id, host_id, status, unique on (cluster_id, name) and (cluster_id, uid).

k8s_namespaces​

Kubernetes namespaces with resource counts.

ColumnTypeDescription
idUUIDPrimary key
cluster_idUUIDFK to clusters.id (ON DELETE CASCADE)
nameTEXTNamespace name
uidTEXTKubernetes UID
statusTEXTNamespace phase (Active, Terminating)
labelsJSONBKubernetes labels
workload_countINTEGERWorkloads in this namespace
pod_countINTEGERPods in this namespace
service_countINTEGERServices in this namespace

Indexes: cluster_id, GIN on labels, unique on (cluster_id, name).

k8s_workloads​

Kubernetes workload resources (Deployments, StatefulSets, DaemonSets and Argo Rollouts; Jobs and CronJobs live in k8s_jobs).

ColumnTypeDescription
idUUIDPrimary key
cluster_idUUIDFK to clusters.id (ON DELETE CASCADE)
namespace_idUUIDFK to k8s_namespaces.id (ON DELETE CASCADE)
kindTEXTResource kind (Deployment, StatefulSet, DaemonSet, Rollout). No CHECK constraint: a new kind needs no migration
nameTEXTWorkload name
uidTEXTKubernetes UID
namespaceTEXTNamespace name (denormalized)
replicas_desiredINTEGERDesired replica count
replicas_readyINTEGERReady replica count
replicas_availableINTEGERAvailable replica count
conditionsJSONBWorkload conditions
strategyTEXTUpdate strategy (RollingUpdate, Recreate, OnDelete; canary or blueGreen for a Rollout)
service_accountTEXTService account name
containersJSONBContainer specs (image, resources, ports)
labelsJSONBKubernetes labels
annotationsJSONBKubernetes annotations
helm_releaseTEXTHelm release name
helm_chartTEXTHelm chart name
helm_chart_versionTEXTHelm chart version

Indexes: cluster_id, namespace_id, kind, GIN on labels, conditional index on helm_release (non-null only), unique on (cluster_id, kind, namespace, name).

Table Relationships​

API Endpoints​

All endpoints require authentication and the assets:read permission, except the cluster metrics endpoint, which requires metrics:read. Responses are scoped to the caller's allowed clients. What the Clusters page and a cluster's Overview draw from these endpoints is described in Kubernetes Clusters: Overview & List.

List Clusters​

GET /api/v1/clusters

Query parameters:

ParameterTypeDescription
client_idUUIDFilter by client
environment_idUUIDFilter by environment
statusstringFilter by cluster status
searchstringName search (case-insensitive substring)
includestringcapacity adds per-cluster capacity and health counts (below). Any other value is a 400.
pageintPage number (default: 1)
per_pageintItems per page (default: 20, maximum 100; larger values are clamped)

The list meta is {total, page, per_page} with no has_more: to read every cluster, keep requesting pages while page * per_page < total, using the per_page the response echoes.

Response:

{
"data": [
{
"id": "a1b2c3d4-...",
"asset_id": "e5f6g7h8-...",
"client_id": "c1d2e3f4-...",
"environment_id": "d4e5f6a7-...",
"name": "k8s-prod-01",
"display_name": "Production Cluster",
"slug": "k8s-prod-01",
"version": "v1.29.3",
"platform": "rke2",
"status": "healthy",
"node_count": 5,
"namespace_count": 12,
"workload_count": 47,
"last_seen_at": "2026-04-05T12:00:00Z"
}
],
"meta": {
"total": 3,
"page": 1,
"per_page": 50
}
}

Get Cluster​

GET /api/v1/clusters/:clusterID

Returns full cluster details including all metadata, cloud context, and counts.

List Capacity (?include=capacity)​

With include=capacity, each list item also carries:

FieldDescription
capacityThe same block, from the same SQL, as the summary's capacity (below)
pods_not_runningPods in any phase other than Running or Succeeded (Pending + Failed + Unknown)
warning_events_1hWarning event objects whose last_seen_at is within the past hour. Kubernetes folds repeats into one object with a count; this counts objects and does not add up count.

The three fields are read in one grouped query for exactly the clusters on the returned page, after the list's tenant scope has been applied. Without the flag they are absent. A 0 is always present when the flag is set, never dropped. A cluster deleted between the list read and the capacity read has the three fields absent, so treat them as optional per item.

Cluster Summary​

GET /api/v1/clusters/:clusterID/summary

A health digest of one cluster, read from PostgreSQL only (the inventory the cluster agent syncs every 5 minutes). All sections are read in one read-only, repeatable-read transaction, so they describe the same snapshot. Lists are capped, and the counts beside them are exact.

FieldDescription
nodes{total, ready, not_ready, with_metrics}. ready means status Ready; not_ready = total − ready (includes Unknown); with_metrics counts nodes linked to a host in the cluster's own environment — the same rule ingest uses to stamp cluster_id, so a stale link to a host moved to another environment does not count. Linkage is not reporting: a linked node whose agent is down still counts
not_ready_nodesNames of not-ready nodes, at most 10
pods{total, running, pending, succeeded, failed, unknown}. unknown = total minus the four named phases, so the five always sum to total
waiting[{reason, count}]: containers currently waiting, by reason (e.g. CrashLoopBackOff), at most 10 reasons. Counts containers, not pods
pending_long[{namespace, name, age_seconds}]: pods Pending for more than 5 minutes, oldest first, at most 10
top_restarters[{namespace, name, workload_kind, workload, restarts}]: pods with the highest restart count (> 0), at most 5
workloads_not_ready[{id, kind, namespace, name, ready, desired}] (id links to the workload page): workloads with fewer ready replicas than desired, largest shortfall first, at most 10
capacitycpu_allocatable_millicores, cpu_requested_millicores, cpu_limits_millicores, cpu_used_millicores, mem_allocatable_bytes, mem_requested_bytes, mem_limits_bytes, mem_used_bytes, pods_allocatable, pods_running. Requests and limits count only pods not in Succeeded/Failed. cpu_used_millicores and mem_used_bytes come from metrics-server (k8s_nodes.usage_*) and are null when no node reports usage — never 0. They sum only the nodes that report usage, so they pair with cpu_allocatable_reporting_millicores / mem_allocatable_reporting_bytes (allocatable over those same nodes). nodes_total, cpu_usage_nodes and mem_usage_nodes give the coverage ("usage from N of M nodes")
warning_events_1hWarning event objects last seen in the past hour (exact)
top_warning_reasons[{reason, count}], at most 5
generated_atThe database time of the snapshot

Every list is [], never null. Authorization follows Get Cluster: an unknown cluster, or one outside the caller's environments, is a 404; assets:read missing on the cluster's environment is a 403. The route accepts an environment-scoped assets:read grant (RequireAnyPermissionInAnyScope) and re-checks the cluster's own environment in the handler.

Cluster Metrics​

GET /api/v1/clusters/:clusterID/metrics/:query?start=…&end=…[&namespace=…][&workload=…&workload_kind=…][&instant=true]

Runs one of ten allow-listed queries in the cluster's VictoriaMetrics tenant. The client picks a name and never sends PromQL. Every selector is confined to the cluster and its environment (cluster_id="…",environment_id="…"). The cluster_id label is added by the backend at ingest to metrics from hosts linked to the cluster's nodes (see docs/standards/metrics.md).

queryPromQL (selector abbreviated sel)Unit
cpu_by_nodemax by (node) (node_cpu_usage_ratio{sel})0–1 ratio of the node's CPU
memory_by_namespacesum by (namespace) (pod_memory_working_set_bytes{sel})bytes
restartssum(clamp_min(last_over_time(container_restarts_total{sel}[W]) - last_over_time(container_restarts_total{sel}[W] offset step), 0)), W = max(step, 5m)restarts within each step
pods_by_phasesum by (phase) (kubelet_pods{sel})pods
cpu_by_workloadsum by (workload_kind, workload) (pod_cpu_usage_ratio{sel, namespace="…", workload!=""}) — requires namespacecores
cpu_by_namespacesum by (namespace) (pod_cpu_usage_ratio{sel})cores
cpu_by_podsum by (pod) (pod_cpu_usage_ratio{selW})cores
memory_by_podsum by (pod) (pod_memory_working_set_bytes{selW})bytes
restarts_by_workloadThe restarts rule over container_restarts_total{selW} (both the current and the offset selector)restarts within each step
replicasmax by (type) (k8s_workload_replicas{selW,host_id=""}) — type is desired or readyreplicas

selW is sel plus namespace="…",workload="…",workload_kind="…": the four per-workload queries require all three parameters. A workload name is unique only per kind within a namespace, so filtering by the kind keeps a Deployment web and a StatefulSet web apart. k8s_workload_replicas is written by the backend at each inventory sync (see docs/standards/metrics.md), so replicas has no data from before the backend version that writes it was deployed. The backend writes it with no host_id, and every agent-sent point carries one (the metrics worker stamps it and never lets an agent unset it), so the host_id="" matcher keeps a node agent from emitting its own k8s_workload_replicas and having it read as the backend's.

restarts is a per-step difference rather than increase(). A new series contributes nothing, so an agent upgrade or new pod does not produce a false spike, and gaps shorter than W are bridged. Each point covers the step ending at its timestamp. Pods without workload labels (older node agents) are left out of cpu_by_workload.

  • Parameters: start and end (RFC 3339, start before end, at most 30 days apart); namespace (a DNS-1123 label), required by cpu_by_workload and the four per-workload queries; workload (a DNS-1123 subdomain, at most 253 characters) and workload_kind (an upper-case letter then at most 62 letters or digits, e.g. Deployment, Rollout), required by the per-workload queries. Each of the three is refused (400) by every query that does not take it, so a caller never believes a filter was applied when it was not. The step is chosen from the range: 1m up to 6h, 5m up to 24h, 15m up to 7d, 1h beyond.
  • instant=true (accepted by cpu_by_namespace and memory_by_namespace only; any other query is a 400, and any value other than true/false is a 400) runs an instant query at end instead: one point per series, capped at 5,000 series with truncated, and step_seconds: 0. The Namespaces tab uses it so that every namespace gets a "used" value.
  • Response: {"data": {"series": [{"labels_key", "labels", "data": [{"time", "avg_value", …}]}], "truncated": false, "step_seconds": 60}}. Series are ranked by peak and capped at 50 with truncated: true. NaN and ±Inf points are dropped as gaps. A cluster whose node agents do not send the series returns "series": [], never null.
  • Errors: an unknown query name, an unknown query parameter, a missing or bad namespace, workload or workload kind, instant on a query that does not take it, or a bad range is a 400. A cluster outside the caller's environments is a 404. metrics:read missing on the cluster's environment is a 403 (assets:read alone is not enough). VictoriaMetrics not configured is a 503, and a VictoriaMetrics failure is a 502 with a generic body. The caller is authorized before the query is validated, so an unauthorized caller learns nothing beyond 404/403.
  • The route uses RequireAnyPermissionInAnyScope(metrics:read), so environment-scoped grants work, and the handler re-checks the cluster's environment.

List Cluster Nodes​

GET /api/v1/clusters/:clusterID/nodes

Returns all nodes in the cluster with capacity, allocatable resources, usage, and host linkage.

Node Capacity​

GET /api/v1/clusters/:clusterID/nodes/capacity

Every node of the cluster for the Nodes tab's capacity table, from PostgreSQL in one query: {"data": {"nodes": [...], "truncated": false, "total": 3}}. Each node carries id, name, status, roles ([], never null), unschedulable, pod_count, cpu_allocatable_millicores, cpu_requested_millicores, cpu_limits_millicores, cpu_used_millicores, mem_allocatable_bytes, mem_requested_bytes, mem_limits_bytes, mem_used_bytes and pods_allocatable.

  • Requests and limits are summed per node over pods not in Succeeded/Failed.
  • cpu_used_millicores / mem_used_bytes are metrics-server usage, null when the node reports none (a legacy stored 0 is also returned as null).
  • At most 1,000 nodes, ordered by name. total is the cluster's node count (computed before the limit), and truncated is total > the nodes returned.

Get Cluster Node​

GET /api/v1/clusters/:clusterID/nodes/:nodeID

{"data": node}, the same record as the node list. The node is looked up by ID and the cluster in the URL: a node of another cluster, or an unknown one, is a 404.

Namespace Usage​

GET /api/v1/clusters/:clusterID/namespaces/usage

{"data": [...]}: every namespace of the cluster (uncapped, ordered by name) with name, workload_count, pods_total, pods_running, restarts, cpu_requested_millicores, cpu_limits_millicores, mem_requested_bytes and mem_limits_bytes. Pods are matched to a namespace by name; requests and limits count pods not in Succeeded/Failed, while pods_total and restarts (the sum of restart_count) count every pod. workload_count is the value the inventory sync stored on the namespace, not a live count. A namespace with no pods returns zeros; the list is [], never null.

List Cluster Namespaces​

GET /api/v1/clusters/:clusterID/namespaces

Returns all namespaces with resource counts.

Get Cluster Workload​

GET /api/v1/clusters/:clusterID/workloads/:workloadID

{"data": workload}, one workload record. Looked up by ID and the cluster in the URL: a workload of another cluster, or an unknown or deleted one, is a 404 ("workload not found").

The node capacity, single-node, namespace usage and single-workload routes follow Get Cluster: RequireAnyPermissionInAnyScope(assets:read) at the router, then in the handler an unknown cluster or one outside the caller's environments is a 404 and assets:read missing on the cluster's environment is a 403.

Search and sort on the resource lists​

The workloads, pods, services, ingresses, PVCs, jobs, HPAs, network policies, resource quotas and endpoint slices lists take two more query parameters:

ParameterDescription
qA case-insensitive substring of the object's name or namespace. It is matched literally: % and _ are escaped, so q=% finds names that contain a percent sign rather than every row. The value is trimmed and may be at most 100 characters. A longer value, a control character or invalid UTF-8 returns 400 invalid_param, with a message saying which.
sortOne column from the list's allow-list (below), optionally prefixed with - for descending. Any other value returns 400 invalid_param, and the error lists the allowed keys. Rows with no value sort last in both directions. Ties are broken by namespace, then name, then ID, so pages never overlap. age ascending is youngest first. With no sort, the list keeps its usual order (namespace, then name; workloads by namespace, kind and name).

Each of q and sort may appear once. A repeated parameter (?sort=name&sort=age) is ambiguous, so it returns 400 invalid_param rather than silently using one of the values.

Listsort keys
workloadsname, namespace, kind, replicas (desired), helm (release), age
podsname, namespace, phase, ready (ready containers), restarts, cpu / memory (requests), node, age (started, else created)
servicesname, namespace, type, cluster_ip (IPv4 addresses numerically, so 10.0.0.2 comes before 10.0.0.10; None, IPv6 and other values follow, ordered as text), endpoints, age
ingressesname, namespace, class, age
pvcsname, namespace, phase, storage_class, capacity, age
jobsname, namespace, kind, schedule, status, completions (succeeded), age
hpasname, namespace, target (target name), min, max, current, age
network-policies, resource-quotasname, namespace, age
endpoint-slicesname, namespace, address_type, service, age

Both parameters are read only after the cluster access checks. A caller who cannot see the cluster still gets 404, and never a 400 that would confirm the cluster exists. The sort key is mapped to a fixed column in the store, which refuses a key it does not know as a second check. The search narrows the existing filters (namespace, kind, phase, labels) and never widens them.

List Cluster Workloads​

GET /api/v1/clusters/:clusterID/workloads

Query parameters:

ParameterTypeDescription
namespacestringFilter by namespace name
kindstringFilter by resource kind
q, sortstringSearch and sort (above)
pageintPage number (default: 1)
per_pageintItems per page (default: 50)

List Cluster Pods​

GET /api/v1/clusters/:clusterID/pods

Query parameters:

ParameterTypeDescription
namespacestringFilter by namespace name
phasestringFilter by pod phase
q, sortstringSearch and sort (above)
pageintPage number (default: 1)
per_pageintItems per page (default: 50)

List Node Pods​

GET /api/v1/clusters/:clusterID/nodes/:nodeID/pods

Returns pods running on a specific node.

List Workload Pods​

GET /api/v1/clusters/:clusterID/workloads/:workloadID/pods

Returns pods belonging to a specific workload.

List Cluster Services​

GET /api/v1/clusters/:clusterID/services

List Cluster Ingresses​

GET /api/v1/clusters/:clusterID/ingresses

List Cluster PVCs​

GET /api/v1/clusters/:clusterID/pvcs

List Cluster Events​

GET /api/v1/clusters/:clusterID/events

Query parameters:

ParameterTypeDescription
typestringFilter by event type (Normal, Warning)
pageintPage number (default: 1)
per_pageintItems per page (default: 50)

List Cluster Jobs​

GET /api/v1/clusters/:clusterID/jobs

List Cluster HPAs​

GET /api/v1/clusters/:clusterID/hpas

List Cluster Network Policies​

GET /api/v1/clusters/:clusterID/network-policies

List Cluster Resource Quotas​

GET /api/v1/clusters/:clusterID/resource-quotas

List Cluster Endpoint Slices​

GET /api/v1/clusters/:clusterID/endpoint-slices

Live Reads (manifest, revisions, logs)​

GET /api/v1/clusters/:clusterID/live/manifest?kind=…&namespace=…&name=…
GET /api/v1/clusters/:clusterID/live/revisions?kind=…&namespace=…&name=…
GET /api/v1/clusters/:clusterID/live/logs?namespace=…&pod=…[&container=…][&tail=…][&previous=true][&since_time=…]

These three routes, with the events timeline, the what-changed timeline and the ArgoCD routes below, make up the read-only operate views; Kubernetes: Operate (read-only) maps them in one page, with the rollout steps. These three routes read the cluster now, through the cluster's own agent (a kube_query NATS request/reply on proxima.system.commands.<agentID>, 15 s timeout). The cluster is chosen by ID, and nothing is stored. Every read runs as the agent's read-only tier, and the caller's email (or proxima-console:<subject type>:<id> for a caller without one) is sent as the audit actor.

  • manifest returns one object, redacted in the agent before it leaves the cluster. managedFields and embedded last-applied copies are removed. Literal env values, probe header values and termination messages become <redacted: N chars>. Value-bearing annotations are redacted. notes says what was done. The rule set is listed in Kubernetes access → Security. kind is one of Deployment, StatefulSet, DaemonSet, Rollout, ReplicaSet, Pod, Service, Ingress, HorizontalPodAutoscaler, Job, CronJob, NetworkPolicy or PersistentVolumeClaim. ConfigMaps and Secrets are never readable. Response: {"data": {"kind", "namespace", "name", "object": {…}, "notes": [], "observed_at"}}.
  • revisions returns a workload's rollout history, newest first and at most 20 (truncated when more exist). Deployments and Rollouts list their ReplicaSets; StatefulSets and DaemonSets list their ControllerRevisions. Each revision carries revision, name, created, replicas, template_hash, containers (name, image) and the redacted pod_template, so any two revisions can be diffed. pod_template is null on the oldest revisions once the reply's 512 KiB budget is spent, and a note says so. kind is Deployment, Rollout, StatefulSet or DaemonSet. Response: {"data": {"kind", "namespace", "name", "revisions": […], "truncated", "notes": [], "observed_at"}}.
  • logs returns one container's recent output as {"data": {"namespace", "pod", "container", "previous", "lines": [{"ts", "text"}], "truncated", "containers"?, "timestamps", "notes": [], "observed_at"}}.
    • Each line's apiserver timestamp is split into ts. A line without one gets "ts": null.

    • text is secret-sanitized before it leaves the backend. This is the same sanitizer the AI chat applies to tool output (chat.SanitizeSecrets). It carries the agent's own scrubber rule set (a port of agent/internal/filewatch, kept in step with it) and adds rules for log-line shapes:

      • key: value anywhere in a line, with the whole value redacted: to the end of the line, or to the closing quote when the value is quoted. key=value anywhere in a line, with a quoted value taken whole.
      • Keys are matched after normalization: camelCase and PascalCase boundaries, - and . all count as _, and case is ignored. So clientSecret, accessToken, DbPassword, x-api-key and DB_PASS are all recognized. The key words are password, passwd, pwd, passphrase, pass, secret, api key, access key, private key, credential(s), token, auth, authorization, bearer, session and cookie.
      • JSON keys in any case, including escaped JSON inside a string ({\"password\":\"…\"}, escaped docker auths), and Ruby/PHP 'password' => '…' pairs;
      • Authorization with any scheme, in every common dump shape: header, JSON, assignment, Go header map (map[Authorization:[Basic …]]), json.Marshal(http.Header) ({"Authorization":["Basic …"]}), Go %#v ([]string{"…"}) and their escaped forms. Also Cookie: and Set-Cookie: headers;
      • an auth scheme word and its credential (Bearer …, Basic …) as one unit wherever they appear, so api_key=Bearer x and --token Bearer x lose the credential, not just the word Bearer;
      • URL-encoded pairs (password%3D…, token%3A%20…);
      • the value of any other key=value treated as text of its own and sanitized again, up to two levels deep. A logfmt msg="login pass=…", a URL whose query carries ?password=…, and an encoded form body lose their nested secrets, while their other content stays;
      • secret command-line flags in both the = and the space form, compound flags included: --password=x, --token x, --db-password x, --client-secret x, --github-token x. A flag counts when its last -/_ segment is a secret word, so --password-file, --token-ttl and --sort-key keep their values. The agent's scrubber has the same rule;
      • PEM private-key blocks, and credentials in URLs of any scheme, including an empty user (rediss://:pw@host);
      • JWTs, Slack xox?- tokens, AWS AKIA… ids, Vault hvs. tokens and sk- keys;
      • a high-entropy backstop.

      The ambiguous keys (token, auth, pass, credential, session, cookie, bearer) keep obvious non-secret values: token: 5, auth: true and cookie: none survive, and mid-line the first word decides, so auth: failed for user bob survives too. Keys that only resemble secret keys survive as well (token_count, tokenCount, auth_method, password_policy, first_pass: true). A differential test holds the sanitizer to never reveal anything the pre-port version hid.

      It is a pattern sanitizer, not a guarantee: a secret in a shape none of the rules knows still passes through. These shapes are known and not handled:

      • the continuation lines of a YAML block scalar (password: | followed by the secret on the next line);
      • doubly escaped JSON ({\\\"password\\\":…});
      • a mid-line key = value with spaces around the = (x password = hunter2). A line-initial one is handled.

      Log lines are sanitized one at a time, so a multi-line secret is only partly covered. The PEM-block rule cannot fire across lines. Body lines of 20 characters or more are caught by the entropy backstop, but a short last line can leak a few bytes.

      The rules err toward hiding. A secret-named key mid-line takes the rest of the line, so some ordinary Kubernetes error text is redacted too. For example, failed to get secret: secrets "foo" not found becomes failed to get secret: [REDACTED]. This is accepted: on this page, hiding too much is the safer failure.

    • tail is 1–300 (default 100).

    • since_time (RFC 3339) asks only for lines at or after that time. This is the page's 5 s "follow" poll. A follow always reads the maximum of 300 lines. The apiserver honors since_time at second precision, so the caller de-duplicates the overlap.

    • truncated is true whenever the read hit its line cap. For a follow that means lines between since_time and the first line returned may be missing.

    • A pod with several containers, or a container name the pod does not have, answers 200 with no lines and containers listing the pod's containers.

Agent versions. manifest and revisions need cluster agent v0.8.0. logs needs v0.7.12, and since_time/timestamps need v0.8.0. An older agent would silently ignore those two fields rather than refuse them. A plain tail on a v0.7.12 agent is still served: it is unstamped, with "timestamps": false, ts: null and a note. A since_time read on it is refused. A gate failure is 409 agent_too_old with details.required_version (for example "v0.8.0") and details.agent_version, and the agent is never asked.

Errors.

StatuscodeWhen
400bad_requestUnknown query parameter; kind outside the allow-list (revisions accepts only its four kinds); namespace or container not a DNS-1123 label; name or pod not a DNS-1123 subdomain; tail outside 1–300; previous not true/false; since_time not RFC 3339
404not_foundUnknown cluster, or a cluster outside the caller's environments
403forbiddenThe route's permission is missing on the cluster's client and environment
409no_agentThe cluster has no cluster agent
409agent_offlineThe cluster's agent is not connected (the message names the last-seen time)
409agent_too_oldSee above
429rate_limitedPer-user budget spent: 30 log reads a minute, and 60 manifest and revisions reads a minute (shared between the two). The budget is counted per backend process, not across replicas, so with N backend replicas a user can make up to N times as many reads
502agent_errorThe agent answered with an error. Its message is passed through, secret-sanitized: not found, not permitted (the read-only tier is not bound, or the Helm chart predates ControllerRevision/Rollout access), Argo Rollouts not installed, object over 512 KiB
502agent_unreachableNo responder on the agent's command subject
503live_unavailableThe backend has no NATS connection
504agent_timeoutNo reply within 15 s

Permissions. manifest and revisions need assets:read, and logs needs logs:read, each on the cluster's client and environment. Each route is gated by RequireAnyPermissionInAnyScope(<permission>), and the handler then applies the same 404/403 check as the other cluster reads with its own permission. assets:read never opens logs.

Events Timeline (14 days)​

GET /api/v1/clusters/:clusterID/events/timeline[?namespace=…][&involved_kind=…][&involved_name=…][&workload=…&workload_kind=…][&type=…][&reason=…][&from=…][&to=…][&limit=…]

The cluster's Kubernetes events from the 14-day copy in VictoriaLogs. List Cluster Events reads Postgres, which keeps 24 hours. The cluster page's Events tab reads this route. The Kubernetes worker forwards every event it syncs to VictoriaLogs as source=k8s-event, in the cluster's client log account. This route reads it back from that same account, so it can only ever read that client's logs. The query is further confined to source=k8s-event, the cluster's ID and its environment, all taken from the stored cluster row.

Why those filters can be trusted. Host logs share the client's log account, and a tailed JSON line's keys become stored fields. Two rules stop a host log from posing as a Kubernetes event of another environment:

  • The fields the server sets on every stored line — _msg, _time, _stream, _stream_id, host_id, environment_id, source, level, unit — can never be set by a log's own keys. A colliding key is kept as field.<key> (vlclient.LogEntry), so {"source":"k8s-event","environment_id":"<another env>"} in an application's log is stored with the host's own source and environment.
  • The logs worker renames the Kubernetes event keys on host logs: k8s_cluster_*, k8s_event_* and k8s_involved_* become field.<key>, and a host log whose source is k8s-event is stored as host:k8s-event. k8s_namespace is left alone, because pod-log enrichment writes it legitimately. Only the Kubernetes worker writes source=k8s-event lines with those keys.

Lines already stored before these rules shipped are not rewritten. A host log written then could still match the timeline until it ages out (14 days).

  • No caller-written LogsQL. The query is assembled in one place (service.BuildK8sEventsTimelineLogsQL) from validated parameters only. Every value is an exact-match filter (field:="value"). The value is quoted with Go's strconv.Quote, whose escapes are exactly the ones LogsQL's string parser undoes. Validation is the first wall and the quoting is the second. A test against a real VictoriaLogs shows that a value that slipped past validation (x" OR *, x | delete, a newline) still matches only itself.

  • Filters (all optional, all exact):

    ParameterRule
    namespaceDNS-1123 label
    involved_kindOne of Pod, Deployment, ReplicaSet, StatefulSet, DaemonSet, Rollout, Job, CronJob, Node, Service, Endpoints, Ingress, HorizontalPodAutoscaler, PersistentVolumeClaim, PersistentVolume, Namespace, NetworkPolicy, PodDisruptionBudget
    involved_nameDNS-1123 subdomain
    typeNormal or Warning
    reason1–64 letters (^[A-Za-z]{1,64}$), e.g. BackOff
    workloadDNS-1123 subdomain; needs workload_kind and namespace; not with involved_name
    workload_kindDeployment, StatefulSet, DaemonSet or Rollout; needs workload

    A parameter sent empty, or any other parameter, is a 400. The involved-object filter is an exact match: involved_kind=Deployment&involved_name=api returns the Deployment's own events (such as ScalingReplicaSet), not its ReplicaSets' or pods'.

  • A workload and what it owns. namespace=shop&workload=api&workload_kind=Deployment returns the events of the Deployment itself, of its ReplicaSets and of its pods — where BackOff, OOMKilling and FailedScheduling land. This is what the workload page's Events tab reads. The store is asked for names equal to api or starting with api- (a quoted exact-prefix filter), and the backend then keeps only the names Kubernetes generates for that kind: api-<hash> ReplicaSets and api-<hash>-<suffix> pods for a Deployment or Rollout, api-<n> for a StatefulSet, api-<suffix> for a DaemonSet (the same rule as the what-changed timeline). So api never includes api-gateway's pods. One residual remains: a DaemonSet whose name suffix uses only pod-template-hash characters, such as api-db, has pods (api-db-x7k2p) that also fit a Deployment api's pod shape (api-v2 does not: v and 2 never occur in a hash). The store's limit is applied before that check, so a truncated workload page can hold fewer than limit events, and its note says older events may have been left out (the row past the limit may have been a sibling the check then dropped). Events of an HPA named like the workload (SuccessfulRescale, FailedGetResourceMetric) are not included: only the workload's own kind, its ReplicaSets and its pods are.

  • Window. from and to are RFC 3339 and both inclusive, to the nanosecond: the exact window is a _time:[from, to] filter in the query, because VictoriaLogs' own start/end request bounds are whole seconds with an exclusive end. An event at exactly to is returned. to defaults to now and from to 24 hours before to. from must be before to, and the window may not exceed 14 days. A window that starts more than 14 days ago gets a note saying nothing before the retention edge can be shown.

  • Limit. limit is 1–1000 (default 200). Events come newest first. truncated is true when more events exist in the window, and a note says so.

  • What a row is. Each inventory sync forwards every event the cluster currently holds. An unchanged event is therefore stored once per sync with identical fields, and those copies are collapsed (uniq by in VictoriaLogs, before the limit). Kubernetes folds repeats of an event into one object with a growing count, so a recurring event appears once per count change. Its time is the event's last-seen time at that sync. The long-term copy does not carry the event's first-seen time or UID.

  • Events written through events.k8s.io (kube-scheduler's FailedScheduling and Scheduled, among others) have no lastTimestamp. From cluster agent v0.8.0 their last-seen time is series.lastObservedTime, else eventTime, else firstTimestamp, and their count is series.count (1 for a single occurrence), so they collapse like any other event. An older agent reports them with no last-seen time; those events are not copied to the 14-day store (VictoriaLogs would stamp every sync's copy with its ingestion time, listing one event once per sync). They still appear in the 24-hour Postgres list, and the response carries a note saying so whenever the cluster's agent is older than v0.8.0 or has reported no version. Copies stored before this fix can still show such an event once per sync, with count: 0, until they age out.

  • Cost. The collapse holds every distinct sighting of the window in VictoriaLogs' memory. It is bounded by VictoriaLogs' own query limits (duration, memory) and by this route's 20 s deadline, not by limit. A 14-day unfiltered window on a busy cluster can therefore fail (504 or 502) rather than truncate; narrow the window or add a filter.

Response:

{"data": {
"events": [{"time": "2026-10-01T11:00:00Z", "namespace": "shop", "involved_kind": "Pod",
"involved_name": "api-6f7-x2k", "type": "Warning", "reason": "BackOff",
"message": "Back-off restarting failed container", "count": 4}],
"truncated": false, "from": "2026-09-30T12:00:00Z", "to": "2026-10-01T12:00:00Z", "notes": []
}}

events and notes are always arrays, never null. An empty events means VictoriaLogs answered and holds no matching events.

Messages are sanitized again. Each message is run through chat.SanitizeSecrets (the live-logs sanitizer, above) before it leaves the backend. The agent scrubs event messages before sending them, but a stored row was scrubbed only by whichever agent version wrote it, and an exec probe's failure message can quote a command line. The what-changed timeline's event items and the 24-hour List Cluster Events route apply the same sanitizer.

Errors.

StatuscodeWhen
400bad_requestUnknown or empty parameter; a filter outside its rule; from/to not RFC 3339; from not before to; window over 14 days; limit outside 1–1000
404not_foundUnknown cluster, or a cluster outside the caller's environments
403forbiddenassets:read missing on the cluster's client and environment
429rate_limited60 timeline reads a minute per user (counted per backend process), with the message "too many event timeline reads"
500internal_errorThe cluster's client log account could not be resolved
502events_backend_unavailableVictoriaLogs answered with an error or could not be reached. This is never turned into an empty list, which would read as "nothing happened"
504events_backend_timeoutVictoriaLogs did not answer within 20 s (under the API server's 30 s write timeout). The read is made once, with no retry, so a slow store is not asked the same heavy query again
503events_unavailableThe backend has no VictoriaLogs configured

A caller that disconnects before the read finishes gets no response; this is logged at debug level, not as a store failure.

Permissions. assets:read on the cluster's client and environment: RequireAnyPermissionInAnyScope(assets:read) on the route, then the same in-handler 404/403 check as the other cluster reads.

What Changed (timeline)​

GET /api/v1/clusters/:clusterID/changes[?namespace=…][&workload=…][&kind=…][&from=…][&to=…]

One list, newest first, of what changed in a cluster, in one namespace, or in one workload. It merges four sources:

kindSourceMatched byPermissionCap
changeRecorded k8s-config rows (workload change records and NetworkPolicy changes)details.cluster_id = this cluster, in its client and environment; then details.namespace, and for a workload details.name plus a workload kindchanges:read200
deployDeploy notifications (source=deployment: ArgoCD syncs, sync failures and other ArgoCD notification types)In the cluster's environment: details.app_name naming an ArgoCD Application mapped to this cluster that concerns the view; otherwise, for apps Console does not know, details.namespace (ArgoCD's destination namespace)changes:read100
eventKubernetes events from the 14-day copy with reason ScalingReplicaSet, Killing, BackOff, FailedScheduling, Evicted or OOMKillingThe cluster, its environment, the namespace, and for a workload the workload or its ReplicaSets and pods (below)assets:read200
alertAlert groups that started firing in the windowThe cluster's environment and the namespace label; for a workload also its own label or a pod label (below)alerts:read100

Each source keeps its newest items up to its cap. truncated.<source> is true when more exist, and a note says so.

Parameters. All are optional, and anything else (or a parameter sent empty or twice) is a 400.

ParameterRule
namespaceDNS-1123 label
workloadDNS-1123 subdomain. Needs namespace, because a workload name is only unique within its namespace
kindDeployment, StatefulSet, DaemonSet or Rollout. Needs workload. Without it, the workload may be any of the four
from, toRFC 3339, both inclusive. to defaults to now and from to 7 days before to. from must be before to, and the window may not exceed 30 days

Response. One merged items list rather than four arrays, because the History tab draws a single timeline. Each item has a kind, so a client can still split or filter.

{"data": {
"items": [
{"id": "event:4f0c1e9a2b7d6c35", "time": "2026-10-01T11:58:00Z", "kind": "event", "title": "BackOff · Pod api-7d9f8c5b4-x2x9k (×4)",
"detail": "Back-off restarting failed container", "type": "Warning", "reason": "BackOff",
"namespace": "shop", "object_kind": "Pod", "object_name": "api-7d9f8c5b4-x2x9k"},
{"id": "change:6a1…", "time": "2026-10-01T11:00:00Z", "kind": "change",
"title": "Deployment shop/api: image api 2.40.3 → 2.41.0", "detail": "Deployment shop/api",
"link": "/changes/6a1…", "severity": "info", "type": "workload_changed",
"namespace": "shop", "object_kind": "Deployment", "object_name": "api",
"changes": [{"field": "containers.api.image", "before": "api:2.40.3", "after": "api:2.41.0"}]},
{"id": "deploy:3c9…", "time": "2026-10-01T10:59:00Z", "kind": "deploy", "title": "ArgoCD sync: shop-api (Synced)",
"detail": "app shop-api · namespace shop · revision 4f1c2d", "link": "/changes/3c9…", "severity": "info", "type": "sync", "namespace": "shop",
"scope": "client"},
{"id": "alert:9b2…", "time": "2026-10-01T10:40:00Z", "kind": "alert", "title": "KubePodCrashLooping",
"detail": "api is crash-looping", "link": "/alerts/9b2…", "severity": "P2", "type": "firing", "namespace": "shop"}
],
"truncated": {"changes": false, "deploys": false, "events": false, "alerts": false},
"sources": {"changes": "ok", "deploys": "ok", "events": "ok", "alerts": "ok"},
"from": "2026-09-24T12:00:00Z", "to": "2026-10-01T12:00:00Z",
"notes": ["Deploys are ArgoCD notifications (source deployment): attributed by their ArgoCD application when Console knows where it deploys, …"]
}}
  • time is a change's observed time (scanned_at, else created_at), a deploy's finish time, an event's last-seen time, or an alert group's first-fired time.
  • id is stable across reads: change:/deploy:/alert: plus the row's UUID, or event: plus a 16-hex-digit hash of the event's content (time, namespace, object, type, reason, count, message). A refetch that brings newer events does not change older items' ids, so UI rows keep their state, and items at the same instant sort by kind, then id.
  • scope is client on a deploy from a webhook source with no environment: such a row is a fact about the whole client, so it is listed for every environment's cluster that has the namespace, and may belong to another environment's cluster. A note gives the count. A client-wide deploy whose app is an ArgoCD Application mapped to this cluster is attributed to it and carries no scope.
  • link is a Console path (/changes/{id} or /alerts/{id}); events have none. actor appears when the source names one (a deploy's provider login).
  • type is the change or deploy event type, the event's Normal/Warning, or the alert group's status.
  • changes carries a recorded change's field-level before/after (names and references only, never values), and changes_truncated is true when the recorder kept only the first 50.
  • items and notes are always arrays, never null. Optional item fields are omitted when empty.

Per-source permissions. The route gate is RequireAnyPermissionInAnyScope(changes:read, alerts:read, assets:read). The handler then loads the cluster and returns 404 unless HasAccessToEnvironment(cluster.client_id, cluster.environment_id). Each source is then checked on its own with the 3-arg Can(perm, client, environment):

  • A source the caller may not read is not queried. sources.<name> is forbidden, and a note names the missing permission. Changes and deploys share changes:read, the same as deploy markers on charts. "Forbidden" is reported separately from "empty" so the UI can say "you can't see alerts here" instead of drawing "nothing happened".
  • Holding none of the three on the cluster's environment is a 403. That includes holding them only on a sibling environment.
  • The store also gets the caller's ScopeFor(<perm>) as a second wall. Every query is pinned to the cluster's client and environment, taken from the stored cluster row.

When VictoriaLogs fails, the rest still answers. A failed event read (store down, 20 s timeout, no VictoriaLogs configured, log account not resolvable) sets sources.events to unavailable and adds a note saying events could not be read. The other sources are returned as usual with a 200. A database failure is a 500. A caller that disconnects gets no response.

Events. The read reuses the events timeline service: the same validated, quoted LogsQL builder, the same client log account, the same 20 s single attempt, and the same notes (retention, old agent). Two internal-only filters were added to it, both validated and quoted the same way as the route's filters: a reasons set (an OR of exact matches, at most 16) and a workload filter (exact name, or an exact <workload>- prefix). The store keeps events for 14 days, so for a longer window only the last 14 days are read, and one note says so ("only the last 14 days of this window show events"); the events reader's own retention note is dropped then, so the two never appear together.

Matching a workload.

  • Changes: the row's details.name equals the workload, and details.kind is the given kind (else any workload kind). A NetworkPolicy that shares the workload's name is not included.
  • Events: the event's object is the workload itself (of an allowed kind), one of its ReplicaSets (<workload>-<hash>, Deployment and Rollout only), or one of its pods. The pod name must have the shape that kind produces: <workload>-<hash>-<5> for Deployment and Rollout, <workload>-<ordinal> for StatefulSet, <workload>-<5> for DaemonSet. The 5-character suffix uses Kubernetes' vowel-free SafeEncodeString alphabet (bcdfghjklmnpqrstvwxz2456789). The <hash> is narrower still: it is the SafeEncoded decimal digits of a 32-bit hash, and SafeEncoding maps the ten digits onto only 4 5 6 7 8 9 b c d f. So a sibling named with a word — Deployment, DaemonSet or Job api-gateway, a Helm hook Job api-migrate — never matches workload api, and neither does DaemonSet api-v2. Residual collision: a sibling whose own name suffix uses only hash characters looks generated. DaemonSet api-db (pods api-db-x7k2p) is indistinguishable by name from a pod of Deployment api with hash db, so its pods' events and alerts appear in api's view. VictoriaLogs filters by prefix, and the backend drops the rows that do not match. That can leave fewer events than the cap while truncated is true.
  • Alerts: the alert's namespace label matches and either its workload label matches (deployment, statefulset, daemonset or rollout for the allowed kinds) or its pod label has the pod-name shape above (a PostgreSQL regular expression built from the validated name, with . escaped).
  • Deploys: by ArgoCD Application first. A deploy whose app_name is an Application mapped to this cluster is shown when the Application concerns the view (the whole cluster; a namespace — its destination namespace or any of its resources; a workload — a resource row naming it). A deploy whose app is an Application Console knows for the client but that maps elsewhere is not shown. A deploy whose app no agent reports falls back to namespace matching, so a workload view then shows every deploy to the workload's namespace (clusterDeployMatch).

Decisions and limits.

  • Alerts without a namespace label are never shown, even in the whole-cluster view. Alert labels are the only link between an alert and Kubernetes. A node or control-plane alert without a namespace cannot be told apart from any other host's alert in the same environment, so including those would fill the cluster timeline with alerts from unrelated VMs. They stay on the Alerts page, and a note says so.
  • Alerts are matched within the cluster's environment: the group's environment_id, or no environment and a host in it. Alert labels carry no reliable cluster identity, so if two clusters in one environment share a namespace name, both clusters' alerts for that namespace appear. When the environment holds more than one cluster, the response carries a note saying so (it does not say how many; a caller reading alerts need not hold assets:read). If counting the environment's clusters fails, the read still succeeds and carries a hedged form of that note ("whether this cluster's environment has other clusters could not be checked …"); the failure is logged. Alert groups that resolved to no environment and no host are not shown.
  • Deploys are ArgoCD notifications (source=deployment): those of the cluster's environment, those mapped to one of its hosts, and client-wide ones (a webhook source with no environment), the same rule as deploy markers on charts. Client-wide ones carry scope: "client" and a note, because one ArgoCD instance serving prod and staging through a single env-less source would otherwise attach prod's shop syncs to the staging cluster's shop without saying so. Deploys must carry details.namespace. GitHub and GitLab write source=github / source=gitlab, which this timeline does not read. A client-wide deploy attributed by its Application is this cluster's for certain and carries no scope. Two clusters in one environment that deploy to the same namespace name both show deploys of apps Console does not know; the shared-environment note applies here too. A hub in another environment: its environment-pinned deploy rows belong to the hub's environment, so they never appear on a destination cluster in a different environment, even when the Application maps to it; only its client-wide rows can (see ArgoCD Drift → What changed).
  • HPA-driven scaling is not a recorded change (see workload change records). It appears as ScalingReplicaSet events instead.
  • No new index. On a test database seeded with 500k change rows and 200k alert groups, the recorded-changes read is a BitmapAnd of idx_change_events_client_effective (client, effective time) and idx_change_events_source, with details.cluster_id a heap filter over one client's windowed rows (3.6 ms for 7 days, 8 ms for a 30-day workload view). Deploys are similar (5.6 ms). The alert read uses idx_alert_groups_status / idx_alert_groups_client on the time range and filters labels over the client's windowed rows (about 20 ms for 30 days). Cost is bounded by client × window.

Errors.

StatuscodeWhen
400bad_requestUnknown, repeated or empty parameter; a value outside its rule; workload without namespace; kind without workload; from/to not RFC 3339; from not before to; window over 30 days
404not_foundUnknown cluster, or a cluster outside the caller's environments
403forbiddenNone of changes:read, alerts:read, assets:read on the cluster's client and environment
429rate_limited60 reads a minute per user (counted per backend process), with the message "too many what-changed reads"
500internal_errorA database read failed

Client-Scoped Clusters​

GET /api/v1/clients/:id/clusters

Returns clusters belonging to the specified client.

Environment-Scoped Clusters​

GET /api/v1/environments/:envID/clusters

Returns clusters in the specified environment.

Node-Host Linkage​

The KubernetesWorker automatically links K8s nodes to existing Proxima hosts when possible. This enables cross-referencing between Kubernetes node data and host-level inventory, metrics, and compliance data.

How It Works​

For each node in a cluster inventory message, the worker attempts to find a matching host in the same environment by:

  1. Hostname match — Compares the node name against hosts.hostname
  2. IP match — Compares node addresses against hosts.ip_address

If a match is found, k8s_nodes.host_id is set to the matching host's ID. This linkage is tracked by the proxima_k8s_node_host_linkage_total Prometheus metric with a result label (linked or unlinked).

PQL Integration​

The cluster.* PQL prefix uses this linkage to find hosts that are K8s nodes:

cluster.name = "k8s-prod-01" AND host.os = "Ubuntu"

This query joins hosts -> k8s_nodes -> clusters, so only hosts that are linked as K8s nodes will match cluster.* filters.

Stale Cluster Detection​

The K8sMaintenanceWorker runs periodically and detects clusters whose collector has stopped reporting:

  • Threshold: If clusters.last_seen_at is older than 15 minutes (configurable), the cluster's agent_status is set to disconnected. The Clusters page then shows the Agent disconnected badge, and the cluster's Overview says Stale instead of giving a health verdict, because every inventory figure is frozen at the last sync.
  • Event cleanup: The maintenance worker also deletes stale K8s events from PostgreSQL (older than the event retention period, default: 24 hours).
  • Signal: the INFO log line marked stale clusters disconnected (with count). There is no metric for this today: observability.K8sStaleClustersMarkedTotal is declared but never registered, so the worker's nil-guard skips it and nothing reaches /metrics.

Cluster Status and Asset Status Mapping​

Cluster health status is mapped to the CMDB asset status when upserting the asset record:

Cluster StatusAsset Status
healthyactive
degradeddegraded
unreachableunreachable
unknownactive

CLI Seeder​

The scripts/k8s-seed.sh script publishes a test Kubernetes NATS message for development and E2E testing. If kubectl is available and connected to a cluster, it reads real cluster data. Otherwise, it uses hardcoded test data.

# Default: internal / development / test-cluster
./scripts/k8s-seed.sh

# Custom slugs
./scripts/k8s-seed.sh acme production k8s-prod

# Custom NATS server
NATS_URL=nats://remote:4222 ./scripts/k8s-seed.sh

Requirements: nats CLI, jq. Optionally kubectl for real cluster data.

Observability​

Prometheus Metrics​

MetricTypeDescription
proxima_k8s_clusters_synced_totalcounterTotal cluster inventory syncs processed
proxima_k8s_sync_duration_secondshistogramTime to process a kubernetes message
proxima_k8s_nodes_synced_totalcounterTotal node upserts across all syncs
proxima_k8s_workloads_synced_totalcounterTotal workload upserts across all syncs
proxima_k8s_node_host_linkage_totalcounterNode-host linkage attempts and results (label: result)

OTel Spans​

The KubernetesWorker creates spans for each processing step:

  • k8s.process_message — Top-level span for the entire message
  • k8s.upsert_cluster — Cluster and asset upsert
  • k8s.sync_nodes — Node sync with host linkage
  • k8s.sync_namespaces — Namespace sync
  • k8s.sync_workloads — Workload sync
  • worker.KubernetesPruneUnseen — Prune phase

Permissions​

Cluster endpoints reuse the existing CMDB permissions:

PermissionGranted ToDescription
assets:readAll system rolesView clusters and sub-resources, including the cluster summary and the 14-day events timeline
metrics:readRoles with metrics accessThe cluster metrics endpoint (Overview charts)
logs:readRoles with log accessThe live container logs endpoint (/live/logs)
changes:read, alerts:readRoles with change / alert accessThe recorded-changes and deploys sources (changes:read) and the alerts source (alerts:read) of the what-changed timeline; assets:read opens its events source
assets:writeSuper Admin, AdminFuture: create and modify clusters

Environment-scoped grants​

Every cluster read route accepts an environment-scoped grant. The routes use RequireAnyPermissionInAnyScope, not RequirePermission: RequirePermission sees only client-wide grants when the request has no client_id, so it would refuse an environment-scoped caller before the handler runs. That gate proves only that the caller holds the permission somewhere. The handlers do the tenant check:

  • GET /api/v1/clusters, GET /api/v1/environments/{envID}/clusters and GET /api/v1/clients/{id}/clusters filter the store query by ac.ScopeFor(assets:read). An environment-scoped grant contributes its environment only, never its whole client. The client route also checks assets:read on the path client and narrows the scope to it.
  • GET /api/v1/clusters/{clusterID} and every /{clusterID}/... read load the cluster, return 404 unless HasAccessToEnvironment(cluster.client_id, cluster.environment_id), and return 403 unless Can(assets:read, client, environment). Child reads query by that cluster's ID. The node and workload pod lists filter on the cluster ID as well as the node or workload ID, so an ID from another cluster returns no rows.

The metrics endpoint follows the same pattern on metrics:read. The what-changed timeline gates its route on any of changes:read, alerts:read, assets:read and checks each of its sources separately: a missing permission empties that source, and only holding none of the three is a 403.

The cluster.* prefix is available in PQL for searching hosts by their associated cluster:

FieldTypeDescriptionExample
cluster.namestringCluster namecluster.name = "k8s-prod-01"
cluster.versionstringK8s versioncluster.version ~ "v1.29*"
cluster.platformstringDistributioncluster.platform = "rke2"
cluster.statusstringHealth statuscluster.status = "healthy"

The cluster.* prefix uses a two-table JOIN chain (k8s_nodes -> clusters), so only hosts that are linked as K8s nodes will match these filters.

See the PQL Reference for full syntax documentation.