Kubernetes: Operate (read-only)
The workload page and the cluster tabs answer the questions an operator asks during an incident without leaving Console and without kubectl: what exactly is deployed, what changed, does it match Git, what is the container saying, and what did Kubernetes report? Every view on this page is read-only. Nothing here can change a cluster.
This page is the map. Each feature is documented in full where it lives; the links below go there.
The views
| View | Where | What it shows | Details |
|---|---|---|---|
| Live manifest | Workload page → YAML | The workload's object as YAML, read from the cluster now and redacted in the agent | YAML tab · Live reads API |
| Revisions | Workload page → YAML | The newest 20 revisions (ReplicaSets or ControllerRevisions), and a diff of any two pod templates | YAML tab · Live reads API |
| ArgoCD drift | Workload page → YAML (card); cluster Overview → Needs attention | Synced / OutOfSync, health, synced → target revision, read in-cluster with no ArgoCD credentials | ArgoCD Drift |
| Live logs | Workload page → Logs | One container's output, live from the pod, followed every 5 s, never stored | Logs tab · Live reads API |
| Events timeline | Cluster → Events; workload page → Events | Kubernetes events from the last 14 days, filtered; a workload's view includes its ReplicaSets and pods | Cluster Events tab · Events timeline API |
| What changed? | Workload page → History | Recorded spec changes, ArgoCD deploys, Kubernetes events and alerts in one list, up to 30 days | History tab · What changed API |
| Workload change records | Changes page, change feed, History tab, L1 triage | One k8s-config change row per workload spec change (image, resources, env var names, config/secret refs, strategy, manual scaling) | Workload spec facts · Workload change records |
Routes, permissions and limits
Every route sits under /api/v1/clusters/{clusterID}/. Each is admitted by RequireAnyPermissionInAnyScope and then checked in the handler on the cluster's own client and environment: 404 when you cannot see the cluster's environment (so the response never confirms the cluster exists), 403 when you can see it but lack the route's permission there.
| Route | Permission | Window / cap | Rate limit (per user, per backend process) |
|---|---|---|---|
GET live/manifest | assets:read | One object, ≤ 512 KiB after redaction | 60 / min, shared with live/revisions |
GET live/revisions | assets:read | Newest 20 revisions, 512 KiB reply budget | (shared with live/manifest) |
GET live/logs | logs:read (assets:read never opens logs) | tail 1–300 lines, 256 KiB | 30 / min |
GET events/timeline | assets:read | Window ≤ 14 days (default 24 h); limit 1–1000 (default 200) | 60 / min |
GET changes | Any of changes:read, alerts:read, assets:read; each source then needs its own | Window ≤ 30 days (default 7 d); caps 200 changes, 100 deploys, 200 events, 100 alerts | 60 / min |
GET argocd/applications | assets:read | 1000 Applications per reporting cluster | 60 / min |
GET workloads/{workloadID}/drift | assets:read | — | 60 / min (its own budget, not shared with argocd/applications) |
The rate limits are counted per backend process, not across replicas: with N backend replicas a user can make up to N times as many reads. They protect the apiserver and VictoriaLogs from a runaway page; the per-request permission checks are what protect the data. A spent budget is 429 rate_limited.
The live reads (live/*) go to the cluster's own agent at the moment you ask. Without a usable agent they answer 409: no_agent, agent_offline (with when it was last seen) or agent_too_old (with details.required_version). The UI turns each of these, and every other failure, into a specific sentence — never an empty panel that reads as "nothing here".
What is never sent or stored
| Data | Rule | Where it is enforced |
|---|---|---|
| Env var values | Never leave the cluster. A live manifest or revision shows each literal env value as <redacted: N chars>; valueFrom references stay because they name a source without revealing it. Spec facts and change records carry env var names only. | Cluster agent (manifest redaction, spec facts) |
| ConfigMap and Secret data | Never read. The live manifest refuses the ConfigMap and Secret kinds outright, and the collector's ClusterRole has no access to either. A workload's references to them are the names written in its own spec. | Cluster agent and its ClusterRole |
| Other value-bearing fields of a manifest | managedFields, embedded last-applied copies (kubectl.kubernetes.io/last-applied-configuration, kapp.k14s.io/original, objectset.rio.cattle.io/applied), probe and lifecycle header values, termination messages and value-bearing annotations are removed or redacted | Cluster agent |
| ArgoCD values | Helm values / valuesObject, parameters and fileParameters, plugin env, kustomize patches and jsonnet variables are never read. The agent decodes allow-listed fields one by one. | Cluster agent (ArgoCD Drift) |
| Repository credentials | The repository URL and destination server lose userinfo, query and fragment. The scrub fails closed: a value whose credentials cannot be told apart from the rest is dropped, not guessed. One documented residual shape remains. The backend runs the same function (api/repourl) again on ingest, so one faulty agent build cannot store a credential. | Cluster agent, then the backend on ingest |
| Kubernetes event messages | Scrubbed by the cluster agent before they are sent, and scrubbed again by the backend (chat.SanitizeSecrets) every time they are read — on the events timeline, in "What changed?" and in the 24-hour events list — because rows already stored were scrubbed only by whichever agent version wrote them. Pattern-based, like the log sanitizer. | Cluster agent, then the backend on read |
| Container logs | Live only. Read from the pod when you look, kept in the tab's memory, and gone when you leave. Never written to Console's database, to VictoriaLogs, or to browser storage. Every line is secret-sanitized by the backend before it reaches the browser. | Backend (sanitizer), frontend (no storage) |
What is not redacted. Container args and command, probe and lifecycle exec commands, and CSI volumeAttributes are shown as written, because operators need them. A secret placed there is visible. Log sanitizing is pattern-based, so a secret in a shape the rules do not know can still appear. See Kubernetes access → Security for the full rule set, and the live reads section for the known gaps of the log sanitizer.
pc kube users see more. The read tier the live reads run under (proxima:kube-readonly) gains get/list on ControllerRevisions and Rollouts. It is shared with pc kube users who hold kubernetes:read, and through kubectl they read those objects unredacted, including historical pod templates with literal env values. Console's own reads redact; kubectl does not.
Logs and events in VictoriaLogs. The 14-day events timeline reads the source=k8s-event copy the backend writes. Host logs share that VictoriaLogs account, so a host log line can no longer set any server-owned field (source, host_id, environment_id, …); a colliding key is stored as field.<key>. See Logs → Reserved Fields.
What is not shown
- HPA-driven scaling is not recorded as a change, so History has no HPA scaling rows; the
ScalingReplicaSetevents and the Replicas chart still show it. Manual scaling is recorded asworkload_scaled. - Workload creations and deletions are not recorded as changes.
- Init container logs are not offered: the agent reads only a pod's regular containers.
- ArgoCD resource rows are reported only for the workload kinds (Deployment, StatefulSet, DaemonSet, Rollout); other kinds count only towards each Application's totals.
- Alerts without a
namespacelabel never appear in History, even for the whole cluster. - GitHub and GitLab changes are not part of History; its deploys are ArgoCD notifications.
- ArgoCD deploys from a hub in another environment do not appear on the destination cluster's History. Deploy rows are read from the cluster's own environment (plus client-wide rows), so a notification recorded in the hub's environment is not read for a destination in a different one, even when the Application maps to this cluster. The drift card still shows that Application, labelled as reported from another environment.
- First seen and the reporting component of an event are not in the 14-day copy.
Rollout notes for operators
Each piece comes from a different component, so each needs its own upgrade. Deploy them in this order: the backend is the consumer of the new inventory fields and must understand them before agents send them.
-
Database migrations.
k8s_workloads_spec_facts(addsk8s_workloads.spec_factsandspec_facts_at) andargocd_applications(addsargocd_reportsandargocd_applications). Their numbers are assigned at merge time (the next free numbers onmainwhen the branch merges), so check the migration file names in the release rather than relying on a number here. They apply with the backend's normalmigrate up. -
Backend and frontend from the same release. The backend decodes spec facts and ArgoCD reports from any agent version; an older agent simply sends neither.
-
Cluster agent image, v0.8.0 or later. The backend gates each feature on the version the cluster agent reports. v0.8.0 is the minimum for each of these:
Feature Without v0.8.0 Live manifest, revisions 409 agent_too_old; the YAML tab names the version it needsLog follow ( since_time, timestamped lines)A plain tail is still served from v0.7.12, unstamped; the Logs tab shows "Limited follow" ArgoCD drift report.state = agent_too_old; drift is unknown, never "not managed"Workload spec facts and change records No facts are sent, so no changes are recorded (no gate needed: absent means unknown) Events without lastTimestamp(FailedScheduling,Scheduled) in the 14-day timelineNot copied to the 14-day store; the timeline adds a note saying so The release that carries these changes must therefore be tagged v0.8.0 or later; an agent reporting a lower version is gated.
-
helm upgradetheproxima-agentchart. Upgrading only the image does not add RBAC rules. The chart adds:- to the cluster agent's own ClusterRole:
listonargoproj.ioapplications— without it ArgoCD reportsforbidden(logged once) and drift is unknown; - to the
proxima:kube-readonlyread tier:get/listonappscontrollerrevisionsandargoproj.iorollouts— without them the StatefulSet/DaemonSet revisions and a Rollout's manifest and revisions answer 502agent_error("not permitted").
The live reads run under that read tier, which the chart installs only with
kubeAccess.enabled=true; without it every live read answers "not permitted". The same setting enablespc kube, so the one cluster, one client rule applies. Clusters installed from the raw manifests ininfra/k8s/cluster-agent/get theapplicationsrule from the updatedclusterrole.yaml. - to the cluster agent's own ClusterRole:
-
Host and node agent image (
agent). Its file and process scrubber was brought to parity with the backend's sanitizer. No page on this list depends on it, but upgrading it closes the same secret shapes on the host paths. -
One-off check for forged
k8s-eventrows, after the deploy. Before this release a JSON line in any tailed host log could be stored as a Kubernetes event of another environment's cluster in the same client's VictoriaLogs account. The fix stops new ones; rows written before it are not rewritten and age out with the 14-day retention. To find any that exist, run in each client's log account:source:="k8s-event" _time:14d | stats by (host_id) count() rowsThe Kubernetes worker writes these rows with
host_idset to the cluster's ID. Anyhost_idthat is not a row inclusters(typically a host's ID) is a forged or spoofed row.That query alone misses rows that also faked
host_id. Before this release a forged line could sethost_idtoo, for example to a real cluster's ID, and then passes the check above. Two more queries cover that case:source:="k8s-event" _time:14d | stats by (host_id, k8s_cluster_id, environment_id) count() rowsOn a genuine row
host_idandk8s_cluster_idare the same cluster ID, andenvironment_idis that cluster's environment. Any group where they differ, or where the environment is not the cluster's, is forged.source:="k8s-event" _time:14d | field_namesThe worker writes exactly these fields:
_msg,_time,host_id,environment_id,source,level,k8s_namespace,k8s_event_type,k8s_event_reason,k8s_involved_kind,k8s_involved_name,k8s_cluster_id,k8s_cluster_slugandk8s_event_count(plus VictoriaLogs' own_streamand_stream_id). Any other field name means at least one row was written by something else; querysource:="k8s-event" _time:14d <that field>:*to list those rows.The residual: a forged row that set
host_id,k8s_cluster_idandenvironment_idconsistently to a real cluster in its real environment, and carried no field the worker does not write, cannot be told apart from a genuine row by these queries. Record what you find; such rows show on the events timeline and in History until they age out. -
Expect one-off cron change pairs after host agents upgrade. The host agent's scrubber now also redacts compound secret flags (
--db-password value). Cron changes are keyed on the scrubbed command line, so a host whose crontab carries such a flag records onecron_removed+cron_addedpair on the first inventory after the upgrade. Nothing changed on the host.
See also
- Kubernetes Clusters — the pages themselves, tab by tab
- Kubernetes Inventory — the routes, the change recorder and the sync
- ArgoCD Drift — collection, destination mapping, trust model
- Kubernetes access → Security — the read tier and manifest redaction
- Cluster Agent — deployment, RBAC and the NATS message limit
- Logs (VictoriaLogs) — reserved fields