Skip to main content

Kubernetes: Operate (read-only)

The workload page and the cluster tabs answer the questions an operator asks during an incident without leaving Console and without kubectl: what exactly is deployed, what changed, does it match Git, what is the container saying, and what did Kubernetes report? Every view on this page is read-only. Nothing here can change a cluster.

This page is the map. Each feature is documented in full where it lives; the links below go there.

The views​

ViewWhereWhat it showsDetails
Live manifestWorkload page → YAMLThe workload's object as YAML, read from the cluster now and redacted in the agentYAML tab · Live reads API
RevisionsWorkload page → YAMLThe newest 20 revisions (ReplicaSets or ControllerRevisions), and a diff of any two pod templatesYAML tab · Live reads API
ArgoCD driftWorkload page → YAML (card); cluster Overview → Needs attentionSynced / OutOfSync, health, synced → target revision, read in-cluster with no ArgoCD credentialsArgoCD Drift
Live logsWorkload page → LogsOne container's output, live from the pod, followed every 5 s, never storedLogs tab · Live reads API
Events timelineCluster → Events; workload page → EventsKubernetes events from the last 14 days, filtered; a workload's view includes its ReplicaSets and podsCluster Events tab · Events timeline API
What changed?Workload page → HistoryRecorded spec changes, ArgoCD deploys, Kubernetes events and alerts in one list, up to 30 daysHistory tab · What changed API
Workload change recordsChanges page, change feed, History tab, L1 triageOne k8s-config change row per workload spec change (image, resources, env var names, config/secret refs, strategy, manual scaling)Workload spec facts · Workload change records

Routes, permissions and limits​

Every route sits under /api/v1/clusters/{clusterID}/. Each is admitted by RequireAnyPermissionInAnyScope and then checked in the handler on the cluster's own client and environment: 404 when you cannot see the cluster's environment (so the response never confirms the cluster exists), 403 when you can see it but lack the route's permission there.

RoutePermissionWindow / capRate limit (per user, per backend process)
GET live/manifestassets:readOne object, ≤ 512 KiB after redaction60 / min, shared with live/revisions
GET live/revisionsassets:readNewest 20 revisions, 512 KiB reply budget(shared with live/manifest)
GET live/logslogs:read (assets:read never opens logs)tail 1–300 lines, 256 KiB30 / min
GET events/timelineassets:readWindow ≤ 14 days (default 24 h); limit 1–1000 (default 200)60 / min
GET changesAny of changes:read, alerts:read, assets:read; each source then needs its ownWindow ≤ 30 days (default 7 d); caps 200 changes, 100 deploys, 200 events, 100 alerts60 / min
GET argocd/applicationsassets:read1000 Applications per reporting cluster60 / min
GET workloads/{workloadID}/driftassets:read—60 / min (its own budget, not shared with argocd/applications)

The rate limits are counted per backend process, not across replicas: with N backend replicas a user can make up to N times as many reads. They protect the apiserver and VictoriaLogs from a runaway page; the per-request permission checks are what protect the data. A spent budget is 429 rate_limited.

The live reads (live/*) go to the cluster's own agent at the moment you ask. Without a usable agent they answer 409: no_agent, agent_offline (with when it was last seen) or agent_too_old (with details.required_version). The UI turns each of these, and every other failure, into a specific sentence — never an empty panel that reads as "nothing here".

What is never sent or stored​

DataRuleWhere it is enforced
Env var valuesNever leave the cluster. A live manifest or revision shows each literal env value as <redacted: N chars>; valueFrom references stay because they name a source without revealing it. Spec facts and change records carry env var names only.Cluster agent (manifest redaction, spec facts)
ConfigMap and Secret dataNever read. The live manifest refuses the ConfigMap and Secret kinds outright, and the collector's ClusterRole has no access to either. A workload's references to them are the names written in its own spec.Cluster agent and its ClusterRole
Other value-bearing fields of a manifestmanagedFields, embedded last-applied copies (kubectl.kubernetes.io/last-applied-configuration, kapp.k14s.io/original, objectset.rio.cattle.io/applied), probe and lifecycle header values, termination messages and value-bearing annotations are removed or redactedCluster agent
ArgoCD valuesHelm values / valuesObject, parameters and fileParameters, plugin env, kustomize patches and jsonnet variables are never read. The agent decodes allow-listed fields one by one.Cluster agent (ArgoCD Drift)
Repository credentialsThe repository URL and destination server lose userinfo, query and fragment. The scrub fails closed: a value whose credentials cannot be told apart from the rest is dropped, not guessed. One documented residual shape remains. The backend runs the same function (api/repourl) again on ingest, so one faulty agent build cannot store a credential.Cluster agent, then the backend on ingest
Kubernetes event messagesScrubbed by the cluster agent before they are sent, and scrubbed again by the backend (chat.SanitizeSecrets) every time they are read — on the events timeline, in "What changed?" and in the 24-hour events list — because rows already stored were scrubbed only by whichever agent version wrote them. Pattern-based, like the log sanitizer.Cluster agent, then the backend on read
Container logsLive only. Read from the pod when you look, kept in the tab's memory, and gone when you leave. Never written to Console's database, to VictoriaLogs, or to browser storage. Every line is secret-sanitized by the backend before it reaches the browser.Backend (sanitizer), frontend (no storage)

What is not redacted. Container args and command, probe and lifecycle exec commands, and CSI volumeAttributes are shown as written, because operators need them. A secret placed there is visible. Log sanitizing is pattern-based, so a secret in a shape the rules do not know can still appear. See Kubernetes access → Security for the full rule set, and the live reads section for the known gaps of the log sanitizer.

pc kube users see more. The read tier the live reads run under (proxima:kube-readonly) gains get/list on ControllerRevisions and Rollouts. It is shared with pc kube users who hold kubernetes:read, and through kubectl they read those objects unredacted, including historical pod templates with literal env values. Console's own reads redact; kubectl does not.

Logs and events in VictoriaLogs. The 14-day events timeline reads the source=k8s-event copy the backend writes. Host logs share that VictoriaLogs account, so a host log line can no longer set any server-owned field (source, host_id, environment_id, …); a colliding key is stored as field.<key>. See Logs → Reserved Fields.

What is not shown​

  • HPA-driven scaling is not recorded as a change, so History has no HPA scaling rows; the ScalingReplicaSet events and the Replicas chart still show it. Manual scaling is recorded as workload_scaled.
  • Workload creations and deletions are not recorded as changes.
  • Init container logs are not offered: the agent reads only a pod's regular containers.
  • ArgoCD resource rows are reported only for the workload kinds (Deployment, StatefulSet, DaemonSet, Rollout); other kinds count only towards each Application's totals.
  • Alerts without a namespace label never appear in History, even for the whole cluster.
  • GitHub and GitLab changes are not part of History; its deploys are ArgoCD notifications.
  • ArgoCD deploys from a hub in another environment do not appear on the destination cluster's History. Deploy rows are read from the cluster's own environment (plus client-wide rows), so a notification recorded in the hub's environment is not read for a destination in a different one, even when the Application maps to this cluster. The drift card still shows that Application, labelled as reported from another environment.
  • First seen and the reporting component of an event are not in the 14-day copy.

Rollout notes for operators​

Each piece comes from a different component, so each needs its own upgrade. Deploy them in this order: the backend is the consumer of the new inventory fields and must understand them before agents send them.

  1. Database migrations. k8s_workloads_spec_facts (adds k8s_workloads.spec_facts and spec_facts_at) and argocd_applications (adds argocd_reports and argocd_applications). Their numbers are assigned at merge time (the next free numbers on main when the branch merges), so check the migration file names in the release rather than relying on a number here. They apply with the backend's normal migrate up.

  2. Backend and frontend from the same release. The backend decodes spec facts and ArgoCD reports from any agent version; an older agent simply sends neither.

  3. Cluster agent image, v0.8.0 or later. The backend gates each feature on the version the cluster agent reports. v0.8.0 is the minimum for each of these:

    FeatureWithout v0.8.0
    Live manifest, revisions409 agent_too_old; the YAML tab names the version it needs
    Log follow (since_time, timestamped lines)A plain tail is still served from v0.7.12, unstamped; the Logs tab shows "Limited follow"
    ArgoCD driftreport.state = agent_too_old; drift is unknown, never "not managed"
    Workload spec facts and change recordsNo facts are sent, so no changes are recorded (no gate needed: absent means unknown)
    Events without lastTimestamp (FailedScheduling, Scheduled) in the 14-day timelineNot copied to the 14-day store; the timeline adds a note saying so

    The release that carries these changes must therefore be tagged v0.8.0 or later; an agent reporting a lower version is gated.

  4. helm upgrade the proxima-agent chart. Upgrading only the image does not add RBAC rules. The chart adds:

    • to the cluster agent's own ClusterRole: list on argoproj.io applications — without it ArgoCD reports forbidden (logged once) and drift is unknown;
    • to the proxima:kube-readonly read tier: get/list on apps controllerrevisions and argoproj.io rollouts — without them the StatefulSet/DaemonSet revisions and a Rollout's manifest and revisions answer 502 agent_error ("not permitted").

    The live reads run under that read tier, which the chart installs only with kubeAccess.enabled=true; without it every live read answers "not permitted". The same setting enables pc kube, so the one cluster, one client rule applies. Clusters installed from the raw manifests in infra/k8s/cluster-agent/ get the applications rule from the updated clusterrole.yaml.

  5. Host and node agent image (agent). Its file and process scrubber was brought to parity with the backend's sanitizer. No page on this list depends on it, but upgrading it closes the same secret shapes on the host paths.

  6. One-off check for forged k8s-event rows, after the deploy. Before this release a JSON line in any tailed host log could be stored as a Kubernetes event of another environment's cluster in the same client's VictoriaLogs account. The fix stops new ones; rows written before it are not rewritten and age out with the 14-day retention. To find any that exist, run in each client's log account:

    source:="k8s-event" _time:14d | stats by (host_id) count() rows

    The Kubernetes worker writes these rows with host_id set to the cluster's ID. Any host_id that is not a row in clusters (typically a host's ID) is a forged or spoofed row.

    That query alone misses rows that also faked host_id. Before this release a forged line could set host_id too, for example to a real cluster's ID, and then passes the check above. Two more queries cover that case:

    source:="k8s-event" _time:14d | stats by (host_id, k8s_cluster_id, environment_id) count() rows

    On a genuine row host_id and k8s_cluster_id are the same cluster ID, and environment_id is that cluster's environment. Any group where they differ, or where the environment is not the cluster's, is forged.

    source:="k8s-event" _time:14d | field_names

    The worker writes exactly these fields: _msg, _time, host_id, environment_id, source, level, k8s_namespace, k8s_event_type, k8s_event_reason, k8s_involved_kind, k8s_involved_name, k8s_cluster_id, k8s_cluster_slug and k8s_event_count (plus VictoriaLogs' own _stream and _stream_id). Any other field name means at least one row was written by something else; query source:="k8s-event" _time:14d <that field>:* to list those rows.

    The residual: a forged row that set host_id, k8s_cluster_id and environment_id consistently to a real cluster in its real environment, and carried no field the worker does not write, cannot be told apart from a genuine row by these queries. Record what you find; such rows show on the events timeline and in History until they age out.

  7. Expect one-off cron change pairs after host agents upgrade. The host agent's scrubber now also redacts compound secret flags (--db-password value). Cron changes are keyed on the scrubbed command line, so a host whose crontab carries such a flag records one cron_removed + cron_added pair on the first inventory after the upgrade. Nothing changed on the host.

See also​