Kubernetes Clusters: Overview, Nodes, Namespaces & Workloads
The Clusters page (/clusters) lists every Kubernetes cluster you can see, with fleet totals and per-cluster capacity. A cluster's Overview tab (/clusters/:clusterId) answers "is this cluster healthy, and if not, what needs attention?" before you look at any resource list. The Nodes tab lists every node's capacity — used, requested and limits against allocatable — as a table, the Namespaces tab compares what each namespace uses with what it requested, and each workload has its own page (/clusters/:clusterId/workloads/:workloadId).
These pages are built from two data sources that refresh at different rates. Knowing which number comes from which source explains most of what you see.
Where the numbers come from
| Source | What it feeds | How fresh |
|---|---|---|
| Inventory (cluster agent → PostgreSQL) | Node, pod, workload and event counts; requests and limits; "Needs attention"; capacity bars; the list's tiles and row cells | One sync every 5 minutes (PROXIMA_K8S_SYNC_INTERVAL). The backend reads the same key and the two must match: the agent gets its value from Helm (clusterAgent.syncInterval), the backend from its own environment, nothing reconciles them, and the cluster-onboarding gates quote the backend's copy. Raise it for the agent alone and the figures on this page, and that gate's wording, both state a cadence that is not the one in use. |
| metrics-server usage (read by the cluster agent, stored on each node row) | "CPU used" and "Memory used" tiles, the "used" layer of the capacity bars | Same 5-minute inventory sync |
| Kubelet metrics (node agent → VictoriaMetrics) | CPU by node, Memory by namespace, Container restarts, the Restarts tile; the Namespaces tab's "used" columns and CPU trend; the workload page's per-pod charts, Restarts, CPU and Memory tiles | Every 60 seconds with the Helm chart's default (nodeAgent.metricsInterval; the agent's own default without the chart is 30 s) |
Replica history (backend → VictoriaMetrics, k8s_workload_replicas) | The workload page's Replicas chart | Written by the backend at each inventory sync (every 5 minutes), from the replica counts it just stored |
The inventory figures come from GET /api/v1/clusters/{id}/summary and GET /api/v1/clusters?include=capacity. Both read PostgreSQL only. The charts come from GET /api/v1/clusters/{id}/metrics/{query}. See Kubernetes Inventory → API Endpoints for the request and response formats.
Kubelet metrics are joined to a cluster through the node's linked host (Node-Host Linkage). The backend adds a cluster_id label to every metric from a host that is linked to a node. A node that is not linked to a host is still counted from the inventory, but its metrics carry no cluster_id and cannot appear in the cluster's charts. The Overview header says so: "metrics from N of M nodes", where N is the number of nodes that are linked to a host. A newly linked node starts appearing in the charts up to 5 minutes later, because the backend caches the lookup.
Cluster Overview
The Overview opens on a 6-hour window. The time picker in its header changes the window for the charts and the Restarts tile. The inventory figures always show the latest sync, whatever the window.
Header
The header gives a verdict and the reason for it. It checks the following in order and stops at the first match:
| Header | When |
|---|---|
| Stale — collector disconnected, last sync … | The cluster agent has not reported for 15 minutes (agent_status = disconnected). Every inventory figure on the page is a frozen picture from the last sync, so the page never calls the cluster Healthy in this state. |
| API unreachable · last seen … | The cluster agent reported that it could not reach the Kubernetes API server. |
| Status unknown | The cluster agent reported no API status. |
| Degraded — … | The API is reachable but something is wrong. The reasons are listed: "N nodes not ready", "N pods Failed", "N containers failing" (containers waiting on CrashLoopBackOff, ImagePullBackOff, ErrImagePull or CreateContainerConfigError). |
| Healthy | The API is reachable, every node is Ready, no pod is Failed and no container is failing. |
The meta line below the verdict shows the client, the environment, the last sync, "metrics from N of M nodes" and "N warning events in 1h".
The badge next to the cluster name uses a narrower vocabulary. It says only what the cluster agent can observe:
| Badge | Meaning |
|---|---|
| API reachable | The cluster agent's last call to the Kubernetes API server succeeded. It is not a health verdict: a cluster whose API answers can still have nodes down. The Overview header is the verdict. |
| API unreachable | The cluster agent could not reach the API server. |
| Agent disconnected | The cluster agent has stopped reporting. Whatever status it last sent is out of date. This badge wins over the other three. |
| Status unknown | The cluster agent sent no status. |
Earlier versions showed the cluster agent's status as Healthy in the page header and on the list. That word only ever meant that the API server answered. The badge now says API reachable, and "Healthy" appears only in the Overview header, where it is derived from the inventory.
Tiles
| Tile | Value | Note |
|---|---|---|
| Nodes ready | Ready nodes / all nodes. A node counts as ready only when its status is Ready. | "all nodes ready", or the names of the not-ready nodes (the tile turns critical with the word "Not ready") |
| Pods | Running pods | "3 CrashLoop · 2 Pending · 1 Failed". CrashLoop counts containers waiting in CrashLoopBackOff; Pending and Failed count pods. With none of those it shows "of N total". |
| CPU used | Used CPU as a % of the allocatable of the nodes that report usage | "12.3 of 72.0 cores". When only some nodes report usage, the percent is taken over those nodes' allocatable and the note adds "usage from N of M nodes" (for example "6.0 of 8.0 cores · usage from 2 of 10 nodes"), so the nodes that send nothing do not dilute the figure. With no metrics-server data it shows "—" and "no metrics-server data", never 0%. |
| Memory used | Used memory as a % of the allocatable of the nodes that report usage | Same as CPU, in GiB |
| Restarts / 6h | Container restarts in the selected window, with a sparkline | See Restarts |
Needs attention
This panel lists what needs a person, worst first, and shows up to 8 rows followed by "+N more". When nothing matches it says "Nothing needs attention". Each row carries an icon and the word Critical or Warning. Where it can, a row links to the resource and opens its drawer.
| Order | Row | Severity | Rule |
|---|---|---|---|
| 1 | Node X is not ready | Critical | Node status is not Ready. The summary names up to 10 nodes; beyond that a "N more nodes not ready" row gives the exact count. |
| 2 | N containers in reason | Critical | Containers currently waiting on CrashLoopBackOff, ImagePullBackOff, ErrImagePull or CreateContainerConfigError. The CrashLoopBackOff row also names the pod with the most restarts. |
| 3 | Pod X Pending > 5 min | Warning | Phase Pending and created more than 5 minutes ago, oldest first, with its age. |
| 4 | Kind X ready/desired ready | Warning | A workload with fewer ready replicas than desired, largest shortfall first. |
| 5 | N pods in Failed phase | Warning | Pods left in the Failed phase. Evicted and failed Job pods stay until they are garbage-collected. |
| 6 | CPU / Memory requests at N% of allocatable | Warning | Requests at 85% of allocatable or more, with the room left for more requests. |
Only pending pods have an age. For the other rows the age column shows "—", because the inventory does not record when a node, container or workload entered that state.
ArgoCD applications. The panel also lists the ArgoCD Applications that concern this cluster (see ArgoCD drift) whose sync status is OutOfSync or whose health is Degraded. Instead of Critical/Warning, such a row carries the status word itself: Degraded (critical, sorted with the critical rows) or OutOfSync (warning). The title reads "ArgoCD app checkout: OutOfSync, Degraded"; the second line names the project, where the app deploys when that is another cluster — or "unmapped destination 'production'" when its destination maps to no Console cluster (listed, not dropped; for ArgoCD's renamed in-cluster cluster set argocd.inClusterNames, after which the app counts as this cluster's) — the hub that reported it, and how many resources are out of sync. Two qualifiers appear as small tags:
- cross-environment — a hub in another environment of the client reports this app;
- stale — the app was reported by this cluster's agent, and this cluster's own ArgoCD report is not current (the agent's role cannot list Applications, the report was oversize or failed), so the row is the last known state. An app reported by a hub is never tagged stale because of this cluster's report: the hub's report is the one that speaks for it.
When this cluster's ArgoCD report is not current — the agent is too old to report ArgoCD, has not sent a report yet, its Helm chart RBAC cannot list applications.argoproj.io, or the report was oversize or failed — one line under the list says so, because apps this cluster would report may be missing or out of date. The same line appears when the report is current but partial: the cluster has more Applications than the agent sends (its cap), or some are not reported because their namespace is excluded by the agent's configuration (the line gives the number). A cluster without ArgoCD installed gets no line: nothing is missing. If the applications cannot be read at all, the line says that instead. ArgoCD rows have no link and no age.
No false all-clear. The green "Nothing needs attention" check appears only when the ArgoCD part is known and complete. While the applications are still loading, or whenever the line above is shown, an empty panel reads "Nothing else needs attention" with a neutral icon, and says whether ArgoCD is still loading or unknown. If the cluster summary fails to load, the panel says "Summary unavailable" and still lists the ArgoCD rows it has.
Capacity
There are three bars: CPU, Memory and Pods. Each bar is measured against the nodes' allocatable capacity (the full track):
- Limits are drawn hatched underneath.
- Requested is drawn over the limits. Requests and limits count only pods that are not
SucceededorFailed. - Used (from metrics-server) is drawn inset on top, so you read "used on top of requested".
- A share above allocatable is drawn to the end of the track and printed as "›100% (overcommitted)". This is common for limits.
- With no metrics-server data the bar draws no used layer and says "used: no metrics-server data". Requests and limits are still drawn.
- When only some nodes report usage, used is a share of those nodes' allocatable (requests and limits stay shares of every node), and the bar says "usage from N of M nodes". The clusters list writes the same case as "used 75% (2 of 10 nodes)".
- The Pods bar shows running pods against allocatable pod slots. Requests and limits do not apply to it.
Charts
The charts need metrics:read on the cluster's environment. Without it, the chart area is replaced by one line: "Charts need metrics:read on this environment".
| Chart | What it shows |
|---|---|
| CPU by node | One line per node, as a % of the node's CPU (node_cpu_usage_ratio, the busiest reading per node). The 5 nodes with the highest peak are coloured; the rest are drawn muted and listed as "+N more nodes". The subtitle names the peak node and time. |
| Memory by namespace | Working-set memory (pod_memory_working_set_bytes) stacked by namespace: the 4 namespaces with the highest peak, plus one "other (K)" band that sums the rest. The subtitle gives the cluster's allocatable memory. |
| Container restarts | Restarts per step across the cluster. See Restarts. |
A response is capped at the 50 series with the highest peak. When the cap applies, the legend says "showing the 50 largest". When no node of the cluster is linked to a host, the charts say "No node of this cluster is linked to a host sending metrics yet."
Restarts
Restarts come from the node agent's container_restarts_total series, which carries the kubelet's restart count for each container, including init containers. Each chart point is the number of restarts within that step: the rise in each container's count over the step, summed across the cluster. The step is 1 minute for windows up to 6 hours, 5 minutes up to 24 hours, 15 minutes up to 7 days and 1 hour beyond that (the endpoint accepts at most 30 days). The first sample of a new container counts nothing, so an agent upgrade or a new pod does not produce a false spike. A gap in a container's samples shorter than 5 minutes (or one step, if longer) is bridged; restarts during a longer gap are lost, never invented.
A gap means missing data, not zero restarts:
- A step with no data is drawn as a gap, never as 0.
- When some steps have data and some do not, the tile reads, for example, "4 / 6h — over 5h 45m of 6h (gaps are missing data)". The sum covers only the steps that have data.
- With no data in the whole window the tile shows "—" and "no restart data in this window", or one of the notices below.
- The total counts only steps inside the window. The chart's first point, at the window's start, covers the step before it and is left out of the total, just as it is left out of the coverage.
The count covers only nodes whose node agent sends the new series. In a mixed fleet where some node agents are upgraded and some are not, the restart total covers only the upgraded nodes, and nothing on the page marks it as partial.
Older node agents
Workload labels, kubelet_pods, container_restarts_total and container_waiting are sent only by node agents at v0.8.0 or later. Older node agents still send every other metric, so CPU by node and Memory by namespace keep working.
When the restarts query comes back empty, at least one node is linked to a host, and CPU by node has data in the window, the Restarts tile and chart say:
Workload breakdown and restart history need a newer node agent (v0.8.0 or later)
The page shows this notice instead of an empty chart that would read as "no restarts". The page does not read node agent versions: it infers the notice from nodes that report metrics (CPU by node, which every node-agent version sends) but none of the new series. Linkage alone is not enough, since a linked node whose agent is down, crash-looping or blocked from NATS sends nothing. If CPU by node is empty too, the tile and both charts say "No node sent metrics in this window" instead: that is an outage of the node agents, not their age. See Node Agent on Kubernetes → Workload labels.
Nodes tab
The Nodes tab is a table with one row per node, read from GET /api/v1/clusters/{id}/nodes/capacity (PostgreSQL only, the latest inventory sync). On a phone the table scrolls sideways inside its own box.
| Column | What it shows |
|---|---|
| Node | The node name, with its roles (for example control-plane, worker) as a muted line under it |
| Status | An icon and a word, never colour alone: Ready; Not ready when the status is NotReady, and Not ready (Unknown) and the like for any other status that is not Ready, so a node that stopped reporting never looks healthy; cordoned when the node is unschedulable (spec.unschedulable). A node can show two, for example "Ready" and "cordoned". |
| CPU, Memory | A capacity bar against the node's allocatable: used (metrics-server) on top of requested, with limits hatched past the requested bar. Under it: used 1.6 · req 2.0 of 4.0 cores (memory in GiB). Requests count only pods that are not Succeeded or Failed, summed per node. With no metrics-server figure the row reads "used —" and draws no used fill — never 0. A node that reports no allocatable reads "allocatable unknown" and draws a hatched empty track, never a division by zero. |
| Pods | Pods on the node of its allocatable pod slots, as a small bar and 22 / 110 |
| Requested | The larger of CPU and memory requested ÷ allocatable, with which one it is (92% Memory) |
Warning marker. When CPU or memory requested is at or above the warning threshold, the Requested column adds a warning icon and "≥85% requested". The default is 85%. The table reads it from this browser's dashboard preferences for the page cluster-metrics:<clusterId>, key node_requested_ratio (its warning value, a 0–1 ratio); a stored null turns the marker off. No screen edits this preference yet, so in practice every cluster uses 85%. The threshold is compared on the whole percentages the table displays, so a node at 84.95% reads "85%" and is marked.
Order. The table opens sorted by Requested, highest first — the nodes closest to full, where the next pod may not schedule. Every column sorts (CPU and Memory by their requested share, not by used; Status puts not ready first, then cordoned); a node with no allocatable goes last whatever the direction. The sorted column's header carries aria-sort, so a screen reader announces the order.
Each bar's accessible value carries the same figures, for example "CPU: used 1.6 cores (40%), requested 2.0 of 4.0 cores (50%), limits 3.0 cores (75%)".
Clicking a row (or Enter on it) opens the node drawer. The drawer shows the full node record (conditions, labels, taints, kubelet version, age), which the table does not carry, so it is read first from GET /api/v1/clusters/{id}/nodes/{nodeID}. If the node was deleted meanwhile, the tab says "Node X no longer exists." instead of opening a stale drawer. A deep link to a node that is not in the table (?resource=<name>&resourceKind=Node) says "Node X is not in this view (deleted)" — or "(deleted or beyond the first 1000)" when the list is truncated — instead of silently opening nothing.
The endpoint returns at most 1,000 nodes, ordered by name. Above that the tab says "showing 1000 of N nodes", with N from the endpoint's total. Above 200 rows the table pages in the browser.
Namespaces tab
The Namespaces tab is a table of every namespace in the cluster: what it uses against what it requested.
| Column | Shows | Source |
|---|---|---|
| Namespace | The name, with a badge when its status is not Active (for example Terminating) | Inventory |
| CPU used / requested | A bar with used drawn over requested, and "0.3 / 1.5 cores". Every row in the column shares one scale: the largest used-or-requested value in the column. | Used: metrics. Requested: inventory |
| Memory used / requested | The same, in GiB | Same |
| Pods | Running / total | Inventory, live from the pod rows |
| Restarts | Container restarts of the current pods over their lifetime | Inventory (restart_count) |
| Workloads | The namespace's workload count as of the last sync. The inventory sync writes it; it is not a live count. | Inventory |
| CPU, 6h | A CPU sparkline over the window, with its peak in the accessible name. The header names the window actually queried ("CPU, 7d" for a link with from=now-7d). | Metrics |
The table sorts by CPU used, highest first; every column sorts. Namespaces with no pods come last, and a missing value sorts after present ones in either direction. Above 200 namespaces the table pages 200 at a time. Clicking a row opens the namespace drawer.
Used comes from instant queries, so every namespace has one. The tab runs cpu_by_namespace and memory_by_namespace with instant=true: one point per namespace at the end of the window, up to 5,000 namespaces (above that, "Used is shown for the 5,000 busiest namespaces; the rest read —"). Requests and limits count only pods that are not Succeeded or Failed; used and requested are joined by namespace name.
The CPU trend covers the top 50. The sparkline is the ordinary range query, which is capped at the 50 namespaces with the highest peak. Namespaces outside the 50 have no sparkline, and the table says "CPU trend shown for the 50 busiest namespaces over the last 6h."
The window is the Overview's default 6 hours; the tab has no time picker. A link that carries its own from/to sets the window instead, and the trend column's label follows it. When used cannot be shown the used columns read "—" and one line says why:
- no
metrics:read: "CPU and memory used need metrics:read on this environment — showing requests only." - a failed metrics call: "CPU and memory used could not be loaded — showing requests only."
A namespace with requests but no used value reads "— / 1.5 cores", never "0 / 1.5". The earlier cards view of this tab has been removed: it could not show the usage bars.
Search and sorting on the cluster page tables
Every table on the cluster page has a search box and sortable column headers. That covers Workloads, Pods, Services, Ingresses, PVCs, Jobs, HPAs, Network Policies, Resource Quotas, Endpoint Slices, Nodes and Namespaces.
- Search matches the name or the namespace, ignoring case. It applies about 300 ms after you stop typing and returns to page 1. A search that matches nothing says so, for example "No workloads match 'web'.", and keeps the box so you can change it.
- Sorting. Click a column header to sort ascending; click it again for descending. The header shows the direction with an arrow and announces it to screen readers (
aria-sort). It works from the keyboard because the header is a button. Changing the sort returns to page 1. Columns without a single stored value to sort by have no sort button: ports, rules, TLS, access modes, selectors, policy types, quota hard/used, endpoints of a slice, and workload containers. Age sorts youngest first. - On the server. The paginated tables search and sort on the server across every page, not just the 50 or 100 rows shown. The backend parameters are listed in Kubernetes inventory → Search and sort on the resource lists.
- In the browser. The Nodes and Namespaces tables receive every row at once (at most 1,000 nodes; every namespace), so their search and sort run in the browser. Nodes search by node name and Namespaces by namespace name. On a cluster with more than 1,000 nodes, the search covers only the first 1,000, and a search that matches none of them says so.
- In the URL. The search and the sort are kept in the URL (
q,sort, for examplesort=-age) together with the view and filters described below, so a link opens the same table.
Links and Back on the cluster page
Everything that decides what a cluster page list shows is held in the URL: the view (?view=workloads) plus that list's filters and position, which are the namespace (ns), the workload kind (kind), the pod phase (phase), the label chips (labels=key:value,key2:value2), the search (q), the sort (sort) and the page (page). As a result:
- Back from a workload page returns to the list as you left it. You get the same view, filters and page in one press. Tabs, the time range, and the Logs pod and container chosen on the workload page replace their history entry rather than adding one, so they never cost an extra Back.
- Changing a filter, the search, the sort or the page does not add a history entry, because it replaces the current one. Choosing another view in the sidebar does add one: Back goes from Workloads to the view you were on before. Clicking the view you are already on adds nothing: it clears that view's filters in place, or does nothing if there are none.
- Switching views keeps the namespace where the new view has a namespace filter, and drops the old view's other filters and its page.
- Any list URL can be shared or bookmarked. It opens the same view with the same filters. A
pagepast the end of the results (for example, from an old link) moves to the last page. - The workload page's breadcrumb leads back to its list. Its namespace crumb opens the Workloads view filtered to that namespace, and its kind crumb adds the kind as well.
Events tab
The cluster's Events tab (?view=events) is a timeline of the cluster's Kubernetes events from the last 14 days. The inventory's own events list keeps only 24 hours; this tab reads the 14-day copy the backend writes to VictoriaLogs (see Kubernetes inventory → Events timeline).
Filters. Window (1 hour, 6 hours, 24 hours — the default — 3, 7 or 14 days), type (Normal or Warning), reason (for example BackOff), namespace, and involved object kind + name. Every filter is an exact match. The selects apply as soon as you change them, on their own — a pending invalid text filter never holds them back, and a select never shows a filter that is not applied; the text filters apply with Enter or Apply, never per keystroke, because each read scans the event store and counts against a budget of 60 requests a minute per user (per backend process). A value that is not valid (a namespace with capitals, a reason with spaces) is explained under the field and never sent. The filters are kept in the URL (ev_ns, ev_type, ev_reason, ev_kind, ev_name, ev_window), so a link opens the same view; an invalid value in a link is ignored and the tab says which. Refresh re-reads the window up to now.
Each event shows its type, reason, involved object, count (×N when it recurred), last seen — relative, with the exact time on hover — and its message, newest first and grouped by day. A Warning carries an icon and the word "Warning", never colour alone. A recurring event appears once per count change. Messages are plain text. The involved object has a copy button (it copies Kind/name). First seen and the reporting component are not in the 14-day copy, so the tab does not show them.
What the tab says when it cannot show everything.
- Truncated — at most 500 events are read at once. When more exist the tab says "Showing the newest N events in this window; older ones were left out. Narrow the filters or the window to see them."
- Notes from the backend are listed under the filters — for example that the window reaches past the 14-day retention, or that FailedScheduling and Scheduled events are missing because the cluster's agent is older than v0.8.0 (those events carry no last-seen time and older agents do not send one).
- Timed out — VictoriaLogs did not answer within 20 seconds. The tab says so and suggests a narrower window or a filter, with a retry button.
- No events in this window appears only when the read succeeded and returned nothing. A failed read, an unavailable event store, a rate limit or a missing permission each has its own sentence, never an empty list. A 404 means the cluster was removed or its environment is not visible to you.
Workload page
Each workload has a page at /clusters/:clusterId/workloads/:workloadId, keyed by the workload's ID, so two workloads with the same name in different namespaces or of different kinds never collide. Clicking a row in the Workloads tab — anywhere on it, or Enter / Space on a focused row — opens this page, the same as clicking the workload's name; Cmd/Ctrl-click or a middle click opens it in a new tab. Workloads have no drawer: a link that names a workload by ?resource=<name>&resourceKind=Workload (an HPA's target in its drawer, an older bookmark) opens this page too. The Overview's "ready/desired" rows in Needs attention link here as well. Rows of kinds without a page of their own — pods, services, ingresses, nodes, namespaces and the rest — still open the drawer.
The page reads the workload with GET /api/v1/clusters/{id}/workloads/{workloadID}. A workload that was deleted, or an ID that belongs to another cluster, is a 404, and the page says "This workload no longer exists or you can't access it" with a link back to the Workloads tab — never another cluster's data.
Argo Rollouts. A Rollout (argoproj.io/v1alpha1) is a workload like any other: it is listed in the Workloads tab with a Rollout kind badge (and can be filtered by kind), counts in Needs attention's ready/desired rows and the namespace workload counts, and has this page. See Argo Rollouts for what has to be deployed for it to appear.
Header. The breadcrumb (Clusters / cluster / namespace / kind), the name, and a status pill that checks, in order: pods waiting on a failure reason ("3 of 6 pods CrashLoopBackOff", the most common reason) → fewer ready than desired ("Degraded — 3 of 6 ready") → desired 0 ("Scaled to zero") → "Ready". Beside the name sits the cluster's API/agent badge (see Header), and the meta line ends with "last sync 5m ago" — when the cluster agent last reported. The meta line also shows the update strategy when the workload has one: RollingUpdate, Recreate or OnDelete for the built-in kinds, and canary or blue-green for a Rollout.
A disconnected collector. The pill, the Ready tile and the pod table are all read from the inventory the cluster agent writes. While the agent is disconnected (or the API is unreachable, or reported no status), that inventory is a frozen picture, so the pill gives no verdict: it reads "Status unknown — collector disconnected, last sync 2d ago" with a warning icon, never "Ready". The Ready tile says "as of last sync 2d ago" and is marked stale, and a warning line above the pod table says the pod states may have changed since. The per-pod charts are unaffected — they read the node agents' metrics, and a gap there already reads as a gap.
A time picker (default 6 hours) sets the window for the charts and the Restarts tile.
Tiles
| Tile | Value | Rules |
|---|---|---|
| Ready | Ready / desired replicas, from the workload row | When short, a warning icon and "N not ready" |
| Restarts / window | Restarts of the workload's containers in the window, with a sparkline | The Overview's rule (see Restarts) confined to this workload: a gap is not zero, partial coverage reads "over Xh of Yh" |
| CPU used | Σ latest CPU ÷ Σ CPU requests, as "% of requests" | Counts only current pods that have both a request and a current value, and says "from k of n pods with a request" when some lack a value, and "N pods have no request" when some set none. No pod sets a request → "—" and "no requests set"; nothing is ever divided by an unset request. |
| Memory used | The worst pod's working set ÷ its own limit, as "% of limit" | Names that pod ("api-7d9 · 612 MiB of 768Mi"). No pod sets a limit → "—" and "no limits set". |
"Latest" is each pod's last point, and only when it is within max(1.5 steps, 3 minutes) of the window's end (or now). A pod whose series stopped reads "—" rather than its last known figure. When the window ends in the past, the tiles say "as of 14:30" (the window's end).
Per-pod charts and request/limit lines
CPU by pod (cores) and Memory by pod (working set, bytes) draw one line per pod (cpu_by_pod, memory_by_pod). The 5 pods with the highest peak are coloured in the fixed palette order; the rest are muted and listed as "+K more pods". The series include pods replaced during the window. A response is capped at the 50 busiest pod series; when the cap applies the legend says "showing the 50 busiest pod series (includes replaced pods)".
Each chart carries a dotted request line and a dashed limit line built from the workload's current pods (not Succeeded or Failed). Requests and limits are summed across a pod's containers, the only granularity the inventory stores.
| The current pods… | Line |
|---|---|
| all set the same value | One line at that value, labelled "request 500m" or "limit 768Mi" |
| set different values, or some set it and some do not | One line at the maximum, labelled "requests vary (250m–500m)" — or "(none–500m)" when some pods set none |
| none set it | No line; the chart subtitle says "no requests set" / "no limits set" |
Replicas history
The Replicas chart draws ready replicas as a step line, with desired replicas as a muted dashed overlay to read it against (replicas query). A step line is used because a smooth curve would draw 2.5 replicas that never existed. The backend writes the series, k8s_workload_replicas, to VictoriaMetrics at each inventory sync, so it has one point every 5 minutes and starts when the backend version that writes it was deployed — there is no history before that. When the window starts before the first point, the subtitle says "history since 09:15" (the first point) rather than drawing zero replicas; with no point at all in the window it says "No replica history in this window — it is recorded on each inventory sync (every 5 minutes)". The write is best-effort: a VictoriaMetrics outage leaves a hole in this chart but never fails the inventory sync (see proxima_k8s_replicas_write_errors_total in docs/standards/metrics.md).
Pod table
One row per pod of the workload (GET /api/v1/clusters/{id}/workloads/{workloadID}/pods, read up to 1,000 pods; beyond that "showing 1,000 of N pods"):
| Column | Shows |
|---|---|
| Pod | The name; a row opens the pod drawer |
| State | Icon and word. A container's waiting reason wins (CrashLoopBackOff, ImagePullBackOff, ErrImagePull and CreateContainerConfigError as danger, any other reason as a warning); otherwise the phase, with "Running · 1/2 ready" when a running pod is not fully ready |
| Restarts | The pod's restart count from the inventory |
| CPU, Memory | The pod's current value from the per-pod charts; "—" when it has none |
| Node | The node the pod runs on |
| Last reason | A terminated container's reason (OOMKilled, Error), otherwise the pod's own status reason (Evicted) |
The cluster agent collects each container's current state and reason, not Kubernetes' lastState. A container that was OOM-killed and is running again has no terminated state any more, so its row shows no last reason. The restart count still rises.
When the per-pod series are empty
The per-pod charts need the workload labels that only node agents at v0.8.0 or later send. When a per-workload series comes back empty while the workload has pods, the page says why, using the same rules as the Overview's Older node agents: no node linked to a host; nodes report but not this series (the node-agent upgrade notice); or "No node sent metrics in this window". A workload with no pods says "This workload has no pods".
Without metrics:read, the charts are replaced by one permission line; the Ready tile, the status pill and the pod table still render.
Tabs
For one page that maps all of the operate views below — their routes, permissions, limits, what is never sent or stored, and the rollout steps — see Kubernetes: Operate (read-only).
The page has tabs: Overview (the header's tiles, charts and pod table above), YAML, Logs, Events and History. The active tab is kept in the URL as ?tab=yaml / ?tab=logs / ?tab=events / ?tab=history, so a shared link opens the same tab; Overview is the default and is never written to the URL, and an unknown ?tab= value opens Overview. The tabs are radix tabs: Tab moves focus into the bar, the arrow keys move between tabs. The time range picker belongs to Overview and is shown only there. Each tab needs a permission on the cluster's own client and environment (see Permissions); a tab you lack it for is hidden rather than shown with a no-permission panel.
YAML tab
The YAML tab answers "what is this workload exactly, what changed between its revisions, and does it match Git?". It reads the cluster live, through the cluster's agent (see Kubernetes inventory → Live Reads), and stores nothing.
ArgoCD card. One line per ArgoCD Application that manages the workload:
✓ Synced · app
shop-checkout· revabc1234→def4567· ✓ Healthy
Sync (Synced / OutOfSync) and health (Healthy, Progressing, Suspended, Degraded, Missing) are each an icon plus a word — never colour alone. They are the workload's own resource row in the Application (an Application can be OutOfSync because of a ConfigMap while this Deployment is Synced); the Application's own counts follow as "Application-wide: 2 of 14 resources OutOfSync". The revision is the synced revision → the target revision (a 40-character SHA is shortened to 7; a tag or branch is shown as-is, once when the two are equal). Below it:
- Another environment. An Application reported by an ArgoCD hub in another environment of the same client carries the label "reported by an ArgoCD in another environment" (with the hub cluster's name when you can read it).
- Repository. The Application's repo URL is a link only when its scheme is
httporhttps(checked case-insensitively, soHTTPS://…is a link); anything else —git@host:org/repo.git,ssh://,javascript:,data:— is plain text. The URL comes from an Application spec in your cluster, which Console does not control, so it never becomes any other kind of link. Links open in a new tab withrel="noopener noreferrer". - Several Applications listing the same workload is ArgoCD's shared-resource conflict, and the card says so (a warning icon and the word "Conflict").
- Not managed. "Not managed by an ArgoCD application Console knows about." — said only when nothing that could prove otherwise is unknown.
- Unknown. When Console cannot conclude "not managed", it lists why, one sentence per reason: this cluster's agent is too old or has not reported; its last report did not succeed — named by the report's own state: the agent's own Kubernetes role cannot list Applications (its Helm chart RBAC, not your Console permission: grant it
listonapplications.argoproj.io), the report was too large to send, or it failed; its last report left out resource lists; it has more Applications than it reports, or skips some namespaces; no Application is reported by or mapped to this cluster (an ArgoCD elsewhere could still deploy here by a server URL Console cannot match); a destination is ambiguous; this cluster's own ArgoCD lists the workload but deploys to a destination Console does not map to any cluster (usually ArgoCD's in-cluster cluster renamed — setargocd.inClusterNamesin the agent's Helm values; the card's note names the destination); an Application's resource list was truncated; or a hub that deploys here last reported a failure or left out resource lists. A managed answer can carry the same reasons ("the check may be incomplete"). - ArgoCD not installed and agent too old (with the version it needs) are sentences of their own, shown in place of an answer when the workload is not managed. "ArgoCD is not installed on this cluster" keeps the caveat beneath it that an ArgoCD elsewhere could still deploy here by a server URL Console cannot match.
See ArgoCD Drift for how Applications are collected and mapped.
Manifest. The workload's manifest as YAML (rendered in the browser from the agent's JSON, keys in the order the agent sent them). The agent redacts it inside the cluster before it is sent: managedFields and the kubectl.kubernetes.io/last-applied-configuration annotation are removed, and every literal env value becomes <redacted: N chars>. Those markers are shown as muted badges, with one line saying the values were redacted in the cluster and never leave it; valueFrom references are kept, since they name a Secret rather than reveal it. The agent's notes are listed under the YAML. Copy redacted YAML copies exactly what is shown, markers included. The view has a bounded height and scrolls; past 1,000 lines it is virtualised, so a manifest at the 512 KiB cap does not freeze the page.
Revisions. The newest 20 revisions, newest first: the revision number (the highest is marked newest and the next previous — "newest", not "current", because in a partitioned StatefulSet rollout the highest revision is the update revision rather than the one most pods run), when it was created (relative, with the exact time on hover), its replica count and its images. A Deployment's and a Rollout's revisions are its ReplicaSets; a StatefulSet's and a DaemonSet's are its ControllerRevisions. Two selects — Compare from and to — pick any two revisions (by default previous → newest), and the diff of their redacted pod templates is shown below in the same one-column diff view as file changes, so it reads the same at phone width. Picking the same revision twice, two identical templates, a revision whose pod template is not in the reply (the agent's notes say why: the reply's size budget, or a ControllerRevision that carries no template), a single revision, or none at all each get their own sentence instead of an empty diff. The diff is bounded: it is computed in the browser with a line-by-line comparison whose cost grows with the product of the two templates' line counts, so a pair past 4,000,000 (for example two templates of about 2,000 lines each) is not compared — the panel says "These pod templates are too large to compare in the browser" rather than freezing the page. It is computed once per pair, not on every refresh of the tab. When older revisions exist, the list says so.
Refresh. Manifest and revisions are agent round trips under a per-user budget of live reads, so they are never polled: each section shows "fetched 14:05:09" (the agent's read time) and a Refresh button. The ArgoCD card is refreshed on the metrics cadence, from Console's stored reports.
Every state is a sentence. No agent, agent offline (with when it was last seen), agent too old (with the version it needs), the cluster no longer available to you (removed, or its environment not visible to you), the agent's own error (for example "Argo Rollouts is not installed", or an object missing from the cluster), timed out, live-read budget used up, refused as invalid, and no permission each render a specific sentence with an icon and a word — never an empty panel that reads as "nothing here". The ones a retry can change offer Try again.
Logs tab
The Logs tab shows one container's output live from the pod, read through the cluster's agent the moment you look. Logs are never stored or shipped: not in Console's database, not in VictoriaLogs, not in your browser's storage. Leave the tab and the lines are gone. (The host log pipeline is a separate thing; container stdout is not part of it.)
Choosing what to read.
- Pod. The workload's pods, from the same pod list as the Overview's pod table. The tab opens on the newest pod that is not Running and Ready (a pending or crash-looping pod is usually why you are here), otherwise on the first pod. Pods that are not ready say so in the picker.
- Container. The pod's containers, in the order the agent reports them. A multi-container pod opens on its first container. When the inventory does not list the containers, the backend answers with the pod's containers and the tab reads the first one. Init containers are not offered: the agent reads only a pod's regular containers.
- Previous instance. Reads the container's last exited instance, which is the crashed one for a crash-looping container. It is a finished stream, so it is read once and not followed. A container that has never restarted has no previous instance, and the agent's message says so.
- Pinned in the URL.
?tab=logs&pod=<pod>&container=<container>opens that stream. The pod the tab opens on is written to the URL as soon as it is chosen, and choosing a pod or container writes them too, so a refresh or a shared link opens the same container, and the pod list refreshing (a pod turning ready, a new pod appearing in a rollout) never moves you to another pod. Only when the pinned pod leaves the workload's pod list does the tab pick the default again, saying "Pod X is gone … Now showing Y". Acontainerin the link is used only for the pod the link names, only when it is a valid container name and one of that pod's containers; otherwise it is dropped from the URL and the pod's default container is read. The tab names the container it could not find only when the link's own pod is open; acontainerin a link with nopod, or whose pod is gone, is dropped silently (a gone pod gets the "Pod X is gone" notice instead). (When the inventory does not list the pod's containers, the backend's own list is shown instead.)
How much, and how often. The labels above the lines state exactly what you are seeing:
| Label | Meaning |
|---|---|
| Live from the pod — not stored | Every line came from the pod just now and is kept only in this tab's memory. |
| Last 300 lines | The first read is the last 300 lines (the agent's cap, also 256 KiB). "(older lines exist above)" when the container has written more than that. |
| Refreshes every 5 s | While following, the tab asks again every 5 seconds for lines since the last line's timestamp and appends only lines it has not shown. The number is the real current interval: 20 s for a while after you hit the rate limit, 30 s while the agent is offline. It is never a WebSocket; each refresh is one ordinary request. |
| Paused | Pause stops the refreshes and keeps the lines; Resume continues from the last line. Switching pod, container or instance starts again unpaused with a fresh buffer. |
| Previous instance — read once, not followed | The previous instance is a finished stream. |
Refresh starts again from the last 300 lines. Up to 5,000 lines are kept in the tab; past that the oldest are dropped, and a notice says how many ("Only the newest 5,000 lines are kept in this tab; 1,234 older lines were dropped from the top").
Gaps. A refresh also reads at most 300 lines. When a busy container wrote more than that in 5 seconds, the lines in between were never read, and the tab inserts a marked row — Gap — some lines may be missing here — instead of splicing the two parts together silently.
Older agents. Following needs cluster agent v0.8.0 or newer, which stamps every line with the apiserver's timestamp. With an older agent the tab says so once ("Limited follow", in place of the backend's own note about it): each refresh re-reads the last 300 lines and replaces what is shown, so lines written between two refreshes beyond those 300 can be missed, and nothing marks where. An agent too old to read logs at all gets the "agent too old" sentence with the version it needs.
Redaction. Every line is secret-sanitised by the Console backend before it reaches your browser. Bearer tokens, key=value secrets, connection-string credentials and JWTs are replaced with [REDACTED]. The tab says so in one line. The redaction is pattern-based, so a secret in a shape it does not recognise can still appear.
Reading the lines.
- Search looks only within the loaded lines; there is no server-side search. Your text is matched literally (characters like
.*(mean themselves) and case-insensitively unless you tick Match case. The count reads "2 of 7 lines"; a screen reader hears the number of matching lines when the search changes, not on every refresh. The current match stays on the same line as new lines arrive or old ones are dropped. Next / Previous, or Enter / Shift+Enter in the box, move between matching lines, and matches are highlighted. - Level words. A line's level is read from a JSON
"level":"…"field, a logfmtlevel=…, a klog prefix (E0930 …) or an upper-case word near the start (ERROR,WARN,INFO,DEBUG, …). It is shown as a word badge, and error and warn also carry an icon, so the level is never shown by colour alone. - Wrap lines and Timestamps (the line's own timestamp, in UTC to the millisecond) are toggles. These two preferences are the only thing the tab saves in your browser.
- Lines are shown as plain text: ANSI colour codes and other control characters are stripped, and nothing in a line is ever interpreted as markup.
- The view sticks to the newest line as lines arrive. Scroll up and it stays where you are, and a Jump to latest button appears. Past 500 lines the list is virtualised, so 5,000 lines scroll smoothly.
- Copy loaded lines (redacted) copies the lines in the tab, as redacted text. There is no download.
Rate limit. Live log reads are limited to 30 requests a minute per user, per backend process. A 5-second follow uses 12, so two following tabs fit, and a third does not. Past the budget the tab shows the live-read budget sentence, backs off to a 20-second refresh ("the next try is in 20 s") and keeps the lines it already has. Every other state (no agent, agent offline, agent too old, the cluster no longer available to you, the agent's own error such as a pod that no longer exists, timed out, refused as invalid, no permission) is a specific sentence with an icon and a word, and a retry button where a retry can help.
Permission. The tab needs logs:read on the cluster's client and environment. It is the first workload tab whose permission differs from the page's own assets:read: a reader with assets:read but without logs:read there does not see the tab, and a ?tab=logs link opens Overview for them.
Workload Events tab
The workload's Events tab is the cluster Events tab read for this workload only. It includes the events of the workload itself and of its ReplicaSets and pods — where BackOff, OOMKilling and FailedScheduling are reported — matched by the names Kubernetes generates for the workload's kind, so api never includes api-gateway's pods. The filters are window (up to 14 days, 24 hours by default), type and reason, kept in the URL like the cluster tab's; an invalid type or reason in a link is ignored and named. Everything else — truncation (at most 300 events at once), notes, the timeout sentence, the empty state — reads as on the cluster tab. A workload whose kind has no owned-object match shows only its own events, and the tab says so.
Permission. assets:read on the cluster's client and environment.
History tab
The History tab answers "what changed?" for this workload, merged newest first from four sources:
| Source | Shows | Needs |
|---|---|---|
| Recorded changes (kind Change) | Spec changes the cluster agent recorded for the workload: image, resources, env var names, config/secret references, strategy, replicas — each as a field list such as "image 2.40.3 → 2.41.0". Values of env vars and secrets are never recorded. | changes:read |
| Deploys (kind Deploy) | ArgoCD deploy notifications attributed to the workload's ArgoCD Application, or matched by namespace when the app is not known to Console | changes:read |
| Events (kind Event) | Kubernetes events with reason ScalingReplicaSet, Killing, BackOff, FailedScheduling, Evicted or OOMKilling (14-day retention) | assets:read |
| Alerts (kind Alert) | Alerts in the cluster's environment whose namespace label — and workload or pod label, when present — match | alerts:read |
Each item has an icon and its kind's word, a title, the time (relative; exact on hover) and its details. An alert's status (Firing, Acked, Resolved, Silenced) and severity use the same badges as the Alerts page, each with an icon. An alert links to its alert page and a change or deploy to its change page; the links are built from the item's id, never from text the backend sent. A deploy from a webhook source with no environment is labelled "client-wide deploy — may belong to another environment's cluster": Console cannot tell which of the client's clusters it went to.
Per-source permissions. The tab opens when you hold any one of changes:read, alerts:read or assets:read on the cluster's client and environment; each source is then checked on its own. A Sources list says, per source, whether it was read, truncated (only the newest kept), forbidden ("you can't see alerts here (needs alerts:read)") or unavailable ("events unavailable right now"). When any source is forbidden or unavailable a banner says the list is incomplete, because a quiet stretch may only mean that part is missing. If no source could be read, the tab says that instead of showing an empty list. An empty list names its gap: "…in the sources you can see" when a source is forbidden, "…in the sources that could be read" when one failed.
Window. 24 hours, 3, 7 (the default), 14 or 30 days, kept in the URL as h_window. Events older than 14 days are gone, and a note says so.
Notes from the backend are listed, including the caps (200 changes, 100 deploys, 200 events, 100 alerts, newest kept), that GitHub and GitLab changes are not part of this timeline, and — when the cluster's environment has other clusters — that alerts and deploys are matched by namespace within the environment, so ones from another cluster with the same namespace may appear.
Not shown. Scaling by a HorizontalPodAutoscaler is not recorded as a change, so there are no HPA scaling rows; the ScalingReplicaSet events still show the replica changes themselves.
Clusters list
The list loads every cluster you can see. It pages through the API 100 at a time, up to 5,000 clusters. Earlier versions stopped at the first 100 and showed a truncation banner.
Fleet tiles
The tiles cover the clusters in the selected project and environment (?client, ?env). The status filter and search narrow only the table. When either is set, the page says "Totals cover all clusters in scope; status and search filter the table only."
| Tile | Value | Note |
|---|---|---|
| Clusters | Number of clusters | "N API reachable · M unreachable/disconnected", plus "K status unknown" when any |
| Nodes ready | Ready / all nodes across the fleet | The clusters with not-ready nodes, by name. The tile turns critical with "Not ready". |
| CPU requested | Requested CPU as a % of allocatable, over the clusters that report capacity | "36.0 of 72.0 cores requested", or "N clusters without capacity data" when some are left out. "—" when none report it. |
| Pods not running | Pods in Pending, Failed or Unknown (any phase other than Running or Succeeded) | "in N clusters" or "all running" |
| Warning events / 1h | Warning event objects last seen in the past hour | See the note below |
Kubernetes folds repeats of the same event into one object with a count. The "warning events in 1h" figures, on the list and on the Overview, count those objects whose last occurrence was within the hour. They do not add up the count field. An event that fired 200 times counts once.
Rows
| Column | Shows |
|---|---|
| Cluster | The name, with environment · version · platform below it |
| API | The API/agent badge (see Header). A warning icon marks an inactive collector. |
| Nodes | "17 / 18 ready" |
| CPU, Memory | A "used on top of requested" bar against allocatable, then "used 41% · requested 68%". With no metrics-server data: "used — · requested 68%". Without capacity data: "—". |
| Pods | "138 running", with "all running" or "N not running". This column sorts by pods not running, so clusters with problems come first. |
| Warnings 1h | Warning event objects in the past hour |
Restarts are not on the list: they need a metrics query per cluster, so they stay on the Overview.
Status filter
The status filter uses the badge vocabulary: API reachable (?status=healthy), API unreachable (?status=unreachable), Agent disconnected (?status=disconnected) and Status unknown (?status=unknown). The URL values predate the new labels and are kept, so existing links still work. ?status=degraded has been removed because no cluster agent ever reports it and it matched nothing. An old link with that value, or any other unrecognised value, now shows every cluster.
Permissions
| What | Needs |
|---|---|
Overview inventory figures (/summary) | assets:read on the cluster's environment |
Overview charts and the Restarts tile (/metrics/{query}) | metrics:read on the cluster's environment. assets:read alone gets a 403 from the metrics endpoint, and the page shows the permission notice. |
Clusters list, including ?include=capacity | assets:read. The list shows only clusters in environments where you hold it. |
| Cluster detail and its inventory tabs (nodes, workloads, pods, services and the rest) | assets:read on the cluster's environment |
Nodes table (/nodes/capacity), the node drawer (/nodes/{nodeID}), Namespaces usage (/namespaces/usage) and the workload page's workload and pods (/workloads/{workloadID}, /workloads/{workloadID}/pods) | assets:read on the cluster's environment |
The Namespaces tab's used columns and trend, and the workload page's charts and CPU, Memory and Restarts tiles (/metrics/{query}) | metrics:read on the cluster's environment |
The workload page's YAML tab: live manifest and revisions (/live/manifest, /live/revisions) and ArgoCD drift (/workloads/{workloadID}/drift) | assets:read on the cluster's environment. The tab is hidden without it — in practice the whole page needs assets:read there, so a reader of the page always sees the tab. Manifest and revisions share a budget of 60 requests a minute per user, per backend process. |
The cluster Events tab and the workload page's Events tab (/events/timeline) | assets:read on the cluster's environment. Rate-limited to 60 requests a minute per user, per backend process. |
The workload page's History tab (/changes) | Any one of changes:read, alerts:read, assets:read on the cluster's environment opens it; recorded changes and deploys then need changes:read, alerts alerts:read, events assets:read, each shown as forbidden without its permission. Rate-limited to 60 requests a minute per user, per backend process. |
Overview Needs attention ArgoCD rows (/argocd/applications) | assets:read on the cluster's environment |
The workload page's Logs tab (/live/logs) | logs:read on the cluster's environment. The tab is hidden without it, even when you can read the rest of the page; the route also answers 403. Rate-limited to 30 requests a minute per user, per backend process. |
Every cluster route accepts an environment-scoped grant, such as assets:read given for one environment only. A user with that grant sees the Clusters page and the clusters in that environment, and nothing from the client's other environments. A client-wide grant covers every environment of the client.
A cluster in an environment you cannot see returns 404, never 403, so the response does not confirm that the cluster exists. A 403 means you can see the environment but do not hold assets:read on it, or hold assets:read nowhere.
A node or workload ID is always looked up together with the cluster in the URL: an ID that belongs to another cluster is a 404, even when you can see that other cluster.
API routes used by these pages
| Route | Returns |
|---|---|
GET /api/v1/clusters/{id}/nodes/capacity | {data: {nodes, truncated, total}}: every node (at most 1,000, by name) with allocatable, requested and limit CPU/memory, metrics-server usage (null when absent, never 0), allocatable pods, pod count, status, roles and unschedulable. total is the cluster's node count, so truncated = total > the nodes returned. |
GET /api/v1/clusters/{id}/nodes/{nodeID} | {data: node}: one node record, as in the node list |
GET /api/v1/clusters/{id}/namespaces/usage | {data: [...]}: every namespace with requested/limit CPU and memory, pods total and running, restarts and workload_count (as of last sync) |
GET /api/v1/clusters/{id}/workloads/{workloadID} | {data: workload}: one workload of this cluster |
GET /api/v1/clusters/{id}/metrics/{query} | New queries cpu_by_namespace, and the per-workload cpu_by_pod, memory_by_pod, restarts_by_workload and replicas, which require namespace, workload and workload_kind. instant=true is accepted by cpu_by_namespace and memory_by_namespace only. |
GET /api/v1/clusters/{id}/live/manifest, /live/revisions, /live/logs | The workload page's YAML and Logs tabs: read through the cluster's agent now, stored nowhere (Live reads) |
GET /api/v1/clusters/{id}/events/timeline | The cluster and workload Events tabs: the 14-day VictoriaLogs copy (Events timeline) |
GET /api/v1/clusters/{id}/changes | The History tab: changes, deploys, events and alerts merged (What changed) |
GET /api/v1/clusters/{id}/argocd/applications, /workloads/{workloadID}/drift | Needs attention's ArgoCD rows and the YAML tab's ArgoCD card (ArgoCD Drift) |
Full request and response formats: Kubernetes Inventory → API Endpoints.
Argo Rollouts
Clusters that deploy with Argo Rollouts instead of Deployments show each Rollout as a workload of kind Rollout. Everything on the workload page works for it, and the inventory export's Kubernetes sheet lists Rollouts next to Deployments. Each piece comes from a different component, so each needs its own upgrade:
| Needs | For | Without it |
|---|---|---|
| Cluster agent image with Rollout support | Rollouts in the Workloads tab, their replica counts, strategy, containers and the Ready tile; their pods linked to them (pod table, request/limit lines) | No Rollouts at all. Older agents also linked a Rollout's pods to a Deployment of the same name that does not exist, so those pods belong to no workload. |
Helm chart upgrade (ClusterRole list on argoproj.io rollouts) | The cluster agent being allowed to list Rollouts | The cluster agent logs one "forbidden to read argo rollouts" warning and skips Rollouts. Upgrading only the image does not add the rule. |
| Node agent image with Rollout support | Per-pod CPU and memory charts, the CPU/Memory/Restarts tiles | Older node agents label a Rollout's pod metrics workload_kind="ReplicaSet", so these series come back empty for the Rollout and the page shows its "nodes report but not this series" notice. |
| Backend and frontend from the same release | The inventory export's Rollout rows (the backend); the Rollout badge colour, the kind filter option and the "canary" / "blue-green" strategy label (the frontend) | Rollouts are still stored and listed — the backend has always accepted any workload kind and written its Replicas history — but the export's Kubernetes sheet lists Deployments only, and the Workloads tab shows Rollouts with a neutral badge and the raw blueGreen strategy |
A failed read removes the Rollouts until the next sync. If the cluster agent cannot read Rollouts in a cycle for a reason other than a missing CRD or a 403 — an API server error or a timeout — that cycle's inventory carries no Rollouts, and the backend's sync prunes every workload it did not see. The Rollouts' rows are deleted and come back on the next successful sync (5 minutes later) with new IDs. In the meantime their pods belong to no workload, and a bookmarked or shared workload page link for a Rollout returns "This workload no longer exists". The built-in kinds behave the same way when their own list fails; Rollouts add one more request (the discovery check) that can fail.
A cluster without the Argo Rollouts CRD is unaffected: the cluster agent checks discovery each cycle and skips Rollouts without logging anything. Details: Cluster Agent → Argo Rollouts.
What does not cover Rollouts yet: the AI assistant's live workload read (k8s_live_workload) works on Deployments, StatefulSets and DaemonSets only, and the agent's scale action on Deployments only. The AI tools that read the inventory snapshot do see Rollouts. A Rollout that takes its template from a Deployment (spec.workloadRef) is shown with no containers. Argo's own rollout status (paused, current canary step) is not collected; the status pill uses the same replica and pod rules as every other kind.
Limitations
- Large clusters undercount above the pod limit. The cluster agent sends at most
PROXIMA_K8S_POD_LIMITpods (HelmclusterAgent.podLimit, default 5000). Above the limit it keeps pods that are not Running first, then the ones with the most restarts, so Running pods are the ones dropped. Pod totals, "running" figures, the Pods capacity bar and CPU/memory requests and limits then undercount. Problem pods (Pending, Failed, restarting) are still reported. Raise the limit for clusters above it, and keep the NATS payload limit in mind. - Mixed node-agent fleets report restarts only from upgraded nodes (see Restarts).
- Nodes not linked to a host are counted from the inventory but are missing from every chart ("metrics from N of M nodes").
- Ages are shown only for pending pods.
- Drawer links match by name. The workload page is keyed by ID, but the drawers' deep links (
?resource=<name>) still match by name: two workloads with the same name in different namespaces open the first match's page. - Truncation. The Nodes table shows at most 1,000 nodes, the Namespaces "used" columns at most 5,000 namespaces and its trend the top 50, the workload page's per-pod charts the 50 busiest pod series and its pod table 1,000 pods. Each says so when it applies.
- The cluster agent's pod limit (
PROXIMA_K8S_POD_LIMIT, above) also undercounts the Nodes table's requested figures, the Namespaces tab's requested, pods and restarts, and the workload page's pod table and request/limit lines, because they all read the synced pods. - Per-pod charts and workload restarts need node agents at v0.8.0 or later. Older agents send no workload labels, so those series are empty (see When the per-pod series are empty).
- A transient Rollout read failure prunes the Rollouts for one sync, and they return with new IDs, so bookmarked workload page links for them break (see Argo Rollouts).
- Replica history starts at deploy of the backend version that writes it, in 5-minute steps.
- Last reason shows only the current container state;
lastStateis not collected.