Skip to main content

ArgoCD Drift

Console shows whether a workload matches what ArgoCD wants it to be: Synced / OutOfSync, the Application that manages it, the target and synced revisions, and the health ArgoCD assesses. The data comes from the cluster agent, which lists ArgoCD Application objects (argoproj.io/v1alpha1) in its own cluster through its own ServiceAccount. Console holds no ArgoCD credentials and never calls the ArgoCD API.

The drift card is one of the read-only operate views mapped in Kubernetes: Operate (read-only), which also lists the rollout steps.

What the agent reads, and what it never reads​

Each inventory sync, the agent asks API discovery whether the cluster serves argoproj.io/v1alpha1 applications, lists them, and reports per Application only:

FieldSource
name, namespacemetadata
projectspec.project
destination (server, name, namespace)spec.destination
sync_statusstatus.sync.status (Synced, OutOfSync, Unknown)
health_statusstatus.health.status (Healthy, Progressing, Degraded, Suspended, Missing, Unknown)
target_revisionspec.source.targetRevision (first source of a multi-source app)
synced_revisionstatus.sync.revision (revisions[0] of a multi-source app)
repo_url, path, chartspec.source — the repository URL with userinfo, query and fragment removed
source_count1 for spec.source, N for spec.sources
resources[]status.resources[] rows of the workload kinds only (Deployment, StatefulSet, DaemonSet, Rollout): group, kind, namespace, name, sync_status, health_status
resource_total, resource_out_of_sync, resource_unhealthyCounts over all status.resources[] rows, every kind (OutOfSync; Degraded or Missing)
resource_namespacesThe distinct namespaces of all rows, sorted, at most 50 (resource_namespaces_truncated)
Deliberate deviation from D9

D9 says "report status.resources[]". The agent lists rows only for the four workload kinds and summarizes the rest (counts and namespaces). Drift is read per workload, the per-app counts and the timeline's namespace view need only the summary, and listing every Service, ConfigMap and Secret of a large hub would push the report past the single inventory message: measured, 100 Applications × 50 resources drop from 723 KiB (every row listed) to 179 KiB (10 workload rows each plus totals).

Values never leave the cluster

An Application can carry secrets: Helm values / valuesObject, parameters and fileParameters, config-management plugin env and parameters, kustomize patches, jsonnet variables, and repository URLs such as https://user:[email protected]/…. None of that is read. The agent decodes the fields above one by one (never the whole spec). It scrubs the repository URL and the destination server URL, and the scrub fails closed: nothing before the userinfo's @ survives, query and fragment never do, and a value whose userinfo cannot be told apart from a query or fragment is dropped (reported empty) rather than guessed.

  • https://user:token@host/repo.git?private_token=x → https://host/repo.git.
  • A URL the Go parser rejects, or with an @ after a ?/# (a /, space, bad escape, # or ? in the password, a bad port) is cut at the last @ first and only then at its query or fragment: https://user:p/ss@host/x and https://user:p#ss@host/x → https://host/x; scp-style git:p#[email protected]:o/r.git → github.com:o/r.git.
  • When an @ follows a ? or # and the text before it does not look like user:password (no :, a /, or a numeric port, as in https://host/x?token=a@b or https://user:1#s@host/x), or when more than one @ follows the ?/#, the whole value is dropped.
  • A / inside a password or token makes the parser read part of the credential as the host and the rest as the path (https://user:8080/s3cret@host/x, https://user:/s3cret@host/x, https://tok/[email protected]/o/r). So when the path contains an @, the value is dropped if the parsed host has a port or an empty port, if a host-like label (with a ., a : or [) follows the @, or if the parsed host is a single label other than localhost or an IP. Scoped package paths such as https://registry.npmjs.org/@scope/pkg and oci://reg.io/@scope/chart still pass; the cost is rare legitimate URLs such as https://host:8443/path/with@sign or a versioned path like host/charts/[email protected]. Every @ in the path is checked, not only the first.
  • Accepted residual: a credential containing both . and / followed by a bare, domain-less host (https://a.b/c@gitea/r) is indistinguishable from a legitimate path such as Azure Artifacts' …/feed@Local/… and is kept. Use fully-qualified repository hosts, or credentials in ArgoCD repository secrets rather than in repoURL.

Annotations, labels, conditions, operation state, history, comparedTo and health messages are not read either.

Three tests hold this (agent/internal/k8scollect/argocd_test.go, argocd_allowlist_test.go): a byte scan with planted values (including unparseable credential URLs); a frozen reflect shape of the report type, so a new field fails until a reviewer accepts it; and a fixture that plants a unique marker (matched case-insensitively, so a case-folded copy is caught too) in every string of an Application saturated across the ArgoCD schema (every source type, sync policy, ignoreDifferences, info, operation, comparedTo, history, operationState.syncResult, conditions, summary, metadata), asserting that only the allow-listed source paths reach the report. The URL scrub itself is repourl.Scrub in the shared api/repourl package, with its case table in api/repourl/repourl_test.go; the backend runs it again on ingest.

Limits: at most 1000 Applications per cluster (sorted by namespace and name; the report is flagged applications_truncated beyond that) and 500 workload rows per Application (flagged resources_truncated; the totals still count every row). Applications and resource rows in an excluded namespace are dropped like every other resource; the report counts the dropped Applications in applications_excluded, so a deliberately partial set never reads as complete.

Report status​

Every report carries a status. Only the first two describe the cluster's set of Applications; every other status means unknown this sync, and Console keeps the last known set (with its own timestamp) instead of reading "no Applications".

StatusMeaning
okListed; the set is complete (up to the cap)
crd_absentThe cluster serves no argoproj.io Applications — ArgoCD is not installed. Silent: no log line
forbiddenThe agent's ClusterRole lacks list on applications.argoproj.io (chart older than this feature). Logged once until it clears
errorDiscovery or the list failed for another reason this sync
oversizeThe Applications were read but dropped to fit the NATS message limit (below)
unavailableThe agent could not build the client it reads Applications with

An agent older than v0.8.0 sends no ArgoCD block at all. Console then touches nothing it stored and the routes say agent_too_old, never "no applications".

Size and the NATS message limit​

The report travels inside the single inventory message (NATS max_payload, 8 MB bundled). Measured: 100 Applications × 50 resources ≈ 179 KiB when 10 of each Application's 50 rows are workloads (the rest summarized); the worst case — every row a workload — is ≈ 723 KiB (about 140 bytes per row; the Applications alone ≈ 39 KiB). When an inventory does not fit, the agent drops optional parts one step at a time, retrying after each, before giving up on the inventory:

  1. workload spec facts;
  2. ArgoCD workload rows (resources_omitted: the Applications and their totals stay, per-workload drift is unknown);
  3. the ArgoCD Applications (status oversize: Console keeps the last known set).

Each sync that needed a step is counted in proxima_collector_inventory_too_large_total{outcome="spec_facts_dropped|argocd_resources_dropped|argocd_apps_dropped|inventory_dropped"} under the deepest step it took.

Helm RBAC​

The cluster agent's ClusterRole gains one rule — list, nothing else (no get, no watch, no write):

- apiGroups: ["argoproj.io"]
resources: ["applications"]
verbs: ["list"]

On a cluster without ArgoCD the rule grants nothing. Upgrading only the agent image on an older chart gives forbidden (one warning in the agent log, and a sentence in the UI); run helm upgrade with the new chart. The pc kube read tier is not widened: this rule is on the collector's own ClusterRole only.

Storage​

TableHolds
argocd_reportsOne row per reporting cluster: last status, reported_at (snapshot time), applications_at (snapshot time of the stored set), applications_truncated, resources_omitted
argocd_applicationsThe last known set per reporting cluster (cluster_id, plus its client_id); resources is JSONB

The backend re-applies the agent's caps on ingest (1000 Applications, 500 workload rows and 50 resource namespaces per Application, setting the matching truncated flag), so a compromised agent cannot use one report to store more than an honest one could. It also re-scrubs repo_url and destination.server with the same function the agent uses (repourl.Scrub in the shared api/repourl package), so the agent's scrub is not the only wall: an agent build with a broken scrub still cannot store a credential.

A report is applied in one transaction that locks the cluster row first (FOR NO KEY UPDATE, the lock order of every Kubernetes sync) and only when its snapshot is strictly newer than the stored report, so a redelivered older inventory cannot rewind the Applications. ok and crd_absent replace the set; any other status updates only argocd_reports.

Resources are JSONB rather than a child table: the set is replaced whole each sync, rows are always read with their Application, and the one lookup by resource ("which Application manages Deployment shop/api") is a containment query a jsonb_path_ops GIN index answers.

The counter proxima_k8s_argocd_reports_total{outcome="applied|status_only|not_newer|error"} counts reports by what happened to them.

Destination mapping​

An Application lives in one cluster (the reporting cluster, where ArgoCD runs) and deploys to another or the same. Console maps its spec.destination to a Console cluster when read — never stored, so a cluster added or renamed later maps without a resync:

DestinationMappingCluster
server: https://kubernetes.default.svc or name: in-clusterin_clusterthe reporting cluster
name: <x>, where <x> is one of the reporting cluster's declared in-cluster namesin_clusterthe reporting cluster
name: <x>, and exactly one Console cluster of the same client has slug <x>mappedthat cluster
name: <x>, and two or more clusters of the client have slug <x>ambiguousnone
name: <x> matching no cluster of the client, or a remote server URL onlyunmappednone
  • Never across clients. A cluster of another client with the same slug is never a candidate: the database query is pinned to the reporting cluster's client, and the mapping checks it again.
  • Not restricted to one environment. A hub ArgoCD commonly runs in one environment (a management or prod cluster) and deploys to the client's clusters in other environments; requiring the same environment would leave every hub destination unmapped. Because slugs are unique per environment, not per client, two environments may share a slug — that is ambiguous, and Console does not pick one.
  • Never guessed from a server URL. A remote destination given only as https://10.0.0.5:6443 is unmapped, even if a cluster's API server URL looks similar.

To map a hub's destinations, name each ArgoCD cluster after the Console cluster's slug.

Renamed in-cluster clusters (argocd.inClusterNames)​

ArgoCD lets you rename the cluster it runs in: its built-in cluster is called in-cluster, but an operator can rename it (to production, say), and Applications then target {name: production, server: ""}. Console never reads ArgoCD's cluster secrets — the agent lists Applications and nothing else — so it cannot learn that production is ArgoCD's own cluster. Unless told, every such Application is unmapped: it is still listed on the reporting cluster (labelled "unmapped destination 'production'") but attributed to no cluster, and a workload it lists has drift unknown with reason unmapped_destination.

Declare the name in the cluster agent's Helm values:

argocd:
inClusterNames:
- production

The chart passes it to the agent as PROXIMA_ARGOCD_IN_CLUSTER_NAMES (comma-separated). The agent validates each name (1–253 characters of letters, digits, ., _, -; at most 20; duplicates removed), logs an invalid one once at startup and ignores it, and sends the list with every ArgoCD report (in_cluster_names, whatever the status). The backend validates them again, stores them on the cluster's report row, and returns them as report.in_cluster_names.

  • The destination is not rewritten. The agent reports spec.destination exactly as ArgoCD states it, so the UI still shows production; only the mapping (in_cluster) changes.
  • Only that cluster's own Applications. A declared name is in-cluster for Applications reported by the cluster that declared it. An Application another cluster reports (a hub) with destination production is classified with that cluster's names, and stays unmapped unless it declares production too. Nothing crosses clients.
  • A declared name cannot be a sibling's slug. If a declared name is the slug of another Console cluster of the client, the declaration contradicts the slug and is void for that name: the slug rule decides (an Application to y still maps to cluster y, which keeps its drift, its list and its visibility — a Helm value can never capture another cluster's Applications or un-hide them for a reader of the declaring cluster only). On the declaring cluster such an Application reads ambiguous (listed only when the caller may read y), its workloads' drift is unknown (ambiguous_destination), and a note says "argocd.inClusterNames declares 'y', which is another Console cluster's slug — fix the Helm value". A declared name equal to the cluster's own slug is fine.
  • Current configuration. The names are not history: the latest report's list applies to all of the cluster's stored Applications, and an agent too old to send the field (v0.8.0) clears it — only the defaults then apply.

Needs a cluster agent, chart and backend newer than v0.8.0 (backend first: an older backend ignores the field).

Trust model​

A cluster agent is trusted for its own cluster: what it reports about its own objects is what Console shows for that cluster. A hub's report about another cluster is a different trust edge: an agent in environment X (the hub) states the sync and health shown on a cluster in environment Y, and its client-wide deploys can be attributed there (environment-pinned ones are not; see What changed). Console accepts that — it is what a hub ArgoCD is — but labels it:

  • every Application in both routes carries cross_environment (always set): true when the reporting cluster and the destination cluster are in different environments of the client. It reveals only that another environment of the client exists. The UI labels such rows ("reported by a cluster in another environment"), also when reported_by is withheld because the reader cannot see the hub;
  • the client boundary is never crossed: an agent of one client cannot attach anything to another client's cluster, whatever destination name it reports.

For the UI: repo_url is agent-supplied text. Render it as a link only when its scheme is http or https; anything else (ssh://, scp-style, or a hostile javascript:) is shown as plain text.

Routes​

Both routes use the Kubernetes pages' gate: assets:read in any scope admits; the handler loads the cluster, answers 404 without access to its environment and 403 without assets:read there (3-arg check). Both are read-only and read only what the agents stored. Each has its own per-user budget of 60 reads a minute, counted per backend process (429 rate_limited); the UI polls each at most once a minute per open view.

List a cluster's Applications​

GET /api/v1/clusters/:clusterID/argocd/applications

Applications reported by this cluster, and Applications reported by any cluster of the same client whose destination maps to this cluster. Per-row visibility:

  • An Application mapped to this cluster is listed. Its reporting hub is named in reported_by only when the caller may read the hub (assets:read on the hub's environment).
  • An Application reported here that deploys to another cluster is listed only when the caller may read that cluster; otherwise it is counted in hidden_count and a note says so. Its resources are the other cluster's workloads, which a reader of this environment alone must not learn.
  • An Application reported here whose destination is unmapped or ambiguous is listed: it is this cluster's own object and carries no other cluster's data. Unmapped ones also get a note naming their destinations and pointing at argocd.inClusterNames.
{"data": {
"applications": [{
"id": "…", "name": "checkout", "namespace": "argocd", "project": "shop",
"destination": {"server": "", "name": "prod-eu", "namespace": "shop"},
"mapping": "mapped",
"destination_cluster": {"id": "…", "environment_id": "…", "name": "prod-eu", "slug": "prod-eu"},
"reported_here": false,
"sync_status": "OutOfSync", "health_status": "Degraded",
"target_revision": "main", "synced_revision": "def456",
"repo_url": "https://git.example.com/shop/deploy.git", "path": "apps/checkout", "chart": "", "source_count": 1,
"resource_count": 14, "resources_truncated": false, "out_of_sync_resources": 1, "unhealthy_resources": 1,
"observed_at": "2026-10-01T09:00:00Z"}],
"report": {"state": "crd_absent", "reported_at": "…", "applications_at": "…",
"applications_truncated": false, "resources_omitted": false, "in_cluster_names": []},
"hidden_count": 0,
"notes": ["ArgoCD is not installed on this cluster (it serves no argoproj.io Applications).",
"1 application reported by another cluster of this client (an ArgoCD hub) deploys to this cluster."]}}

report describes this cluster's own agent report: state is the agent status, or agent_too_old (with required_version v0.8.0 and agent_version) or not_reported; in_cluster_names are the names its agent declared (never null). The Overview "Needs attention" list reads sync_status / health_status from here.

A workload's drift​

The workload page's YAML tab renders this answer as its ArgoCD card; see Kubernetes Clusters → YAML tab.

GET /api/v1/clusters/:clusterID/workloads/:workloadID/drift

The workload must belong to the cluster (404 otherwise). The response lists the Applications mapped to this cluster whose resources include the workload — matched by API group (apps; argoproj.io for a Rollout), kind, namespace and name — with the workload's own resource row:

{"data": {
"workload": {"kind": "Deployment", "namespace": "shop", "name": "checkout-api"},
"matches": [{"application": {"name": "checkout", "sync_status": "OutOfSync", "health_status": "Degraded",
"target_revision": "main", "synced_revision": "def456", "…": "…"},
"resource": {"group": "apps", "kind": "Deployment", "namespace": "shop", "name": "checkout-api",
"sync_status": "OutOfSync", "health_status": "Degraded"}}],
"report": {"state": "ok", "…": "…"},
"notes": []}}
  • state is managed (matches found), not_managed, or unknown; unknown_reasons lists every reason "not managed" cannot be concluded, and not_managed is answered only when it is empty:

    ReasonMeaning
    agent_too_old, not_reportedThis cluster's agent sent no ArgoCD report (older than v0.8.0 / not yet)
    report_not_currentThis cluster's latest report is forbidden, error, oversize or unavailable
    resources_omittedIts latest report dropped workload rows to fit the message
    applications_truncatedIt has more Applications than the agent reports
    applications_excludedApplications in excluded namespaces were not reported
    no_applications_visibleNo Application is reported by or mapped to this cluster — whether ArgoCD is installed here with none visible (ok, zero apps) or not installed (crd_absent). Both get this same answer: a hub may target the cluster by a server URL Console cannot map, so "not managed" needs a visible ArgoCD that deploys here and does not list the workload
    ambiguous_destinationAn Application lists the workload and names this cluster's slug, which more than one cluster of the client carries
    unmapped_destinationAn Application reported by this cluster lists the workload, but its destination maps to no Console cluster (a renamed in-cluster cluster not declared in argocd.inClusterNames, a server URL, or a cluster Console does not know). It may well deploy here, so "not managed" cannot be concluded; a note per Application names its destination
    resources_truncatedAn Application deploying here hit the per-app row cap
    hub_report_not_currentA hub that maps Applications here last reported something other than ok (or never)
    hub_resources_omittedSuch a hub's latest report dropped workload rows

    unknown_reasons may be non-empty with managed too: another Application could also claim it.

  • Two or more matches are ArgoCD's shared-resource conflict; a note says so.

  • An Application whose destination maps elsewhere never counts, even if it names a same-named object. An Application reported by this cluster whose destination maps nowhere (unmapped), is ambiguous (whatever name it carries), or is a declared name that conflicts with a sibling's slug never counts as a match either — but if it lists the workload the answer is unknown (unmapped_destination / ambiguous_destination), and if its resource list was truncated it adds resources_truncated; either way, never not_managed. A hub's ambiguous Application naming this cluster's slug is treated the same way.

  • Limitation — hubs. An unmapped Application reported by another cluster (a hub whose ArgoCD calls this cluster by a name that is not its slug, or by a server URL) is not a candidate for this cluster at all: nothing ties it here, so it cannot block not_managed. The no_applications_visible reason covers only the case where nothing at all is reported by or mapped to this cluster. Name a hub's ArgoCD clusters after the Console slugs.

What changed: deploys by Application​

The what-changed timeline attributes ArgoCD deploy notifications (details.app_name) by Application: a deploy whose Application is mapped to the cluster (and, in a namespace or workload view, concerns it) is shown whatever namespace it names, and is not flagged scope: "client"; a deploy of an Application known to map elsewhere is no longer shown merely because its namespace name matches. Only Applications that resolve to a cluster (in_cluster — including by a declared in-cluster name — or mapped) take part, and only when their name is unique among all the client's Applications, resolved or not — a deploy row carries a bare name, so a name two ArgoCD instances share (even if one of them deploys by an unmappable server URL) is not attributed. A deploy row that names an environment must also be from the hub's own environment to be that app's deploy (a hub in environment X whose webhook source is pinned to X therefore never attaches its deploys to a destination cluster in Y).

Deploys from a hub in another environment are not on the destination's History

The timeline reads deploy rows of the cluster's own environment and client-wide rows (a webhook source with no environment) only. So when the hub runs in environment X and the destination cluster is in Y, the hub's environment-pinned deploy notifications never appear on the destination cluster's History — even though the drift card shows that Application, labelled as reported from another environment. Only a hub whose webhook source is client-wide (no environment) has its deploys attributed across environments. This is the conservative choice: a cluster's History keeps to change records of its own environment (plus client-wide ones), the rule every other source follows. ::: Everything else — apps no agent reports, unmapped and ambiguous destinations, shared names — keeps the namespace match and the client-wide hedge.