UNCLASSIFIED - NO CUI

Stop the policy engine being restarted while it waits for the data-syncer

Problem

The package's own default upgrade and ambient upgrade tests fail before any upgrade is attempted, on the baseline install of the target branch (anchore-enterprise@4.1.2-bb.0):

Helm install failed for release anchore/anchore-enterprise with chart anchore-enterprise@4.1.2-bb.0:
  failed early due to stalled resources: [Deployment/anchore/...-policy status: 'Failed']

Cluster events show why:

Warning  Unhealthy  .../policy   Liveness probe failed: HTTP probe failed with statuscode: 500
Normal   Killing    .../policy   Container upstream-policyengine failed liveness probe, will be restarted

The policy engine blocks in preflight until the data-syncer reports ready (Preflight checks failed with error: Data-syncer service not yet ready) and answers /health with a 500 for as long as it waits. The upstream chart's global liveness probe allows initialDelaySeconds: 120 plus failureThreshold: 6 at periodSeconds: 10, so roughly 180s — the kubelet restarts the container part-way through startup. Each restart begins the preflight again, the Deployment never goes Available inside its 600s progressDeadlineSeconds, kstatus reports it as Failed, and Flux abandons the release. tests/test-values.yaml deliberately sets install.remediation.retries: 0, so there is no retry and wait-resources runs to its timeout.

Change

Raise upstream.probes.liveness.failureThreshold from 6 to 60, giving 120s + 600s.

Readiness is deliberately untouched, which is the argument for being this generous: a service that is still starting is still pulled out of its Service endpoints within 30s and takes no traffic, so liveness only has to catch a genuine hang.

There is no narrower fix available in the chart. enterprise.common.livenessProbe reads the global .Values.probes.* for every service, there is no per-service override, and the only startupProbe in the chart belongs to the cloudsql-proxy sidecar. This is the same in enterprise 4.1.x and 4.2.0.

Validation

Rendered with helm template; failureThreshold: 60 reaches every Anchore service container with the remaining liveness keys still inherited from the subchart, and readiness unchanged:

ae-anchore-enterprise-policy -> upstream-policyengine
  liveness : {..., 'initialDelaySeconds': 120, 'periodSeconds': 10, 'failureThreshold': 60}
  readiness: {..., 'periodSeconds': 10, 'failureThreshold': 3}

Not yet exercised against a test pipeline — that is what this MR's run will show.

Residual risk worth flagging: progressDeadlineSeconds is the Kubernetes default of 600s and the chart does not expose it. Removing the restarts should get the policy engine to Ready well inside that, but if a run still trips it, the follow-ups are a wait-for-datasyncer initContainer on upstream.policyEngine (mirroring the existing wait-for-catalog on dataSyncer in tests/test-values.yaml) or a kustomize postRenderer.

Does not fix the sso upgrade cypress failure seen in !436 — that was already addressed on main by cdd8886d. Once this merges, rebasing !436 should pick up both.


Complete MR checklist

Assignee

  • Followed upgrade instructions outlined in docs/DEVELOPMENT_MAINTENANCE.md
  • Update Docs with new/updated steps as needed
  • Tested and Validated Changes made with supporting info like logs or screenshots from test pipelines

Add supporting info below

Validated by chart render only (above); pipeline validation pending this MR's run.

Reviewer only

  • Tested and Validated changes

Upgrade Notices

N/A — values-only change to the package's liveness probe configuration. No action is required of deployers on upgrade; existing releases pick up the wider threshold on the next reconcile.


Closes #302 (closed)

Edited by Ben Lang

Merge request reports

Loading
Loading