Stop the policy engine being restarted while it waits for the data-syncer
Problem
The package's own default upgrade and ambient upgrade tests fail before any upgrade is
attempted, on the baseline install of the target branch (anchore-enterprise@4.1.2-bb.0):
Helm install failed for release anchore/anchore-enterprise with chart anchore-enterprise@4.1.2-bb.0:
failed early due to stalled resources: [Deployment/anchore/...-policy status: 'Failed']Cluster events show why:
Warning Unhealthy .../policy Liveness probe failed: HTTP probe failed with statuscode: 500
Normal Killing .../policy Container upstream-policyengine failed liveness probe, will be restartedThe policy engine blocks in preflight until the data-syncer reports ready
(Preflight checks failed with error: Data-syncer service not yet ready) and answers
/health with a 500 for as long as it waits. The upstream chart's global liveness probe
allows initialDelaySeconds: 120 plus failureThreshold: 6 at periodSeconds: 10, so
roughly 180s — the kubelet restarts the container part-way through startup. Each restart
begins the preflight again, the Deployment never goes Available inside its 600s
progressDeadlineSeconds, kstatus reports it as Failed, and Flux abandons the release.
tests/test-values.yaml deliberately sets install.remediation.retries: 0, so there is no
retry and wait-resources runs to its timeout.
Change
Raise upstream.probes.liveness.failureThreshold from 6 to 60, giving 120s + 600s.
Readiness is deliberately untouched, which is the argument for being this generous: a service that is still starting is still pulled out of its Service endpoints within 30s and takes no traffic, so liveness only has to catch a genuine hang.
There is no narrower fix available in the chart. enterprise.common.livenessProbe reads the
global .Values.probes.* for every service, there is no per-service override, and the only
startupProbe in the chart belongs to the cloudsql-proxy sidecar. This is the same in
enterprise 4.1.x and 4.2.0.
Validation
Rendered with helm template; failureThreshold: 60 reaches every Anchore service container
with the remaining liveness keys still inherited from the subchart, and readiness unchanged:
ae-anchore-enterprise-policy -> upstream-policyengine
liveness : {..., 'initialDelaySeconds': 120, 'periodSeconds': 10, 'failureThreshold': 60}
readiness: {..., 'periodSeconds': 10, 'failureThreshold': 3}Not yet exercised against a test pipeline — that is what this MR's run will show.
Residual risk worth flagging: progressDeadlineSeconds is the Kubernetes default of 600s and
the chart does not expose it. Removing the restarts should get the policy engine to Ready well
inside that, but if a run still trips it, the follow-ups are a wait-for-datasyncer
initContainer on upstream.policyEngine (mirroring the existing wait-for-catalog on
dataSyncer in tests/test-values.yaml) or a kustomize postRenderer.
Related
Does not fix the sso upgrade cypress failure seen in !436 — that was already addressed on
main by cdd8886d. Once this merges, rebasing !436 should pick up both.
Complete MR checklist
Assignee
- Followed upgrade instructions outlined in docs/DEVELOPMENT_MAINTENANCE.md
- Update Docs with new/updated steps as needed
- Tested and Validated Changes made with supporting info like logs or screenshots from test pipelines
Add supporting info below
Validated by chart render only (above); pipeline validation pending this MR's run.
Reviewer only
- Tested and Validated changes
Upgrade Notices
N/A — values-only change to the package's liveness probe configuration. No action is required of deployers on upgrade; existing releases pick up the wider threshold on the next reconcile.
Closes #302 (closed)