GitOps promotion gates¶
The operator reaches a verdict on every reconcile and writes it to
status.contractStatus. Neither Flux nor Argo CD reads that field on its own, so
a Kustomization holding a NonCompliant Pacto reports healthy and a violated
contract cannot turn an Argo Application red. Both tools have an extension point
for exactly this. This page is the two snippets that close the gap.
Nothing here changes the operator, the CRD or your contracts. It is configuration you add to the delivery tool you already run.
What the gate buys you¶
The operator observes; it never writes to your workloads. By the time a verdict exists the pods are already serving. What these snippets buy is that the next step stops: the Kustomization goes unready, the dependent one never starts, the Argo Application goes Degraded and whatever alerting you already point at unhealthy resources picks it up.
For catching a breaking change before it lands, the tool is
pacto impact in the pull request. A promotion gate is
the second line, not the first.
Flux¶
Flux decides whether a resource is healthy using kstatus, which recognises three
condition names: Ready, Reconciling and Stalled. Pacto publishes
ContractValid, RuntimeObserved and ReadinessSatisfied, so kstatus discards
all three unread and the object falls through to Current. The gate looks
configured and gates nothing.
spec.healthCheckExprs (kustomize-controller v1.5.0, flux2 v2.5.0 and later)
replaces the guess with a CEL expression:
# Flux promotion gate driven by the Pacto contract verdict.
#
# This file is published verbatim in integrations/kubernetes/docs/gitops.md and
# applied verbatim by tests/acceptance/kind/gitops-flux.sh. Edit it and you edit
# both; the documented snippet cannot drift from the tested one.
---
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: orders-staging
namespace: flux-system
spec:
interval: 1m
# Must exceed the operator's stabilization window (default 2m) plus one
# requeue interval, or a real violation arrives as an ambiguous timeout
# instead of a clean failure.
timeout: 5m
prune: true
wait: true
sourceRef:
kind: OCIRepository
name: orders
path: ./staging
healthCheckExprs:
- apiVersion: pacto.trianalab.io/v1alpha1
kind: Pacto
failed: status.contractStatus in ['NonCompliant', 'Invalid']
current: status.contractStatus in ['Compliant', 'Reference', 'Warning']
---
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: orders-production
namespace: flux-system
spec:
interval: 1m
timeout: 5m
prune: true
wait: true
# Production does not start until staging is healthy, and healthy now
# includes the contract verdict. This line is the promotion gate.
dependsOn:
- name: orders-staging
sourceRef:
kind: OCIRepository
name: orders
path: ./production
sourceRef is whatever you already use — GitRepository, OCIRepository or
Bucket. The gate does not care where the manifests came from, only what the
operator says about them once they are applied.
Four things about that snippet are worth knowing, each checked against
fluxcd/pkg runtime/cel/status_evaluator.go:
- There is no
inProgressexpression, deliberately. When no expression matches, Flux falls through to in-progress. SoUnknown,NotEvaluatedand any verdict added in a later release hold the deploy and time out rather than going falsely green. The gate fails closed. - There is no hand-written generation guard, deliberately. Flux already
compares
status.observedGenerationtometadata.generationbefore any expression runs. Writing your own is redundant, and it throws on an object that has no status yet. Warningsits in the passing set. Move that one word tofailedif you want contract warnings to block promotions. That is the whole knob.timeoutmust exceed the operator's stabilization window. Set it lower and a real violation reaches you as an ambiguous timeout instead of a clean failure. See Timing below.
A Pacto that has just been created has no status at all, and the expression
errors while that is true. Flux treats the error as not-yet-healthy and keeps
polling, so the first few seconds of a fresh apply are noisy in the logs and
harmless in the outcome.
That snippet is not an illustration. tests/acceptance/kind/gitops-flux.sh
applies that exact file to a kind cluster running Flux, publishes a contract
that contradicts the workload that ships, and asserts the dependent
Kustomization's own manifest never reaches the cluster — then corrects the
contract and asserts it does.
HelmRelease gained the same field in helm-controller v1.5.0 / flux2 v2.8.0. The
snippet above should transfer unchanged, but it is not covered here.
Argo CD¶
Argo picks a health check from a fixed list of built-in kinds and returns nothing for everything else; the roll-up into Application health starts at Healthy and ignores that nothing. A Pacto object is therefore not unhealthy to Argo but invisible, and nothing reports the missing check.
A resource health customization in argocd-cm supplies it:
# Argo CD health customization that makes a violated Pacto contract turn its
# Application red.
#
# This file is published verbatim in integrations/kubernetes/docs/gitops.md.
# Apply it as a MERGE PATCH, never as a plain `kubectl apply` -- argocd-cm holds
# Argo's own keys and a full apply would drop them:
#
# kubectl -n argocd patch configmap argocd-cm --type merge \
# --patch-file argocd-cm-pacto-health.yaml
#
# The data key must be typed exactly. Argo silently ignores a key it does not
# recognise, which looks identical to having configured nothing at all.
apiVersion: v1
kind: ConfigMap
metadata:
name: argocd-cm
namespace: argocd
data:
resource.customizations.health.pacto.trianalab.io_Pacto: |
-- Argo runs this with the Lua string library disabled.
-- Only concatenation, comparison, ipairs and tostring() are available.
-- Set to "Degraded" to make contract warnings block a promotion.
local warningHealth = "Healthy"
local hs = {}
if obj.status == nil or obj.status.contractStatus == nil then
hs.status = "Progressing"
hs.message = "waiting for the first contract evaluation"
return hs
end
if obj.status.observedGeneration ~= obj.metadata.generation then
hs.status = "Progressing"
hs.message = "contract not yet evaluated at this generation"
return hs
end
local s = obj.status.contractStatus
if s == "Compliant" or s == "Reference" then
hs.status = "Healthy"
hs.message = "contract satisfied"
return hs
end
if s == "Warning" then
hs.status = warningHealth
hs.message = "contract satisfied with warnings"
return hs
end
if s == "NonCompliant" then
hs.status = "Degraded"
hs.message = "contract violated"
if obj.status.findings ~= nil then
for _, f in ipairs(obj.status.findings) do
if f.severity == "error" then
hs.message = "contract violated: " .. tostring(f.code) .. ": " .. tostring(f.message)
break
end
end
end
return hs
end
if s == "Invalid" then
hs.status = "Degraded"
hs.message = "contract is invalid and could not be evaluated"
return hs
end
-- Unknown, NotEvaluated, and anything added later.
hs.status = "Progressing"
hs.message = "contract verdict pending: " .. tostring(s)
return hs
Apply it as a merge patch: argocd-cm holds Argo's own configuration, and a
plain kubectl apply would drop it:
kubectl -n argocd patch configmap argocd-cm --type merge \
--patch-file argocd-cm-pacto-health.yaml
kubectl -n argocd rollout restart statefulset/argocd-application-controller
The restart is not optional. The application controller reads health
customizations into its resource cache at startup, and that cached verdict is
what it compares to decide whether a changed object needs re-examining. Until it
has the customization, a Pacto's health is cached as nothing, every verdict the
operator writes compares equal to the last, and the Application only catches up
on the next periodic resync. On an install without argocd-server
— Argo's core install — a restart is the only way in: hot reload of argocd-cm
needs server.secretkey, which only argocd-server creates.
- Argo has no built-in generation check, so the script does its own. Be
precise about what that proves:
observedGenerationhere is the Pacto object's own generation, so it says the operator has seen the current contract — not the current workload. - The findings loop puts the reason in the Argo UI instead of a bare red dot.
- Nothing maps to Argo's
Unknown. It ranks worse thanDegradedin the roll-up, so it would mask genuinely broken workloads in the same Application, and it does not fire the on-degraded trigger. Anything unrecognised becomesProgressingand times out. - The Lua runs with the string library disabled. Concatenation, comparison,
ipairsandtostring()are available;string.format,s:gsub()and the rest are not, and reaching for one fails at runtime rather than at load.
That snippet is not an illustration either. tests/acceptance/kind/gitops-argocd.sh
runs that exact file twice. Once with no cluster, through
argocd admin settings resource-overrides health, which puts every contract
status through Argo's Lua sandbox — including states a running cluster passes
through too quickly to catch, like a verdict that has not caught up with the
contract. Then inside a kind cluster running Argo CD, where an Application must
go Degraded naming the finding while the contract is violated and back to
Healthy once corrected.
Check that it took¶
Argo ignores a data key it does not recognise, and an ignored key looks exactly
like having configured nothing. Confirm the customization is live before you rely
on it:
# The key must read back exactly, group and kind included.
kubectl -n argocd get cm argocd-cm \
-o jsonpath='{.data.resource\.customizations\.health\.pacto\.trianalab\.io_Pacto}'
Then ask Argo what it makes of a real object. argocd admin settings evaluates
the customization against files on disk, in the same Lua sandbox the controller
uses, so it answers without waiting for a sync:
kubectl -n argocd get cm argocd-cm -o yaml > /tmp/argocd-cm.yaml
kubectl -n <namespace> get pacto <name> -o yaml > /tmp/pacto.yaml
argocd admin settings resource-overrides health /tmp/pacto.yaml \
--argocd-cm-path /tmp/argocd-cm.yaml
A Compliant Pacto prints STATUS: Healthy and MESSAGE: contract satisfied.
A key that did not take prints Health script is not configured for
'pacto.trianalab.io/Pacto' instead — and prints it while exiting 0, so read
the output rather than the exit code if you wire this into a check.
Both checks read the ConfigMap, not the controller. They will pass on an instance whose controller started before the patch and still has no idea the customization exists. The one that answers that question is the Application itself: it should change health within seconds of a verdict changing, not minutes.
Timing¶
The gate is only as fresh as the verdict behind it, and verdicts do not all land at the same speed.
| Change | When it becomes NonCompliant |
|---|---|
| A mismatch — workload, persistence or configuration conformance | The first reconcile after the workload is observed |
| An absence — a missing interface, capability, dependency, Secret or ConfigMap | After the stabilization window (default two minutes) plus one requeue interval |
That split is why timeout has a floor. A five-minute timeout against the
default two-minute window leaves room for the window, one requeue and the apply
itself.
There is one gap worth naming. kstatus will not call a Deployment current until
its controller writes observedGeneration back, and that same write is the watch
event that queues the Pacto reconcile. So the re-check is guaranteed queued
before Flux can first see the workload as current. It is not guaranteed
finished. The gap is one reconcile.
Every row above is about how fast the verdict lands. When the GitOps tool looks
is a separate question, and only Argo has an answer worth knowing. It re-examines
a Pacto when the health status the customization returns changes — Healthy to
Degraded and back shows up in about a second. A change that lands on the same
status does not: one NonCompliant reason replaced by another is still
Degraded, so the Application keeps showing the old message until the next
periodic resync, two to five minutes out. The red dot is prompt. The wording
behind it is not always.
Limits¶
- Argo health customizations are instance-global. They live in
argocd-cm, so one Argo serving several teams cannot give one team a blockingWarningand another a passing one. - Neither tool can refuse an artifact for what is inside it. Flux's only pre-apply gate is a signature check. "Reject this because its contract breaks three consumers" is not something a health gate can express — that is a pull request check.
- The verdict is about the contract, not the rollout. A Pacto reporting
Compliantsays the running workload matches its declared contract. It says nothing about request errors, saturation or anything else your normal progressive-delivery signals cover. status.lastReconciledAtcannot be used as a freshness gate. Flux hands the expression the custom resource and nothing else; Argo hands the Lua onlyobj. Neither has a clock to compare it against.
Related¶
- Troubleshooting — the events the
operator emits when a verdict changes, and why a
Countabove 1 means the status actually flapped. - Limitations — what the operator declines to judge, and why
those cases read
Unknownrather thanNonCompliant. - CRD reference — the full
statusschema the expressions above read from.