Skip to content
anecticoDocsDashboard
Browse documentation
Reference

Evaluation release gates

Declare release objectives, evaluate admissible evidence, and interpret permanent decisions.

On this page

Start with your first evaluation for a runnable workflow. See API and permissions for endpoint and credential requirements.

Release gates — deciding whether a release may ship

A release gate binds a candidate release to evaluation evidence and produces a pass or a refusal with the reason, kept for ever. It is the thing that turns a comparison into a decision.

A gate is fixed when you create it — its evidence, its objectives, its confidence level. A changed objective is a new gate, because a gate whose bar can move under it is not a gate.

The two kinds of evidence, which are different designs

evidence_kind What it compares Paired?
EXPERIMENT One arm of a declared experiment against that experiment's own declared baseline, on the same frozen cases. Yes — the same case is scored under both configurations.
CANARY Two windows of live traffic that you declare: one before the candidate shipped, one after. No. The runs in the two windows are different runs, so the interval carries the full case-to-case variance.

You do not name the comparison arm for experiment evidence — the experiment's own declared baseline is used. Choosing the arm afterwards is how a flattering pair gets picked.

For a canary, the window boundary is your declaration of when the candidate shipped. Everything else that changed at that boundary is confounded with the release, and every stored canary comparison says so in words alongside its numbers, because a narrow interval and a causal claim look identical on a dashboard.

Objectives

Each objective names an evaluator at an exact version, the metric it must produce, what to compare, and the line to clear:

basis The number compared against the threshold
CANDIDATE_MEAN The candidate's own average. An absolute floor or ceiling; needs no baseline.
DIFFERENCE_MEAN The point estimate of candidate minus baseline.
DIFFERENCE_INTERVAL The conservative bound of the interval around that difference — the low bound for a floor, the high bound for a ceiling. The strong form: the whole interval must clear the line, so a candidate that passes on noise does not.

min_samples is required and zero is refused. A gate that will pass on a single observation is something you have to ask for out loud.

min_judge_agreement is optional and puts a floor on how far the objective's evaluator must agree with human judgments — see Calibration. It is not a way to opt out: an objective bound to a model judge with no calibration at all is inadmissible whatever this field says, and a gate whose only evidence comes from that judge answers INADMISSIBLE rather than passing. Leaving it at zero asks only that a calibration exist.

The rule applies to model judges alone. A deterministic or reference-trajectory evaluator is reproducible from its own definition and is calibrated against nobody, so nothing is asked of it.

Calibration is resolved when the gate is evaluated, not when it is created. A gate written before its judge was calibrated passes once the judge is; a judge whose calibration was superseded by an edited judgment set stops passing until it is recomputed.

An objective is checked against your evaluator registry when the gate is created, including that the evaluator you named actually produces the metric you named — so a gate that could only ever measure nothing is refused before it exists rather than after a release waited on it.

Verdicts: only PASS passes

Verdict What it means Where to look
PASS Every objective was measured on admissible evidence and met.
FAIL At least one objective was measured and missed. The objective's numbers.
PENDING Created, never evaluated. Not a pass. Evaluate it.
NO_EVIDENCE An objective had nothing admissible to measure, or less than it declared it needs. Your sampling — nothing ran.
INADMISSIBLE Results exist for an objective and none of them may be used — incomplete provenance, an execution record that cannot say how the score was produced, a job that did not complete, or a model judge that has not been calibrated. Your evaluator reference — things ran, and cannot be used.
RELEASE_MISMATCH Every objective met, but the candidate runs did not all observably declare the release the gate names. Not a pass. The decision's release_binding: it distinguishes another release, no declared release, partial instrumentation, and an unmeasured subject shape.

The last two are kept apart deliberately: "nothing ran" and "what ran cannot be used" are different problems with different fixes, and one empty answer cannot tell them apart.

Per objective you also get observed_present, which is the only way to tell "the value was 0" from "there was no value". A score of 0.0 is a legitimate measurement, so the number can never be that discriminator — never infer absence from observed_value. inadmissible_samples counts results that exist and may not be used; they are reported rather than filtered away, for the same reason.

Evaluating one

The first call reaches a verdict over the evidence available then and appends one permanent decision. You may supply an idempotency_key; repeating the same gate, caller and key returns that existing decision with replayed: true — even if new results arrived between the calls. If you omit the key, Anectico derives one stable default decision identity from the organization, project, gate and authenticated caller. Repeating the omitted-key call therefore replays that first decision too; it never silently regains at-least-once append behavior. The original decided_at, sample counts and outcomes show which evidence snapshot you received.

Use a new explicit key when you deliberately want to evaluate newly arrived evidence. That creates a new decision beside the first, and the whole history stays readable. Evidence is intentionally not part of the retry identity: otherwise a response-loss retry could silently become a contradictory second decision just because a result landed in between.

A gate never runs an evaluator and never spends your evaluation budget. That is deliberate — a gate that produced its own evidence on demand could never report NO_EVIDENCE, which is the one answer it exists to make reachable. Score first, then gate.

If the evidence store cannot be read, you get an error and nothing is stored. "We could not look" is never recorded as "we looked and found nothing".

A release label is a claim that the evidence must prove

release is the label you declare when creating the gate. Evaluation reconciles it against the release actually carried by every admissible candidate run, deduplicated across objectives; baseline runs are expected to belong to the earlier comparison arm and do not participate in this check.

Every decision stores release_binding with the observed release names, released, unreleased, and unmeasured subject counts, a reason, and one of MATCHED, MISMATCH, UNDECLARED, PARTIAL, UNMEASURED, or NO_EVIDENCE. Only MATCHED may accompany PASS. If the objective numbers meet but the label does not, the verdict is RELEASE_MISMATCH rather than a green decision carrying a warning. If an objective already missed, the verdict remains the more direct FAIL, but the same binding discrepancy is still recorded.

“No release declared” is a first-class answer. A gate evaluated over those runs returns binding UNDECLARED and cannot pass; it does not error, silently discard those runs, or pretend they belong to the gate's label.