Product experiments
Two-arm, fixed 50/50, application-server experiments with one fixed-horizon result — how your agent defines, launches, stops and reads one, and how your server delivers it.
On this page
A product experiment assigns each identified person to one of two fixed 50/50 arms, your application server registers the first exposure, and a captured event you name up front becomes the trusted binary outcome. The result is computed once, at a fixed horizon: no peeking while the run is live, no automatic rollout, and a minimum sample the platform enforces before drawing any inference.
A refused launch preview returns status: error over MCP, leading with refused: <reason>
and retaining every refusal code. The CLI prints the full verdict and exits 1, including
JSON output. Previewing never launches the experiment.
Your agent defines, previews, launches and stops the experiment through MCP or the CLI, and reads its counts and result. Your own application server delivers it: it assigns people, registers exposures and records outcomes.
Assignment needs a customer identifier Anectico resolved to an identified person. Identify it first. Declaring another person's person ID without that mapping does not enroll them. Attribution is as reliable as your application's own identification of customers; an intake credential does not authenticate an end user.
This is not AI Evaluation. AI Evaluation experiments (under a separate surface) compare prompt/model trials against human or judge labels. Product experiments never reuse feature flags either: an experiment's arm payloads live only in the frozen launched revision, so flag decisions and local flag evaluation cannot allocate or reveal a treatment. If you need to kill a bad experiment immediately, stop it — there is no flag toggle for this.
Ask your agent
"Draft an experiment called
checkout-copythat compares the long and the short checkout copy on all identified people. The outcome isorder_completedwithin one hour. Enroll for two weeks from next Monday. Run the launch preview and tell me every refusal before you launch."
"How is
checkout-copydoing? Give me the counts per arm. If it has finished, give me the result and the decision."
| Job | MCP tool or action | CLI command |
|---|---|---|
| List a project's experiments, or read one | list_experiments, get_experiment |
anectico experiments list, anectico experiments get |
| Create a draft | create_experiment |
anectico experiments create |
| Replace a draft's definition | update_experiment_draft |
anectico experiments update-draft |
| Check launch admission without launching | preview_experiment_launch |
anectico experiments preview-launch |
| Launch (previewed first, then confirmed) | launch_experiment |
anectico experiments launch |
| Stop a draft or a run | stop_experiment |
anectico experiments stop |
| Read the per-arm ledger counts | get_experiment_counts |
anectico experiments counts |
| Read the fixed-horizon result | get_experiment_result |
anectico experiments result |
The reads are MCP read actions: your agent finds them with list_read_actions and runs them with
execute_read_action. The writes are write actions: it finds them with list_write_actions and
runs them with execute_internal_action. Launch and stop are previewed first and applied only with
the returned confirmation token. Management needs experiments:read for every call, plus
experiments:write (drafts) or experiments:launch (launch and stop); over MCP the key also needs
mcp:read (reads) or mcp:write (writes). Delivery uses a different key (see
Delivery).
What your agent gets back
- From
get_experiment: the definition (frozen once launched), the state, the stop reason and the invalidation evidence. - From
get_experiment_counts: counts only, no inference — per arm, how many were assigned and exposed, and the outcome receipts by state. - From
get_experiment_result: descriptive per-arm counts until the run is final; then the immutable snapshot, the rates and, for afinalresult, the difference, its interval and the decision (see Reading a result). - From
preview_experiment_launch: every launch refusal found, without launching anything (see Launch admission).
Open the proof
An experiment has no proof page, and its tools return no link. Your agent reports the definition, the counts and the result in its answer, and cites the tool result.
Lifecycle
| State | Meaning |
|---|---|
draft |
Editable. Assigns nobody. |
running |
Launched; assignment is accepted inside the enrollment window. |
observing |
Enrollment closed; outcomes are still arriving or settling. |
final |
Analyzed once, from an immutable snapshot. |
insufficient_evidence |
Analyzed once; the floor, minimum sample or missing evidence withheld inference. |
canceled |
Stopped by a person. Terminal — never yields a result. |
invalidated |
Required evidence became unusable (identity change of an exposed person, the declared audience became unreadable, or an exposed person was erased). Terminal — never yields a result, and withdraws one that was already final. |
A launched definition is frozen: arms, eligibility, the outcome selector, the enrollment window and the horizon can never change after launch. To change any of them, stop the run and create a new draft.
Define, preview and launch
control_payload and treatment_payload are JSON values. Objects, arrays, strings, numbers,
booleans and null are accepted. A draft is replaced as a whole, at the revision your agent read:
pass it as --expected-revision in the CLI, or as expected_revision to update_experiment_draft.
If someone else changed the draft first, the call is refused with a 409 conflict and saves
nothing. Read the draft again, reapply your edits to the latest revision and retry.
anectico experiments create --key checkout-copy --body '{
"name": "Shorter checkout copy",
"hypothesis": "Shorter copy converts more people",
"control_payload": "long copy",
"treatment_payload": "short copy",
"eligibility": {"kind": "EXPERIMENT_ELIGIBILITY_KIND_ALL_RESOLVED_PEOPLE"},
"primary_outcome": {"event": "order_completed", "source_id": "<server-capture source id>"},
"conversion_horizon_seconds": 3600,
"enrollment_start": "2026-11-01T00:00:00Z",
"enrollment_end": "2026-11-15T00:00:00Z",
"minimum_exposed_per_arm": 100,
"direction": "EXPERIMENT_DIRECTION_INCREASE"
}'
anectico experiments preview-launch checkout-copy --project PROJECT_UUID
anectico experiments launch checkout-copy --expected-revision 1 --yes --project PROJECT_UUID
eligibility.kind is either ALL_RESOLVED_PEOPLE or one exact AUDIENCE_GENERATION (a saved
audience id plus its exact generation — a saved audience is described in
Events, cohorts and audiences). primary_outcome.source_id must be an
active server-capture ingestion source of the project — a browser, mobile or import source is
refused. enrollment_start/enrollment_end are half-open and must span 7 to 30 days;
conversion_horizon_seconds is 1 second to 30 days; minimum_exposed_per_arm is at least 100.
Launch admission
preview-launch evaluates every requirement on actual instants and returns every refusal it
finds, without launching anything:
| Code | Meaning |
|---|---|
EXPERIMENT_HORIZON_LIMIT |
The horizon or enrollment window is outside its bound. |
UNSUPPORTED_EXPERIMENT_MODE |
Eligibility or direction is not a supported shape. |
EXPERIMENT_ENROLLMENT_IN_PAST |
enrollment_start is more than 5 minutes in the past. |
EXPERIMENT_RETENTION_INSUFFICIENT |
An evidence family's retention would expire before the finalization deadline (enrollment_end + horizon + 24h), or the organization has no configured retention policy. |
EXPERIMENT_SOURCE_UNTRUSTED |
The primary outcome's (or guardrail's) source_id is not an active server-capture source. |
EXPERIMENT_AUDIENCE_UNAVAILABLE |
The declared audience generation, or the launcher's own authority over it, is not readable. |
EXPERIMENT_AUDIENCE_EXPIRES_BEFORE_FINALIZATION |
The audience generation expires before the finalization deadline; no replacement generation is ever substituted. |
launch re-evaluates admission at --expected-revision and refuses with the same codes if the
draft still fails; a refusal never launches anything. Launching a revision that has already moved
(edited or already launched) is a 409 conflict — re-read the draft and retry.
Delivery: an APPLICATION key only
Assignment, exposure registration and outcome recording are a separate, server-to-server surface.
They accept only an APPLICATION-purpose API key — never a management token, never a
server_capture intake key — and the project and organization come from that key, never from the
request body:
anectico apikey create --purpose application --scope experiments:deliver --project PROJECT_UUID --name checkout-experiments
The plaintext secret (an_...) is shown once. It is a server-side secret: it must never reach
a browser, a mobile app, or any client-side code. Use it from your own backend, with the Node
helper below or the equivalent in your own server language.
assign → expose → capture → record-outcome
import { experiments, ExperimentDeliveryError } from "@anectico/sdk/node";
import { AnecticoClient } from "@anectico/sdk";
const delivery = experiments({
applicationKey: process.env.ANECTICO_EXPERIMENTS_KEY!, // the application key above
baseUrl: "https://api.example.com",
});
const analytics = new AnecticoClient({
endpoint: "https://api.example.com",
apiKey: process.env.ANECTICO_CAPTURE_API_KEY!, // a server_capture key
});
async function checkout(distinctId: string) {
// 1. Assign. Stable per canonical person: aliases, retries and concurrent
// calls all return the same assignment.
const assignment = await delivery.assign("checkout-copy", distinctId);
if (!assignment.enrolled) return renderDefault();
// 2. Expose, BEFORE any treatment-dependent work. idempotencyKey makes a
// retry after a network failure safe.
const exposure = await delivery.expose(
"checkout-copy",
assignment.assignmentId!,
`checkout-copy:${assignment.assignmentId}`,
);
render(assignment.arm, assignment.payload);
// 3. Capture the qualifying event with captureAndAck — it resolves with the
// exact message_id the platform durably queued. capture() alone cannot
// give you that id to record an outcome against.
const captured = await analytics.captureAndAck("order_completed", { amount: 4200 });
// 4. Record the outcome, binding that captured event to the exposure.
try {
await delivery.recordOutcome(
"checkout-copy",
exposure.exposureId,
"<the primary_outcome.source_id from the definition above>",
captured.messageId,
);
} catch (error) {
if (error instanceof ExperimentDeliveryError && !error.retryable) throw error; // do not retry
// any other failure: retry is safe with the same arguments
}
}
Assignment, exposure and outcome timestamps are all stamped by the platform's own clock at the
first durable write — never by a client timestamp or an ingestion-host clock. Retrying expose or
recordOutcome with the same idempotency key (or the same exposure/event pair) is always safe and
returns the original result (replayed: true).
Two refusals no retry can fix
EXPERIMENT_SUBJECT_UNRESOLVED (assignment) and EXPERIMENT_OUTCOME_CONFLICT (outcome recording)
are never retryable, with the same idempotency key or a new one:
EXPERIMENT_SUBJECT_UNRESOLVED— thedistinct_idyou named does not resolve to an identified person in this project. Identify the person first (see Events, cohorts and audiences); an anonymous visitor is never assigned.EXPERIMENT_OUTCOME_CONFLICT— the captured event you named is already bound to a different exposure of this launched revision. Every captured event can back at most one outcome receipt.
Receipt-timed outcomes and the 24-hour arrival grace
recordOutcome stores a provisional receipt immediately; a background reconciliation pass
settles it against the captured event within (exposed_at, exposed_at + conversion_horizon) —
strictly after exposure, strictly before the horizon ends. An outcome state is never a failure
signal by itself:
| State | Meaning |
|---|---|
provisional |
Recorded; not yet reconciled. |
reconciled |
The captured event matched project, trusted source, the declared event name and the exposed person's identity, strictly inside the timing window. Counts as a success. |
refused |
The captured event contradicts the receipt (wrong timing, wrong person, untrusted origin, or a second success once one is already counted). |
missing |
No qualifying captured event arrived within 24 hours of the receipt. This withholds inference — it is never counted as a failure. |
The 24-hour arrival grace is fixed and cannot be shortened or lengthened per run.
Check that delivery works
While the experiment runs, ask your agent for the counts (get_experiment_counts, or
anectico experiments counts KEY). The assigned count rises after your server calls assign, the
exposed count rises after expose, and each recorded outcome appears as provisional and later as
reconciled. If no qualifying captured event arrives within 24 hours of a receipt, its outcome
becomes missing. Then check that the captured event carries the declared name, the trusted source
and the exposed person's identity.
Identity changes invalidate a run
An experiment's ledger freezes a person's aliases at assignment time as evidence, never as identifiers. If an identity event (a merge, split or correction) later changes an exposed person's identity, the run is invalidated immediately — a treatment effect measured against a person who has since become someone else in the platform's identity graph cannot be trusted. An assignment that was never exposed is only marked; exposing it afterward is refused and then invalidates the run. An identity change that touches nobody assigned to the run changes nothing.
Erasing a person invalidates the runs that exposed them
Erasing a person removes their assignment, their exposure and the
outcomes recorded for them from every experiment. A run that had exposed them counted them, so
it is invalidated with reason PERSON_ERASED, and the invalidation names the erasure operation,
never the person. This is the one cause that reaches a run that is already final: its result
counted someone who is no longer there, so the run becomes invalidated and its result stops being
served. Reading it returns the descriptive per-arm counts, which no longer include the person.
A run that had only assigned the person, and never exposed them, is not affected: the assignment is removed and the run carries on. The person cannot be assigned again.
A run that takes its eligible people from a saved audience is also invalidated if that audience contained the erased person: the audience is a fixed list and is withdrawn with them.
Stop
anectico experiments stop checkout-copy --reason "bad copy" --yes --project PROJECT_UUID
Stopping cancels a draft or a run immediately: assignment, exposure registration and outcome recording all stop at once. A canceled run is terminal and never yields a result. Stopping an already-canceled experiment returns it unchanged; stopping an invalidated or final run is refused — both are already terminal.
Reading a result
anectico experiments result checkout-copy --project PROJECT_UUID
While running or observing, and for a canceled or invalidated run, a result carries
descriptive per-arm counts only — assigned, exposed, successes, missing evidence, refused —
never a rate, a test statistic, an interval or a decision. The REST response may omit
arms[].conversion_rate or return it as null; both mean unavailable, never zero. Peeking at an in-flight run cannot be
trusted the way the one fixed-horizon analysis can.
After the finalization deadline (enrollment_end + conversion_horizon + 24h), the platform
computes the result once, from an immutable ledger snapshot (its digest and per-arm counts are
in the result, for exact reproducibility), and finalizes to one of two states:
final— a valid run that met the floor with no missing outcome evidence. The result carries, per arm, the conversion rate; the Miettinen–Nurminen score-test difference (p_treatment − p_control), its z-statistic, p-value and 95% confidence interval; a sample-ratio check; and a decision.insufficient_evidence— analyzed, but inference is withheld.reasonsnames why:
| Reason | Meaning |
|---|---|
FLOOR_NOT_MET |
Fewer than 100 exposed, or fewer than 5 successes, or fewer than 5 failures, in some arm — the inference floor. |
MINIMUM_SAMPLE_NOT_MET |
The run's own declared minimum_exposed_per_arm (which may ask for more than the floor) was not met. |
MISSING_OUTCOME_EVIDENCE |
At least one exposure's outcome evidence is still missing (see the outcome states above) — the result waits rather than analyzing an incomplete ledger. |
DEGENERATE_TABLE |
Reserved for a 2×2 outcome table with no meaningful test (no successes or all successes in both arms, or complete separation). Any such table also has fewer than 5 successes or failures in some arm, so a run in that state reports FLOOR_NOT_MET first. |
NUMERICAL_NONCONVERGENCE |
The score-test inversion did not converge. |
SAMPLE_RATIO_MISMATCH |
The exposed 50/50 split is statistically unlikely (p < 0.001) — a randomization or delivery bug is suspected, so no decision is made even though inference was otherwise computable. sample_ratio is still reported. |
The decision is a read of the interval, never an automatic rollout
decision is set only on a final result with no sample-ratio mismatch, and reads the 95%
confidence interval of p_treatment − p_control against the experiment's declared direction:
treatment_better— the interval excludes 0 on the side the declared direction prefers.treatment_worse— the interval excludes 0 on the other side.no_detectable_difference— the interval contains 0. This is not evidence of no effect; the floor establishes a minimum sample, not statistical power.
Anectico never ships the treatment automatically, whatever the decision says — reading and acting on it is always a decision your team makes.
Guardrail: descriptive, never decides
An experiment may optionally declare a guardrail selector (a second event to watch, purely
descriptive). Its result reports per-arm exposure and event rates when it can be computed
(status: COMPUTED), or nothing at all when there were too many guardrail events to safely
attribute (status: LOOKUP_LIMIT_EXCEEDED) — a result is still final either way. A guardrail
never affects the decision.
Reads and writes at a glance
| Surface | Reads | Writes |
|---|---|---|
| CLI | anectico experiments list|get|counts|result |
create|update-draft|preview-launch|launch|stop |
| MCP | list_experiments, get_experiment, get_experiment_counts, get_experiment_result, preview_experiment_launch |
create_experiment, update_experiment_draft, launch_experiment (confirm), stop_experiment (confirm) |
| REST | GET .../experiments, GET .../experiments/{key}, GET .../experiments/{key}/counts, GET .../experiments/{key}/result |
POST .../experiments, PUT .../experiments/{key}/draft, POST .../experiments/{key}/launch-preview, POST .../experiments/{key}/launch, POST .../experiments/{key}/stop |
Management requires experiments:read for every call, plus experiments:write (drafts) or
experiments:launch (launch, stop). Delivery requires the application key's own
experiments:deliver scope, which no other credential can ever hold.
See also: Authentication for the application
key purpose, Permissions for the scope list, and the
Node.js SDK for captureAndAck and the experiments() delivery client.