Skip to content
Console
Browse documentation
Guide

Product experiments

Two-arm, fixed 50/50, application-server experiments with one fixed-horizon result — how your agent defines, launches, stops and reads one, and how your server delivers it.

On this page

A product experiment assigns each identified person to one of two fixed 50/50 arms, your application server registers the first exposure, and a captured event you name up front becomes the trusted binary outcome. The result is computed once, at a fixed horizon: no peeking while the run is live, no automatic rollout, and a minimum sample the platform enforces before drawing any inference.

A refused launch preview returns status: error over MCP, leading with refused: <reason> and retaining every refusal code. The CLI prints the full verdict and exits 1, including JSON output. Previewing never launches the experiment.

Your agent defines, previews, launches and stops the experiment through MCP or the CLI, and reads its counts and result. Your own application server delivers it: it assigns people, registers exposures and records outcomes.

Assignment needs a customer identifier Anectico resolved to an identified person. Identify it first. Declaring another person's person ID without that mapping does not enroll them. Attribution is as reliable as your application's own identification of customers; an intake credential does not authenticate an end user.

This is not AI Evaluation. AI Evaluation experiments (under a separate surface) compare prompt/model trials against human or judge labels. Product experiments never reuse feature flags either: an experiment's arm payloads live only in the frozen launched revision, so flag decisions and local flag evaluation cannot allocate or reveal a treatment. If you need to kill a bad experiment immediately, stop it — there is no flag toggle for this.

Ask your agent

"Draft an experiment called checkout-copy that compares the long and the short checkout copy on all identified people. The outcome is order_completed within one hour. Enroll for two weeks from next Monday. Run the launch preview and tell me every refusal before you launch."

"How is checkout-copy doing? Give me the counts per arm. If it has finished, give me the result and the decision."

Job MCP tool or action CLI command
List a project's experiments, or read one list_experiments, get_experiment anectico experiments list, anectico experiments get
Create a draft create_experiment anectico experiments create
Replace a draft's definition update_experiment_draft anectico experiments update-draft
Check launch admission without launching preview_experiment_launch anectico experiments preview-launch
Launch (previewed first, then confirmed) launch_experiment anectico experiments launch
Stop a draft or a run stop_experiment anectico experiments stop
Read the per-arm ledger counts get_experiment_counts anectico experiments counts
Read the fixed-horizon result get_experiment_result anectico experiments result

The reads are MCP read actions: your agent finds them with list_read_actions and runs them with execute_read_action. The writes are write actions: it finds them with list_write_actions and runs them with execute_internal_action. Launch and stop are previewed first and applied only with the returned confirmation token. Management needs experiments:read for every call, plus experiments:write (drafts) or experiments:launch (launch and stop); over MCP the key also needs mcp:read (reads) or mcp:write (writes). Delivery uses a different key (see Delivery).

What your agent gets back

  • From get_experiment: the definition (frozen once launched), the state, the stop reason and the invalidation evidence.
  • From get_experiment_counts: counts only, no inference — per arm, how many were assigned and exposed, and the outcome receipts by state.
  • From get_experiment_result: descriptive per-arm counts until the run is final; then the immutable snapshot, the rates and, for a final result, the difference, its interval and the decision (see Reading a result).
  • From preview_experiment_launch: every launch refusal found, without launching anything (see Launch admission).

Open the proof

An experiment has no proof page, and its tools return no link. Your agent reports the definition, the counts and the result in its answer, and cites the tool result.

Lifecycle

State Meaning
draft Editable. Assigns nobody.
running Launched; assignment is accepted inside the enrollment window.
observing Enrollment closed; outcomes are still arriving or settling.
final Analyzed once, from an immutable snapshot.
insufficient_evidence Analyzed once; the floor, minimum sample or missing evidence withheld inference.
canceled Stopped by a person. Terminal — never yields a result.
invalidated Required evidence became unusable (identity change of an exposed person, the declared audience became unreadable, or an exposed person was erased). Terminal — never yields a result, and withdraws one that was already final.

A launched definition is frozen: arms, eligibility, the outcome selector, the enrollment window and the horizon can never change after launch. To change any of them, stop the run and create a new draft.

Define, preview and launch

control_payload and treatment_payload are JSON values. Objects, arrays, strings, numbers, booleans and null are accepted. A draft is replaced as a whole, at the revision your agent read: pass it as --expected-revision in the CLI, or as expected_revision to update_experiment_draft. If someone else changed the draft first, the call is refused with a 409 conflict and saves nothing. Read the draft again, reapply your edits to the latest revision and retry.

anectico experiments create --key checkout-copy --body '{
  "name": "Shorter checkout copy",
  "hypothesis": "Shorter copy converts more people",
  "control_payload": "long copy",
  "treatment_payload": "short copy",
  "eligibility": {"kind": "EXPERIMENT_ELIGIBILITY_KIND_ALL_RESOLVED_PEOPLE"},
  "primary_outcome": {"event": "order_completed", "source_id": "<server-capture source id>"},
  "conversion_horizon_seconds": 3600,
  "enrollment_start": "2026-11-01T00:00:00Z",
  "enrollment_end": "2026-11-15T00:00:00Z",
  "minimum_exposed_per_arm": 100,
  "direction": "EXPERIMENT_DIRECTION_INCREASE"
}'

anectico experiments preview-launch checkout-copy --project PROJECT_UUID
anectico experiments launch checkout-copy --expected-revision 1 --yes --project PROJECT_UUID

eligibility.kind is either ALL_RESOLVED_PEOPLE or one exact AUDIENCE_GENERATION (a saved audience id plus its exact generation — a saved audience is described in Events, cohorts and audiences). primary_outcome.source_id must be an active server-capture ingestion source of the project — a browser, mobile or import source is refused. enrollment_start/enrollment_end are half-open and must span 7 to 30 days; conversion_horizon_seconds is 1 second to 30 days; minimum_exposed_per_arm is at least 100.

Launch admission

preview-launch evaluates every requirement on actual instants and returns every refusal it finds, without launching anything:

Code Meaning
EXPERIMENT_HORIZON_LIMIT The horizon or enrollment window is outside its bound.
UNSUPPORTED_EXPERIMENT_MODE Eligibility or direction is not a supported shape.
EXPERIMENT_ENROLLMENT_IN_PAST enrollment_start is more than 5 minutes in the past.
EXPERIMENT_RETENTION_INSUFFICIENT An evidence family's retention would expire before the finalization deadline (enrollment_end + horizon + 24h), or the organization has no configured retention policy.
EXPERIMENT_SOURCE_UNTRUSTED The primary outcome's (or guardrail's) source_id is not an active server-capture source.
EXPERIMENT_AUDIENCE_UNAVAILABLE The declared audience generation, or the launcher's own authority over it, is not readable.
EXPERIMENT_AUDIENCE_EXPIRES_BEFORE_FINALIZATION The audience generation expires before the finalization deadline; no replacement generation is ever substituted.

launch re-evaluates admission at --expected-revision and refuses with the same codes if the draft still fails; a refusal never launches anything. Launching a revision that has already moved (edited or already launched) is a 409 conflict — re-read the draft and retry.

Delivery: an APPLICATION key only

Assignment, exposure registration and outcome recording are a separate, server-to-server surface. They accept only an APPLICATION-purpose API key — never a management token, never a server_capture intake key — and the project and organization come from that key, never from the request body:

anectico apikey create --purpose application --scope experiments:deliver --project PROJECT_UUID --name checkout-experiments

The plaintext secret (an_...) is shown once. It is a server-side secret: it must never reach a browser, a mobile app, or any client-side code. Use it from your own backend, with the Node helper below or the equivalent in your own server language.

assign → expose → capture → record-outcome

import { experiments, ExperimentDeliveryError } from "@anectico/sdk/node";
import { AnecticoClient } from "@anectico/sdk";

const delivery = experiments({
  applicationKey: process.env.ANECTICO_EXPERIMENTS_KEY!, // the application key above
  baseUrl: "https://api.example.com",
});
const analytics = new AnecticoClient({
  endpoint: "https://api.example.com",
  apiKey: process.env.ANECTICO_CAPTURE_API_KEY!, // a server_capture key
});

async function checkout(distinctId: string) {
  // 1. Assign. Stable per canonical person: aliases, retries and concurrent
  //    calls all return the same assignment.
  const assignment = await delivery.assign("checkout-copy", distinctId);
  if (!assignment.enrolled) return renderDefault();

  // 2. Expose, BEFORE any treatment-dependent work. idempotencyKey makes a
  //    retry after a network failure safe.
  const exposure = await delivery.expose(
    "checkout-copy",
    assignment.assignmentId!,
    `checkout-copy:${assignment.assignmentId}`,
  );
  render(assignment.arm, assignment.payload);

  // 3. Capture the qualifying event with captureAndAck — it resolves with the
  //    exact message_id the platform durably queued. capture() alone cannot
  //    give you that id to record an outcome against.
  const captured = await analytics.captureAndAck("order_completed", { amount: 4200 });

  // 4. Record the outcome, binding that captured event to the exposure.
  try {
    await delivery.recordOutcome(
      "checkout-copy",
      exposure.exposureId,
      "<the primary_outcome.source_id from the definition above>",
      captured.messageId,
    );
  } catch (error) {
    if (error instanceof ExperimentDeliveryError && !error.retryable) throw error; // do not retry
    // any other failure: retry is safe with the same arguments
  }
}

Assignment, exposure and outcome timestamps are all stamped by the platform's own clock at the first durable write — never by a client timestamp or an ingestion-host clock. Retrying expose or recordOutcome with the same idempotency key (or the same exposure/event pair) is always safe and returns the original result (replayed: true).

Two refusals no retry can fix

EXPERIMENT_SUBJECT_UNRESOLVED (assignment) and EXPERIMENT_OUTCOME_CONFLICT (outcome recording) are never retryable, with the same idempotency key or a new one:

  • EXPERIMENT_SUBJECT_UNRESOLVED — the distinct_id you named does not resolve to an identified person in this project. Identify the person first (see Events, cohorts and audiences); an anonymous visitor is never assigned.
  • EXPERIMENT_OUTCOME_CONFLICT — the captured event you named is already bound to a different exposure of this launched revision. Every captured event can back at most one outcome receipt.

Receipt-timed outcomes and the 24-hour arrival grace

recordOutcome stores a provisional receipt immediately; a background reconciliation pass settles it against the captured event within (exposed_at, exposed_at + conversion_horizon) — strictly after exposure, strictly before the horizon ends. An outcome state is never a failure signal by itself:

State Meaning
provisional Recorded; not yet reconciled.
reconciled The captured event matched project, trusted source, the declared event name and the exposed person's identity, strictly inside the timing window. Counts as a success.
refused The captured event contradicts the receipt (wrong timing, wrong person, untrusted origin, or a second success once one is already counted).
missing No qualifying captured event arrived within 24 hours of the receipt. This withholds inference — it is never counted as a failure.

The 24-hour arrival grace is fixed and cannot be shortened or lengthened per run.

Check that delivery works

While the experiment runs, ask your agent for the counts (get_experiment_counts, or anectico experiments counts KEY). The assigned count rises after your server calls assign, the exposed count rises after expose, and each recorded outcome appears as provisional and later as reconciled. If no qualifying captured event arrives within 24 hours of a receipt, its outcome becomes missing. Then check that the captured event carries the declared name, the trusted source and the exposed person's identity.

Identity changes invalidate a run

An experiment's ledger freezes a person's aliases at assignment time as evidence, never as identifiers. If an identity event (a merge, split or correction) later changes an exposed person's identity, the run is invalidated immediately — a treatment effect measured against a person who has since become someone else in the platform's identity graph cannot be trusted. An assignment that was never exposed is only marked; exposing it afterward is refused and then invalidates the run. An identity change that touches nobody assigned to the run changes nothing.

Erasing a person invalidates the runs that exposed them

Erasing a person removes their assignment, their exposure and the outcomes recorded for them from every experiment. A run that had exposed them counted them, so it is invalidated with reason PERSON_ERASED, and the invalidation names the erasure operation, never the person. This is the one cause that reaches a run that is already final: its result counted someone who is no longer there, so the run becomes invalidated and its result stops being served. Reading it returns the descriptive per-arm counts, which no longer include the person.

A run that had only assigned the person, and never exposed them, is not affected: the assignment is removed and the run carries on. The person cannot be assigned again.

A run that takes its eligible people from a saved audience is also invalidated if that audience contained the erased person: the audience is a fixed list and is withdrawn with them.

Stop

anectico experiments stop checkout-copy --reason "bad copy" --yes --project PROJECT_UUID

Stopping cancels a draft or a run immediately: assignment, exposure registration and outcome recording all stop at once. A canceled run is terminal and never yields a result. Stopping an already-canceled experiment returns it unchanged; stopping an invalidated or final run is refused — both are already terminal.

Reading a result

anectico experiments result checkout-copy --project PROJECT_UUID

While running or observing, and for a canceled or invalidated run, a result carries descriptive per-arm counts only — assigned, exposed, successes, missing evidence, refused — never a rate, a test statistic, an interval or a decision. The REST response may omit arms[].conversion_rate or return it as null; both mean unavailable, never zero. Peeking at an in-flight run cannot be trusted the way the one fixed-horizon analysis can.

After the finalization deadline (enrollment_end + conversion_horizon + 24h), the platform computes the result once, from an immutable ledger snapshot (its digest and per-arm counts are in the result, for exact reproducibility), and finalizes to one of two states:

  • final — a valid run that met the floor with no missing outcome evidence. The result carries, per arm, the conversion rate; the Miettinen–Nurminen score-test difference (p_treatment − p_control), its z-statistic, p-value and 95% confidence interval; a sample-ratio check; and a decision.
  • insufficient_evidence — analyzed, but inference is withheld. reasons names why:
Reason Meaning
FLOOR_NOT_MET Fewer than 100 exposed, or fewer than 5 successes, or fewer than 5 failures, in some arm — the inference floor.
MINIMUM_SAMPLE_NOT_MET The run's own declared minimum_exposed_per_arm (which may ask for more than the floor) was not met.
MISSING_OUTCOME_EVIDENCE At least one exposure's outcome evidence is still missing (see the outcome states above) — the result waits rather than analyzing an incomplete ledger.
DEGENERATE_TABLE Reserved for a 2×2 outcome table with no meaningful test (no successes or all successes in both arms, or complete separation). Any such table also has fewer than 5 successes or failures in some arm, so a run in that state reports FLOOR_NOT_MET first.
NUMERICAL_NONCONVERGENCE The score-test inversion did not converge.
SAMPLE_RATIO_MISMATCH The exposed 50/50 split is statistically unlikely (p < 0.001) — a randomization or delivery bug is suspected, so no decision is made even though inference was otherwise computable. sample_ratio is still reported.

The decision is a read of the interval, never an automatic rollout

decision is set only on a final result with no sample-ratio mismatch, and reads the 95% confidence interval of p_treatment − p_control against the experiment's declared direction:

  • treatment_better — the interval excludes 0 on the side the declared direction prefers.
  • treatment_worse — the interval excludes 0 on the other side.
  • no_detectable_difference — the interval contains 0. This is not evidence of no effect; the floor establishes a minimum sample, not statistical power.

Anectico never ships the treatment automatically, whatever the decision says — reading and acting on it is always a decision your team makes.

Guardrail: descriptive, never decides

An experiment may optionally declare a guardrail selector (a second event to watch, purely descriptive). Its result reports per-arm exposure and event rates when it can be computed (status: COMPUTED), or nothing at all when there were too many guardrail events to safely attribute (status: LOOKUP_LIMIT_EXCEEDED) — a result is still final either way. A guardrail never affects the decision.

Reads and writes at a glance

Surface Reads Writes
CLI anectico experiments list|get|counts|result create|update-draft|preview-launch|launch|stop
MCP list_experiments, get_experiment, get_experiment_counts, get_experiment_result, preview_experiment_launch create_experiment, update_experiment_draft, launch_experiment (confirm), stop_experiment (confirm)
REST GET .../experiments, GET .../experiments/{key}, GET .../experiments/{key}/counts, GET .../experiments/{key}/result POST .../experiments, PUT .../experiments/{key}/draft, POST .../experiments/{key}/launch-preview, POST .../experiments/{key}/launch, POST .../experiments/{key}/stop

Management requires experiments:read for every call, plus experiments:write (drafts) or experiments:launch (launch, stop). Delivery requires the application key's own experiments:deliver scope, which no other credential can ever hold.

See also: Authentication for the application key purpose, Permissions for the scope list, and the Node.js SDK for captureAndAck and the experiments() delivery client.