# Product experiments

> Two-arm, fixed 50/50, application-server experiments with one fixed-horizon result — how your agent defines, launches, stops and reads one, and how your server delivers it.

Canonical page: https://anectico.com/docs/investigate/product-experiments/


A product experiment assigns each identified person to one of two fixed 50/50 arms, your
application server registers the first exposure, and a captured event you name up front becomes
the trusted binary outcome. The result is computed once, at a fixed horizon: no peeking while the
run is live, no automatic rollout, and a minimum sample the platform enforces before drawing any
inference.

A refused launch preview returns `status: error` over MCP, leading with `refused: <reason>`
and retaining every refusal code. The CLI prints the full verdict and exits 1, including
JSON output. Previewing never launches the experiment.

Your agent defines, previews, launches and stops the experiment through MCP or the CLI, and reads
its counts and result. Your own application server delivers it: it assigns people, registers
exposures and records outcomes.

Assignment needs a customer identifier Anectico resolved to an identified person.
Identify it first. Declaring another person's person ID without that mapping
does not enroll them. Attribution is as reliable as your application's own
identification of customers; an intake credential does not authenticate an end user.

**This is not AI Evaluation.** AI Evaluation experiments (under a separate surface) compare
prompt/model trials against human or judge labels. Product experiments never reuse feature flags
either: an experiment's arm payloads live only in the frozen launched revision, so flag decisions
and local flag evaluation cannot allocate or reveal a treatment. If you need to kill a bad
experiment immediately, stop it — there is no flag toggle for this.

## Ask your agent

> "Draft an experiment called `checkout-copy` that compares the long and the short checkout copy
> on all identified people. The outcome is `order_completed` within one hour. Enroll for two
> weeks from next Monday. Run the launch preview and tell me every refusal before you launch."

> "How is `checkout-copy` doing? Give me the counts per arm. If it has finished, give me the
> result and the decision."

| Job | MCP tool or action | CLI command |
| --- | --- | --- |
| List a project's experiments, or read one | `list_experiments`, `get_experiment` | `anectico experiments list`, `anectico experiments get` |
| Create a draft | `create_experiment` | `anectico experiments create` |
| Replace a draft's definition | `update_experiment_draft` | `anectico experiments update-draft` |
| Check launch admission without launching | `preview_experiment_launch` | `anectico experiments preview-launch` |
| Launch (previewed first, then confirmed) | `launch_experiment` | `anectico experiments launch` |
| Stop a draft or a run | `stop_experiment` | `anectico experiments stop` |
| Read the per-arm ledger counts | `get_experiment_counts` | `anectico experiments counts` |
| Read the fixed-horizon result | `get_experiment_result` | `anectico experiments result` |

The reads are MCP read actions: your agent finds them with `list_read_actions` and runs them with
`execute_read_action`. The writes are write actions: it finds them with `list_write_actions` and
runs them with `execute_internal_action`. Launch and stop are previewed first and applied only with
the returned confirmation token. Management needs `experiments:read` for every call, plus
`experiments:write` (drafts) or `experiments:launch` (launch and stop); over MCP the key also needs
`mcp:read` (reads) or `mcp:write` (writes). Delivery uses a different key (see
[Delivery](#delivery-an-application-key-only)).

## What your agent gets back

- From `get_experiment`: the definition (frozen once launched), the state, the stop reason and the
  invalidation evidence.
- From `get_experiment_counts`: counts only, no inference — per arm, how many were assigned and
  exposed, and the outcome receipts by state.
- From `get_experiment_result`: descriptive per-arm counts until the run is final; then the
  immutable snapshot, the rates and, for a `final` result, the difference, its interval and the
  decision (see [Reading a result](#reading-a-result)).
- From `preview_experiment_launch`: every launch refusal found, without launching anything (see
  [Launch admission](#launch-admission)).

## Open the proof

An experiment has no proof page, and its tools return no link. Your agent reports the definition,
the counts and the result in its answer, and cites the tool result.

## Lifecycle

| State | Meaning |
| --- | --- |
| `draft` | Editable. Assigns nobody. |
| `running` | Launched; assignment is accepted inside the enrollment window. |
| `observing` | Enrollment closed; outcomes are still arriving or settling. |
| `final` | Analyzed once, from an immutable snapshot. |
| `insufficient_evidence` | Analyzed once; the floor, minimum sample or missing evidence withheld inference. |
| `canceled` | Stopped by a person. Terminal — never yields a result. |
| `invalidated` | Required evidence became unusable (identity change of an exposed person, the declared audience became unreadable, or an exposed person was erased). Terminal — never yields a result, and withdraws one that was already final. |

A launched definition is **frozen**: arms, eligibility, the outcome selector, the enrollment
window and the horizon can never change after launch. To change any of them, stop the run and
create a new draft.

## Define, preview and launch

`control_payload` and `treatment_payload` are JSON values. Objects, arrays, strings, numbers,
booleans and `null` are accepted. A draft is replaced as a whole, at the revision your agent read:
pass it as `--expected-revision` in the CLI, or as `expected_revision` to `update_experiment_draft`.
If someone else changed the draft first, the call is refused with a `409` conflict and saves
nothing. Read the draft again, reapply your edits to the latest revision and retry.

```bash
anectico experiments create --key checkout-copy --body '{
  "name": "Shorter checkout copy",
  "hypothesis": "Shorter copy converts more people",
  "control_payload": "long copy",
  "treatment_payload": "short copy",
  "eligibility": {"kind": "EXPERIMENT_ELIGIBILITY_KIND_ALL_RESOLVED_PEOPLE"},
  "primary_outcome": {"event": "order_completed", "source_id": "<server-capture source id>"},
  "conversion_horizon_seconds": 3600,
  "enrollment_start": "2026-11-01T00:00:00Z",
  "enrollment_end": "2026-11-15T00:00:00Z",
  "minimum_exposed_per_arm": 100,
  "direction": "EXPERIMENT_DIRECTION_INCREASE"
}'

anectico experiments preview-launch checkout-copy --project PROJECT_UUID
anectico experiments launch checkout-copy --expected-revision 1 --yes --project PROJECT_UUID
```

`eligibility.kind` is either `ALL_RESOLVED_PEOPLE` or one exact `AUDIENCE_GENERATION` (a saved
audience id plus its exact generation — a saved audience is described in
[Events, cohorts and audiences](/docs/investigate/events-and-cohorts)). `primary_outcome.source_id` must be an
active **server-capture** ingestion source of the project — a browser, mobile or import source is
refused. `enrollment_start`/`enrollment_end` are half-open and must span 7 to 30 days;
`conversion_horizon_seconds` is 1 second to 30 days; `minimum_exposed_per_arm` is at least 100.

### Launch admission

`preview-launch` evaluates every requirement on actual instants and returns every refusal it
finds, without launching anything:

| Code | Meaning |
| --- | --- |
| `EXPERIMENT_HORIZON_LIMIT` | The horizon or enrollment window is outside its bound. |
| `UNSUPPORTED_EXPERIMENT_MODE` | Eligibility or direction is not a supported shape. |
| `EXPERIMENT_ENROLLMENT_IN_PAST` | `enrollment_start` is more than 5 minutes in the past. |
| `EXPERIMENT_RETENTION_INSUFFICIENT` | An evidence family's retention would expire before the finalization deadline (`enrollment_end + horizon + 24h`), or the organization has no configured retention policy. |
| `EXPERIMENT_SOURCE_UNTRUSTED` | The primary outcome's (or guardrail's) `source_id` is not an active server-capture source. |
| `EXPERIMENT_AUDIENCE_UNAVAILABLE` | The declared audience generation, or the launcher's own authority over it, is not readable. |
| `EXPERIMENT_AUDIENCE_EXPIRES_BEFORE_FINALIZATION` | The audience generation expires before the finalization deadline; no replacement generation is ever substituted. |

`launch` re-evaluates admission at `--expected-revision` and refuses with the same codes if the
draft still fails; a refusal never launches anything. Launching a revision that has already moved
(edited or already launched) is a `409` conflict — re-read the draft and retry.

## Delivery: an APPLICATION key only

Assignment, exposure registration and outcome recording are a separate, server-to-server surface.
They accept **only** an APPLICATION-purpose API key — never a management token, never a
`server_capture` intake key — and the project and organization come from that key, never from the
request body:

```bash
anectico apikey create --purpose application --scope experiments:deliver --project PROJECT_UUID --name checkout-experiments
```

The plaintext secret (`an_...`) is shown once. It is a **server-side secret**: it must never reach
a browser, a mobile app, or any client-side code. Use it from your own backend, with the Node
helper below or the equivalent in your own server language.

### assign → expose → capture → record-outcome

```typescript
import { experiments, ExperimentDeliveryError } from "@anectico/sdk/node";
import { AnecticoClient } from "@anectico/sdk";

const delivery = experiments({
  applicationKey: process.env.ANECTICO_EXPERIMENTS_KEY!, // the application key above
  baseUrl: "https://api.example.com",
});
const analytics = new AnecticoClient({
  endpoint: "https://api.example.com",
  apiKey: process.env.ANECTICO_CAPTURE_API_KEY!, // a server_capture key
});

async function checkout(distinctId: string) {
  // 1. Assign. Stable per canonical person: aliases, retries and concurrent
  //    calls all return the same assignment.
  const assignment = await delivery.assign("checkout-copy", distinctId);
  if (!assignment.enrolled) return renderDefault();

  // 2. Expose, BEFORE any treatment-dependent work. idempotencyKey makes a
  //    retry after a network failure safe.
  const exposure = await delivery.expose(
    "checkout-copy",
    assignment.assignmentId!,
    `checkout-copy:${assignment.assignmentId}`,
  );
  render(assignment.arm, assignment.payload);

  // 3. Capture the qualifying event with captureAndAck — it resolves with the
  //    exact message_id the platform durably queued. capture() alone cannot
  //    give you that id to record an outcome against.
  const captured = await analytics.captureAndAck("order_completed", { amount: 4200 });

  // 4. Record the outcome, binding that captured event to the exposure.
  try {
    await delivery.recordOutcome(
      "checkout-copy",
      exposure.exposureId,
      "<the primary_outcome.source_id from the definition above>",
      captured.messageId,
    );
  } catch (error) {
    if (error instanceof ExperimentDeliveryError && !error.retryable) throw error; // do not retry
    // any other failure: retry is safe with the same arguments
  }
}
```

Assignment, exposure and outcome timestamps are all stamped by the platform's own clock at the
first durable write — never by a client timestamp or an ingestion-host clock. Retrying `expose` or
`recordOutcome` with the same idempotency key (or the same exposure/event pair) is always safe and
returns the original result (`replayed: true`).

### Two refusals no retry can fix

`EXPERIMENT_SUBJECT_UNRESOLVED` (assignment) and `EXPERIMENT_OUTCOME_CONFLICT` (outcome recording)
are never retryable, with the same idempotency key or a new one:

- **`EXPERIMENT_SUBJECT_UNRESOLVED`** — the `distinct_id` you named does not resolve to an
  identified person in this project. Identify the person first (see
  [Events, cohorts and audiences](/docs/investigate/events-and-cohorts)); an anonymous visitor is never assigned.
- **`EXPERIMENT_OUTCOME_CONFLICT`** — the captured event you named is already bound to a
  *different* exposure of this launched revision. Every captured event can back at most one
  outcome receipt.

### Receipt-timed outcomes and the 24-hour arrival grace

`recordOutcome` stores a **provisional** receipt immediately; a background reconciliation pass
settles it against the captured event within (`exposed_at`, `exposed_at + conversion_horizon`) —
strictly after exposure, strictly before the horizon ends. An outcome state is never a failure
signal by itself:

| State | Meaning |
| --- | --- |
| `provisional` | Recorded; not yet reconciled. |
| `reconciled` | The captured event matched project, trusted source, the declared event name and the exposed person's identity, strictly inside the timing window. Counts as a success. |
| `refused` | The captured event contradicts the receipt (wrong timing, wrong person, untrusted origin, or a second success once one is already counted). |
| `missing` | No qualifying captured event arrived within 24 hours of the receipt. **This withholds inference — it is never counted as a failure.** |

The 24-hour arrival grace is fixed and cannot be shortened or lengthened per run.

### Check that delivery works

While the experiment runs, ask your agent for the counts (`get_experiment_counts`, or
`anectico experiments counts KEY`). The assigned count rises after your server calls `assign`, the
exposed count rises after `expose`, and each recorded outcome appears as `provisional` and later as
`reconciled`. If no qualifying captured event arrives within 24 hours of a receipt, its outcome
becomes `missing`. Then check that the captured event carries the declared name, the trusted source
and the exposed person's identity.

## Identity changes invalidate a run

An experiment's ledger freezes a person's aliases at assignment time as evidence, never as
identifiers. If an identity event (a merge, split or correction) later changes an **exposed**
person's identity, the run is invalidated immediately — a treatment effect measured against a
person who has since become someone else in the platform's identity graph cannot be trusted. An
assignment that was never exposed is only marked; exposing it afterward is refused and *then*
invalidates the run. An identity change that touches nobody assigned to the run changes nothing.

## Erasing a person invalidates the runs that exposed them

[Erasing a person](/docs/manage/erase-a-person) removes their assignment, their exposure and the
outcomes recorded for them from every experiment. A run that had **exposed** them counted them, so
it is invalidated with reason `PERSON_ERASED`, and the invalidation names the erasure operation,
never the person. This is the one cause that reaches a run that is already `final`: its result
counted someone who is no longer there, so the run becomes `invalidated` and its result stops being
served. Reading it returns the descriptive per-arm counts, which no longer include the person.

A run that had only **assigned** the person, and never exposed them, is not affected: the
assignment is removed and the run carries on. The person cannot be assigned again.

A run that takes its eligible people from a saved audience is also invalidated if that audience
contained the erased person: the audience is a fixed list and is withdrawn with them.

## Stop

```bash
anectico experiments stop checkout-copy --reason "bad copy" --yes --project PROJECT_UUID
```

Stopping cancels a draft or a run immediately: assignment, exposure registration and outcome
recording all stop at once. A canceled run is terminal and **never** yields a result. Stopping an
already-canceled experiment returns it unchanged; stopping an invalidated or final run is refused
— both are already terminal.

## Reading a result

```bash
anectico experiments result checkout-copy --project PROJECT_UUID
```

While `running` or `observing`, and for a `canceled` or `invalidated` run, a result carries
**descriptive per-arm counts only** — assigned, exposed, successes, missing evidence, refused —
never a rate, a test statistic, an interval or a decision. The REST response may omit
`arms[].conversion_rate` or return it as `null`; both mean unavailable, never zero. Peeking at an in-flight run cannot be
trusted the way the one fixed-horizon analysis can.

After the finalization deadline (`enrollment_end + conversion_horizon + 24h`), the platform
computes the result **once**, from an immutable ledger snapshot (its digest and per-arm counts are
in the result, for exact reproducibility), and finalizes to one of two states:

- **`final`** — a valid run that met the floor with no missing outcome evidence. The result
  carries, per arm, the conversion rate; the Miettinen–Nurminen score-test difference
  (`p_treatment − p_control`), its z-statistic, p-value and 95% confidence interval; a sample-ratio
  check; and a decision.
- **`insufficient_evidence`** — analyzed, but inference is withheld. `reasons` names why:

| Reason | Meaning |
| --- | --- |
| `FLOOR_NOT_MET` | Fewer than 100 exposed, or fewer than 5 successes, or fewer than 5 failures, in some arm — the inference floor. |
| `MINIMUM_SAMPLE_NOT_MET` | The run's own declared `minimum_exposed_per_arm` (which may ask for more than the floor) was not met. |
| `MISSING_OUTCOME_EVIDENCE` | At least one exposure's outcome evidence is still missing (see the outcome states above) — the result waits rather than analyzing an incomplete ledger. |
| `DEGENERATE_TABLE` | Reserved for a 2×2 outcome table with no meaningful test (no successes or all successes in both arms, or complete separation). Any such table also has fewer than 5 successes or failures in some arm, so a run in that state reports `FLOOR_NOT_MET` first. |
| `NUMERICAL_NONCONVERGENCE` | The score-test inversion did not converge. |
| `SAMPLE_RATIO_MISMATCH` | The exposed 50/50 split is statistically unlikely (p < 0.001) — a randomization or delivery bug is suspected, so no decision is made even though inference was otherwise computable. `sample_ratio` is still reported. |

### The decision is a read of the interval, never an automatic rollout

`decision` is set **only** on a `final` result with no sample-ratio mismatch, and reads the 95%
confidence interval of `p_treatment − p_control` against the experiment's declared `direction`:

- **`treatment_better`** — the interval excludes 0 on the side the declared direction prefers.
- **`treatment_worse`** — the interval excludes 0 on the other side.
- **`no_detectable_difference`** — the interval contains 0. This is *not* evidence of no effect;
  the floor establishes a minimum sample, not statistical power.

Anectico never ships the treatment automatically, whatever the decision says — reading and acting
on it is always a decision your team makes.

### Guardrail: descriptive, never decides

An experiment may optionally declare a guardrail selector (a second event to watch, purely
descriptive). Its result reports per-arm exposure and event rates when it can be computed
(`status: COMPUTED`), or nothing at all when there were too many guardrail events to safely
attribute (`status: LOOKUP_LIMIT_EXCEEDED`) — a result is still `final` either way. A guardrail
never affects the decision.

## Reads and writes at a glance

| Surface | Reads | Writes |
| --- | --- | --- |
| CLI | `anectico experiments list\|get\|counts\|result` | `create\|update-draft\|preview-launch\|launch\|stop` |
| MCP | `list_experiments`, `get_experiment`, `get_experiment_counts`, `get_experiment_result`, `preview_experiment_launch` | `create_experiment`, `update_experiment_draft`, `launch_experiment` (confirm), `stop_experiment` (confirm) |
| REST | `GET .../experiments`, `GET .../experiments/{key}`, `GET .../experiments/{key}/counts`, `GET .../experiments/{key}/result` | `POST .../experiments`, `PUT .../experiments/{key}/draft`, `POST .../experiments/{key}/launch-preview`, `POST .../experiments/{key}/launch`, `POST .../experiments/{key}/stop` |

Management requires `experiments:read` for every call, plus `experiments:write` (drafts) or
`experiments:launch` (launch, stop). Delivery requires the application key's own
`experiments:deliver` scope, which no other credential can ever hold.

See also: [Authentication](/docs/reference/authentication#common-key-recipes) for the application
key purpose, [Permissions](/docs/reference/permissions) for the scope list, and the
[Node.js SDK](/docs/reference/javascript-sdk) for `captureAndAck` and the `experiments()` delivery client.
