# Evaluator configuration

> Configure evaluator kinds, exact versions, subjects, sampling, and paid-evaluation budgets.

Canonical page: https://anectico.com/docs/reference/evaluators/


Start with [your first evaluation](/docs/agents/quality-evaluations) for a runnable workflow.
See [API and permissions](/docs/reference/evaluation-api) for endpoint and credential requirements.


## The three pieces

| Piece | What it is |
|---|---|
| **Evaluator** | A named, versioned rule that turns a run into a number and a label. |
| **Sampling rule** | Which runs an evaluator sees, at what rate, with what per-tick ceiling. |
| **Result** | One score, with the full record of what produced it. |

## Evaluator versions are exact, and never float

An evaluator has an id and a list of **versions**. A version is immutable: changing a rule creates a
new version and leaves the old one exactly as it was.

Every reference must name an existing version explicitly; there is no "latest" alias.
References are validated when written, so a missing version fails the request immediately.
Pin the same version when comparing releases.

## Creating an evaluator

```bash
curl -X POST https://app.anectico.com/api/v1/evaluation/evaluators \
  -H "Authorization: Bearer $ANECTICO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "evaluator_id": "task_completion",
    "display_name": "Task completion",
    "metric_name": "task_completion",
    "kind": "EVALUATOR_KIND_DETERMINISTIC",
    "config_json": "{\"aggregation\":\"all\",\"assertions\":[{\"field\":\"run.status\",\"op\":\"equals\",\"value\":\"AGENT_RUN_STATUS_COMPLETED\"},{\"field\":\"run.error_steps\",\"op\":\"lte\",\"number\":0}]}"
  }'
```

The response carries the evaluator and its version 1. Append a new version with
`POST /api/v1/evaluation/evaluators/{evaluatorId}/versions`; there is no update or patch, by design.

### Evaluator kinds

**`EVALUATOR_KIND_DETERMINISTIC`** — declarative assertions over a run's observable fields.

```json
{
  "aggregation": "all",
  "assertions": [
    {"field": "run.status",       "op": "equals",   "value": "AGENT_RUN_STATUS_COMPLETED"},
    {"field": "run.error_steps",  "op": "lte",      "number": 0},
    {"field": "steps.tool_names", "op": "contains", "value": "search_docs"}
  ]
}
```

`aggregation` is `all` (a gate: every assertion must hold, the score is 1 or 0) or `fraction` (the
share of assertions that held).

| Field | Type | Operators |
|---|---|---|
| `run.status` | one value | `equals`, `not_equals` |
| `run.error_steps`, `run.step_count`, `run.tool_calls`, `run.model_calls`, `run.duration_ms` | number | `lte`, `gte` |
| `steps.outcomes`, `steps.tool_names`, `steps.canonical_operations` | set | `contains`, `not_contains` |
| `session.turn_count`, `session.run_count`, `session.duration_ms`, `session.cost_nanos` | number | `lte`, `gte` |

The `session.*` fields describe a whole conversation and are answerable **only** for
`SUBJECT_KIND_AGENT_SESSION` — see [What you can name as the subject](/docs/reference/evaluators#what-you-can-name-as-the-subject).

A deterministic evaluator can instead return one native measurement. A measurement config contains
`measurement` and no `assertions` or `aggregation`:

```json
{"measurement":"run.cost_nanos"}
```

| Measurement | Unit | When it is measurable |
|---|---|---|
| `run.duration_ms` | milliseconds | The recorded run reports a duration. |
| `run.cost_nanos` | billionths of a US dollar | The run has client-reported or calculated model cost. A missing or unpriced cost is not zero and produces no score. |
| `safety.error_steps` | count | The run contains explicitly identified safety-check steps with reported outcomes. No safety check, or an unknown outcome, is not a safe zero and produces no score. |
| `session.turn_count` | count | The session's turns carry turn ids. A session with none is not a session of nought turns; it produces no score. |
| `session.run_count` | count | Always measurable for a session that exists. |
| `session.duration_ms` | milliseconds | The time between the session's first and last recorded activity. A session with one moment of activity measures `0`, which is a real answer. |
| `session.cost_nanos` | billionths of a US dollar | The session's runs reported model cost. Nothing on a session says whether an unpriced conversation was free, so a zero is treated as unmeasured and produces no score. |

Use the assertion form for **reliability** (for example, completed status and zero error steps), a
reference-trajectory or judge evaluator for **quality**, and the native measurements for **safety,
latency, and cost**. Experiments and release gates treat each as the evaluator's real numeric result;
they do not fill missing dimensions with defaults.

**`EVALUATOR_KIND_REFERENCE_TRAJECTORY`** — did the agent do the right things **in the right order**.

```json
{
  "step_key": "canonical_operation",
  "match": "ordered_subsequence",
  "reference": [
    "CANONICAL_OPERATION_INVOKE_AGENT",
    "CANONICAL_OPERATION_RETRIEVAL",
    "CANONICAL_OPERATION_CHAT"
  ]
}
```

`ordered_subsequence` (the default) scores the longest in-order match divided by the reference
length; extra intervening steps do not reduce the score. `exact` requires an identical sequence.
`step_key` may be `canonical_operation` or `tool_name`.

**`EVALUATOR_KIND_LLM_JUDGE`** — a model scores the run against a rubric you write.

Use a judge for criteria such as groundedness or relevance that cannot be expressed as deterministic
assertions. It reads the selected evidence and scores it against your rubric.

```json
{
  "instructions": "Be strict. Prefer a low score when you are unsure.",
  "criteria": [
    {"name": "grounded",  "description": "every claim in the answer traces to the retrieved context", "weight": 2},
    {"name": "addressed", "description": "the answer resolves the question that was asked"}
  ],
  "evidence": ["run_summary", "model_transcript"],
  "max_evidence_chars": 12000
}
```

| Field | Requirement | Meaning |
|---|---|---|
| `criteria` | required, 1–8 | Each needs a `name` and a `description`. `weight` defaults to 1. |
| `evidence` | required | What the judge is shown. See below. |
| `instructions` | optional | A preamble, rendered ahead of the criteria. |
| `model` | optional | A specific model id. Omit to use the default. |
| `max_evidence_chars` | optional, 500–40000 | Evidence budget. Defaults to 12000. |

Rubric fields are validated; unknown fields are rejected, not ignored.

The result shape is fixed — one score, one label, one sentence of reasoning — and is not something a
rubric sets.

#### What a judge is shown

`evidence` is a declaration, and it is the one that decides what leaves your project:

| Selector | What the judge sees | Content |
|---|---|---|
| `run_summary` | Status, step counts, duration. | — |
| `step_sequence` | The ordered operations, outcomes and tool **names**. | — |
| `model_transcript` | Recorded prompt and response bodies. | `model_transcript` |
| `tool_calls` | Recorded tool and MCP call arguments and results. | `tool_arguments` |

The first two carry no message bodies at all — a tool *name* is a symbol, not the arguments it was
called with — so a rubric built only from them shows a model nothing you recorded.

#### Your content policy decides whether the evidence may be sent

Sending recorded content to an evaluation model is a transfer, and it is governed by your project's
[content policy](/docs/manage/content-policy) at the `judge_transfer` boundary. The policy is
consulted **before** anything is sent, never after.

Two properties are worth stating exactly, because both are easy to get wrong and neither is visible
from the outside:

- **The question asked of your policy is derived from what the evidence actually contains**, not from
  what the rubric might want. A rubric that selects `tool_calls`, running against a run that recorded
  no tool arguments, asks nothing about `tool_arguments` — so a project that permits transcripts and
  refuses tool arguments is evaluated normally on those runs.
- **But a kind of content your policy actually withheld is never quietly dropped.** When the content
  your rubric asked for was removed by the policy, the transfer is put to the policy again for
  exactly that kind, and a refusal is a refusal for the whole evaluation — `403 Forbidden` — even
  when the rubric's other selectors still had something to show. A judge scoring "was the answer
  grounded in the transcript" over a run summary alone is scoring something you did not ask about.
  If the removal was not your policy's decision (a narrower credential, for instance), the
  evaluation proceeds and the model is **told** which kinds of content were withheld, by name, so it
  cannot read an absence as evidence that nothing was there.
- **A refusal is terminal and it is recorded.** The run is not sent, is never scored 0, and is not
  retried in a loop. It is recorded as withheld, naming the kind of content, the boundary, and which
  of three things decided: your policy said so, the project has no policy yet, or the policy could
  not be read. "Nobody judged this" and "nobody sampled this" stay different facts.

The **read** and `judge_transfer` cells are independent for automatic quality checks. A run may be
hidden from product reads and still evaluated when `judge_transfer` allows its content; denying
`judge_transfer` still prevents the transfer whatever the read cell says.

## Why some rules are refused

Evaluator configuration is validated when you write it, and some rules are refused even though they
are well-formed. Each refusal names its reason. The reason is always the same one:

> **A field that nothing fills is not a field you can measure.**

A run's operation, outcome and similar values carry a placeholder when your instrumentation did not
report them. A rule written against that placeholder does not measure your agent — it measures how
completely you are instrumented, and reports the answer as a quality score. Worse, a trajectory
written against it scores *higher* the less instrumented the run is, inverting the metric entirely.

So you cannot assert on `OUTCOME_UNKNOWN`, on `CANONICAL_OPERATION_UNKNOWN`, or on fields no
instrumentation populates. If you need a value that is refused, the fix is to emit it — see
[Instrumenting LLM and agent calls](/docs/instrument/llm-calls).

The same principle applies at run time. If a run carries nothing the evaluator can measure, it is
recorded as **unmeasured** and retried later, never scored zero. A zero would say your agent failed;
the truth is that nobody measured it. A judge is held to the same rule twice over: a run carrying
none of the evidence its rubric selected is not sent to a model at all, and a model that comes back
with something unusable is retried rather than recorded as a failing score.

## Sampling production traffic

```bash
curl -X PUT https://app.anectico.com/api/v1/evaluation/sampling-rules \
  -H "Authorization: Bearer $ANECTICO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "evaluator": {"evaluator_id": "task_completion", "version": 1},
    "subject_kind": "SUBJECT_KIND_AGENT_RUN",
    "sampling_rate": 0.1,
    "per_tick_limit": 50,
    "enabled": true
  }'
```

- **`sampling_rate`** (0–1) is a **true cost cap**. Selection is derived from a stable hash of the
  run, so the same run gets the same answer every time and the unselected share is never evaluated at
  all. It is not a per-cycle coin flip, which would eventually select everything.
- **`per_tick_limit`** caps how many evaluations start per cycle. Zero disables the rule, as does
  `enabled: false` or a rate of 0.

A rule is evaluated only where you create one. Nothing is evaluated by default.

## Evaluating one run on demand

```bash
curl -X POST https://app.anectico.com/api/v1/evaluation/evaluate \
  -H "Authorization: Bearer $ANECTICO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "evaluator": {"evaluator_id": "task_completion", "version": 1},
    "subject_kind": "SUBJECT_KIND_AGENT_RUN",
    "subject_ref": "0af7651916cd43dd8448eb211c80319c"
  }'
```

This is the same engine, with sampling skipped — naming a run *is* selecting it. Asking twice for the
same run and evaluator version returns the stored answer with `already_evaluated: true`; it does not
re-run and does not spend again.

### Set your client's timeout above 30 seconds

The call is **synchronous**: it reads the subject, and for an LLM-judge evaluator it waits for the
model. A judge call routinely takes 10–30 seconds and is allowed up to five minutes. Clients whose
default request timeout is 30 seconds — which is most of them — will hang up on a judge that is still
working, and the evaluation is then abandoned rather than stored.

Give this one endpoint its own timeout of **six minutes** — a little above the five the server
allows, so that what you read is the server's answer rather than your own client giving up — or use
a deterministic evaluator (which answers in under a second) when you need a fast answer.

If the five minutes do run out, you get **`504 Gateway Timeout`** with a message, not a dropped
connection. The `anectico` CLI already applies its own longer timeout to `anectico evals evaluate`
and `anectico evals experiments run`; every other command keeps the short default.

### What you can name as the subject

`subject_kind` says what `subject_ref` is, and it is also the label the stored result carries. Three
kinds can be evaluated today:

| `subject_kind` | `subject_ref` is | Notes |
|---|---|---|
| `SUBJECT_KIND_AGENT_RUN` | the run (trace) id | The whole run. This is what sampling rules select. |
| `SUBJECT_KIND_AGENT_STEP` | `"<run_id>/<span_id>"` | One step of a run — the `span_id` is the one the run graph reports for that step. The score is about that step: the evaluator sees only it, not its siblings. |
| `SUBJECT_KIND_AGENT_SESSION` | the session key | The whole conversation: every run that shares that session key. Only the `session.*` fields can score it, and only a rubric that asks for **no recorded content** can be run over it — see below. |

`SUBJECT_KIND_AGENT_TURN` is part of the result vocabulary — a stored result can carry it — but
cannot be submitted here: a turn cannot be looked up from its id alone. `SUBJECT_KIND_DATASET_ITEM`
is a batch subject and is evaluated by running an experiment over a dataset version, not through this
endpoint. Naming either returns **`501 Not Implemented`**. A result is never labelled with a subject
that was not read: an unsupported kind is refused rather than scored from the run behind it.

Which fields an evaluator can assert on depends on what you named, and each family is answerable for
its own subject only:

- **`run.*`** describe the whole run and are answerable for `SUBJECT_KIND_AGENT_RUN` and for dataset
  items. Asking one about a **step** or a **session** returns `400 Bad Request` ("carries nothing
  this evaluator can measure") rather than reporting the run's number under another subject's name.
- **`steps.*`** and `safety.error_steps` are answerable for a run and for a step, and for a step
  subject they describe that step alone.
- **`session.*`** describe a whole conversation and are answerable **only** for
  `SUBJECT_KIND_AGENT_SESSION`. Asking one about a single **run** of that session returns
  `400 Bad Request` for the same reason in the other direction — a run's own duration or cost is not
  the conversation's.

### What a judge reads for each subject

A judge that asked for `model_transcript` or `tool_calls` has to be shown recorded content, and what
that content **is** depends on the subject:

- **A run** — that run's own recorded content.
- **A step** — the content of the run the step belongs to, with the step itself as the subject the
  score is filed under.
- **A session** — nothing. A conversation is a set of runs and has no transcript of its own, and
  there is no answer to "show the judge the session" that would not mean choosing runs on your
  behalf. Rather than pick one, a content-asking rubric over a session is refused with
  `400 Bad Request` naming the kind. A rubric built only from `run_summary` and `step_sequence` is
  not refused for content — it reads none — but a session rollup carries neither a run summary nor a
  step sequence, so it has nothing to show a judge either and answers "carries nothing this
  evaluator can measure". Score a conversation with the `session.*` fields, which read the rollup
  itself.

| Response | Meaning |
|---|---|
| `404 Not Found` | Nothing here matches what you named — the run, the step, or the evaluator version. Look the identifier up again; retrying it unchanged will not help. A run that was just produced may not be readable for a few seconds, and the same call succeeds once it is. |
| `504 Gateway Timeout` | Either the subject could not be read within its own budget, or the whole evaluation ran past the five minutes the call is allowed. **This is not a `404`**: it says the subject could not be *read*, never that it does not exist, so retry the identifier you already have rather than looking for a new one. No result is stored. A judge that had already started may still have been charged, and the retry is free once a verdict exists — asking again for the same run and evaluator version returns the stored answer. |
| `501 Not Implemented` | This `subject_kind` cannot be evaluated on demand. See the table above; retrying will not help. |
| `400 Bad Request` | The `subject_ref` cannot name a subject of that kind — a step reference must be `"<run_id>/<span_id>"`. |
| `409 Conflict` | Already being evaluated. Retry shortly. |
| `429 Too Many Requests` | The daily evaluation budget is spent. |
| `400 Bad Request` | The run carries nothing this evaluator can measure. |
| `403 Forbidden` | Your content policy withholds content your rubric asked for from a judge. The message names the kind of content and the boundary, never the content itself. |
| `400 Bad Request` | A judge whose rubric asks for recorded content cannot be run over this `subject_kind`. Today that is a **session**: a conversation has no single transcript to show, and the answer names the kind rather than reading one of its runs. Score the session with `session.*` fields, or judge its runs individually. |
| `400 Bad Request` | Judging is not configured or its credentials were permanently refused. An operator must repair the judge setup before retrying. |
| `503 Service Unavailable` | Judging is temporarily unavailable because of a timeout, throttle, transport failure, or upstream outage. Retry is appropriate. |

## The spend budget

Evaluators that call a model cost money, so the daily ceiling is a **hard** limit rather than a
throttle: it is enforced transactionally, and it does not open up when anything is degraded.

```bash
curl -X PUT https://app.anectico.com/api/v1/evaluation/budget \
  -H "Authorization: Bearer $ANECTICO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"daily_paid_evaluation_cap": 500}'
```

`GET /api/v1/evaluation/budget` returns the cap and today's consumption.

Two things to know:

- **An organization with no configured budget runs no paid evaluations.** The safe state is the
  default; you opt in to spending. A judge in an organization with no budget never runs.
- **Evaluators that do not call a model never consume budget.** `EVALUATOR_KIND_DETERMINISTIC` and
  `EVALUATOR_KIND_REFERENCE_TRAJECTORY` are in that category, so they keep working after a paid
  budget is exhausted. `EVALUATOR_KIND_LLM_JUDGE` is not: it is charged one unit per run evaluated,
  once — a retried delivery of a score that was already produced never charges again.

The budget is charged **after** your content policy permits the transfer, so a run your policy
withholds costs nothing.
