Evaluation API and permissions
Look up evaluation endpoints, scopes, and the supported CLI and MCP surface.
On this page
Start with your first evaluation for a runnable workflow.
Permissions
| Scope | Grants |
|---|---|
evals:read |
List and read evaluators, sampling rules, the budget, results, the ordered observed-release series, datasets, experiments — including an experiment's comparison — and release gates with their full decision history. Also replay one frozen case in a sandbox: it stores nothing, and it returns strictly less than reading the case does. |
evals:write |
Create evaluators and versions, manage sampling rules, set the budget, evaluate on demand, create datasets and promote reviewed cases or remove cases, create experiments and their trials, run an experiment, and create a release gate or evaluate one. Evaluating a gate is a write because it permanently records whether a release is admitted or held — a key minted only to look at gates must not be able to create it. |
agents:content:read |
Additionally see the explanation field on results, when the project's content policy also allows it. |
scores:write |
Record or replace your own human judgment of a run. It is a member capability rather than an administrative one, deliberately: rating a run is something anyone reviewing the product does. |
scores:read |
Read the human judgments recorded for runs. |
A project-scoped key can only reach its own project; a request naming another project returns
404 Not Found, the same answer an evaluator that does not exist returns.
Endpoints
| Method | Path |
|---|---|
GET |
/api/v1/evaluation/evaluators |
POST |
/api/v1/evaluation/evaluators |
GET |
/api/v1/evaluation/evaluators/{evaluatorId} |
POST |
/api/v1/evaluation/evaluators/{evaluatorId}/versions |
GET |
/api/v1/evaluation/evaluators/{evaluatorId}/versions/{version} |
GET |
/api/v1/evaluation/sampling-rules |
PUT |
/api/v1/evaluation/sampling-rules |
DELETE |
/api/v1/evaluation/sampling-rules/{ruleId} |
GET |
/api/v1/evaluation/budget |
PUT |
/api/v1/evaluation/budget |
POST |
/api/v1/evaluation/evaluate |
GET |
/api/v1/evaluation/results |
GET |
/api/v1/evaluation/release-series |
POST |
/api/v1/agent-runs/{runId}/annotations |
GET |
/api/v1/agent-runs/{runId}/annotations |
GET |
/api/v1/evaluation/annotations |
GET |
/api/v1/evaluation/datasets |
POST |
/api/v1/evaluation/datasets |
GET |
/api/v1/evaluation/datasets/{datasetId} |
GET |
/api/v1/evaluation/datasets/{datasetId}/versions/{version}/items |
POST |
/api/v1/evaluation/datasets/{datasetId}/promotions |
POST |
/api/v1/evaluation/datasets/{datasetId}/removals |
POST |
/api/v1/evaluation/datasets/{datasetId}/versions/{version}/items/{itemId}/replays |
GET |
/api/v1/evaluation/experiments |
POST |
/api/v1/evaluation/experiments |
GET |
/api/v1/evaluation/experiments/{experimentId} |
POST |
/api/v1/evaluation/experiments/{experimentId}/trials |
POST |
/api/v1/evaluation/experiments/{experimentId}/run |
GET |
/api/v1/evaluation/experiments/{experimentId}/comparison |
GET |
/api/v1/evaluation/gates |
POST |
/api/v1/evaluation/gates |
GET |
/api/v1/evaluation/gates/{gateId} |
POST |
/api/v1/evaluation/gates/{gateId}/evaluations |
Promoting cases and removing them are both POSTs that seal a new version, rather than a PUT
or a DELETE on an existing one — nothing about a sealed version ever changes. Evaluating a gate is
a POST onto /evaluations for the same reason: a new explicit idempotency key appends a decision
rather than editing one. Its optional JSON body is
{ "idempotency_key": "your-stable-decision-key" }; an empty body selects the stable default
decision intent.
Replaying a case is a POST for a different reason: it carries the sandbox plan, not because
anything is written. A replay stores nothing and reproduces its answer every time, which is why
it needs only evals:read.
GET /api/v1/evaluation/gates accepts an optional ?release= to narrow to one release label.
All accept an optional ?project_id= query parameter; a project-scoped key is pinned to its own
project regardless. /budget is the exception: it is org-wide and reads no project.
From the command line, and from an agent
The API is the complete surface. The CLI mirrors it; MCP deliberately exposes a narrower subset.
The CLI mirrors the endpoints one-for-one under anectico evals:
anectico evals datasets list
anectico evals datasets items refunds --version 3
anectico evals experiments comparison exp-refunds
anectico evals release-series --evaluator task_completion --evaluator-version 1 --metric task_completion
anectico evals gates list --release 2026.08.3
anectico evals gates evaluate gate-1 --idempotency-key release-2026-08-26
anectico evals datasets replay refunds case-17 --version 3 --file plan.json
anectico evals annotations add $RUN_ID --metric helpfulness --label good --comment "answered the question"
anectico evals annotations list $RUN_A $RUN_B $RUN_C
anectico evals annotations add omits the number unless you pass --value, so a --label-only
judgment is recorded as "not scored" rather than as a zero. list names every run you give it in one
request.
Structured declarations — an experiment, a gate, a promotion — are supplied as JSON with --body or
--file (--file - reads standard input), and are sent to the API exactly as you wrote them. Reads
print a table on a terminal and JSON when piped; --output json forces it. anectico evals also
covers evaluators, rules, budget, evaluate and results. Run anectico evals --help, or see
the CLI reference.
An MCP agent can read this surface through the read gateway: list_eval_datasets,
get_eval_dataset, list_eval_experiments, get_eval_experiment, compare_eval_experiment,
list_release_gates and get_release_gate, all under evals:read. They are reached through
list_read_actions and execute_read_action rather than appearing as tools of their own.
Three writes are available through list_write_actions and execute_internal_action.
create_release_gate and evaluate_release_gate require evals:write; both preview first because
gates and decisions are permanent and have no delete, void, or correction path. Creation requires a
caller-owned gate_id, so its retry is refused by the unique ID instead of minting a duplicate.
Evaluation through MCP requires idempotency_key; reusing it safely returns the original decision,
while a new key asks for a fresh evidence snapshot. The API and CLI also accept omission and use the
server-derived default intent. Both actions are gatewayed and add no first-class
host tool.
submit_annotation, under scores:write, is also reached through
list_write_actions and execute_internal_action. It previews first and applies on a second call
with the returned confirm_token, because re-judging replaces the caller's previous judgment of that
run. Its receipt reports the metric, the number and the label; it does not repeat the rationale back.
Three things an agent cannot do over MCP, deliberately: it cannot read a dataset version's frozen cases — those can carry recorded content, and the read that returns them is available over the API and the CLI only — it cannot read the ordered release series yet, and it cannot change the evaluation configuration. Creating a dataset, promoting cases, and declaring or running an experiment remain API and CLI operations.