# Ingestion sources and pipelines

> Manage intake sources, bind keys to them, and validate, activate or roll back a native-intake source's transform pipeline.

Canonical page: https://anectico.com/docs/manage/ingestion-pipelines/


An **ingestion source** is the intake credential surface a native-capture or OTLP API key binds to.

Your agent creates and manages sources and pipelines through MCP or the CLI; the Console has no page
for them. What the browser SDK may capture is a privacy control, so it is in the Console: open
**Privacy** (`/privacy`) and choose the **Browser capture** tab. See
[Browser capture settings](#browser-capture-settings). A source's **purpose** — `browser_capture`,
`mobile_capture`, `server_capture` or `native_ingest` — and its project are fixed for the source's
whole life; nothing can change either after creation. Every telemetry-sending key names the source
it binds to (`source_id` when the key is created — see
[authentication and API keys](/docs/reference/authentication)), never the other way around: a source
does not list "its" keys, because a source outlives any single key bound to it.

An invalid validation or program preview returns `status: error` over MCP and a summary
starting `invalid: <reason>`. The CLI prints the full diagnostics and exits 1, including
JSON output. Individual sample refusals in a completed preview are separate outcomes.
An invalid program cannot receive a confirmation token to create a revision.

## Ask your agent

> "Create a native ingestion source called `collector`, then check this pipeline program against my
> sample log lines. Do not activate it yet."

> "What did the active pipeline on `collector` drop or refuse in the last 48 hours?"

| Job | MCP tool or action | CLI command |
| --- | --- | --- |
| Create, list, read, rename or revoke a source | `create_ingestion_source`, `list_ingestion_sources`, `get_ingestion_source`, `update_ingestion_source`, `revoke_ingestion_source` | `anectico sources create`, `list`, `update <source-id>`, `revoke <source-id> --yes` |
| Check a pipeline program | `validate_pipeline_program` | `anectico sources pipeline validate <source-id>` |
| Preview a program on sample records | `preview_pipeline` | `anectico sources pipeline preview <source-id>` |
| Store, list or read a pipeline revision | `create_pipeline_revision`, `list_pipeline_revisions`, `get_pipeline_revision` | `anectico sources pipeline create <source-id>`, `anectico sources pipeline list <source-id>` |
| Activate or roll back a revision | `activate_pipeline_revision` | `anectico sources pipeline activate <source-id> <revision>` |
| Read how the active pipeline behaves | `get_pipeline_diagnostics` | `anectico sources pipeline diagnostics <source-id>` |
| Read an activation operation | `list_source_operations`, `get_source_operation` | `anectico sources operations list <source-id>` |
| Read or change the browser capture settings | `get_capture_settings`, `update_capture_settings` | `anectico capture-settings get`, `set`, `enable-autocapture`, `disable-autocapture` |

Reading needs `ingestion:pipelines:read`; creating, renaming, revoking, storing, activating and
changing capture settings also need `ingestion:pipelines:write`. The MCP write tools preview first
and apply on a second call with a `confirm_token`; see [MCP tools](/docs/reference/mcp-tools).

## What your agent gets back

A source read returns its purpose, project, state, revision and catalog epoch. It never returns a key's
secret. A pipeline validation or preview returns diagnostics and the output records intake would
store, and it saves no program. Diagnostics return counts only. A capture-settings read returns the
settings and their revision. None of these results has a proof page.

## Default sources

Every project starts with one stable **default** source per purpose, created automatically the
first time a key is minted for that purpose without an explicit `source_id`. You do not create it
yourself, and its list entry is marked as the default. Revoking a default source is permanent in a way
explicit sources are not: the project's default for that purpose is never recreated, so every future
key that would have used it implicitly must instead name an explicitly created source with
`source_id`.

## Create, rename and revoke

Creating a source asks for a purpose and a display name (description is optional) and returns the
new source immediately — creation does not affect any existing key. Renaming replaces the name and
description together and requires the exact revision you loaded, so a save that raced a concurrent
edit is refused rather than silently overwriting it; reload and retry.

**Revoke** is immediate and permanent — there is no un-revoke. The moment a source is revoked, every
key bound to it stops authenticating, independent of that key's own expiry or scopes. The refusal
answer is explicit and identical in shape across every transport a key can use:

| Transport | Refusal |
| --- | --- |
| Product capture (`POST /api/v1/capture`) | `403 forbidden`, header `X-Anectico-Refusal-Reason: INGRESS_SOURCE_REVOKED` |
| OTLP over HTTP | `403`, message `intake source revoked` |
| OTLP over gRPC | `PermissionDenied`, message `intake source revoked` |

Revocation never deletes anything already ingested; it only stops future admission through that
source's keys.

## Pipelines (native-intake sources only)

A **pipeline** is an optional transform program attached to a `native_ingest` source; other purposes
do not support one at all (`pipelines-unsupported`). It rewrites a small, fixed set of log-record
fields on the way in — `body`, `severity_text`, and any `attributes.<key>` or `resource.<key>` value
— and cannot touch record identity, timestamps, trace context, or any platform-reserved
`anectico.*` key, whether reading or writing.

The reserved prefix is checked without regard to case in both attributes and resource fields.
This includes the source of `extract`, `parse_json` and `parse_key_value`, and every `when.field`
condition. Naming a reserved key returns `PROTECTED_FIELD`: the program can only be saved as a
draft, and preview returns `INVALID_PROGRAM` without running any samples. Reserved keys already
present on a record are preserved when a pipeline transforms other fields.

**Pipeline revisions are immutable and versioned.** Creating one from a program document never edits
an earlier revision; it either stores a new one or, if the document is byte-for-byte identical to an
existing revision, returns that revision unchanged. A revision that fails structural checks is still
stored, as a **draft** you cannot activate — its `diagnostics` name exactly what to fix — while a
structurally sound one is **validated** and can be activated. Only one revision is ever **active** at
a time.

### The program document (schema version 1)

```json
{
  "schema_version": 1,
  "steps": [
    { "op": "redact", "field": "attributes.card_number" },
    { "op": "rename", "from": "attributes.msg", "to": "attributes.message" }
  ]
}
```

At most 32 steps, nested at most 8 levels deep, in a document of at most 64 KiB. Each step is one
operator plus its operands, and an optional `when` condition (`field`, plus exactly one of `equals`,
`matches` — an RE2 pattern of at most 256 bytes — or `exists`) that skips the step when it is false:

| Operator | Operands | Effect |
| --- | --- | --- |
| `parse_json` | `from`, `into` | Parse a field as JSON into an attribute/resource container |
| `parse_key_value` | `from`, `into`, optional `pair_separator`, `value_separator` | Parse `key=value` pairs into a container |
| `rename` | `from`, `to` | Move a field to a new key |
| `remove` | `field` | Delete a field |
| `extract` | `from`, `pattern`, `into`, `type` (`string`, `int`, `float`, `bool` or `timestamp`) | Pull a regex capture into a typed field |
| `set` | `field`, `value` | Write a fixed string, number or boolean |
| `redact` | `field`, optional `pattern` | Replace a field's value; an optional pattern limits which occurrences |
| `drop` | `when` | Discard the record entirely when the condition holds |

Field keys are at most 128 bytes and a `set` value at most 1024 bytes. A program that exceeds any of
these limits, uses an unknown operator, or names a protected field returns the corresponding
diagnostic code (`TRANSFORM_PROGRAM_LIMIT`, `UNKNOWN_OPERATOR`, `PROTECTED_FIELD`, `INVALID_FIELD`,
`INVALID_REGEX`, `INVALID_PROGRAM` or `UNSUPPORTED_SCHEMA_VERSION`) rather than an activatable draft.

### Processing order

A record is redacted for known-sensitive content twice: once on the way in, before any pipeline step
runs, and once more after every step has run, over anything a step wrote or moved. The first pass
already replaces the value of a quoted credential member such as `{"password": "…"}` inside a log
line. The second pass protects what only a step can create: a value moved under a sensitive name
(a `rename` onto `api_key`), or content lifted out of text in a shape the first pass does not
recognize. There is no way to write a pipeline that skips the second pass.

### Drop vs. refuse

A pipeline can remove a record from storage in two different ways, and they are reported
differently:

- **`drop`** is an explicit, intentional step: the program said to discard records matching a
  condition. A dropped record is still acknowledged to the sender — it is not an error — but is
  counted and never stored.
- **A refusal** happens when a step's input doesn't fit its expectations, or a program's output
  would break a limit or a platform invariant (see [reason codes](#pipeline-reason-codes) below). A
  refused record is rejected through the OTLP partial-success mechanism, the same way a
  malformed record would be rejected without any pipeline at all, and is never stored.

Both are visible in an OTLP export response's partial-success fields: `rejected_log_records` counts
only refusals (a drop is never "rejected"), and `error_message` names what happened in closed,
count-only text such as `"2 log records refused by the intake pipeline (TRANSFORM_PARSE_FAILED=2); 1
log records dropped by pipeline rules"` — never record content. A batch that only dropped records
(nothing refused) still carries a message, with `rejected_log_records` at `0`. If the pipeline
revision pinned for a request cannot be loaded at all, the whole request is refused as retryable:
`503` with `Retry-After` over HTTP, `Unavailable` over gRPC, message `intake pipeline unavailable` —
nothing in that batch is admitted, and the export should be retried unchanged.

### Pipeline reason codes

| Code | Meaning |
| --- | --- |
| `TRANSFORM_DROPPED` | An explicit `drop` step matched. |
| `TRANSFORM_PARSE_FAILED` | `parse_json`/`parse_key_value` source was missing type-correctness, or the JSON did not decode to one object. |
| `TRANSFORM_TYPE_MISMATCH` | `extract` matched text that could not convert to the requested type, or its source was not a scalar. |
| `TRANSFORM_DEPTH_LIMIT` | A `parse_json` step produced a value nested more than 8 levels deep. |
| `TRANSFORM_EXPANSION_LIMIT` | The record ended up with more than 256 attributes or 256 resource keys, or its measured size exceeded the output bound. |
| `TRANSFORM_WORK_LIMIT` | The record's steps and conditions together fed more than 1 MiB to parsing or pattern matching. |
| `TRANSFORM_PROTECTED_FIELD` | A step tried to add, remove or change a platform-reserved key. |
| `TRANSFORM_INVALID_OUTPUT` | The record no longer validates (e.g. an empty body, an unrecognized severity, or a key that cannot be stored as written). |

A dropped or refused record is counted under exactly one of these codes; see
[diagnostics](#diagnostics) for the per-revision, per-reason totals.

### Preview

**Preview** runs 1–10 sample records — each with `body`, `severity_text`, `attributes` and
`resource`, at most 16 KiB encoded **each** (not a total across samples) — through either a stored
revision or an unsaved program document, with the exact same compiler, runtime, redaction and limits
real intake uses. Samples and their outputs are never stored or logged. Each result names its
sample's outcome (unchanged, transformed, dropped or refused), the reason code and the step that
produced it, and the output record intake would store (for an unchanged or transformed sample).
Previewing an unsaved program that does not itself validate returns `INVALID_PROGRAM` with the same
structural diagnostics `validate` would, and runs nothing. Use preview to check a program's actual
effect on real-shaped events before creating a revision, or before activating one you already
created.

### Diagnostics

**Diagnostics** reports how a source's active-intake pipeline is actually behaving: accepted,
transformed, dropped and refused record counts per pipeline revision over a recent window (1–168
hours, default 24), broken down by reason code, for records received through this source. It also
reports **freshness** — how recently anything was admitted through this source, which pipeline
revision and catalog state that admission used, and whether that state is still the source's current
one. Freshness is absent when nothing was admitted in the window. Counts reach this read within
about 15 seconds of admission. This is metadata only: no event content, ever.

## Activate and roll back

**Activate** makes a stored, validated revision the source's live pipeline; picking an earlier
revision than the one currently active is how you roll back, and revision `0` removes the active
pipeline entirely. Every activation (including a rollback) is a compare-and-swap on the source's
**catalog epoch**: you must supply the exact epoch you last read, and a stale value is refused
rather than silently applied over a change you have not seen — reload the source and retry with its
current epoch. The one exception is asking for the revision that is **already active**: that
changes nothing, so it succeeds with `already_active: true` whatever epoch you send. A retried
activation whose first attempt already went through therefore reports success, not a conflict.

A successful activation returns an operation in state **draining**: admission already in flight
under the old pipeline is allowed to finish for about fifteen seconds before the switch is
considered safe. Reading that operation again after the drain window reports **complete** — nothing
else needs to happen for it to advance; simply read it again. **Complete** means every subsequent
admission through this source uses the new pipeline, not that historical data was reprocessed.

## Browser capture settings

Each project has **capture settings**: what its browsers are allowed to collect automatically.
They are a ceiling on [autocapture](/docs/reference/javascript-sdk#opt-in-autocapture) that you
change here, without redeploying your site. A site that starts autocapture with
`remoteConfig: true` reads them and applies the more private of them and its own code, option by
option. The settings can turn capture off, mask more, block more and lower limits; they can never
make a site collect more than its code configured.

| Setting | Default | Meaning |
| --- | --- | --- |
| `enabled` | `false` | The kill switch. While `false`, autocapture is off in every browser that reads these settings. |
| `dom_events` | click, change, submit | The events browsers may observe. |
| `mask_all_text` | `false` | Send no visible text or free-text attributes. |
| `mask_selectors` | none | CSS selectors whose events are sent without text or attributes. At most 50, each at most 512 characters. |
| `block_selectors` | none | CSS selectors that send no event at all. Same limits. |
| `capture_pointer` | `false` | Allow click position data. |
| `rate_limits` | 300 overall, 120 clicks, 60 changes, 30 submits, 10 rage clicks | Events per minute per browser, 0 to 100000. `0` blocks that kind. |

A project that has never stored settings is at **revision 0** with the defaults above, so
autocapture is off until you turn it on. Every change advances the revision, and a change must
name the revision it was based on: if someone else changed the settings in between, yours is
refused with a conflict (`409`) instead of overwriting theirs. Read the settings again and retry.

**In the Console.** Open **Privacy** (`/privacy`) and choose the **Browser capture** tab, for the
project selected in the header. See the [Console guide](/docs/manage/console#privacy). The tab
shows the current settings, the revision, and when and by whom they were last changed. A project
that has never stored settings reads "Not configured yet — autocapture is off". Everyone in the
workspace can view the tab; changing it needs the same access as other ingestion settings, and
without it the values are shown with no edit controls.

- **Autocapture** is the kill switch. Turning it off stops automatic collection in every browser
  that reads these settings, whatever the site configured. Saving a change to this switch, in
  either direction, asks you to confirm first.
- **DOM events** are the events browsers may observe: click, change and submit.
- **Mask all text** sends no visible text or free-text attributes. **Capture pointer position**
  adds where on the page a click happened, which [heatmaps](/docs/investigate/heatmaps) need to place clicks; it is off by default.
- **Mask selectors** send an event without its text or attributes. **Block selectors** send no
  event at all for anything inside a matching element. Each list takes up to 50 CSS selectors of
  at most 512 bytes. The tab checks that every selector is valid CSS and that none repeats,
  because a browser that cannot read one selector turns autocapture off.
- **Rate limits** are events per minute per browser, 0 to 100,000, for all events, clicks,
  changes, submits and rage clicks. All five are required.

If someone else saved a change after you opened the tab, your save is refused. The tab tells
you, keeps your unsaved edits on screen next to the latest saved values, and pauses saving until
you choose **Discard my edits and load the latest**. Any other refusal is shown exactly as it was
returned and keeps your edits.

**Turn autocapture off (the kill switch).**

```bash
anectico capture-settings get
anectico capture-settings disable-autocapture   # changes only "enabled"; everything else is kept
anectico capture-settings enable-autocapture
```

A change is served within 30 seconds. A page loaded after that follows it at once; a page that is
already open follows it at its next refresh, five minutes by default. If a browser cannot read the
settings at all, autocapture in that browser stays off.

**Change the other settings.** `set` replaces the whole autocapture section, so start from what
`get` returned:

```bash
anectico capture-settings set --expected-revision 3 --body '{
  "enabled": true,
  "dom_events": ["AUTOCAPTURE_DOM_EVENT_CLICK", "AUTOCAPTURE_DOM_EVENT_SUBMIT"],
  "mask_all_text": false,
  "mask_selectors": ["[data-account]"],
  "block_selectors": ["#billing"],
  "capture_pointer": false,
  "rate_limits": {"global": 300, "click": 120, "change": 60, "submit": 30, "rageclick": 10}
}'
```

**REST.** `GET /api/v1/projects/{projectId}/capture-settings` returns
`{"settings": {"project_id", "revision", "autocapture", "updated_at", "updated_by"}}`.
`PUT` on the same path takes `{"expected_revision": "<revision>", "autocapture": {...}}` with the
whole section as above and returns the stored settings at their new revision. `dom_events` uses
the names `AUTOCAPTURE_DOM_EVENT_CLICK`, `AUTOCAPTURE_DOM_EVENT_CHANGE` and
`AUTOCAPTURE_DOM_EVENT_SUBMIT`; all five rate limits are required. A setting outside its limits is
refused with `400` naming the field.

```bash
curl -X PUT "https://app.anectico.com/api/v1/projects/$PROJECT_ID/capture-settings" \
  -H "Authorization: Bearer $ANECTICO_TOKEN" -H "Content-Type: application/json" \
  -d @capture-settings.json
```

**MCP.** `get_capture_settings` reads the settings and their revision.
`update_capture_settings` previews and then confirms a change; it takes `expected_revision` and
only the fields you want to change, so the kill switch is
`{"expected_revision": "3", "enabled": false}`.

**Page snapshots.** The same settings carry a second section, `page_snapshot`, which governs the
masked page copy a browser may take for a
[heatmap's backdrop](/docs/investigate/heatmaps#the-page-behind-the-grid). It is off by default and
masks all text by default. Change it under **Heatmap page snapshots** on the same tab,
with `anectico capture-settings enable-page-snapshot`, `disable-page-snapshot` (its kill switch) or
`set-page-snapshot`, with MCP's `update_capture_settings`
(`{"expected_revision": "3", "page_snapshot": {"enabled": true}}`), or by sending `page_snapshot`
in the `PUT` above; an update that leaves it out keeps it as it is. Unlike autocapture, this switch is
also enforced by the server: while it is off, a copy that a browser still sends is refused.

Capture settings govern what a browser collects automatically. The mobile SDKs collect only what
your app sends explicitly, so there is nothing for these settings to switch off there, and
server-side capture is not affected by them.

## Permissions

`ingestion:pipelines:read` (any workspace member) lists and reads sources, pipeline revisions and
operations, runs validate/preview, and reads intake diagnostics — metadata and counts only, never a
key's secret or captured event content. `ingestion:pipelines:write` (workspace admin/owner) is required in addition to create,
rename or revoke a source, to store or activate a pipeline revision, or to change capture
settings (reading them needs only the read scope). Neither scope grants intake
itself (that is `ingest:write` / `analytics:write`) or key creation (`api_key:write`). See the
[permission scope reference](/docs/reference/permissions).

## REST, MCP and CLI

The full REST contract is in [Ingestion sources and pipelines](/docs/reference/rest-api#ingestion-sources-and-pipelines).
MCP exposes the same surface as reads (`list_ingestion_sources`, `get_ingestion_source`,
`list_pipeline_revisions`, `get_pipeline_revision`, `validate_pipeline_program`, `preview_pipeline`,
`get_pipeline_diagnostics`, `list_source_operations`, `get_source_operation`) plus
preview-then-confirm writes (`create_ingestion_source`, `update_ingestion_source`,
`revoke_ingestion_source`, `create_pipeline_revision`, `activate_pipeline_revision`) — see the
[MCP tool reference](/docs/reference/mcp-tools).

The CLI's `sources` command group covers the same operations:

```bash
anectico sources list --include-revoked
anectico sources create --purpose native_ingest --name collector --idempotency-key checkout-collector
anectico sources update <source-id> --expected-revision 1 --name renamed-collector
anectico sources revoke <source-id> --yes

anectico sources pipeline validate <source-id> --file program.json
anectico sources pipeline create <source-id> --file program.json
anectico sources pipeline list <source-id>
anectico sources pipeline preview <source-id> <revision> --sample '{"body":"sample"}'
anectico sources pipeline preview <source-id> --program-file draft.json --samples-file samples.json
anectico sources pipeline diagnostics <source-id> --window-hours 48
anectico sources pipeline activate <source-id> <revision> --expected-epoch 1 --wait

anectico sources operations list <source-id>
```

`preview` takes exactly one of a stored `<revision>` or `--program-file` (an unsaved program
document), and samples from one or more `--sample '<json object>'` or a single `--samples-file`
holding a JSON array of them — never both forms of either at once.

`anectico sources revoke` is destructive and requires `--yes`. `anectico sources pipeline activate
--wait` blocks until the resulting operation leaves draining and prints the terminal read, instead of
the immediate response, so a script does not need its own poll loop.
