# External monitoring

> Probe a public HTTP endpoint on a schedule and page through your existing escalation policies.

Canonical page: https://anectico.com/docs/respond/external-monitoring/


A monitor watches one external thing on a schedule and pages through your escalation policies the
same way any other alert does. It answers a narrower question than a rule over your own telemetry:
is this endpoint reachable, this certificate still valid, this domain still registered, or this job
still running on schedule — right now, independent of anything your own services report about
themselves. Four kinds share the same lifecycle, routing, response gate and maintenance windows:

| Kind | Watches | Regions | Typical interval |
| --- | --- | --- | --- |
| **HTTP** | A request/response check against a public endpoint | Yes, a majority decides | Seconds to minutes |
| **TLS** | A certificate chain's expiry | Yes, a majority decides | Hours |
| **Domain** | A domain's registration expiry (RDAP) | No — one lookup covers every region | Hours |
| **Heartbeat** | Pings from a scheduled job, never a probe | No — the job pings the platform | The job's own schedule |

A monitor's **kind cannot change** after creation; create a new monitor instead. The sections below
cover HTTP first (the original kind, and the most configuration surface), then
[TLS certificate expiry](#tls-certificate-expiry),
[domain registration expiry](#domain-registration-expiry), [heartbeats](#heartbeats) and
[maintenance windows](#maintenance-windows), which apply to every kind.

Your agent creates and manages monitors through MCP or the CLI. There is no monitor page in the
Console. If a monitor's check needs a secret in a request header, you add that secret as a
connection in the Console first.

## Ask your agent

> "Monitor `https://status.example.com/health` every minute from every region. Page the platform
> escalation policy if it fails twice in a row."

> "Why did the checkout monitor fail last night? Show the regional results of the failed rounds."

| Job | MCP tool or action | CLI command |
| --- | --- | --- |
| Create a monitor (`http`, `tls`, `domain` or `heartbeat`) | `create_monitor` | `anectico monitors create` |
| List monitors or read one | `list_monitors`, `get_monitor` | `anectico monitors list`, `anectico monitors get <monitor-id>` |
| Change a monitor | `update_monitor` | `anectico monitors update <monitor-id>` |
| Pause or resume | `pause_monitor`, `resume_monitor` | `anectico monitors pause <monitor-id>`, `anectico monitors resume <monitor-id>` |
| Run one test round now | `test_monitor` | `anectico monitors test <monitor-id>` |
| Resume paging after a responder ended the response | `rearm_monitor` | `anectico monitors rearm <monitor-id>` |
| Delete a monitor | `delete_monitor` | `anectico monitors delete <monitor-id> --yes` |
| Read rounds, job history and the audit trail | `list_monitor_rounds`, `list_monitor_history`, `list_monitor_audit` | `anectico monitors rounds`, `history`, `audit` |
| List probe regions | `list_probe_regions` | `anectico monitors regions` |
| Read heartbeat pings, rotate its token | `list_heartbeat_pings`, `rotate_heartbeat_token` | `anectico monitors heartbeat pings <monitor-id>`, `anectico monitors heartbeat rotate <monitor-id> --yes` |
| Print the heartbeat snippets | none | `anectico monitors heartbeat snippet <monitor-id>` |
| Create, read, change or end a maintenance window | `create_maintenance_window`, `list_maintenance_windows`, `get_maintenance_window`, `update_maintenance_window`, `end_maintenance_window` | `anectico monitors maintenance create`, `list`, `get`, `update`, `end` |

The scopes are in [Scopes](#scopes). The MCP write tools preview first and apply on a second call
with a `confirm_token`; see [MCP tools](/docs/reference/mcp-tools).

## What your agent gets back

A monitor read returns its kind, state, configuration revision, owning service, routing status,
diagnostics, health and, for a heartbeat, its token prefix and the run it is waiting for. See
[Reading a monitor's state](#reading-a-monitors-state) for every field. A round read returns the
verdict, the coverage and each region's result and timing. A heartbeat's ping token is returned once,
by the call that creates the monitor or rotates the token.

## Open the proof

A monitor has no proof page. When its failure opens an incident, the incident has one at
`/view/incident/:id`: see [Manage incidents and on-call](/docs/respond/incidents-and-on-call) and
[Proof pages](/docs/agents/proof-pages).

## Create an HTTP monitor

Ask your agent to create the monitor. An HTTP monitor
needs:

- a **name**;
- a **check**: `GET` or `HEAD`, an absolute `http://` or `https://` URL on a public host, and how long
  to wait for a response (1-10 seconds, default 10);
- an **interval** between rounds (at least 60 seconds, default 60) and which **regions** probe it (at
  least one; leave this unset to use every region available to your organization). A round is
  decided by a majority of the monitor's regions — 1 of 1 with a single region, 2 of 3 with three;
- how many consecutive **failed rounds** open an occurrence and how many consecutive **healthy
  rounds** close it (1-10 each, default 2); and
- **routing**: either the escalation policy that pages when the check fails and a severity
  (`info`, `warning` or `critical`, default `critical`), or the **service** from your
  [service catalog](/docs/respond/incidents-and-on-call#set-up-the-response-catalog) that owns the
  monitor (see [Monitors owned by a service](#monitors-owned-by-a-service)).

A monitor can also be created **paused**, so you can finish wiring up routing before it starts
probing.

### What "reachable and expected" means

By default a round succeeds on any `2xx` response. You can widen or narrow that with one or more
accepted status ranges (for example `200-299` or a single code like `204`), and add up to 10 response
checks against the first 256 KiB of the body — a `GET` check only, since a `HEAD` response has no
body:

- **contains** or **does not contain** a piece of text; or
- a **JSON path** (a dotted path of object keys and array indexes, for example `status` or
  `checks.0.state`) **equals** an exact JSON string, number, boolean or `null`.

Accepted status ranges, response body checks and redirect behavior are part of the monitor's
configuration; set them with the API or the CLI flags `--expect`, `--assert-contains` and
`--assert-json`. A paused monitor stays paused after an edit.

### Targets that are refused

A monitor cannot target a private, loopback, link-local or cloud metadata address, `localhost`, a
single-label host name, or a name ending in `.local`, `.internal`, `.localhost`, `.home.arpa`,
`.lan` or `.intranet` — these only ever resolve inside a private network, and a monitor exists to
watch what the public internet sees. A name that legitimately cannot be resolved is accepted (a DNS
outage is exactly the kind of thing a monitor should report); a name that resolves to one of those
addresses is not. The accepted port range is 80, 443, or 1024-65535.

### Secrets go through Connections, never through the monitor

A monitor can send up to 10 request headers, but it never accepts a header **value** directly — only
a reference to an active, secret-holding connection of kind `api_key` or `basic_auth`. Add that
connection first. A secret is a person's job, so you add it yourself: in the Console, open
**Connections and notifications** (`/connections`); see
[Connect tools and notification delivery](/docs/manage/connections-and-notifications). Then ask your
agent to point the header at the connection's id by name (for example an `Authorization` header
referencing a stored bearer token). Do not paste the secret into the chat with your agent. Framing, routing and
hop-by-hop headers (`Host`, `Content-Length`, `Connection`, `Proxy-Authorization` and similar) are
refused: a secret can add a credential, never redirect or reshape the request.

## TLS certificate expiry

A `tls` monitor connects to a public DNS host and port, completes a real TLS handshake verified
against the public trust store, and measures the earliest expiry among the verified certificate
chain (the trust anchor itself is excluded) — so if an intermediate certificate expires before the
leaf does, that earlier date is what the monitor reports. It needs:

- a **host** — a public DNS name (no literal IP address), refused under the same policy as an HTTP
  target: no private, loopback, link-local or metadata address, `localhost`, single-label names, or
  `.local`/`.internal`/`.localhost`/`.home.arpa`/`.lan`/`.intranet` suffixes;
- a **port** — 443 (default), 465, 636, 853, 993, 995, or 1024-65535;
- **warning_days** — fail the check once the chain is within this many days of expiring, 1-180
  (default 21); and
- an **interval** (1-24 hours, 3600-86400 seconds, default 6 hours) and **regions**, exactly like an
  HTTP monitor — a certificate check still runs from your configured regions and is still decided
  by a majority of them.

Every failure is one of three closed outcomes: **expiring** (within `warning_days` of expiry),
**expired**, or **invalid** — any other verification failure (untrusted issuer, wrong host name, not
yet valid). All three are failures a client of the endpoint would also see. If the target cannot be
reached at all, or the handshake fails before any certificate is presented, that region's result is
**unknown coverage** — never a failure, and never an invented date. A completed check's evidence
(measured expiry, days remaining, subject and issuer common names, the leaf certificate's SHA-256
fingerprint, and whether the chain verified) is attached to that region's round result.

## Domain registration expiry

A `domain` monitor reads a domain's registration expiry from its authoritative registry over RDAP,
located automatically through the public IANA bootstrap registry — you never choose or configure a
registry yourself. It needs:

- a **domain** — stored as its registrable domain under a public suffix (for example
  `status.example.co.uk` is stored and checked as `example.co.uk`; a bare public suffix like `com`,
  or a name under a *private* suffix such as `example.github.io`, is refused); and
- **warning_days** — fail once the registration is within this many days of expiring, 1-365 (default
  30). Its interval is 6-24 hours (21600-86400 seconds, default 24 hours), and it takes **no
  regions**: one lookup, run once by the platform, covers every region a monitor of another kind
  would need several of.

A domain check only ever fails on an expiring or expired registration date it actually read. Every
other registry answer is **unknown coverage**, never a date and never a failure, with a closed
reason: no RDAP service exists for that registry, the registry redacts the expiration date, the
record carries no expiration event at all, the registry has no record of the domain, or the registry
is rate-limiting or failing lookups (retried with growing backoff, up to six hours between
attempts). A registration date the platform already knows survives a registry that starts failing
or throttling afterward — the last successful lookup keeps being judged for up to 72 hours before it
too becomes unknown, so a brief registry outage never manufactures either a false pass or a false
alarm.

## Heartbeats

A `heartbeat` monitor never probes anything: instead, the job it watches pings the platform, and the
monitor expects that ping on a schedule. There is no target to configure and no regions — the
schedule itself is the check.

### Schedule: interval or cron

Set exactly one of:

- **interval_seconds** — expect a successful ping at most this long after the previous one settled,
  60 seconds to 7 days; or
- **cron** — a five-field cron expression (`minute hour day-of-month month day-of-week`, the
  `@daily`/`@hourly`/`@weekly`/`@monthly`/`@yearly` shorthands accepted) evaluated in **timezone**
  (an IANA zone; default UTC).

**grace_seconds** (60-86400, default 300) is how long after the expected time a successful ping may
still arrive — set it to comfortably cover the job's own runtime if it reports a **start** ping, not
just success or failure.

Expectations begin at **activation** (creation, or the moment a paused heartbeat is resumed): a job
that has never pinged is judged missed at its very first expected run plus grace, not held open
indefinitely. Editing the schedule, like editing any monitor's configuration, restarts the schedule
from the moment of the edit rather than continuing to judge against the old one.

Pausing discards unfinished expectations and marks unused accepted pings as ignored.
Completed history stays available. Resuming starts fresh expectations; a completion paired with a
start from before the pause is late and cannot satisfy the resumed heartbeat. Report a new run to
satisfy its new expectation. Repeating pause or resume when already in that state changes nothing.

Across a daylight-saving change, a `cron` schedule's wall-clock time that is skipped entirely is
expected once, at the moment the gap ends; one that occurs twice is expected once, at its first
occurrence — neither change manufactures a missed run or a duplicate one.

If the platform itself falls behind by more than an hour (a genuine backlog, not a single missed
run), it does not replay every expected run that piled up: it skips ahead to the run closest to now
and resumes judging from there.

### Reporting a run

The job reports itself with `GET` or `POST` requests to the ping endpoint, sending its token in the
`X-Anectico-Heartbeat-Token` header, with an empty request body and run details in the query string.
A `token` query parameter or any non-empty body is refused with `400`, even if the header contains a
valid token; no ping is recorded. A token never belongs in a URL, where it ends up in access logs
and shell history. Report:

- `event=start` when the run begins (only needed if you want overlapping or long runs tracked
  individually);
- `event=success` or `event=fail` when it ends; or
- an **exit code** (0-255) instead of an explicit event — `0` implies success, anything else implies
  failure, and one that contradicts an explicit `event` is refused.

An optional **run id** (up to 64 characters) pairs a `start` with its later `success`/`fail`; without
one, a completion pairs with the latest unfinished start that also had none. A **repeated** start or
completion of the same run id changes nothing and is reported back as a duplicate — safe to retry a
ping that may not have arrived. **Overlapping runs** (a new run starts before the previous one
finished) are tracked individually: the expected run is satisfied by a success once every run
started within its window has finished, and any one of them failing fails it. A completion that
arrives for a run whose expected window has already been judged is **late** — it cannot retroactively
satisfy that run, and it does not spuriously satisfy the next one either. A ping that arrives before
its round's window has opened (a manual run between scheduled times) is recorded but **ignored** —
it settles nothing. A **paused** heartbeat's pings are accepted but ignored and not stored; pings are
capped at 30 accepted per minute per monitor.

Every ping — its event, run id, exit code, timestamp, whether it was a duplicate, and how its round
eventually counted it (`counted`, `late`, `ignored`) — is visible on the monitor's ping history.

### Tokens and rotation

A heartbeat monitor's ping token (`anhb_` followed by 43 characters) is shown **exactly once**: in
the response that creates the monitor, or the response that rotates it. It is not recoverable after
that — hold it as a secret in whatever runs the job, the same as any other credential, and never
paste it into a crontab, a script body, or a ticket. If it is lost, **rotate** it: the previous token
stops working the instant a new one is issued, so update the job in the same step.

### Snippets

These two shapes cover most jobs. Neither one embeds a real token — replace `$TOKEN` with the value
you received, held as a secret, and substitute your own script:

A crontab line that reports only success or failure, with no run id:

```bash
0 2 * * * /job.sh && curl -fsS -m 10 --retry 3 -X POST -H "X-Anectico-Heartbeat-Token: $TOKEN" "https://<api>/api/v1/heartbeats/ping"
```

A wrapper that reports a start, then the outcome, paired by run id (recommended once a run can
overlap the next, or you want `RUN_UNFINISHED` distinguished from a plain miss):

```bash
RID=$(date +%s)
curl -fsS -m 10 --retry 3 -X POST -H "X-Anectico-Heartbeat-Token: $TOKEN" "https://<api>/api/v1/heartbeats/ping?event=start&run_id=$RID"
/job.sh
curl -fsS -m 10 --retry 3 -X POST -H "X-Anectico-Heartbeat-Token: $TOKEN" "https://<api>/api/v1/heartbeats/ping?exit_code=$?&run_id=$RID"
```

`anectico monitors heartbeat snippet <monitor-id>` prints both, and
`anectico monitors heartbeat rotate <monitor-id> --yes` replaces a lost token.

## Maintenance windows

A maintenance window suppresses **paging**, not checking: every monitor in its scope keeps running
its schedule and every round is still recorded truthfully — health still moves to failing, an
already-open episode still receives its updates — but a failure that would otherwise open a **new**
episode opens none while the window is active. This is deliberately different from **pausing** a
monitor, which stops it from running or recording anything at all; use a maintenance window when you
still want the history, and pause when you do not want the check to run.

A window covers either **every monitor in a project** (including ones created while it is active) or
an explicit list of up to 200 monitors, between a start (which may be in the past — the window is
then active immediately) and an end (in the future, at most 14 days after the start, itself at most
a year ahead). A project may hold at most 50 upcoming-and-active windows at once.

**When the window ends**, any monitor in its scope that is still failing — and has no open episode
of its own, because maintenance suppressed opening one — pages **once**, at that moment, rather than
waiting for its next scheduled failed round (which, for a domain or heartbeat check running once a
day, could be a day away). A monitor whose response was separately ended by a responder (see
[When a responder resolves a still-failing monitor](#when-a-responder-resolves-a-still-failing-monitor))
does not page when the window ends either — that override stands until a qualified recovery or an
explicit rearm, exactly as it would with no maintenance window involved. Two overlapping windows
covering the same monitor: ending the earlier one pages nothing while the later one still covers it;
ending the one that covers it last is what pages.

**Editing a window** depends on its state: an **upcoming** one accepts any change to its scope,
bounds or time zone; an **active** one accepts only its **name** and its **end** (you can shorten or
extend an active window, or rename it, but not change what it covers); a **completed** or
**canceled** one accepts no change at all. **Ending** a window early cancels an upcoming one
outright, or ends an active one right now (with the paging behavior above). A window, once ended,
cannot be restarted — create a new one.

## How a round is decided

This section covers HTTP and TLS monitors, whose rounds run from configured regions. A domain
round has a single platform-run lookup in place of regional votes, and a heartbeat round is judged
from pings rather than regions (see [Heartbeats](#heartbeats)); both still drive the same
consecutive failure/recovery counters and response gate described below.

Every round probes every one of a monitor's regions independently, and the round's verdict is
decided by a **majority** of those regions — 2 of 3 with three regions, 1 of 1 with a single
region, 3 of 5 with five. A region that never reports a valid result (it was never reached, or the
probe itself could not produce a clean observation) is **unknown coverage**: it never counts as
healthy, and it never, by itself, confirms an outage.

- If a majority of regions report a failure, the round is **failed**.
- If a majority report success, the round is **healthy**.
- Anything else — including an even split, or too few regions reporting at all — is **unknown**,
  and an unknown round moves nothing: it neither opens nor updates a failing episode, and it never
  counts toward recovery.

A round's **coverage** tells you how much of that majority you can trust: `full` (every region
reported), `degraded` (enough regions reported to decide, but not all of them), or `unknown` (too
few did — the round decided nothing).

A single failed round does not page. It takes **failure_rounds** consecutive failed rounds (of the
monitor's current configuration) to open one failing episode — one alert, no matter how many rounds
that episode goes on to repeat. Recovery is symmetric: **recovery_rounds** consecutive healthy
rounds close it. Editing a monitor's check or schedule restarts both counters from zero under the
new configuration, though an already-open episode stays open and is simply updated by the next
failed round rather than opening a second one.

## Regional results and timing

Every round's regional results are readable through the API, MCP (`list_monitor_rounds`) and the CLI
(`anectico monitors rounds <monitor-id>`): each region's outcome (`success`, `failure`, `invalid`, or `missing` when
no result arrived in time), the response status code, a closed error class when it did not succeed
(for example `DNS_FAILURE`, `TLS_FAILURE`, `CONNECT_FAILURE`, `TIMEOUT`, or `STATUS_MISMATCH` when
the response came back but not with an accepted status), and per-phase timing in milliseconds (DNS,
connect, TLS, time to first byte, and the total). A phase that did not happen for that request (a
reused connection, plain HTTP) reads as zero, not as missing data.

## Probe regions

A monitor's regions come from your organization's probe region catalog. Each region is served by
one or more regional probe executors, and the catalog itself — separate from any one monitor — is
readable from the same surfaces: `serving` (an executor polled recently), `stale` (one was seen
before, but not recently), or `absent` (never seen). A region that is not currently serving still
belongs to a monitor's configuration; its rounds simply produce unknown coverage for that region
until an executor resumes polling it.

## Routing and what pages

A monitor pages through exactly one route, chosen when you create it and changeable later:

- **Its own escalation policy.** The monitor creates and owns a dedicated alert rule the moment it
  is created. You never edit that rule directly; changing the monitor's escalation policy or
  severity changes the rule.
- **A service that owns it.** The monitor has no alert rule of its own and pages through the
  service's — see [Monitors owned by a service](#monitors-owned-by-a-service).

**An active monitor needs routing coverage**: an escalation policy that exists, is active, and has at
least one level — its own, or its owning service's. Creating or resuming a monitor without that
coverage is refused, naming which part is missing (`missing_policy`, `policy_inactive` or
`policy_has_no_levels`). Create the monitor paused if you want to finish setting up the policy
first, or fix the policy and try again.

### Monitors owned by a service

Give a monitor a `service_key` and the [catalog service](/docs/respond/incidents-and-on-call#set-up-the-response-catalog)
with that key owns it. That one choice decides who owns the monitor's failures everywhere:

- **Paging** goes through the service's escalation policy at the service's severity. The monitor
  has no routing of its own — sending an escalation policy or severity together with a
  `service_key` is refused (`MONITOR_ROUTING_CONFLICT`), and re-routing an owned monitor directly
  is refused too (`MONITOR_SERVICE_OWNED`): change the service's routing instead.
- **Incidents** opened for the monitor's failures carry the service's
  [ownership snapshot](/docs/respond/incidents-and-on-call#ownership-and-checklist) — owner,
  schedule, escalation policy and runbook checklist — exactly like an incoming alert routed to that
  service.
- **Response reports** attribute those incidents to the service's key, so filtering a report by the
  service includes the monitor's failures.

The service must exist in the monitor's project, be live, and have its alert rule bound
(`MONITOR_SERVICE_INVALID` otherwise, or `MONITOR_SERVICE_NOT_ROUTABLE` while the service is still
being set up). Only one rule ever pages for the failure, so the monitor can never page twice.

**Moving a monitor between routes.** Setting a `service_key` on a monitor that had its own routing
removes its own alert rule. Clearing it (an empty `service_key`, sent together with the new routing)
gives the monitor its own rule again. Moving to another service is the same change with the new key.
If the monitor is failing when it moves, the failure moves with it: the alert on the old route
closes with resolution `withdrawn` — never as recovered, because the check has not recovered — and
the new route is paged at once rather than at the next failed round. A move whose old rule could
not be removed yet answers `MONITOR_RELEASE_PENDING`; the monitor already pages through its new
service, and repeating the same change finishes the removal.

If the owning service is **deleted**, the monitor cannot page until you give it another service or
its own routing; its routing then reads `unverified`, naming the deleted service.

Every failing occurrence becomes one alert, however many rounds repeat the same failure; recovery
closes it the same way an ordinary alert closes. From there, acknowledging, resolving, and
escalation behave exactly like any other alert — see
[Create useful alerts](/docs/respond/alerts) and
[Manage incidents and on-call](/docs/respond/incidents-and-on-call).

## When a responder resolves a still-failing monitor

Every failing episode is one alert, and a monitor's alert behaves like any other: it can be
acknowledged, resolved, and escalated the normal way (see
[Create useful alerts](/docs/respond/alerts)). Resolving it **while the check is still failing** —
directly, or by resolving an incident linked to it — is a deliberate override: the responder is
telling the system "I've handled this, stop paging me about it." From that point, further failed
rounds keep updating the monitor's evidence (so the history stays honest), but **paging is held**
until either:

- a **qualified recovery** — enough consecutive healthy rounds to close the episode the normal
  way; or
- an explicit **rearm** — telling the monitor to start paging again immediately, even though the
  check has not recovered. Re-arming requires the monitor's current response-gate version (shown on
  every read) so a rearm sent against a stale copy of the monitor is refused rather than silently
  reapplied; the very next failed round then opens a new episode and pages once.

Two related actions do **not** rearm a monitor on their own: **reopening an incident** linked to the
alert is a change to the incident, not a resumption of paging for the check itself, and a
**silence** ending (by expiry or by being removed early) resumes normal alerting exactly the way it
would for any other alert, without needing a rearm.

## Pause, resume and delete

**Pausing** a monitor stops scheduling new rounds immediately; any rounds already in flight are
cancelled. Pausing does **not** touch an alert that is already firing — it keeps firing until it is
resolved, or until the monitor is resumed and recovers on its own. **Resuming** restarts the
schedule from now, and re-checks routing coverage exactly like creating an active monitor does — a
monitor cannot resume into a state where nothing would page.

**Deleting** a monitor is a tombstone: scheduling stops, its dedicated alert rule is removed, and its
full history and audit trail remain readable. A monitor owned by a service has no rule to remove:
if it is failing, its open alert on the service's rule closes with resolution `withdrawn`, which
closes the incident linked to it. A deleted monitor's name is not held back for reuse.
Deleting the monitor's **project** (or organization) is different: its monitors and their history
are erased, not kept as tombstones, and no further checks run.
Both an edit and a delete require you to state the exact revision you are changing (shown on every
read), so a change from a stale copy of the monitor is refused rather than silently overwriting a
concurrent one.

## Test a monitor on demand

A test (`test_monitor`, or `anectico monitors test <monitor-id>`) runs one bounded, extra round right
now, against every one of the monitor's regions, without waiting for the schedule. By default a test
round's result **never pages**. Ask for **page on failure** (the CLI flag is `--page`) when you want a
failing test to page like a real round would. Only one test round can be open per monitor at a time;
wait for it to finish (or expire) before starting another. A paging test against a monitor whose
routing is not yet covered is refused for the same reason an active monitor would be. Read its outcome
the same way you read any other round: `monitors rounds <id> --kind test` (CLI),
`GET .../rounds?kind=test` (API), or `list_monitor_rounds` (MCP). The round can take a moment, so read
again until it is complete.

A test of a `domain` monitor whose last registry lookup failed asks the registry again
rather than repeating the failure it already recorded — at most once a minute per domain, however
many tests are started, so a test within a minute of the last lookup reports that lookup's result.
A registration date that is already known, or any other answer the registry gave (a redacted date,
no record, no RDAP service, rate limiting), is reported as it stands without asking again.

A heartbeat has no check to run on demand — there is nothing for a test to probe — so it is
refused with a closed reason (`MONITOR_TEST_UNSUPPORTED`); test it by sending it a ping instead.

## Reading a monitor's state

Every monitor read includes:

- its **kind** (`http`, `tls`, `domain` or `heartbeat`) and current **state** (`active`, `paused` or
  a deleted tombstone) and configuration **revision**;
- its **owning service** (`service_key`, empty when it has its own routing) and the alert rule it
  pages through — its own, or the owning service's (then the rule's `binding_state` is `none`: the
  monitor has no rule of its own);
- that rule's **routing status** — `covered`, or a reason it is not (`missing_policy`,
  `policy_inactive`, `policy_has_no_levels`, or `unverified` when it could not be checked);
- **diagnostics**: why it is not currently producing results, as one or more closed reason codes,
  plus counts of jobs waiting and jobs that expired unrun in the last 24 hours;
- **health**: its current state (`unknown`, `healthy` or `failing`), how many consecutive
  failed/healthy rounds have been counted under its current configuration, the last reduced round's
  verdict and coverage, the open episode's id and alert (if any), and its response-gate state and
  version (see [When a responder resolves a still-failing monitor](#when-a-responder-resolves-a-still-failing-monitor));
- for a **heartbeat**, its token prefix (never the token itself), when it was last rotated, the most
  recent accepted ping, and the run it is currently waiting for (expected time and deadline); and
- the active **maintenance window** covering it, if any (id, name, end time).

**Executor status** says whether the monitor's regions are currently being probed: `available`
(every region), `degraded` (enough regions for a majority, not all — reason `EXECUTOR_DEGRADED`) or
`unavailable` (too few to decide a round — reason `EXECUTOR_UNAVAILABLE`). A region that produces no
result is **unknown coverage**: it never counts as healthy and never, by itself, as an outage. A
domain or heartbeat monitor has no regions of its own — its executor status reflects the platform's
own worker instead, so it is still accurate when that worker is down. The other reason codes explain
the rest: `MONITOR_PAUSED` (paused, not scheduling), `MONITOR_DELETED` (tombstoned),
`MONITOR_BINDING_PENDING` (its alert rule has not finished being created or removed yet; retrying
the create, resume or delete completes it), `ROUTING_NOT_COVERED` (its escalation policy cannot
currently page), `RESPONSE_ENDED_WHILE_UNHEALTHY` (paging is held; see above), and
`MAINTENANCE_ACTIVE` (a maintenance window currently covers it).

A monitor's **history** lists every round and test job, newest first, each with its state
(admitted and waiting, claimed, completed, fenced by a later edit or pause, or expired) and, for a
fenced or expired job, why. A monitor's **rounds** list is the reduced view of that same history:
one entry per round with its verdict, coverage and each region's result (see
[Regional results and timing](#regional-results-and-timing)). Both are kept for 90 days and survive
pausing or deleting the monitor.

## Scopes

| Scope | Grants |
| --- | --- |
| `monitoring:read` | List and read monitors, their configuration, routing, diagnostics, health, history, rounds and audit; list probe regions; list a heartbeat's pings; list and read maintenance windows. Required by every monitor operation, including the writes below. |
| `monitoring:write` | Create, edit, pause, resume, test-run and re-arm monitors; rotate a heartbeat's ping token; create, edit and end maintenance windows. |
| `monitoring:delete` | Delete monitors. |

Because a monitor's routing lives on an alert rule, creating, resuming or re-routing a monitor also
requires `alerts:rules:write`, and deleting one also requires `alerts:rules:delete`. Giving a monitor
an owning service reads that service, so it also requires `incidents:read`; moving a monitor that
has its own rule to a service removes that rule, so it also requires `alerts:rules:delete`. A header
that references a connection also requires `connections:read`. See
[permission scopes](/docs/reference/permissions) for who holds each scope by default.

## API, MCP and CLI

REST is documented at [REST API reference § External monitoring](/docs/reference/rest-api#external-monitoring).
The `list_monitors`, `get_monitor`, `list_monitor_history`, `list_monitor_audit`,
`list_monitor_rounds`, `list_probe_regions`, `list_heartbeat_pings`, `list_maintenance_windows` and
`get_maintenance_window` reads and the `create_monitor`, `update_monitor`, `pause_monitor`,
`resume_monitor`, `test_monitor`, `delete_monitor`, `rearm_monitor`, `rotate_heartbeat_token`,
`create_maintenance_window`, `update_maintenance_window` and `end_maintenance_window` actions are
available over MCP (`create_monitor` and `update_monitor` take `service_key`, and `update_monitor`
takes `clear_service` with an escalation policy; every write previews before it applies; a heartbeat's ping token is delivered
only in the one-time text of a `create_monitor` or `rotate_heartbeat_token` response, never in the
structured content a later read could echo back).

```bash
anectico monitors create --name "Checkout API" \
  --url https://status.example.com/health --method GET \
  --expect 200-299 --assert-json status=\"ok\" \
  --interval 60 --failure-rounds 2 --recovery-rounds 2 \
  --escalation-policy 6f6b1a2e-9c3d-4b8e-8e2a-1f7b6a5d4c3b --severity critical

# Owned by a catalog service: pages through the service, and its incidents carry the service's ownership
anectico monitors create --name "Checkout API" --url https://status.example.com/health --service checkout
anectico monitors update <monitor-id> --revision 2 --service checkout
anectico monitors update <monitor-id> --revision 2 --clear-service \
  --escalation-policy 6f6b1a2e-9c3d-4b8e-8e2a-1f7b6a5d4c3b

anectico monitors list
anectico monitors get <monitor-id>
anectico monitors test <monitor-id> --page
anectico monitors rounds <monitor-id> --kind scheduled
anectico monitors history <monitor-id> --kind scheduled
anectico monitors regions
anectico monitors rearm <monitor-id> --gate-version 3
anectico monitors delete <monitor-id> --revision 1 --yes

# TLS certificate expiry
anectico monitors create --name "Shop cert" --kind tls --host shop.example.com \
  --warning-days 21 --escalation-policy 6f6b1a2e-9c3d-4b8e-8e2a-1f7b6a5d4c3b

# Domain registration expiry
anectico monitors create --name "example.com registration" --kind domain \
  --domain example.com --warning-days 30 --escalation-policy 6f6b1a2e-9c3d-4b8e-8e2a-1f7b6a5d4c3b

# Heartbeat, and its token / snippet / ping-history commands
anectico monitors create --name "Nightly ETL" --kind heartbeat --cron "0 2 * * *" \
  --timezone America/New_York --grace 600 --escalation-policy 6f6b1a2e-9c3d-4b8e-8e2a-1f7b6a5d4c3b
anectico monitors heartbeat snippet <monitor-id>
anectico monitors heartbeat pings <monitor-id>
anectico monitors heartbeat rotate <monitor-id> --yes

# Maintenance windows
anectico monitors maintenance create --name "db upgrade" --scope monitors --monitor <monitor-id> \
  --starts-at 2026-10-01T09:00:00Z --ends-at 2026-10-01T11:00:00Z
anectico monitors maintenance list --state active
anectico monitors maintenance update <window-id> --revision 1 --ends-at 2026-10-01T12:00:00Z
anectico monitors maintenance end <window-id> --revision 2 --yes
```

- [Create useful alerts](/docs/respond/alerts)
- [Manage incidents and on-call](/docs/respond/incidents-and-on-call)
- [Response and availability reports](/docs/respond/response-reports)
- [Connect tools and notification delivery](/docs/manage/connections-and-notifications)
