External monitoring
Probe a public HTTP endpoint on a schedule and page through your existing escalation policies.
On this page
A monitor watches one external thing on a schedule and pages through your escalation policies the same way any other alert does. It answers a narrower question than a rule over your own telemetry: is this endpoint reachable, this certificate still valid, this domain still registered, or this job still running on schedule — right now, independent of anything your own services report about themselves. Four kinds share the same lifecycle, routing, response gate and maintenance windows:
| Kind | Watches | Regions | Typical interval |
|---|---|---|---|
| HTTP | A request/response check against a public endpoint | Yes, a majority decides | Seconds to minutes |
| TLS | A certificate chain's expiry | Yes, a majority decides | Hours |
| Domain | A domain's registration expiry (RDAP) | No — one lookup covers every region | Hours |
| Heartbeat | Pings from a scheduled job, never a probe | No — the job pings the platform | The job's own schedule |
A monitor's kind cannot change after creation; create a new monitor instead. The sections below cover HTTP first (the original kind, and the most configuration surface), then TLS certificate expiry, domain registration expiry, heartbeats and maintenance windows, which apply to every kind.
Your agent creates and manages monitors through MCP or the CLI. There is no monitor page in the Console. If a monitor's check needs a secret in a request header, you add that secret as a connection in the Console first.
Ask your agent
"Monitor
https://status.example.com/healthevery minute from every region. Page the platform escalation policy if it fails twice in a row."
"Why did the checkout monitor fail last night? Show the regional results of the failed rounds."
| Job | MCP tool or action | CLI command |
|---|---|---|
Create a monitor (http, tls, domain or heartbeat) |
create_monitor |
anectico monitors create |
| List monitors or read one | list_monitors, get_monitor |
anectico monitors list, anectico monitors get <monitor-id> |
| Change a monitor | update_monitor |
anectico monitors update <monitor-id> |
| Pause or resume | pause_monitor, resume_monitor |
anectico monitors pause <monitor-id>, anectico monitors resume <monitor-id> |
| Run one test round now | test_monitor |
anectico monitors test <monitor-id> |
| Resume paging after a responder ended the response | rearm_monitor |
anectico monitors rearm <monitor-id> |
| Delete a monitor | delete_monitor |
anectico monitors delete <monitor-id> --yes |
| Read rounds, job history and the audit trail | list_monitor_rounds, list_monitor_history, list_monitor_audit |
anectico monitors rounds, history, audit |
| List probe regions | list_probe_regions |
anectico monitors regions |
| Read heartbeat pings, rotate its token | list_heartbeat_pings, rotate_heartbeat_token |
anectico monitors heartbeat pings <monitor-id>, anectico monitors heartbeat rotate <monitor-id> --yes |
| Print the heartbeat snippets | none | anectico monitors heartbeat snippet <monitor-id> |
| Create, read, change or end a maintenance window | create_maintenance_window, list_maintenance_windows, get_maintenance_window, update_maintenance_window, end_maintenance_window |
anectico monitors maintenance create, list, get, update, end |
The scopes are in Scopes. The MCP write tools preview first and apply on a second call
with a confirm_token; see MCP tools.
What your agent gets back
A monitor read returns its kind, state, configuration revision, owning service, routing status, diagnostics, health and, for a heartbeat, its token prefix and the run it is waiting for. See Reading a monitor's state for every field. A round read returns the verdict, the coverage and each region's result and timing. A heartbeat's ping token is returned once, by the call that creates the monitor or rotates the token.
Open the proof
A monitor has no proof page. When its failure opens an incident, the incident has one at
/view/incident/:id: see Manage incidents and on-call and
Proof pages.
Create an HTTP monitor
Ask your agent to create the monitor. An HTTP monitor needs:
- a name;
- a check:
GETorHEAD, an absolutehttp://orhttps://URL on a public host, and how long to wait for a response (1-10 seconds, default 10); - an interval between rounds (at least 60 seconds, default 60) and which regions probe it (at least one; leave this unset to use every region available to your organization). A round is decided by a majority of the monitor's regions — 1 of 1 with a single region, 2 of 3 with three;
- how many consecutive failed rounds open an occurrence and how many consecutive healthy rounds close it (1-10 each, default 2); and
- routing: either the escalation policy that pages when the check fails and a severity
(
info,warningorcritical, defaultcritical), or the service from your service catalog that owns the monitor (see Monitors owned by a service).
A monitor can also be created paused, so you can finish wiring up routing before it starts probing.
What "reachable and expected" means
By default a round succeeds on any 2xx response. You can widen or narrow that with one or more
accepted status ranges (for example 200-299 or a single code like 204), and add up to 10 response
checks against the first 256 KiB of the body — a GET check only, since a HEAD response has no
body:
- contains or does not contain a piece of text; or
- a JSON path (a dotted path of object keys and array indexes, for example
statusorchecks.0.state) equals an exact JSON string, number, boolean ornull.
Accepted status ranges, response body checks and redirect behavior are part of the monitor's
configuration; set them with the API or the CLI flags --expect, --assert-contains and
--assert-json. A paused monitor stays paused after an edit.
Targets that are refused
A monitor cannot target a private, loopback, link-local or cloud metadata address, localhost, a
single-label host name, or a name ending in .local, .internal, .localhost, .home.arpa,
.lan or .intranet — these only ever resolve inside a private network, and a monitor exists to
watch what the public internet sees. A name that legitimately cannot be resolved is accepted (a DNS
outage is exactly the kind of thing a monitor should report); a name that resolves to one of those
addresses is not. The accepted port range is 80, 443, or 1024-65535.
Secrets go through Connections, never through the monitor
A monitor can send up to 10 request headers, but it never accepts a header value directly — only
a reference to an active, secret-holding connection of kind api_key or basic_auth. Add that
connection first. A secret is a person's job, so you add it yourself: in the Console, open
Connections and notifications (/connections); see
Connect tools and notification delivery. Then ask your
agent to point the header at the connection's id by name (for example an Authorization header
referencing a stored bearer token). Do not paste the secret into the chat with your agent. Framing, routing and
hop-by-hop headers (Host, Content-Length, Connection, Proxy-Authorization and similar) are
refused: a secret can add a credential, never redirect or reshape the request.
TLS certificate expiry
A tls monitor connects to a public DNS host and port, completes a real TLS handshake verified
against the public trust store, and measures the earliest expiry among the verified certificate
chain (the trust anchor itself is excluded) — so if an intermediate certificate expires before the
leaf does, that earlier date is what the monitor reports. It needs:
- a host — a public DNS name (no literal IP address), refused under the same policy as an HTTP
target: no private, loopback, link-local or metadata address,
localhost, single-label names, or.local/.internal/.localhost/.home.arpa/.lan/.intranetsuffixes; - a port — 443 (default), 465, 636, 853, 993, 995, or 1024-65535;
- warning_days — fail the check once the chain is within this many days of expiring, 1-180 (default 21); and
- an interval (1-24 hours, 3600-86400 seconds, default 6 hours) and regions, exactly like an HTTP monitor — a certificate check still runs from your configured regions and is still decided by a majority of them.
Every failure is one of three closed outcomes: expiring (within warning_days of expiry),
expired, or invalid — any other verification failure (untrusted issuer, wrong host name, not
yet valid). All three are failures a client of the endpoint would also see. If the target cannot be
reached at all, or the handshake fails before any certificate is presented, that region's result is
unknown coverage — never a failure, and never an invented date. A completed check's evidence
(measured expiry, days remaining, subject and issuer common names, the leaf certificate's SHA-256
fingerprint, and whether the chain verified) is attached to that region's round result.
Domain registration expiry
A domain monitor reads a domain's registration expiry from its authoritative registry over RDAP,
located automatically through the public IANA bootstrap registry — you never choose or configure a
registry yourself. It needs:
- a domain — stored as its registrable domain under a public suffix (for example
status.example.co.ukis stored and checked asexample.co.uk; a bare public suffix likecom, or a name under a private suffix such asexample.github.io, is refused); and - warning_days — fail once the registration is within this many days of expiring, 1-365 (default 30). Its interval is 6-24 hours (21600-86400 seconds, default 24 hours), and it takes no regions: one lookup, run once by the platform, covers every region a monitor of another kind would need several of.
A domain check only ever fails on an expiring or expired registration date it actually read. Every other registry answer is unknown coverage, never a date and never a failure, with a closed reason: no RDAP service exists for that registry, the registry redacts the expiration date, the record carries no expiration event at all, the registry has no record of the domain, or the registry is rate-limiting or failing lookups (retried with growing backoff, up to six hours between attempts). A registration date the platform already knows survives a registry that starts failing or throttling afterward — the last successful lookup keeps being judged for up to 72 hours before it too becomes unknown, so a brief registry outage never manufactures either a false pass or a false alarm.
Heartbeats
A heartbeat monitor never probes anything: instead, the job it watches pings the platform, and the
monitor expects that ping on a schedule. There is no target to configure and no regions — the
schedule itself is the check.
Schedule: interval or cron
Set exactly one of:
- interval_seconds — expect a successful ping at most this long after the previous one settled, 60 seconds to 7 days; or
- cron — a five-field cron expression (
minute hour day-of-month month day-of-week, the@daily/@hourly/@weekly/@monthly/@yearlyshorthands accepted) evaluated in timezone (an IANA zone; default UTC).
grace_seconds (60-86400, default 300) is how long after the expected time a successful ping may still arrive — set it to comfortably cover the job's own runtime if it reports a start ping, not just success or failure.
Expectations begin at activation (creation, or the moment a paused heartbeat is resumed): a job that has never pinged is judged missed at its very first expected run plus grace, not held open indefinitely. Editing the schedule, like editing any monitor's configuration, restarts the schedule from the moment of the edit rather than continuing to judge against the old one.
Pausing discards unfinished expectations and marks unused accepted pings as ignored. Completed history stays available. Resuming starts fresh expectations; a completion paired with a start from before the pause is late and cannot satisfy the resumed heartbeat. Report a new run to satisfy its new expectation. Repeating pause or resume when already in that state changes nothing.
Across a daylight-saving change, a cron schedule's wall-clock time that is skipped entirely is
expected once, at the moment the gap ends; one that occurs twice is expected once, at its first
occurrence — neither change manufactures a missed run or a duplicate one.
If the platform itself falls behind by more than an hour (a genuine backlog, not a single missed run), it does not replay every expected run that piled up: it skips ahead to the run closest to now and resumes judging from there.
Reporting a run
The job reports itself with GET or POST requests to the ping endpoint, sending its token in the
X-Anectico-Heartbeat-Token header, with an empty request body and run details in the query string.
A token query parameter or any non-empty body is refused with 400, even if the header contains a
valid token; no ping is recorded. A token never belongs in a URL, where it ends up in access logs
and shell history. Report:
event=startwhen the run begins (only needed if you want overlapping or long runs tracked individually);event=successorevent=failwhen it ends; or- an exit code (0-255) instead of an explicit event —
0implies success, anything else implies failure, and one that contradicts an expliciteventis refused.
An optional run id (up to 64 characters) pairs a start with its later success/fail; without
one, a completion pairs with the latest unfinished start that also had none. A repeated start or
completion of the same run id changes nothing and is reported back as a duplicate — safe to retry a
ping that may not have arrived. Overlapping runs (a new run starts before the previous one
finished) are tracked individually: the expected run is satisfied by a success once every run
started within its window has finished, and any one of them failing fails it. A completion that
arrives for a run whose expected window has already been judged is late — it cannot retroactively
satisfy that run, and it does not spuriously satisfy the next one either. A ping that arrives before
its round's window has opened (a manual run between scheduled times) is recorded but ignored —
it settles nothing. A paused heartbeat's pings are accepted but ignored and not stored; pings are
capped at 30 accepted per minute per monitor.
Every ping — its event, run id, exit code, timestamp, whether it was a duplicate, and how its round
eventually counted it (counted, late, ignored) — is visible on the monitor's ping history.
Tokens and rotation
A heartbeat monitor's ping token (anhb_ followed by 43 characters) is shown exactly once: in
the response that creates the monitor, or the response that rotates it. It is not recoverable after
that — hold it as a secret in whatever runs the job, the same as any other credential, and never
paste it into a crontab, a script body, or a ticket. If it is lost, rotate it: the previous token
stops working the instant a new one is issued, so update the job in the same step.
Snippets
These two shapes cover most jobs. Neither one embeds a real token — replace $TOKEN with the value
you received, held as a secret, and substitute your own script:
A crontab line that reports only success or failure, with no run id:
0 2 * * * /job.sh && curl -fsS -m 10 --retry 3 -X POST -H "X-Anectico-Heartbeat-Token: $TOKEN" "https://<api>/api/v1/heartbeats/ping"
A wrapper that reports a start, then the outcome, paired by run id (recommended once a run can
overlap the next, or you want RUN_UNFINISHED distinguished from a plain miss):
RID=$(date +%s)
curl -fsS -m 10 --retry 3 -X POST -H "X-Anectico-Heartbeat-Token: $TOKEN" "https://<api>/api/v1/heartbeats/ping?event=start&run_id=$RID"
/job.sh
curl -fsS -m 10 --retry 3 -X POST -H "X-Anectico-Heartbeat-Token: $TOKEN" "https://<api>/api/v1/heartbeats/ping?exit_code=$?&run_id=$RID"
anectico monitors heartbeat snippet <monitor-id> prints both, and
anectico monitors heartbeat rotate <monitor-id> --yes replaces a lost token.
Maintenance windows
A maintenance window suppresses paging, not checking: every monitor in its scope keeps running its schedule and every round is still recorded truthfully — health still moves to failing, an already-open episode still receives its updates — but a failure that would otherwise open a new episode opens none while the window is active. This is deliberately different from pausing a monitor, which stops it from running or recording anything at all; use a maintenance window when you still want the history, and pause when you do not want the check to run.
A window covers either every monitor in a project (including ones created while it is active) or an explicit list of up to 200 monitors, between a start (which may be in the past — the window is then active immediately) and an end (in the future, at most 14 days after the start, itself at most a year ahead). A project may hold at most 50 upcoming-and-active windows at once.
When the window ends, any monitor in its scope that is still failing — and has no open episode of its own, because maintenance suppressed opening one — pages once, at that moment, rather than waiting for its next scheduled failed round (which, for a domain or heartbeat check running once a day, could be a day away). A monitor whose response was separately ended by a responder (see When a responder resolves a still-failing monitor) does not page when the window ends either — that override stands until a qualified recovery or an explicit rearm, exactly as it would with no maintenance window involved. Two overlapping windows covering the same monitor: ending the earlier one pages nothing while the later one still covers it; ending the one that covers it last is what pages.
Editing a window depends on its state: an upcoming one accepts any change to its scope, bounds or time zone; an active one accepts only its name and its end (you can shorten or extend an active window, or rename it, but not change what it covers); a completed or canceled one accepts no change at all. Ending a window early cancels an upcoming one outright, or ends an active one right now (with the paging behavior above). A window, once ended, cannot be restarted — create a new one.
How a round is decided
This section covers HTTP and TLS monitors, whose rounds run from configured regions. A domain round has a single platform-run lookup in place of regional votes, and a heartbeat round is judged from pings rather than regions (see Heartbeats); both still drive the same consecutive failure/recovery counters and response gate described below.
Every round probes every one of a monitor's regions independently, and the round's verdict is decided by a majority of those regions — 2 of 3 with three regions, 1 of 1 with a single region, 3 of 5 with five. A region that never reports a valid result (it was never reached, or the probe itself could not produce a clean observation) is unknown coverage: it never counts as healthy, and it never, by itself, confirms an outage.
- If a majority of regions report a failure, the round is failed.
- If a majority report success, the round is healthy.
- Anything else — including an even split, or too few regions reporting at all — is unknown, and an unknown round moves nothing: it neither opens nor updates a failing episode, and it never counts toward recovery.
A round's coverage tells you how much of that majority you can trust: full (every region
reported), degraded (enough regions reported to decide, but not all of them), or unknown (too
few did — the round decided nothing).
A single failed round does not page. It takes failure_rounds consecutive failed rounds (of the monitor's current configuration) to open one failing episode — one alert, no matter how many rounds that episode goes on to repeat. Recovery is symmetric: recovery_rounds consecutive healthy rounds close it. Editing a monitor's check or schedule restarts both counters from zero under the new configuration, though an already-open episode stays open and is simply updated by the next failed round rather than opening a second one.
Regional results and timing
Every round's regional results are readable through the API, MCP (list_monitor_rounds) and the CLI
(anectico monitors rounds <monitor-id>): each region's outcome (success, failure, invalid, or missing when
no result arrived in time), the response status code, a closed error class when it did not succeed
(for example DNS_FAILURE, TLS_FAILURE, CONNECT_FAILURE, TIMEOUT, or STATUS_MISMATCH when
the response came back but not with an accepted status), and per-phase timing in milliseconds (DNS,
connect, TLS, time to first byte, and the total). A phase that did not happen for that request (a
reused connection, plain HTTP) reads as zero, not as missing data.
Probe regions
A monitor's regions come from your organization's probe region catalog. Each region is served by
one or more regional probe executors, and the catalog itself — separate from any one monitor — is
readable from the same surfaces: serving (an executor polled recently), stale (one was seen
before, but not recently), or absent (never seen). A region that is not currently serving still
belongs to a monitor's configuration; its rounds simply produce unknown coverage for that region
until an executor resumes polling it.
Routing and what pages
A monitor pages through exactly one route, chosen when you create it and changeable later:
- Its own escalation policy. The monitor creates and owns a dedicated alert rule the moment it is created. You never edit that rule directly; changing the monitor's escalation policy or severity changes the rule.
- A service that owns it. The monitor has no alert rule of its own and pages through the service's — see Monitors owned by a service.
An active monitor needs routing coverage: an escalation policy that exists, is active, and has at
least one level — its own, or its owning service's. Creating or resuming a monitor without that
coverage is refused, naming which part is missing (missing_policy, policy_inactive or
policy_has_no_levels). Create the monitor paused if you want to finish setting up the policy
first, or fix the policy and try again.
Monitors owned by a service
Give a monitor a service_key and the catalog service
with that key owns it. That one choice decides who owns the monitor's failures everywhere:
- Paging goes through the service's escalation policy at the service's severity. The monitor
has no routing of its own — sending an escalation policy or severity together with a
service_keyis refused (MONITOR_ROUTING_CONFLICT), and re-routing an owned monitor directly is refused too (MONITOR_SERVICE_OWNED): change the service's routing instead. - Incidents opened for the monitor's failures carry the service's ownership snapshot — owner, schedule, escalation policy and runbook checklist — exactly like an incoming alert routed to that service.
- Response reports attribute those incidents to the service's key, so filtering a report by the service includes the monitor's failures.
The service must exist in the monitor's project, be live, and have its alert rule bound
(MONITOR_SERVICE_INVALID otherwise, or MONITOR_SERVICE_NOT_ROUTABLE while the service is still
being set up). Only one rule ever pages for the failure, so the monitor can never page twice.
Moving a monitor between routes. Setting a service_key on a monitor that had its own routing
removes its own alert rule. Clearing it (an empty service_key, sent together with the new routing)
gives the monitor its own rule again. Moving to another service is the same change with the new key.
If the monitor is failing when it moves, the failure moves with it: the alert on the old route
closes with resolution withdrawn — never as recovered, because the check has not recovered — and
the new route is paged at once rather than at the next failed round. A move whose old rule could
not be removed yet answers MONITOR_RELEASE_PENDING; the monitor already pages through its new
service, and repeating the same change finishes the removal.
If the owning service is deleted, the monitor cannot page until you give it another service or
its own routing; its routing then reads unverified, naming the deleted service.
Every failing occurrence becomes one alert, however many rounds repeat the same failure; recovery closes it the same way an ordinary alert closes. From there, acknowledging, resolving, and escalation behave exactly like any other alert — see Create useful alerts and Manage incidents and on-call.
When a responder resolves a still-failing monitor
Every failing episode is one alert, and a monitor's alert behaves like any other: it can be acknowledged, resolved, and escalated the normal way (see Create useful alerts). Resolving it while the check is still failing — directly, or by resolving an incident linked to it — is a deliberate override: the responder is telling the system "I've handled this, stop paging me about it." From that point, further failed rounds keep updating the monitor's evidence (so the history stays honest), but paging is held until either:
- a qualified recovery — enough consecutive healthy rounds to close the episode the normal way; or
- an explicit rearm — telling the monitor to start paging again immediately, even though the check has not recovered. Re-arming requires the monitor's current response-gate version (shown on every read) so a rearm sent against a stale copy of the monitor is refused rather than silently reapplied; the very next failed round then opens a new episode and pages once.
Two related actions do not rearm a monitor on their own: reopening an incident linked to the alert is a change to the incident, not a resumption of paging for the check itself, and a silence ending (by expiry or by being removed early) resumes normal alerting exactly the way it would for any other alert, without needing a rearm.
Pause, resume and delete
Pausing a monitor stops scheduling new rounds immediately; any rounds already in flight are cancelled. Pausing does not touch an alert that is already firing — it keeps firing until it is resolved, or until the monitor is resumed and recovers on its own. Resuming restarts the schedule from now, and re-checks routing coverage exactly like creating an active monitor does — a monitor cannot resume into a state where nothing would page.
Deleting a monitor is a tombstone: scheduling stops, its dedicated alert rule is removed, and its
full history and audit trail remain readable. A monitor owned by a service has no rule to remove:
if it is failing, its open alert on the service's rule closes with resolution withdrawn, which
closes the incident linked to it. A deleted monitor's name is not held back for reuse.
Deleting the monitor's project (or organization) is different: its monitors and their history
are erased, not kept as tombstones, and no further checks run.
Both an edit and a delete require you to state the exact revision you are changing (shown on every
read), so a change from a stale copy of the monitor is refused rather than silently overwriting a
concurrent one.
Test a monitor on demand
A test (test_monitor, or anectico monitors test <monitor-id>) runs one bounded, extra round right
now, against every one of the monitor's regions, without waiting for the schedule. By default a test
round's result never pages. Ask for page on failure (the CLI flag is --page) when you want a
failing test to page like a real round would. Only one test round can be open per monitor at a time;
wait for it to finish (or expire) before starting another. A paging test against a monitor whose
routing is not yet covered is refused for the same reason an active monitor would be. Read its outcome
the same way you read any other round: monitors rounds <id> --kind test (CLI),
GET .../rounds?kind=test (API), or list_monitor_rounds (MCP). The round can take a moment, so read
again until it is complete.
A test of a domain monitor whose last registry lookup failed asks the registry again
rather than repeating the failure it already recorded — at most once a minute per domain, however
many tests are started, so a test within a minute of the last lookup reports that lookup's result.
A registration date that is already known, or any other answer the registry gave (a redacted date,
no record, no RDAP service, rate limiting), is reported as it stands without asking again.
A heartbeat has no check to run on demand — there is nothing for a test to probe — so it is
refused with a closed reason (MONITOR_TEST_UNSUPPORTED); test it by sending it a ping instead.
Reading a monitor's state
Every monitor read includes:
- its kind (
http,tls,domainorheartbeat) and current state (active,pausedor a deleted tombstone) and configuration revision; - its owning service (
service_key, empty when it has its own routing) and the alert rule it pages through — its own, or the owning service's (then the rule'sbinding_stateisnone: the monitor has no rule of its own); - that rule's routing status —
covered, or a reason it is not (missing_policy,policy_inactive,policy_has_no_levels, orunverifiedwhen it could not be checked); - diagnostics: why it is not currently producing results, as one or more closed reason codes, plus counts of jobs waiting and jobs that expired unrun in the last 24 hours;
- health: its current state (
unknown,healthyorfailing), how many consecutive failed/healthy rounds have been counted under its current configuration, the last reduced round's verdict and coverage, the open episode's id and alert (if any), and its response-gate state and version (see When a responder resolves a still-failing monitor); - for a heartbeat, its token prefix (never the token itself), when it was last rotated, the most recent accepted ping, and the run it is currently waiting for (expected time and deadline); and
- the active maintenance window covering it, if any (id, name, end time).
Executor status says whether the monitor's regions are currently being probed: available
(every region), degraded (enough regions for a majority, not all — reason EXECUTOR_DEGRADED) or
unavailable (too few to decide a round — reason EXECUTOR_UNAVAILABLE). A region that produces no
result is unknown coverage: it never counts as healthy and never, by itself, as an outage. A
domain or heartbeat monitor has no regions of its own — its executor status reflects the platform's
own worker instead, so it is still accurate when that worker is down. The other reason codes explain
the rest: MONITOR_PAUSED (paused, not scheduling), MONITOR_DELETED (tombstoned),
MONITOR_BINDING_PENDING (its alert rule has not finished being created or removed yet; retrying
the create, resume or delete completes it), ROUTING_NOT_COVERED (its escalation policy cannot
currently page), RESPONSE_ENDED_WHILE_UNHEALTHY (paging is held; see above), and
MAINTENANCE_ACTIVE (a maintenance window currently covers it).
A monitor's history lists every round and test job, newest first, each with its state (admitted and waiting, claimed, completed, fenced by a later edit or pause, or expired) and, for a fenced or expired job, why. A monitor's rounds list is the reduced view of that same history: one entry per round with its verdict, coverage and each region's result (see Regional results and timing). Both are kept for 90 days and survive pausing or deleting the monitor.
Scopes
| Scope | Grants |
|---|---|
monitoring:read |
List and read monitors, their configuration, routing, diagnostics, health, history, rounds and audit; list probe regions; list a heartbeat's pings; list and read maintenance windows. Required by every monitor operation, including the writes below. |
monitoring:write |
Create, edit, pause, resume, test-run and re-arm monitors; rotate a heartbeat's ping token; create, edit and end maintenance windows. |
monitoring:delete |
Delete monitors. |
Because a monitor's routing lives on an alert rule, creating, resuming or re-routing a monitor also
requires alerts:rules:write, and deleting one also requires alerts:rules:delete. Giving a monitor
an owning service reads that service, so it also requires incidents:read; moving a monitor that
has its own rule to a service removes that rule, so it also requires alerts:rules:delete. A header
that references a connection also requires connections:read. See
permission scopes for who holds each scope by default.
API, MCP and CLI
REST is documented at REST API reference § External monitoring.
The list_monitors, get_monitor, list_monitor_history, list_monitor_audit,
list_monitor_rounds, list_probe_regions, list_heartbeat_pings, list_maintenance_windows and
get_maintenance_window reads and the create_monitor, update_monitor, pause_monitor,
resume_monitor, test_monitor, delete_monitor, rearm_monitor, rotate_heartbeat_token,
create_maintenance_window, update_maintenance_window and end_maintenance_window actions are
available over MCP (create_monitor and update_monitor take service_key, and update_monitor
takes clear_service with an escalation policy; every write previews before it applies; a heartbeat's ping token is delivered
only in the one-time text of a create_monitor or rotate_heartbeat_token response, never in the
structured content a later read could echo back).
anectico monitors create --name "Checkout API" \
--url https://status.example.com/health --method GET \
--expect 200-299 --assert-json status=\"ok\" \
--interval 60 --failure-rounds 2 --recovery-rounds 2 \
--escalation-policy 6f6b1a2e-9c3d-4b8e-8e2a-1f7b6a5d4c3b --severity critical
# Owned by a catalog service: pages through the service, and its incidents carry the service's ownership
anectico monitors create --name "Checkout API" --url https://status.example.com/health --service checkout
anectico monitors update <monitor-id> --revision 2 --service checkout
anectico monitors update <monitor-id> --revision 2 --clear-service \
--escalation-policy 6f6b1a2e-9c3d-4b8e-8e2a-1f7b6a5d4c3b
anectico monitors list
anectico monitors get <monitor-id>
anectico monitors test <monitor-id> --page
anectico monitors rounds <monitor-id> --kind scheduled
anectico monitors history <monitor-id> --kind scheduled
anectico monitors regions
anectico monitors rearm <monitor-id> --gate-version 3
anectico monitors delete <monitor-id> --revision 1 --yes
# TLS certificate expiry
anectico monitors create --name "Shop cert" --kind tls --host shop.example.com \
--warning-days 21 --escalation-policy 6f6b1a2e-9c3d-4b8e-8e2a-1f7b6a5d4c3b
# Domain registration expiry
anectico monitors create --name "example.com registration" --kind domain \
--domain example.com --warning-days 30 --escalation-policy 6f6b1a2e-9c3d-4b8e-8e2a-1f7b6a5d4c3b
# Heartbeat, and its token / snippet / ping-history commands
anectico monitors create --name "Nightly ETL" --kind heartbeat --cron "0 2 * * *" \
--timezone America/New_York --grace 600 --escalation-policy 6f6b1a2e-9c3d-4b8e-8e2a-1f7b6a5d4c3b
anectico monitors heartbeat snippet <monitor-id>
anectico monitors heartbeat pings <monitor-id>
anectico monitors heartbeat rotate <monitor-id> --yes
# Maintenance windows
anectico monitors maintenance create --name "db upgrade" --scope monitors --monitor <monitor-id> \
--starts-at 2026-10-01T09:00:00Z --ends-at 2026-10-01T11:00:00Z
anectico monitors maintenance list --state active
anectico monitors maintenance update <window-id> --revision 1 --ends-at 2026-10-01T12:00:00Z
anectico monitors maintenance end <window-id> --revision 2 --yes