Skip to main content
GET /core/v1/metrics?range=1h|6h|24h|7d reports Core’s own health: its process, execution queue and slots, PostgreSQL and background jobs. It requires the Core key (Core administration API). range is the only parameter, sent at most once; it defaults to 1h. An empty, repeated or unsupported value, or any other parameter, returns 400 invalid_request. When Core cannot read its metrics, the route returns 503 core_metrics_unavailable. When only some measurements fail, the response is still 200 with service.status set to degraded and each missing value set to null. The response never contains database or native error text, credentials, bodies, resource IDs or tenant labels.

Time and missing data

range in the response has UTC RFC 3339 start and end and resolution_seconds. end is the most recent complete bucket boundary; the interval is [start, end), so the current partial bucket is never included.
  • Execution slots, connected daemons, the connection pool, Go heap and goroutines are read when the request arrives. Process CPU, RSS and limits, queue counts and database size come from a sample Core takes every 30 seconds; a sample older than 60 seconds is not reported as current.
  • Each series bucket reports the highest value observed in it, not every intermediate peak. Missing observations and the process’s partial first bucket are null.
  • Samples and rejection counts live in memory for seven days, plus two hours of padding for bucket alignment. A restart loses them; Core does not backfill. Turn history comes from PostgreSQL and survives restarts.
  • execution.unavailable is null when the interval starts before this process began observing, and zero for a fully observed interval without rejections.
  • An empty queue has a count of zero and a null oldest age. No started Turns or no successful ping samples give null percentiles, not zero latency. Percentiles of periodic pings (p50, p95) are linearly interpolated; a request never triggers a ping.

Fields

The response has object: "core.metrics", range, service, execution, database, jobs and process. Every numeric value and service.execution_owner is nullable; each series always lists every complete bucket of the range. Core resolves its own cgroup, including nested and subtree mounts, and reports that cgroup’s limits, not the host’s or an ancestor’s. Unreadable or malformed values are null. Non-Linux builds report null CPU, RSS and memory limit, with GOMAXPROCS as the CPU limit. Missing process measurements alone do not make the service degraded.

Background jobs

jobs lists scheduler, runtime_sampler, history_cleanup and audit_cleanup. Each has status (unknown before its first run, ok, failing, or stopped when its loop is disabled or ended), last_run_at (when the last pass finished), processed and failed, both describing the last pass.
  • The scheduler’s processed counts the Turn and Environment work it selected in its last poll.
  • The Runtime sampler’s processed and failed count the targets its last sweep observed and failed to observe.
  • A cleanup job’s processed counts the rows it removed.
  • For the scheduler and the cleanup jobs, a failed pass sets processed to null and failed to 1; failed never estimates lost rows or failed Turns.