⌁ Checkpulse

Incidents and silences

Incidents

An incident opens when a check goes down and closes when the job next reports success. Each incident records its cause:

CauseWhen
missedNo ping arrived before the grace period ended.
timeoutA run started but did not finish within the grace period.
failedThe job sent /fail.
exit status NThe job reported a non-zero exit status.
failure keywordA success ping's output matched one of the check's failure keywords.

On the check page, and through the API and MCP, people and agents can:

  • Acknowledge an incident, optionally with a note. This shows others it is being handled and stops reminder alerts.
  • Add notes with findings and actions, so the history explains what happened.

Each check shows uptime and total downtime over the last 7, 30, and 90 days.

Silences

A silence holds a check's alerts for a period you choose (up to 7 days) while monitoring continues:

  • If the check goes down during the silence, the incident is recorded but no alert is sent.
  • When the silence ends, an alert is sent if the check is still down.
  • If the job recovers during the silence, nobody is alerted.

Silences suit planned maintenance and deployments better than pausing, because they end on their own. Ending a silence early sends any alert it held back.

Pausing

Pausing stops monitoring until the check is resumed, either by hand or, unless the check requires a manual resume, by its next ping. A resumed check waits for its next ping like a new check.

Status badges

Every project has public badge URLs that show only overall status (up, late, or down), for the whole project or for one tag:

https://checkpulse.foo/badge/<badge key>.svg
https://checkpulse.foo/badge/<badge key>/<tag>.svg

Replace .svg with .json for {"status": "up", "total": 4, "down": 0}, or .shields for a Shields.io endpoint. Rotate the badge key in project settings to retire old URLs.

Prometheus metrics

GET /api/v1/metrics returns the project's checks in Prometheus text format: checkpulse_check_up, checkpulse_check_started, checkpulse_check_last_ping_timestamp_seconds, checkpulse_tag_up, checkpulse_checks_total, and checkpulse_checks_down_total. Scrape it with a read-only API key:

scrape_configs:
  - job_name: checkpulse
    scheme: https
    metrics_path: /api/v1/metrics
    authorization:
      credentials: cpk_…
    static_configs:
      - targets: ["checkpulse.example"]

Put your Checkpulse host name in targets, without https://.

View as Markdown