Incidents and silences
Incidents
An incident opens when a check goes down and closes when the job next reports success. Each incident records its cause:
| Cause | When |
|---|---|
| missed | No ping arrived before the grace period ended. |
| timeout | A run started but did not finish within the grace period. |
| failed | The job sent /fail. |
| exit status N | The job reported a non-zero exit status. |
| failure keyword | A success ping's output matched one of the check's failure keywords. |
On the check page, and through the API and MCP, people and agents can:
- Acknowledge an incident, optionally with a note. This shows others it is being handled and stops reminder alerts.
- Add notes with findings and actions, so the history explains what happened.
Each check shows uptime and total downtime over the last 7, 30, and 90 days.
Silences
A silence holds a check's alerts for a period you choose (up to 7 days) while monitoring continues:
- If the check goes down during the silence, the incident is recorded but no alert is sent.
- When the silence ends, an alert is sent if the check is still down.
- If the job recovers during the silence, nobody is alerted.
Silences suit planned maintenance and deployments better than pausing, because they end on their own. Ending a silence early sends any alert it held back.
Pausing
Pausing stops monitoring until the check is resumed, either by hand or, unless the check requires a manual resume, by its next ping. A resumed check waits for its next ping like a new check.
Status badges
Every project has public badge URLs that show only overall status (up, late, or down), for the whole project or for one tag:
https://checkpulse.foo/badge/<badge key>.svg
https://checkpulse.foo/badge/<badge key>/<tag>.svg
Replace .svg with .json for {"status": "up", "total": 4, "down": 0}, or .shields for a Shields.io endpoint. Rotate the badge key in project settings to retire old URLs.
Prometheus metrics
GET /api/v1/metrics returns the project's checks in Prometheus text format: checkpulse_check_up, checkpulse_check_started, checkpulse_check_last_ping_timestamp_seconds, checkpulse_tag_up, checkpulse_checks_total, and checkpulse_checks_down_total. Scrape it with a read-only API key:
scrape_configs:
- job_name: checkpulse
scheme: https
metrics_path: /api/v1/metrics
authorization:
credentials: cpk_…
static_configs:
- targets: ["checkpulse.example"]
Put your Checkpulse host name in targets, without https://.