# Incidents and silences

## Incidents

An incident opens when a check goes down and closes when the job next reports success. Each incident records its cause:

| Cause | When |
|---|---|
| missed | No ping arrived before the grace period ended. |
| timeout | A run started but did not finish within the grace period. |
| failed | The job sent `/fail`. |
| exit status N | The job reported a non-zero exit status. |
| failure keyword | A success ping's output matched one of the check's failure keywords. |

On the check page, and through the API and MCP, people and agents can:

- **Acknowledge** an incident, optionally with a note. This shows others it is being handled and stops reminder alerts.
- **Add notes** with findings and actions, so the history explains what happened.

Each check shows uptime and total downtime over the last 7, 30, and 90 days.

## Silences

A silence holds a check's alerts for a period you choose (up to 7 days) while monitoring continues:

- If the check goes down during the silence, the incident is recorded but no alert is sent.
- When the silence ends, an alert is sent if the check is still down.
- If the job recovers during the silence, nobody is alerted.

Silences suit planned maintenance and deployments better than pausing, because they end on their own. Ending a silence early sends any alert it held back.

## Pausing

Pausing stops monitoring until the check is resumed, either by hand or, unless the check requires a manual resume, by its next ping. A resumed check waits for its next ping like a new check.

## Status badges

Every project has public badge URLs that show only overall status (up, late, or down), for the whole project or for one tag:

```text
https://checkpulse.foo/badge/<badge key>.svg
https://checkpulse.foo/badge/<badge key>/<tag>.svg
```

Replace `.svg` with `.json` for `{"status": "up", "total": 4, "down": 0}`, or `.shields` for a [Shields.io endpoint](https://shields.io/badges/endpoint-badge). Rotate the badge key in project settings to retire old URLs.

## Prometheus metrics

`GET /api/v1/metrics` returns the project's checks in Prometheus text format: `checkpulse_check_up`, `checkpulse_check_started`, `checkpulse_check_last_ping_timestamp_seconds`, `checkpulse_tag_up`, `checkpulse_checks_total`, and `checkpulse_checks_down_total`. Scrape it with a read-only API key:

```yaml
scrape_configs:
  - job_name: checkpulse
    scheme: https
    metrics_path: /api/v1/metrics
    authorization:
      credentials: cpk_…
    static_configs:
      - targets: ["checkpulse.example"]
```

Put your Checkpulse host name in `targets`, without `https://`.
