# Monitoring and retention

## Health

| Endpoint | Answers |
|---|---|
| `GET /healthz` | `200` while the process serves requests. |
| `GET /readyz` | `200` when the database answers, `503` otherwise. |

Both are open without a token. In containers use `shipyard -healthcheck` — see [Run with Docker](/docker/).

## Prometheus metrics

Set `observability.metrics: true` and scrape `/metrics`:

```yaml
scrape_configs:
  - job_name: shipyard
    static_configs:
      - targets: ["127.0.0.1:8080"]
    authorization:
      credentials_file: /etc/prometheus/shipyard.token
```

`shipyard_runs{state="…"}` reports stored runs per state: `queued`, `running`, `cancelling`, `succeeded`, `failed`, `cancelled`. Leave out `authorization` when no API token is configured.

## Alertmanager

Shipyard receives grouped notifications on `POST /webhooks/alertmanager`:

```yaml
# alertmanager.yml
receivers:
  - name: shipyard
    webhook_configs:
      - url: http://127.0.0.1:8080/webhooks/alertmanager
        http_config:
          authorization:
            credentials_file: /etc/alertmanager/shipyard.token
```

```yaml
# .shipyard.yaml
alerts:
  token_env: ALERTMANAGER_TOKEN
  project_label: shipyard_project
  enqueue_diagnostics: true
```

Each alert needs the label `shipyard_project` with the ID of an enabled project. Shipyard records firing and resolved alerts; with `enqueue_diagnostics` a firing alert also queues one diagnostic run. A repeated notification for the same alert does not queue a second run. A batch with an unknown project is rejected as a whole and nothing is recorded.

## Traces

Set `observability.otel_endpoint` to an OTLP/HTTP collector, e.g. `http://127.0.0.1:4318`. Shipyard exports a span per run and per pipeline stage. Export failures never block runs.

## Service logs

Under systemd, Shipyard's own log goes to the journal:

```sh
journalctl -u shipyard -f
journalctl -u shipyard --since today -o short-iso
```

Journal size and retention are host settings (`SystemMaxUse` in `journald.conf`). Run logs are separate — they live in `logging.directory` and follow Shipyard's retention.

## Retention

Shipyard cleans up at worker start and every hour:

- run logs older than `retention.logs_days` are deleted; the run stays in the history marked *logs expired*;
- finished runs older than `retention.completed_runs_days` are deleted;
- audit entries and alert records older than `retention.audit_days` are deleted.

Queued and running work is never touched. Each run's log is capped at `logging.max_run_size_mb`, and single log messages at 16 KiB. When the log directory reaches `retention.max_disk_size_mb`, new runs are refused with `507` until retention frees space — newer logs are never deleted to make room.

The database file itself is not counted against the budget. Back it up with the service stopped, or with `sqlite3 shipyard.db ".backup backup.db"` while it runs.
