Monitoring and retention
Health
| Endpoint | Answers |
|---|---|
GET /healthz |
200 while the process serves requests. |
GET /readyz |
200 when the database answers, 503 otherwise. |
Both are open without a token. In containers use shipyard -healthcheck — see Run with Docker.
Prometheus metrics
Set observability.metrics: true and scrape /metrics:
scrape_configs:
- job_name: shipyard
static_configs:
- targets: ["127.0.0.1:8080"]
authorization:
credentials_file: /etc/prometheus/shipyard.token
shipyard_runs{state="…"} reports stored runs per state: queued, running, cancelling, succeeded, failed, cancelled. Leave out authorization when no API token is configured.
Alertmanager
Shipyard receives grouped notifications on POST /webhooks/alertmanager:
# alertmanager.yml
receivers:
- name: shipyard
webhook_configs:
- url: http://127.0.0.1:8080/webhooks/alertmanager
http_config:
authorization:
credentials_file: /etc/alertmanager/shipyard.token
# .shipyard.yaml
alerts:
token_env: ALERTMANAGER_TOKEN
project_label: shipyard_project
enqueue_diagnostics: true
Each alert needs the label shipyard_project with the ID of an enabled project. Shipyard records firing and resolved alerts; with enqueue_diagnostics a firing alert also queues one diagnostic run. A repeated notification for the same alert does not queue a second run. A batch with an unknown project is rejected as a whole and nothing is recorded.
Traces
Set observability.otel_endpoint to an OTLP/HTTP collector, e.g. http://127.0.0.1:4318. Shipyard exports a span per run and per pipeline stage. Export failures never block runs.
Service logs
Under systemd, Shipyard's own log goes to the journal:
journalctl -u shipyard -f
journalctl -u shipyard --since today -o short-iso
Journal size and retention are host settings (SystemMaxUse in journald.conf). Run logs are separate — they live in logging.directory and follow Shipyard's retention.
Retention
Shipyard cleans up at worker start and every hour:
- run logs older than
retention.logs_daysare deleted; the run stays in the history marked logs expired; - finished runs older than
retention.completed_runs_daysare deleted; - audit entries and alert records older than
retention.audit_daysare deleted.
Queued and running work is never touched. Each run's log is capped at logging.max_run_size_mb, and single log messages at 16 KiB. When the log directory reaches retention.max_disk_size_mb, new runs are refused with 507 until retention frees space — newer logs are never deleted to make room.
The database file itself is not counted against the budget. Back it up with the service stopped, or with sqlite3 shipyard.db ".backup backup.db" while it runs.