Monitoring and retention

Health

Endpoint Answers
GET /healthz 200 while the process serves requests.
GET /readyz 200 when the database answers, 503 otherwise.

Both are open without a token. In containers use shipyard -healthcheck — see Run with Docker.

Prometheus metrics

Set observability.metrics: true and scrape /metrics:

scrape_configs:
  - job_name: shipyard
    static_configs:
      - targets: ["127.0.0.1:8080"]
    authorization:
      credentials_file: /etc/prometheus/shipyard.token

shipyard_runs{state="…"} reports stored runs per state: queued, running, cancelling, succeeded, failed, cancelled. Leave out authorization when no API token is configured.

Alertmanager

Shipyard receives grouped notifications on POST /webhooks/alertmanager:

# alertmanager.yml
receivers:
  - name: shipyard
    webhook_configs:
      - url: http://127.0.0.1:8080/webhooks/alertmanager
        http_config:
          authorization:
            credentials_file: /etc/alertmanager/shipyard.token
# .shipyard.yaml
alerts:
  token_env: ALERTMANAGER_TOKEN
  project_label: shipyard_project
  enqueue_diagnostics: true

Each alert needs the label shipyard_project with the ID of an enabled project. Shipyard records firing and resolved alerts; with enqueue_diagnostics a firing alert also queues one diagnostic run. A repeated notification for the same alert does not queue a second run. A batch with an unknown project is rejected as a whole and nothing is recorded.

Traces

Set observability.otel_endpoint to an OTLP/HTTP collector, e.g. http://127.0.0.1:4318. Shipyard exports a span per run and per pipeline stage. Export failures never block runs.

Service logs

Under systemd, Shipyard's own log goes to the journal:

journalctl -u shipyard -f
journalctl -u shipyard --since today -o short-iso

Journal size and retention are host settings (SystemMaxUse in journald.conf). Run logs are separate — they live in logging.directory and follow Shipyard's retention.

Retention

Shipyard cleans up at worker start and every hour:

  • run logs older than retention.logs_days are deleted; the run stays in the history marked logs expired;
  • finished runs older than retention.completed_runs_days are deleted;
  • audit entries and alert records older than retention.audit_days are deleted.

Queued and running work is never touched. Each run's log is capped at logging.max_run_size_mb, and single log messages at 16 KiB. When the log directory reaches retention.max_disk_size_mb, new runs are refused with 507 until retention frees space — newer logs are never deleted to make room.

The database file itself is not counted against the budget. Back it up with the service stopped, or with sqlite3 shipyard.db ".backup backup.db" while it runs.