3 Health metrics and logs
Tiago Águeda edited this page 2026-09-22 15:51:02 +02:00

Health, metrics and logs

Three addresses answer machines about the instance itself, rather than about anybody's job search: /healthz for whatever restarts Postulo when it stops, /metrics for Prometheus, and /logs for a log collector. They live outside the API because they are not a person's: no API token opens them, and none of them is scoped to an account.

Address Off by default Guarded by Answers
GET /healthz no, always on nothing JSON: the application and the database are up
GET /metrics yes POSTULO_METRICS_TOKEN, optional Prometheus text: counts, nothing about anybody
GET /logs yes POSTULO_LOGS_TOKEN, required One JSON record per line

/healthz

For a container's health check, a load balancer or an uptime monitor. It asks the database for a connection and nothing more, so it is cheap to call often.

{"status": "ok", "database": "ok", "version": "0.2.1"}

200 when the database answered. 503 with {"status": "error", "database": "unavailable"} when it did not. The image's own health check calls it over plain HTTP on 127.0.0.1:8000.

/metrics

Prometheus metrics, in the text exposition format (text/plain; version=0.0.4), read from the database at the moment of the scrape and never cached.

Switched on with POSTULO_METRICS_ENABLED=true. Off, the address is a 404, so a stranger cannot tell it exists.

The token is optional, on purpose. A metric is a count of things on the instance and carries nothing about anybody, so it may reasonably be left open on a private network. With POSTULO_METRICS_TOKEN empty, anybody who can reach the instance can read the numbers, and Server settings says so. With a token set, a request needs Authorization: Bearer <token> or it gets 401.

Metric Labels What it counts
postulo_info version, python, django Always 1; the labels say what is running
postulo_database_reachable 1 when the database answered, 0 when it did not. When it is 0, it is the last metric in the response.
postulo_migrations_applied 1 when every migration has been applied, 0 when some are outstanding
postulo_records kind: people, applications, listings, companies, documents How many of each kind of record exist
postulo_pending kind: document_copies, captures, suggestions, reminders Work waiting to be done or looked at. reminders counts the ones that have fallen due, not every reminder anybody has set for the months ahead.
postulo_overdue kind: reminders Work that is not merely waiting but late: due for a quarter of an hour and still not announced. On a healthy instance this is 0.
postulo_failures kind: document_copies, syncs Things that tried, did not succeed and are still in that state. This is the one to alert on.
postulo_scheduler_last_pass_timestamp_seconds When the scheduler last finished a pass, as a Unix time. 0 means it has never finished one here — which is also what an instance running no scheduler reports.
postulo_plugins state: enabled, disabled Installed plugins, by whether they are switched on

Every metric is a gauge. A scrape configuration:

scrape_configs:
  - job_name: postulo
    scheme: https
    metrics_path: /metrics
    scrape_interval: 1m
    authorization:
      credentials: YOUR_METRICS_TOKEN   # leave the block out when no token is set
    static_configs:
      - targets: ["postulo.example.org"]

Noticing that the scheduler has stopped

The scheduler is a loop in a container, and the way it fails is quietly: nothing errors, reminders simply stop arriving, and the first person to notice is the one who missed a deadline. Two rules cover it — the first says the loop has stopped, the second says it is running but not getting through, which is the one a broken notifier produces.

groups:
  - name: postulo
    rules:
      - alert: PostuloSchedulerStopped
        # Also fires on an instance that has no scheduler at all, which is worth knowing
        # once. Drop the rule, or run one, rather than silencing it.
        expr: time() - postulo_scheduler_last_pass_timestamp_seconds > 1800
        for: 10m
        annotations:
          summary: No scheduler pass for half an hour

      - alert: PostuloRemindersOverdue
        expr: postulo_overdue{kind="reminders"} > 0
        for: 30m
        annotations:
          summary: Reminders have fallen due and nobody has been told

      - alert: PostuloDeliveriesFailing
        expr: postulo_failures > 0
        for: 1h
        annotations:
          summary: "{{ $labels.kind }} have been failing for an hour"

If you run the scheduler from cron rather than with --loop, set the first threshold to comfortably more than the gap between runs.

The scheduler service in the Compose files has a healthcheck of its own that reads the same heartbeat from the data volume, so docker compose ps shows it as unhealthy once it has stopped going round. Before 0.4.0 it inherited the image's healthcheck, which curls the web port that this container does not serve — so it was always unhealthy and said nothing.

/logs

The log, for a collector somewhere else on your network — Grafana Alloy, Vector, Promtail — that can poll a URL more easily than it can have log shipping arranged off the host. Anybody who can read the container's standard output should carry on doing that.

Switched on with POSTULO_LOGS_ENDPOINT_ENABLED=true; off, it is a 404. A token is required: a log entry names connections, companies and applications, and an open log endpoint is a data leak with a URL. With the endpoint on and POSTULO_LOGS_TOKEN empty, it refuses with 503 and writes an error to the log saying why, rather than publishing anything because a variable was forgotten. It reads the file Postulo keeps in POSTULO_LOG_DIR, so with that empty there is nothing to serve.

curl -H "Authorization: Bearer YOUR_LOGS_TOKEN" \
  "https://postulo.example.org/logs?since=2026-09-10T08:00:00&level=WARNING"
Parameter Meaning
limit How many records, 200 by default and 1000 at most
level The lowest level wanted: DEBUG, INFO, WARNING, ERROR or CRITICAL
since Only records after this ISO-8601 time — the last time the collector already has

The answer is application/x-ndjson, oldest first, the order a collector appends in: one JSON object per line, with time, level, logger and message, plus whatever fields that record carried -- request_id among them, whenever the line was written while answering a request, during a scheduler pass (pass-…) or by a background errand (errand-<number>). The same id is in the response's X-Request-ID header and at the end of gunicorn's access line, so the three logs can be laid side by side. It is answered once and not streamed. A collector polls, and a stream would hold one of the server's few workers open for as long as the collector wanted.

What /metrics and /logs share

  • Too often is 429, with Retry-After in seconds. /metrics and /logs allow POSTULO_ENDPOINT_RATE requests per calling address, 120/h by default. The limit is per address because a shared token, not an account, is what opens them. A scrape every minute is half of it.
  • A wrong or missing token is 401, with WWW-Authenticate: Bearer (realm="postulo-metrics" or realm="postulo-logs"), and tokens are compared in constant time.
  • Nothing is cached (Cache-Control: no-store).

Every variable is in Configuration. /manifest.webmanifest is not one of these: it is what a phone reads when Postulo is installed on its home screen.