The health check cannot fail in production: SSL redirect answers it before the view does #82
Labels
No labels
accessibility
authentication
breaking change
bug
documentation
enhancement
interface
internationalisation
observability
security
tier
1
tier
2
tier
3
tier/4
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
Postulo/postulo#82
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Observation
Found while deploying 0.2.0 behind traefik. With
POSTULO_SSL_REDIRECTon — which is theproduction default — the container's health check stops checking anything.
docker/Dockerfile:prod.py:There is no
SECURE_REDIRECT_EXEMPT, soSecurityMiddlewareanswers that plain-HTTPrequest with a 301 to
https://127.0.0.1:8000/healthz— before any view runs, and beforeanything touches the database.
And
curl -fonly fails on 4xx and 5xx. Verified rather than assumed:Empty body, exit 0. The check passes.
What that means
The health check reports healthy whenever
SecurityMiddlewareis loaded, which is always.It would report healthy with the database gone, the migrations unapplied, every view
raising — anything the
healthzview was written to detect:That 503 is unreachable in production. So is the restart that
restart: unless-stoppedplusa failing health check would have produced. Every production deployment of the shipped
image has a liveness probe that cannot fail.
It is invisible for the usual reason: a health check that always passes looks exactly like a
healthy service.
The fix
With a test that asserts
/healthzreturns 200 under production settings with the redirecton, and that everything else still redirects — the second half matters, since an exemption
pattern that is too broad would be worse than the bug.
Worth checking
/metricsat the same time: a Prometheus scraper reaching it over plain HTTPinside the network would be redirected too, and a scraper follows redirects less politely
than a browser does.
Workaround in the meantime
Set
POSTULO_SSL_REDIRECT=falseand let the reverse proxy redirect at its entrypoint, whichis what traefik does anyway. That is what the ragnar deployment does, with a comment saying
why, and it is a perfectly good production posture — but it should be a choice rather than
the only configuration in which the health check works.
Classification
Bug. Not a security hole: the redirect still happens for real traffic, and nothing is
exposed. What is lost is the ability to notice that the application has stopped working.