A route that cannot deliver is not a way back in #152

Closed
opened 2026-09-09 11:06:22 +00:00 by tiagoagueda · 2 comments
Owner

Observation

Found while reading the mail plugin for the OAuth work, and worth separating because it is
true today and wants fixing whether or not that lands.

What exists

recovery_routes() decides whether anybody could get back into their account, and
refuse_switching_off() reads it to lock the mail transport while it is the last way in.
That lock is the whole of #104 and it is load-bearing.

What it actually checks is whether a transport is selected:

transport = selected()
if transport is not None and transport.name != without:
    routes.append("email")

Selected, not working. An instance whose SMTP password is wrong, whose relay has been
switched off, whose host no longer resolves, or whose credentials were revoked, still counts
email as a route — and the lock stays shut on the strength of a route that delivers
nothing. The refusal message tells an administrator they may not switch off the transport
because it is protecting accounts it is not, in fact, protecting.

The interface has the missing piece already: Server settings → Email has a Send a test
message
button, and its result is shown to the administrator and then forgotten.

What this asks for

A route that counts because it delivers, not because it is configured.

Worth being careful about

Nothing may send mail to answer this question. Evaluating a lock must not have a side
effect, and refuse_switching_off() is called while rendering a page. The honest source is
the last known outcome — recorded when the test button is pressed and when real mail is
sent — rather than a probe at decision time.

Stale is not the same as broken, and the difference has to be chosen. A transport that
worked last week and has not been used since is probably fine; one that failed an hour ago is
probably not. Whatever rule is picked has to fail in the safe direction: unknown counts as
working
, because a lock that opens on a shrug is worse than one that stays shut on an
optimistic guess. This issue makes the lock more honest, never more eager to open.

The administrator should be able to see why. If the lock opens because mail has been
failing, the page saying so is more useful than the lock quietly changing state — and if it
stays shut, the reason is already worded and needs only the extra fact.

OAuth is what makes this urgent rather than tidy. A password stays right until somebody
changes it; a refresh token expires on its own. #150 introduces credentials that go stale
without anybody touching them, on the path that carries password resets.

The same question will be asked of every future route. SMS (#143), an
administrator-issued link, a passkey: recovery_routes() returns a list precisely so routes
can be added, and "does this one actually work" is a question each of them will have to
answer in its own way. Deciding the shape now is cheaper than three times later.

## Observation Found while reading the mail plugin for the OAuth work, and worth separating because it is true today and wants fixing whether or not that lands. ## What exists `recovery_routes()` decides whether anybody could get back into their account, and `refuse_switching_off()` reads it to lock the mail transport while it is the last way in. That lock is the whole of #104 and it is load-bearing. What it actually checks is whether a transport is **selected**: ```python transport = selected() if transport is not None and transport.name != without: routes.append("email") ``` Selected, not working. An instance whose SMTP password is wrong, whose relay has been switched off, whose host no longer resolves, or whose credentials were revoked, still counts `email` as a route — and the lock stays shut on the strength of a route that delivers nothing. The refusal message tells an administrator they may not switch off the transport because it is protecting accounts it is not, in fact, protecting. The interface has the missing piece already: *Server settings → Email* has a **Send a test message** button, and its result is shown to the administrator and then forgotten. ## What this asks for A route that counts because it delivers, not because it is configured. ## Worth being careful about **Nothing may send mail to answer this question.** Evaluating a lock must not have a side effect, and `refuse_switching_off()` is called while rendering a page. The honest source is the **last known outcome** — recorded when the test button is pressed and when real mail is sent — rather than a probe at decision time. **Stale is not the same as broken, and the difference has to be chosen.** A transport that worked last week and has not been used since is probably fine; one that failed an hour ago is probably not. Whatever rule is picked has to fail in the safe direction: **unknown counts as working**, because a lock that opens on a shrug is worse than one that stays shut on an optimistic guess. This issue makes the lock *more* honest, never more eager to open. **The administrator should be able to see why.** If the lock opens because mail has been failing, the page saying so is more useful than the lock quietly changing state — and if it stays shut, the reason is already worded and needs only the extra fact. **OAuth is what makes this urgent rather than tidy.** A password stays right until somebody changes it; a refresh token expires on its own. #150 introduces credentials that go stale without anybody touching them, on the path that carries password resets. **The same question will be asked of every future route.** SMS (#143), an administrator-issued link, a passkey: `recovery_routes()` returns a list precisely so routes can be added, and "does this one actually work" is a question each of them will have to answer in its own way. Deciding the shape now is cheaper than three times later.
tiagoagueda added this to the 0.3.0 milestone 2026-09-09 11:06:22 +00:00
Author
Owner

#153 is the first thing that needs this rather than merely benefiting from it. It offers an administrator a switch that is only safe when the instance's mail works — the maintainer's own wording — and there is currently no way to ask that question. Offering sign-in by email on an instance whose SMTP is quietly broken produces a sign-in page promising something it cannot do, to somebody who may have no other way in.

So the last-known-outcome record this issue proposes has a second reader: not only the lock deciding whether a transport may be switched off, but a settings page deciding whether an option may be offered at all.

#153 is the first thing that needs this rather than merely benefiting from it. It offers an administrator a switch that is only safe *when the instance's mail works* — the maintainer's own wording — and there is currently no way to ask that question. Offering sign-in by email on an instance whose SMTP is quietly broken produces a sign-in page promising something it cannot do, to somebody who may have no other way in. So the last-known-outcome record this issue proposes has a second reader: not only the lock deciding whether a transport may be switched off, but a settings page deciding whether an option may be offered at all.
Author
Owner

A route now counts because it delivers, not because it is configured.

The outcome of every send is recorded on the way past, at the one point every message already passes through — so the Send a test message button counts as evidence without anything extra being wired up, and nothing opens a connection to answer the question. That mattered: the lock is evaluated while rendering a page.

Three failures in a row, not one. A relay that refuses a single address has told us about that address rather than about itself, and an instance that has never sent anything counts as working. Both fail towards the lock staying shut, because a lock that opens on a shrug is worse than one that stays shut on an optimistic guess.

When it does open, the page carries the fact. Mail failing while it is the only route means nobody who forgets a password can get back in — true whether the lock is shut or not, since those accounts are stranded by the relay and not by the setting. So Server settings → Email names how many people that is, and the lock gets out of the way of an administrator installing something that would fix it.

recovery_routes() now returns routes that separate exists from delivers, which is the shape #143, #144 and an administrator-issued link will each need.

Also fixed on the way past: PluggableBackend now owns its fail_silently rather than inheriting it. Django 7.0 removes it from BaseEmailBackend, and reading the inherited attribute already warns — on exactly the path that matters, the one where a send has failed.

22 tests in tests/test_mail_health.py, a wiki section under Configuration → Email, and nine strings in all 39 European catalogues.

Shipped in 037f7f1 on 0.3.0, with main kept level.

A route now counts because it delivers, not because it is configured. The outcome of every send is recorded on the way past, at the one point every message already passes through — so the **Send a test message** button counts as evidence without anything extra being wired up, and nothing opens a connection to answer the question. That mattered: the lock is evaluated while rendering a page. **Three failures in a row, not one.** A relay that refuses a single address has told us about that address rather than about itself, and an instance that has never sent anything counts as working. Both fail towards the lock staying shut, because a lock that opens on a shrug is worse than one that stays shut on an optimistic guess. **When it does open, the page carries the fact.** Mail failing while it is the only route means nobody who forgets a password can get back in — true whether the lock is shut or not, since those accounts are stranded by the relay and not by the setting. So *Server settings → Email* names how many people that is, and the lock gets out of the way of an administrator installing something that would fix it. `recovery_routes()` now returns routes that separate *exists* from *delivers*, which is the shape #143, #144 and an administrator-issued link will each need. Also fixed on the way past: `PluggableBackend` now owns its `fail_silently` rather than inheriting it. Django 7.0 removes it from `BaseEmailBackend`, and reading the inherited attribute already warns — on exactly the path that matters, the one where a send has failed. 22 tests in `tests/test_mail_health.py`, a wiki section under *Configuration → Email*, and nine strings in all 39 European catalogues. Shipped in `037f7f1` on `0.3.0`, with `main` kept level.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Reference
Postulo/postulo#152
No description provided.