A route that cannot deliver is not a way back in #152
Labels
No labels
accessibility
authentication
breaking change
bug
documentation
enhancement
interface
internationalisation
observability
security
tier
1
tier
2
tier
3
tier/4
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Blocks
Reference
Postulo/postulo#152
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Observation
Found while reading the mail plugin for the OAuth work, and worth separating because it is
true today and wants fixing whether or not that lands.
What exists
recovery_routes()decides whether anybody could get back into their account, andrefuse_switching_off()reads it to lock the mail transport while it is the last way in.That lock is the whole of #104 and it is load-bearing.
What it actually checks is whether a transport is selected:
Selected, not working. An instance whose SMTP password is wrong, whose relay has been
switched off, whose host no longer resolves, or whose credentials were revoked, still counts
emailas a route — and the lock stays shut on the strength of a route that deliversnothing. The refusal message tells an administrator they may not switch off the transport
because it is protecting accounts it is not, in fact, protecting.
The interface has the missing piece already: Server settings → Email has a Send a test
message button, and its result is shown to the administrator and then forgotten.
What this asks for
A route that counts because it delivers, not because it is configured.
Worth being careful about
Nothing may send mail to answer this question. Evaluating a lock must not have a side
effect, and
refuse_switching_off()is called while rendering a page. The honest source isthe last known outcome — recorded when the test button is pressed and when real mail is
sent — rather than a probe at decision time.
Stale is not the same as broken, and the difference has to be chosen. A transport that
worked last week and has not been used since is probably fine; one that failed an hour ago is
probably not. Whatever rule is picked has to fail in the safe direction: unknown counts as
working, because a lock that opens on a shrug is worse than one that stays shut on an
optimistic guess. This issue makes the lock more honest, never more eager to open.
The administrator should be able to see why. If the lock opens because mail has been
failing, the page saying so is more useful than the lock quietly changing state — and if it
stays shut, the reason is already worded and needs only the extra fact.
OAuth is what makes this urgent rather than tidy. A password stays right until somebody
changes it; a refresh token expires on its own. #150 introduces credentials that go stale
without anybody touching them, on the path that carries password resets.
The same question will be asked of every future route. SMS (#143), an
administrator-issued link, a passkey:
recovery_routes()returns a list precisely so routescan be added, and "does this one actually work" is a question each of them will have to
answer in its own way. Deciding the shape now is cheaper than three times later.
#153 is the first thing that needs this rather than merely benefiting from it. It offers an administrator a switch that is only safe when the instance's mail works — the maintainer's own wording — and there is currently no way to ask that question. Offering sign-in by email on an instance whose SMTP is quietly broken produces a sign-in page promising something it cannot do, to somebody who may have no other way in.
So the last-known-outcome record this issue proposes has a second reader: not only the lock deciding whether a transport may be switched off, but a settings page deciding whether an option may be offered at all.
A route now counts because it delivers, not because it is configured.
The outcome of every send is recorded on the way past, at the one point every message already passes through — so the Send a test message button counts as evidence without anything extra being wired up, and nothing opens a connection to answer the question. That mattered: the lock is evaluated while rendering a page.
Three failures in a row, not one. A relay that refuses a single address has told us about that address rather than about itself, and an instance that has never sent anything counts as working. Both fail towards the lock staying shut, because a lock that opens on a shrug is worse than one that stays shut on an optimistic guess.
When it does open, the page carries the fact. Mail failing while it is the only route means nobody who forgets a password can get back in — true whether the lock is shut or not, since those accounts are stranded by the relay and not by the setting. So Server settings → Email names how many people that is, and the lock gets out of the way of an administrator installing something that would fix it.
recovery_routes()now returns routes that separate exists from delivers, which is the shape #143, #144 and an administrator-issued link will each need.Also fixed on the way past:
PluggableBackendnow owns itsfail_silentlyrather than inheriting it. Django 7.0 removes it fromBaseEmailBackend, and reading the inherited attribute already warns — on exactly the path that matters, the one where a send has failed.22 tests in
tests/test_mail_health.py, a wiki section under Configuration → Email, and nine strings in all 39 European catalogues.Shipped in
037f7f1on0.3.0, withmainkept level.