Two tests pass only on the machine they were written on #76
Labels
No labels
accessibility
authentication
breaking change
bug
documentation
enhancement
interface
internationalisation
observability
security
tier
1
tier
2
tier
3
tier/4
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
Postulo/postulo#76
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Observation
#75 fixed the dependency audit, which was one reason CI failed. It was not the only one.
With the audit corrected, the run still failed, and the log shows two tests that pass on
Windows and fail on Linux — which is to say, two tests that had never passed anywhere except
the machine they were written on.
1. The production-policy test depended on a gitignored file
Two faults stacked on each other:
postulo.config.settings.baseis already imported by the time any test runs — thetest settings import it — so
SECRET_KEYwas resolved long beforesetdefaultran.prod'sfrom .base import *gets the cached module, withSECRET_KEYstill empty, andthe guard fires.
basecallsenviron.Env.read_env(REPO_DIR / ".env"), and theauthor's
.env— which is gitignored — supplies a key. So what the test asserteddepended on whether the person running it happened to have a file that is not in the
repository.
Fixed by importing
prodin a subprocess with an explicit, scrubbed environment andreading the settings back as JSON. Immune to import order, and
read_envis stubbed outinside the probe so a developer's
.envcannot change the answer. A second test now assertsthe guard itself — that an instance with no key refuses to start rather than inventing one —
which is worth stating outright rather than discovering by accident.
2. The brand check compared bytes that two machines do not agree on
--checkrebuilt every image and compared the bytes. That cannot work across machines:Pillow's platform wheels take different paths through LANCZOS, so the resampled pixels
differ by one here and there. The tell is in the list —
favicon-16,-32and-48matched and everything from 64 pixels upward did not, which is resampling, not the
encoder.
Fixed by recording digests instead.
assets/brand/brand.jsonholds the digest of thesource, the digest of each derived file as committed, and the size-and-padding table they
were built from.
--checkcompares those and rebuilds nothing, so it answers the questionworth asking — were these built from this source? — rather than does this machine round
the same way as the last one? It still fails on all three things that matter: the mark
changed without a rebuild, a derived file edited or missing, and the recipe changed. Each
verified by breaking it deliberately.
Why this went unnoticed
Neither is exotic. Both would have been caught the first time CI ran on a machine that was
not the author's — and CI has been running on exactly such a machine, and failing, thirty-
eight times. The failures were invisible because the logs are not reachable: this
Forgejo instance has no
actions/runs/{id}/jobsroute, noactions/jobs/{id}/logs, noartifacts endpoint, and the web UI's JSON returns
500: task ... does not existeven forruns that succeeded. Diagnosing this needed the log pasted in by hand.
That is arguably the more important finding. A build whose failures cannot be read is a
build nobody reads, and thirty-eight red runs is what that looks like. Worth its own issue
against the instance's
ACTIONS.LOG_RETENTION_DAYSor log storage.Classification
Bug. Two tests that asserted nothing reliable, and a suite that was green only where it was
written.