Two tests pass only on the machine they were written on #76

Closed
opened 2026-09-07 10:08:17 +00:00 by tiagoagueda · 0 comments
Owner

Observation

#75 fixed the dependency audit, which was one reason CI failed. It was not the only one.
With the audit corrected, the run still failed, and the log shows two tests that pass on
Windows and fail on Linux — which is to say, two tests that had never passed anywhere except
the machine they were written on.

FAILED tests/security/test_requests.py::test_the_production_settings_are_what_the_policy_says
FAILED tests/test_brand.py::test_the_derived_images_match_the_source

1. The production-policy test depended on a gitignored file

os.environ.setdefault("POSTULO_SECRET_KEY", "x" * 64)
prod = importlib.import_module("postulo.config.settings.prod")
django.core.exceptions.ImproperlyConfigured: POSTULO_SECRET_KEY must be set

Two faults stacked on each other:

  • postulo.config.settings.base is already imported by the time any test runs — the
    test settings import it — so SECRET_KEY was resolved long before setdefault ran.
    prod's from .base import * gets the cached module, with SECRET_KEY still empty, and
    the guard fires.
  • It passed locally because base calls environ.Env.read_env(REPO_DIR / ".env"), and the
    author's .env — which is gitignored — supplies a key. So what the test asserted
    depended on whether the person running it happened to have a file that is not in the
    repository.

Fixed by importing prod in a subprocess with an explicit, scrubbed environment and
reading the settings back as JSON. Immune to import order, and read_env is stubbed out
inside the probe so a developer's .env cannot change the answer. A second test now asserts
the guard itself — that an instance with no key refuses to start rather than inventing one —
which is worth stating outright rather than discovering by accident.

2. The brand check compared bytes that two machines do not agree on

Out of date; run `uv run python scripts/brand.py`:
  apple-touch-icon.png  icon-192.png  icon-512.png  logo-64.png  logo-256.png  favicon.ico

--check rebuilt every image and compared the bytes. That cannot work across machines:
Pillow's platform wheels take different paths through LANCZOS, so the resampled pixels
differ by one here and there. The tell is in the list — favicon-16, -32 and -48
matched and everything from 64 pixels upward did not
, which is resampling, not the
encoder.

Fixed by recording digests instead. assets/brand/brand.json holds the digest of the
source, the digest of each derived file as committed, and the size-and-padding table they
were built from. --check compares those and rebuilds nothing, so it answers the question
worth asking — were these built from this source? — rather than does this machine round
the same way as the last one?
It still fails on all three things that matter: the mark
changed without a rebuild, a derived file edited or missing, and the recipe changed. Each
verified by breaking it deliberately.

Why this went unnoticed

Neither is exotic. Both would have been caught the first time CI ran on a machine that was
not the author's — and CI has been running on exactly such a machine, and failing, thirty-
eight times. The failures were invisible because the logs are not reachable: this
Forgejo instance has no actions/runs/{id}/jobs route, no actions/jobs/{id}/logs, no
artifacts endpoint, and the web UI's JSON returns 500: task ... does not exist even for
runs that succeeded. Diagnosing this needed the log pasted in by hand.

That is arguably the more important finding. A build whose failures cannot be read is a
build nobody reads, and thirty-eight red runs is what that looks like. Worth its own issue
against the instance's ACTIONS.LOG_RETENTION_DAYS or log storage.

Classification

Bug. Two tests that asserted nothing reliable, and a suite that was green only where it was
written.

## Observation #75 fixed the dependency audit, which was one reason CI failed. It was not the only one. With the audit corrected, the run still failed, and the log shows two tests that pass on Windows and fail on Linux — which is to say, two tests that had never passed anywhere except the machine they were written on. ``` FAILED tests/security/test_requests.py::test_the_production_settings_are_what_the_policy_says FAILED tests/test_brand.py::test_the_derived_images_match_the_source ``` ## 1. The production-policy test depended on a gitignored file ```python os.environ.setdefault("POSTULO_SECRET_KEY", "x" * 64) prod = importlib.import_module("postulo.config.settings.prod") ``` ``` django.core.exceptions.ImproperlyConfigured: POSTULO_SECRET_KEY must be set ``` Two faults stacked on each other: * `postulo.config.settings.base` is **already imported** by the time any test runs — the test settings import it — so `SECRET_KEY` was resolved long before `setdefault` ran. `prod`'s `from .base import *` gets the cached module, with `SECRET_KEY` still empty, and the guard fires. * It passed locally because `base` calls `environ.Env.read_env(REPO_DIR / ".env")`, and the author's `.env` — which is **gitignored** — supplies a key. So what the test asserted depended on whether the person running it happened to have a file that is not in the repository. **Fixed** by importing `prod` in a subprocess with an explicit, scrubbed environment and reading the settings back as JSON. Immune to import order, and `read_env` is stubbed out inside the probe so a developer's `.env` cannot change the answer. A second test now asserts the guard itself — that an instance with no key refuses to start rather than inventing one — which is worth stating outright rather than discovering by accident. ## 2. The brand check compared bytes that two machines do not agree on ``` Out of date; run `uv run python scripts/brand.py`: apple-touch-icon.png icon-192.png icon-512.png logo-64.png logo-256.png favicon.ico ``` `--check` rebuilt every image and compared the bytes. That cannot work across machines: Pillow's platform wheels take different paths through LANCZOS, so the resampled pixels differ by one here and there. The tell is in the list — **`favicon-16`, `-32` and `-48` matched and everything from 64 pixels upward did not**, which is resampling, not the encoder. **Fixed** by recording digests instead. `assets/brand/brand.json` holds the digest of the source, the digest of each derived file as committed, and the size-and-padding table they were built from. `--check` compares those and rebuilds nothing, so it answers the question worth asking — *were these built from this source?* — rather than *does this machine round the same way as the last one?* It still fails on all three things that matter: the mark changed without a rebuild, a derived file edited or missing, and the recipe changed. Each verified by breaking it deliberately. ## Why this went unnoticed Neither is exotic. Both would have been caught the first time CI ran on a machine that was not the author's — and CI has been running on exactly such a machine, and failing, thirty- eight times. The failures were invisible because **the logs are not reachable**: this Forgejo instance has no `actions/runs/{id}/jobs` route, no `actions/jobs/{id}/logs`, no artifacts endpoint, and the web UI's JSON returns `500: task ... does not exist` even for runs that succeeded. Diagnosing this needed the log pasted in by hand. **That is arguably the more important finding.** A build whose failures cannot be read is a build nobody reads, and thirty-eight red runs is what that looks like. Worth its own issue against the instance's `ACTIONS.LOG_RETENTION_DAYS` or log storage. ## Classification Bug. Two tests that asserted nothing reliable, and a suite that was green only where it was written.
tiagoagueda added this to the 0.3.0 milestone 2026-09-07 10:08:17 +00:00
tiagoagueda modified the milestone from 0.3.0 to 0.2.0 2026-09-07 11:44:16 +00:00
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
Postulo/postulo#76
No description provided.