A dev image channel, so a feature can be run before it is released #190
Labels
No labels
accessibility
authentication
breaking change
bug
documentation
enhancement
interface
internationalisation
observability
security
tier
1
tier
2
tier
3
tier/4
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
Postulo/postulo#190
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
What is wanted
An image built from a development branch, published on its own channel, so a feature can be
run somewhere real before it is in a release. Independent of the release images: a dev build
must never be what
docker pullgives somebody who asked for Postulo.Most of it already exists
image.ymlbuilds the image, scans it with both Trivy and Grype, and pushes themulti-architecture manifest only if the scan passes — the native build is scanned first
on purpose, because "a gate after the push would be a report about something already
published". It signs in to Forgejo's registry with
REGISTRY_USERandREGISTRY_TOKEN, andkeeps an SBOM.
So this is a second trigger and a second tag scheme, not new machinery. What it is not is
a small change, for one reason.
The decision, left open: what triggers it
image.ymlisworkflow_dispatchonly, andCONTRIBUTING.mdsays why in as many words:A dev channel that builds on every push removes exactly that mitigation. The
root-equivalent label stops being reachable only by a human and becomes reachable by a push.
That is fine while one person pushes to a protected branch and stops being fine the day
somebody else's pull request can reach it — and the point of a dev channel is usually that
other people are involved.
Three shapes, none chosen here:
A.
workflow_dispatchwith a branch input. The same button, pointed at a ref instead ofa tag. No new exposure whatsoever; the mitigation above survives word for word. The cost is
that somebody presses it, so the image is as fresh as the last time anyone remembered.
B.
pushto one named protected branch. Automatic, which is the point. Needs theworkflow to refuse everything else explicitly —
pull_requestnever, and anifon the refrather than trusting the trigger list — and needs that branch to stay protected for as long
as the runner carries the label. The exposure is real but bounded, and it is bounded by a
branch protection rule rather than by the workflow.
C. Cron. A nightly
:devfrom the branch tip. No push-triggered path to the label atall, so the mitigation survives; the image is up to a day old and builds on nights when
nothing changed.
The security question is the same in each: who can cause code to run as root on the runner.
A, C and B-with-protection all answer it; B-without-protection does not.
Settled regardless of which
:latestmust not move. The tag step computes${image}:${version},${image}:${version%.*}and${image}:latestunconditionally. A dev build reaching:latestmakesdocker pull postuloa dev image. Dev needs its own namespace —:dev,and something reproducible beside it like
:0.3.0-dev.<short-sha>, since a floating tagnobody can pin is not much use for reporting a bug against.
image. A gate skipped "just for dev" is not a gate, and #155 and #157 were both found in an
image that had been built and published without one.
built from a moving branch is an image nobody can reproduce". A dev image is built from a
moving branch by definition, so the commit has to be recorded in the tag or a label on the
image, or the same reasoning bites in a worse place.
Decide what keeps a dev tag alive before there are two hundred of them.
are not supported, and they may break a database in ways a release will not.
Check this first
No image has ever been built by CI.
image.ymlneeds a runner advertisingdocker, andas of #81 the only runner —
ouranos— advertisedubuntu-latest,ubuntu-24.04andubuntu-22.04and nothing else, which is why the workflow has never run.CONTRIBUTING.md§ Giving a runner the
dockerlabel has the recipe and the three host requirements thateach fail confusingly when missing (node, the
dockergroup, QEMU binfmt).Whether the label was ever added cannot be read from outside the instance; Site
administration → Actions → Runners shows what is advertised. If it was not, this issue is
the first thing that would ever have built an image, and it should expect to find the
usual first-run problems rather than assume the path works.
Checked on the host: the
dockerlabel was never added, and adding it is not one lineRead off
ouranosdirectly. The runner is containerised —code.forgejo.org/forgejo/runner:6, on Alpine — deployed fromstacks/ouranos/forgejo/docker-compose.yml, and itsconfig.ymldeclares exactly threelabels:
No
docker. Soimage.ymlhas never run and still cannot, which confirms what this issueassumed rather than leaving it assumed.
Four things that change the shape of the work:
The daemon is already reachable. The runner mounts
/var/run/docker.sockand carriesgroup_add: ${DOCKER_GID:-983}. Nothing about host access needs arranging — and note therunner container therefore already holds root-equivalent access to the daemon. The label
decides whether a workflow job can reach it, not whether the machine is exposed.
The runner container has git and nothing else. No
node, nodockerCLI. Underdocker:hosta job runs inside this container, soactions/checkout@v4— a JavaScriptaction — and every
dockercommand inimage.ymlwould both fail. Adding the label on itsown converts "queues for ever" into "fails confusingly", which is the outcome
CONTRIBUTING.mdwarns about by name.config.ymlis rewritten on every deploy. Therunner-registerservice writes it from aheredoc in the compose file, which says so itself: "change it in this file, not in the
volume." Editing the volume would survive until the next deploy and no longer.
capacity: 2. A dev-image build competes with CI for one of two slots, which matters morefor a channel that builds often than for a release image built by hand.
The two routes
A —
docker:host, asCONTRIBUTING.mddocumentsAdd
- docker:hostto the compose heredoc, and give the runner container the two things itlacks:
apk add nodejs docker-cli, either baked into a small custom image or as a stepwrapping the daemon command.
dockerlabel evertouch the socket.
runner:6.CONTRIBUTING.mdwere written for a runner installed onthe host; for a containerised one they become "the runner image needs node and the docker
CLI", and QEMU binfmt is still registered on the host kernel and shared. That section
wants a correction either way.
B —
docker:docker://catthehacker/ubuntu:act-latestA container label rather than host mode. Verified on the host: that image carries
/usr/bin/dockerand node 24, so nothing needs building.container.optionsis global — it would hand the socket to every CI job, includingany future
pull_request. That gives away exactly the isolation this issue is trying toprotect.
container.valid_volumes: ["/var/run/docker.sock"]plus an explicitmount declared in
image.yml, so the socket reaches only the workflow that asks for it ina file that is reviewable. That is a repository change as well as a host change.
The recommendation, for whenever this is decided
A. It keeps the one property everything else rests on, and a small custom image is a
smaller cost than a wider blast radius. B is quicker today and spends isolation that was
deliberately arranged.
Neither has been done. The label is not added.
Done while looking
FORGEJO-ACTIONS-TASK-2122_WORKFLOW-CI_JOB-browserhad been running for two days — anorphaned job container from roughly 120 tasks ago, holding a container and its resources with
no job behind it. Removed; no task containers are running now. Worth knowing that the runner
can leave these behind, since nothing reaps them.
Unblocked, not resolved — and by a better answer than either route above
Checked on
ouranos. A second runner instance now exists and is healthy:registered with
--scope Postulo/postulo. Soruns-on: dockerschedules now, and theprerequisite this issue was parked on is gone.
It is not route A or route B. It is better than both, for a reason neither of them had:
That is the flaw in route B stated exactly, and the fix is not to narrow the mount but to
narrow what can reach the runner at all. A second runner scoped to one repository is
enforced by Forgejo, rather than by the convention that nothing schedules onto a label by
itself — which is what route A and
CONTRIBUTING.mdboth rest on. The blast radius is aproperty of the registration now, not of everybody remembering.
It also sidesteps the thing that made route A expensive:
catthehacker/ubuntu:act-latestalready carries node 24, git, the docker CLI and buildx, so no custom runner image.
CONTRIBUTING.mdis now wrong, and the compose file says soThe compose carries a NOTE against following it:
§ Giving a runner the
dockerlabel tells the next person to do the thing that does notwork here, and its three host prerequisites (node, the
dockergroup, QEMU binfmt) describea host-installed runner. That section needs rewriting, and the reasoning above is what it
should say. Worth its own commit.
What is still open on this issue
Everything the issue is actually about. The runner was the prerequisite:
image.ymlis stillworkflow_dispatchwith a release-tag input. There is no dev channel.only this repository can reach, which makes B's exposure much smaller than when it was
written.
:latestmust not move; the tag step still computes it unconditionally.One thing to note before the first build:
image.ymlhas never run, so it should expect theusual first-run problems rather than assume the path works now that a runner answers.
The channel works. Its first scan stopped the push, which is the channel working.
dev-image.ymllanded and runs on every push tomain. The last run built the image,scanned it with both scanners, and then refused to publish:
That is the gate doing exactly what #156 built it for. Nothing is wrong with the workflow.
Getting there took four fixes, none of them in the new workflow
All four were in the release path, none had ever been exercised, and each was found only by
running the thing:
${GITHUB_SERVER_URL#https://}.GITHUB_SERVER_URLhere ishttp://server:3000— the runner's route to Forgejo — sothe strip left the scheme on, and the pushing daemon runs on the host and cannot resolve
a compose name regardless. Now a
REGISTRY_HOSTrepository variable with a fallback.image.ymlhad the identical line and would have failed the same way on the firstrelease anybody tried to publish. Fixed in
7fd3a6d4f.scan-image.shandcheck-image.shwererecorded
100644, so./scripts/scan-image.sh— howCONTRIBUTING.mdtells a person torun it and how both workflows call it — was Permission denied on any fresh clone.
e8cc7c2bd.aquasec/trivy:0.68.0is not a tag and never was;the scanners had only ever run on a machine that already had one pulled.
9ab5587cd.scan-image.shpassed-v "$OUT:/out", and abind mount is resolved by the daemon, against the host filesystem — so running inside a
container with the socket mounted in, the scanner wrote where the caller could not read.
Trivy writes to stdout now, which grype already did.
1f42054ae.Two of those needed a file to be read, not run, so
tests/test_shell_scripts.pynowdoes: every tracked script parses under
bash -n, is recorded executable, and has ashebang. Both faults were reintroduced to check the tests catch them.
b3b8339e3.What the gate actually found, and the decision it needs
Debian: 0. Every OS finding — sqlite, systemd, ncurses, util-linux, perl, zlib — is
fix_deferred,will_not_fixor has no fix, and--ignore-unfixedexcluded all of them.That is the design working: "six unfixable CRITICALs is the normal state of a Debian base
image".
Python: 2, and both are the same thing —
msgpack1.1.2, GHSA-6v7p-g79w-8964, HIGH,fixed in 1.2.1.
It is not ours.
msgpackis not inuv.lock, not inpyproject.toml, and not in thevirtual environment. It lives at:
It is the copy pip vendors for its own HTTP cache, arriving with
python:3.14-slim-bookworm. No dependency bump can move it.So the choice is not "upgrade a package". It is one of:
uvdeliberately —"
plugins/installing.pyprefers it 'where the image put it', falling back to pip" — anduvis present and on PATH. If the fallback is genuinely never taken when uv is there,dropping pip removes the finding, shrinks the image and cuts attack surface. But
installing.pystill has that fallback, and removing pip turns a documented path into acrash on an image where uv somehow is not usable. That wants checking, not assuming.
control and it will drift back.
parser, and pip only runs when an administrator installs a plugin.
I have not chosen. Removing pip touches plugin installation and that is not a call to make
while proving a CI channel works.
So, where this issue stands
The channel is built, running, and correct. It publishes nothing yet because the first thing
it scanned had a fixable HIGH — which is the gate, not a fault. Whichever way the msgpack
question is answered, the next push publishes
:devand:<version>-dev.<sha>.Still open from the original body: retention (nothing prunes old dev tags), and
:latestremains untouched by design.
Built, published, and running on ragnar
3ac8448c7published:devand:0.2.1-dev.3ac8448, for amd64 and arm64 —digest
sha256:3425488e…. The gate passed with nothing to report once pip left the image.ragnar is deployed from it and healthy:
on both
http://100.66.30.95:8000/healthzand https://postulo.tiagoagueda.com/healthz.compose.ymlpulls now instead of building — that is roughly ten minutes of Pi build timeper deploy gone. Config and a verified database backup were taken first
(
…bak-20260912-1614), anddeploy.logrecords the tag, the pinned tag, the digest andthe commit, so a rollback has something to name.
Two more faults, which makes six
Both in the release path, both would have hit the first real release:
pythonin the runtime stage is/app/.venv/bin/python— PATH puts the environment first — and that one never hadpip, so
python -m pip uninstallfailed before removing anything. Removed by path now,checked against
/usr/local/bin/python. Worth keeping: this meansinstaller()'s pipbranch could never have worked in the container anyway, because
sys.executableis theenvironment's python.
GITHUB_REPOSITORYisPostulo/postuloand aDocker repository name may not have one, so the tag was refused after a full
two-architecture build.
image.ymlcomputes it identically.Changed from what this issue proposed
The dev image is multi-arch after all. The workflow said native-only, arguing that
anybody needing another architecture wanted a release. That was wrong here: the instance
these images exist to be run on is a Raspberry Pi, so an amd64-only dev image is one nobody
can deploy. QEMU binfmt was already registered on the runner, so it cost time and nothing
else.
Still open
0.x.y-dev.<sha>tags, and there will be one per push.:latestremains untouched, as intended.Retention is in:
1683088a8dev-image.ymlends withscripts/prune-dev-images.pynow. The newest five pinned devtags stay (
DEV_IMAGES_KEPT); the rest go, and so do the per-architecture manifests onlythey referenced -- a manifest nothing names still holds its layers, so deleting a tag alone
would free nothing. Releases,
latest,dev, and anything untagged the run did not itselforphan are never candidates; that last case is what a push in flight from
image.ymllookslike, which is why the script does not sweep "everything unreferenced".
What was exercised where
registry in
tests/test_prune_dev_images.py(9 tests).(
postulo/prune-test, two multi-architecture pushes from ouranos, since removed):the listing, the
/v2/tokenexchange, reading an index, andDELETEof a taggedversion and of
sha256:versions -- the older tag went with its four manifests, thekept build's four survived. Then the package was deleted.
REGISTRY_TOKEN:nothing to prune: 3 tags, none past the newest 5 dev tags-- which is the listingworking with the CI credential, and nothing doomed yet.
The one thing still unproven
Whether
REGISTRY_TOKENmay delete a package version. It can push, and the listingworked with it, but delete needs
write:packageand nothing has asked it to delete yet.The first time it will is the sixth pinned tag -- four pushes to
mainfrom now. Afailing prune fails the run on purpose, so if the scope is missing it will be a red run
with
prune failed: DELETE … -> 403in the log, and the fix is the token's scope, not thecode. Worth glancing at that run.
Also corrected in the same commit:
CONTRIBUTING.mdstill said the dev image was builtfor the native architecture only.
The one thing left unproven is proven
The sixth pinned tag arrived with
789f0d9d4(run 327), and the prune step made its firstreal deletion with
REGISTRY_TOKEN:Those four are exactly the manifests only the first build referenced -- its two
architectures and two attestations -- and nothing that a remaining tag names was touched.
The registry holds
dev, five pinned tags and 18 manifests. So the token may delete,retention works unattended, and there is nothing left open on this issue.
One consequence worth knowing: ragnar's
deploy.lognames0.2.1-dev.3ac8448, which isthe tag that just went. A rollback to it by tag is no longer possible from the registry;
ragnar's own daemon still holds that image. The next deploy moves the rollback window along.