A dev image channel, so a feature can be run before it is released #190

Closed
opened 2026-09-12 11:11:49 +00:00 by tiagoagueda · 6 comments
Owner

What is wanted

An image built from a development branch, published on its own channel, so a feature can be
run somewhere real before it is in a release. Independent of the release images: a dev build
must never be what docker pull gives somebody who asked for Postulo.

Most of it already exists

image.yml builds the image, scans it with both Trivy and Grype, and pushes the
multi-architecture manifest only if the scan passes
— the native build is scanned first
on purpose, because "a gate after the push would be a report about something already
published"
. It signs in to Forgejo's registry with REGISTRY_USER and REGISTRY_TOKEN, and
keeps an SBOM.

So this is a second trigger and a second tag scheme, not new machinery. What it is not is
a small change, for one reason.

The decision, left open: what triggers it

image.yml is workflow_dispatch only, and CONTRIBUTING.md says why in as many words:

What that label costs, stated plainly. A job on it runs as the runner's user with
Docker, and Docker is root on that machine. The mitigation is that nothing schedules
onto this label by itself
: image.yml is workflow_dispatch only, so the only way to
reach it is somebody pressing the button. Do not put the label on a runner that also
serves pull_request from people who are not you.

A dev channel that builds on every push removes exactly that mitigation. The
root-equivalent label stops being reachable only by a human and becomes reachable by a push.
That is fine while one person pushes to a protected branch and stops being fine the day
somebody else's pull request can reach it — and the point of a dev channel is usually that
other people are involved.

Three shapes, none chosen here:

A. workflow_dispatch with a branch input. The same button, pointed at a ref instead of
a tag. No new exposure whatsoever; the mitigation above survives word for word. The cost is
that somebody presses it, so the image is as fresh as the last time anyone remembered.

B. push to one named protected branch. Automatic, which is the point. Needs the
workflow to refuse everything else explicitly — pull_request never, and an if on the ref
rather than trusting the trigger list — and needs that branch to stay protected for as long
as the runner carries the label. The exposure is real but bounded, and it is bounded by a
branch protection rule rather than by the workflow.

C. Cron. A nightly :dev from the branch tip. No push-triggered path to the label at
all, so the mitigation survives; the image is up to a day old and builds on nights when
nothing changed.

The security question is the same in each: who can cause code to run as root on the runner.
A, C and B-with-protection all answer it; B-without-protection does not.

Settled regardless of which

  • :latest must not move. The tag step computes ${image}:${version},
    ${image}:${version%.*} and ${image}:latest unconditionally. A dev build reaching
    :latest makes docker pull postulo a dev image. Dev needs its own namespace — :dev,
    and something reproducible beside it like :0.3.0-dev.<short-sha>, since a floating tag
    nobody can pin is not much use for reporting a bug against.
  • The scan gate applies. A dev image somebody runs against their real applications is an
    image. A gate skipped "just for dev" is not a gate, and #155 and #157 were both found in an
    image that had been built and published without one.
  • Reproducibility. The release path checks out the tag deliberately: "a release image
    built from a moving branch is an image nobody can reproduce"
    . A dev image is built from a
    moving branch by definition, so the commit has to be recorded in the tag or a label on the
    image, or the same reasoning bites in a worse place.
  • Retention. A build per push fills a registry, and Forgejo does not prune on its own.
    Decide what keeps a dev tag alive before there are two hundred of them.
  • Say what the channel is for. A README line: dev images are for trying a feature, they
    are not supported, and they may break a database in ways a release will not.

Check this first

No image has ever been built by CI. image.yml needs a runner advertising docker, and
as of #81 the only runner — ouranos — advertised ubuntu-latest, ubuntu-24.04 and
ubuntu-22.04 and nothing else, which is why the workflow has never run. CONTRIBUTING.md
§ Giving a runner the docker label has the recipe and the three host requirements that
each fail confusingly when missing (node, the docker group, QEMU binfmt).

Whether the label was ever added cannot be read from outside the instance; Site
administration → Actions → Runners
shows what is advertised. If it was not, this issue is
the first thing that would ever have built an image
, and it should expect to find the
usual first-run problems rather than assume the path works.

## What is wanted An image built from a development branch, published on its own channel, so a feature can be run somewhere real before it is in a release. Independent of the release images: a dev build must never be what `docker pull` gives somebody who asked for Postulo. ## Most of it already exists `image.yml` builds the image, **scans it with both Trivy and Grype, and pushes the multi-architecture manifest only if the scan passes** — the native build is scanned first on purpose, because *"a gate after the push would be a report about something already published"*. It signs in to Forgejo's registry with `REGISTRY_USER` and `REGISTRY_TOKEN`, and keeps an SBOM. So this is a second trigger and a second tag scheme, not new machinery. What it is *not* is a small change, for one reason. ## The decision, left open: what triggers it `image.yml` is `workflow_dispatch` only, and `CONTRIBUTING.md` says why in as many words: > **What that label costs, stated plainly.** A job on it runs as the runner's user with > Docker, and **Docker is root on that machine**. The mitigation is that **nothing schedules > onto this label by itself**: `image.yml` is `workflow_dispatch` only, so the only way to > reach it is somebody pressing the button. **Do not put the label on a runner that also > serves `pull_request` from people who are not you.** A dev channel that builds on every push **removes exactly that mitigation**. The root-equivalent label stops being reachable only by a human and becomes reachable by a push. That is fine while one person pushes to a protected branch and stops being fine the day somebody else's pull request can reach it — and the point of a dev channel is usually that other people are involved. Three shapes, none chosen here: **A. `workflow_dispatch` with a branch input.** The same button, pointed at a ref instead of a tag. No new exposure whatsoever; the mitigation above survives word for word. The cost is that somebody presses it, so the image is as fresh as the last time anyone remembered. **B. `push` to one named protected branch.** Automatic, which is the point. Needs the workflow to refuse everything else explicitly — `pull_request` never, and an `if` on the ref rather than trusting the trigger list — and needs that branch to stay protected for as long as the runner carries the label. The exposure is real but bounded, and it is bounded by a branch protection rule rather than by the workflow. **C. Cron.** A nightly `:dev` from the branch tip. No push-triggered path to the label at all, so the mitigation survives; the image is up to a day old and builds on nights when nothing changed. The security question is the same in each: *who can cause code to run as root on the runner*. A, C and B-with-protection all answer it; B-without-protection does not. ## Settled regardless of which - **`:latest` must not move.** The tag step computes `${image}:${version}`, `${image}:${version%.*}` and `${image}:latest` unconditionally. A dev build reaching `:latest` makes `docker pull postulo` a dev image. Dev needs its own namespace — `:dev`, and something reproducible beside it like `:0.3.0-dev.<short-sha>`, since a floating tag nobody can pin is not much use for reporting a bug against. - **The scan gate applies.** A dev image somebody runs against their real applications is an image. A gate skipped "just for dev" is not a gate, and #155 and #157 were both found in an image that had been built and published without one. - **Reproducibility.** The release path checks out the tag deliberately: *"a release image built from a moving branch is an image nobody can reproduce"*. A dev image is built from a moving branch by definition, so the commit has to be recorded in the tag or a label on the image, or the same reasoning bites in a worse place. - **Retention.** A build per push fills a registry, and Forgejo does not prune on its own. Decide what keeps a dev tag alive before there are two hundred of them. - **Say what the channel is for.** A README line: dev images are for trying a feature, they are not supported, and they may break a database in ways a release will not. ## Check this first **No image has ever been built by CI.** `image.yml` needs a runner advertising `docker`, and as of #81 the only runner — `ouranos` — advertised `ubuntu-latest`, `ubuntu-24.04` and `ubuntu-22.04` and nothing else, which is why the workflow has never run. `CONTRIBUTING.md` § *Giving a runner the `docker` label* has the recipe and the three host requirements that each fail confusingly when missing (node, the `docker` group, QEMU binfmt). Whether the label was ever added cannot be read from outside the instance; *Site administration → Actions → Runners* shows what is advertised. **If it was not, this issue is the first thing that would ever have built an image**, and it should expect to find the usual first-run problems rather than assume the path works.
Author
Owner

Checked on the host: the docker label was never added, and adding it is not one line

Read off ouranos directly. The runner is containerised —
code.forgejo.org/forgejo/runner:6, on Alpine — deployed from
stacks/ouranos/forgejo/docker-compose.yml, and its config.yml declares exactly three
labels:

labels:
  - ubuntu-latest:docker://catthehacker/ubuntu:act-latest
  - ubuntu-24.04:docker://catthehacker/ubuntu:act-24.04
  - ubuntu-22.04:docker://catthehacker/ubuntu:act-22.04

No docker. So image.yml has never run and still cannot, which confirms what this issue
assumed rather than leaving it assumed.

Four things that change the shape of the work:

The daemon is already reachable. The runner mounts /var/run/docker.sock and carries
group_add: ${DOCKER_GID:-983}. Nothing about host access needs arranging — and note the
runner container therefore already holds root-equivalent access to the daemon. The label
decides whether a workflow job can reach it, not whether the machine is exposed.

The runner container has git and nothing else. No node, no docker CLI. Under
docker:host a job runs inside this container, so actions/checkout@v4 — a JavaScript
action — and every docker command in image.yml would both fail. Adding the label on its
own converts "queues for ever" into "fails confusingly", which is the outcome
CONTRIBUTING.md warns about by name.

config.yml is rewritten on every deploy. The runner-register service writes it from a
heredoc in the compose file, which says so itself: "change it in this file, not in the
volume."
Editing the volume would survive until the next deploy and no longer.

capacity: 2. A dev-image build competes with CI for one of two slots, which matters more
for a channel that builds often than for a release image built by hand.

The two routes

A — docker:host, as CONTRIBUTING.md documents

Add - docker:host to the compose heredoc, and give the runner container the two things it
lacks: apk add nodejs docker-cli, either baked into a small custom image or as a step
wrapping the daemon command.

  • Keeps the property the whole arrangement rests on: only jobs on the docker label ever
    touch the socket.
  • Costs a custom image to build and keep current with runner:6.
  • The three host prerequisites in CONTRIBUTING.md were written for a runner installed on
    the host; for a containerised one they become "the runner image needs node and the docker
    CLI"
    , and QEMU binfmt is still registered on the host kernel and shared. That section
    wants a correction either way.

B — docker:docker://catthehacker/ubuntu:act-latest

A container label rather than host mode. Verified on the host: that image carries
/usr/bin/docker and node 24, so nothing needs building.

  • No custom image, works as soon as the label is added.
  • But the socket has to get into the job container, and forgejo-runner's
    container.options is global — it would hand the socket to every CI job, including
    any future pull_request. That gives away exactly the isolation this issue is trying to
    protect.
  • The narrower form is container.valid_volumes: ["/var/run/docker.sock"] plus an explicit
    mount declared in image.yml, so the socket reaches only the workflow that asks for it in
    a file that is reviewable. That is a repository change as well as a host change.

The recommendation, for whenever this is decided

A. It keeps the one property everything else rests on, and a small custom image is a
smaller cost than a wider blast radius. B is quicker today and spends isolation that was
deliberately arranged.

Neither has been done. The label is not added.

Done while looking

FORGEJO-ACTIONS-TASK-2122_WORKFLOW-CI_JOB-browser had been running for two days — an
orphaned job container from roughly 120 tasks ago, holding a container and its resources with
no job behind it. Removed; no task containers are running now. Worth knowing that the runner
can leave these behind, since nothing reaps them.

## Checked on the host: the `docker` label was never added, and adding it is not one line Read off `ouranos` directly. The runner is **containerised** — `code.forgejo.org/forgejo/runner:6`, on **Alpine** — deployed from `stacks/ouranos/forgejo/docker-compose.yml`, and its `config.yml` declares exactly three labels: labels: - ubuntu-latest:docker://catthehacker/ubuntu:act-latest - ubuntu-24.04:docker://catthehacker/ubuntu:act-24.04 - ubuntu-22.04:docker://catthehacker/ubuntu:act-22.04 No `docker`. So `image.yml` has never run and still cannot, which confirms what this issue assumed rather than leaving it assumed. Four things that change the shape of the work: **The daemon is already reachable.** The runner mounts `/var/run/docker.sock` and carries `group_add: ${DOCKER_GID:-983}`. Nothing about host access needs arranging — and note the runner container therefore *already* holds root-equivalent access to the daemon. The label decides whether a **workflow job** can reach it, not whether the machine is exposed. **The runner container has git and nothing else.** No `node`, no `docker` CLI. Under `docker:host` a job runs *inside this container*, so `actions/checkout@v4` — a JavaScript action — and every `docker` command in `image.yml` would both fail. **Adding the label on its own converts "queues for ever" into "fails confusingly", which is the outcome `CONTRIBUTING.md` warns about by name.** **`config.yml` is rewritten on every deploy.** The `runner-register` service writes it from a heredoc in the compose file, which says so itself: *"change it in this file, not in the volume."* Editing the volume would survive until the next deploy and no longer. **`capacity: 2`.** A dev-image build competes with CI for one of two slots, which matters more for a channel that builds often than for a release image built by hand. ## The two routes ### A — `docker:host`, as `CONTRIBUTING.md` documents Add `- docker:host` to the compose heredoc, and give the runner container the two things it lacks: `apk add nodejs docker-cli`, either baked into a small custom image or as a step wrapping the daemon command. - Keeps the property the whole arrangement rests on: **only jobs on the `docker` label ever touch the socket.** - Costs a custom image to build and keep current with `runner:6`. - The three host prerequisites in `CONTRIBUTING.md` were written for a runner installed on the host; for a containerised one they become *"the runner image needs node and the docker CLI"*, and QEMU binfmt is still registered on the host kernel and shared. That section wants a correction either way. ### B — `docker:docker://catthehacker/ubuntu:act-latest` A container label rather than host mode. Verified on the host: that image carries **`/usr/bin/docker` and node 24**, so nothing needs building. - No custom image, works as soon as the label is added. - **But the socket has to get into the job container**, and forgejo-runner's `container.options` is **global** — it would hand the socket to *every* CI job, including any future `pull_request`. That gives away exactly the isolation this issue is trying to protect. - The narrower form is `container.valid_volumes: ["/var/run/docker.sock"]` plus an explicit mount declared in `image.yml`, so the socket reaches only the workflow that asks for it in a file that is reviewable. That is a repository change as well as a host change. ### The recommendation, for whenever this is decided **A.** It keeps the one property everything else rests on, and a small custom image is a smaller cost than a wider blast radius. B is quicker today and spends isolation that was deliberately arranged. Neither has been done. The label is **not** added. ## Done while looking `FORGEJO-ACTIONS-TASK-2122_WORKFLOW-CI_JOB-browser` had been running for two days — an orphaned job container from roughly 120 tasks ago, holding a container and its resources with no job behind it. Removed; no task containers are running now. Worth knowing that the runner can leave these behind, since nothing reaps them.
Author
Owner

Unblocked, not resolved — and by a better answer than either route above

Checked on ouranos. A second runner instance now exists and is healthy:

runner: ouranos-docker, with version: v6.4.0, with labels: [docker], declared successfully

labels:
  - docker:docker://catthehacker/ubuntu:act-latest
container:
  docker_host: "automount"

registered with --scope Postulo/postulo. So runs-on: docker schedules now, and the
prerequisite this issue was parked on is gone.

It is not route A or route B. It is better than both, for a reason neither of them had:

container.docker_host is per RUNNER INSTANCE, not per label. Setting it to automount
mounts /var/run/docker.sock into every job container that runner starts — and docker is
root on ouranos.

That is the flaw in route B stated exactly, and the fix is not to narrow the mount but to
narrow what can reach the runner at all. A second runner scoped to one repository is
enforced by Forgejo, rather than by the convention that nothing schedules onto a label by
itself — which is what route A and CONTRIBUTING.md both rest on. The blast radius is a
property of the registration now, not of everybody remembering.

It also sidesteps the thing that made route A expensive: catthehacker/ubuntu:act-latest
already carries node 24, git, the docker CLI and buildx, so no custom runner image.

CONTRIBUTING.md is now wrong, and the compose file says so

The compose carries a NOTE against following it:

Postulo's CONTRIBUTING.md says to declare docker:host — do NOT follow that here. host
runs the job inside the RUNNER container, and this runner image is Alpine with no node and
no docker CLI, so actions/checkout (a JavaScript action) fails before anything reaches
the daemon. That recipe assumes a runner installed on the host.

§ Giving a runner the docker label tells the next person to do the thing that does not
work here, and its three host prerequisites (node, the docker group, QEMU binfmt) describe
a host-installed runner. That section needs rewriting, and the reasoning above is what it
should say. Worth its own commit.

What is still open on this issue

Everything the issue is actually about. The runner was the prerequisite:

  • image.yml is still workflow_dispatch with a release-tag input. There is no dev channel.
  • The trigger is still undecided — A, B or C in the issue body, now against a runner that
    only this repository can reach, which makes B's exposure much smaller than when it was
    written.
  • :latest must not move; the tag step still computes it unconditionally.
  • Retention, and recording the commit in a dev tag, are untouched.

One thing to note before the first build: image.yml has never run, so it should expect the
usual first-run problems rather than assume the path works now that a runner answers.

## Unblocked, not resolved — and by a better answer than either route above Checked on `ouranos`. A **second runner instance** now exists and is healthy: runner: ouranos-docker, with version: v6.4.0, with labels: [docker], declared successfully labels: - docker:docker://catthehacker/ubuntu:act-latest container: docker_host: "automount" registered with `--scope Postulo/postulo`. So `runs-on: docker` schedules now, and the prerequisite this issue was parked on is gone. **It is not route A or route B.** It is better than both, for a reason neither of them had: > `container.docker_host` is per RUNNER INSTANCE, not per label. Setting it to `automount` > mounts /var/run/docker.sock into every job container that runner starts — and docker is > root on ouranos. That is the flaw in route B stated exactly, and the fix is not to narrow the mount but to narrow *what can reach the runner at all*. A second runner scoped to one repository is enforced by Forgejo, rather than by the convention that nothing schedules onto a label by itself — which is what route A and `CONTRIBUTING.md` both rest on. The blast radius is a property of the registration now, not of everybody remembering. It also sidesteps the thing that made route A expensive: `catthehacker/ubuntu:act-latest` already carries node 24, git, the docker CLI and buildx, so no custom runner image. ### `CONTRIBUTING.md` is now wrong, and the compose file says so The compose carries a NOTE against following it: > Postulo's CONTRIBUTING.md says to declare `docker:host` — do NOT follow that here. `host` > runs the job inside the RUNNER container, and this runner image is Alpine with no node and > no docker CLI, so `actions/checkout` (a JavaScript action) fails before anything reaches > the daemon. That recipe assumes a runner installed on the host. § *Giving a runner the `docker` label* tells the next person to do the thing that does not work here, and its three host prerequisites (node, the `docker` group, QEMU binfmt) describe a host-installed runner. **That section needs rewriting**, and the reasoning above is what it should say. Worth its own commit. ### What is still open on this issue Everything the issue is actually about. The runner was the prerequisite: - `image.yml` is still `workflow_dispatch` with a release-tag input. There is no dev channel. - **The trigger is still undecided** — A, B or C in the issue body, now against a runner that only this repository can reach, which makes B's exposure much smaller than when it was written. - `:latest` must not move; the tag step still computes it unconditionally. - Retention, and recording the commit in a dev tag, are untouched. One thing to note before the first build: `image.yml` has never run, so it should expect the usual first-run problems rather than assume the path works now that a runner answers.
tiagoagueda referenced this issue from a commit 2026-09-12 13:15:08 +00:00
Author
Owner

The channel works. Its first scan stopped the push, which is the channel working.

dev-image.yml landed and runs on every push to main. The last run built the image,
scanned it with both scanners, and then refused to publish:

Fixable findings at HIGH,CRITICAL or above.

That is the gate doing exactly what #156 built it for. Nothing is wrong with the workflow.

Getting there took four fixes, none of them in the new workflow

All four were in the release path, none had ever been exercised, and each was found only by
running the thing:

  1. The registry address. Both workflows computed ${GITHUB_SERVER_URL#https://}.
    GITHUB_SERVER_URL here is http://server:3000 — the runner's route to Forgejo — so
    the strip left the scheme on, and the pushing daemon runs on the host and cannot resolve
    a compose name regardless. Now a REGISTRY_HOST repository variable with a fallback.
    image.yml had the identical line and would have failed the same way on the first
    release anybody tried to publish. Fixed in 7fd3a6d4f.
  2. The shell scripts were not executable. scan-image.sh and check-image.sh were
    recorded 100644, so ./scripts/scan-image.sh — how CONTRIBUTING.md tells a person to
    run it and how both workflows call it — was Permission denied on any fresh clone.
    e8cc7c2bd.
  3. The pinned Trivy did not exist. aquasec/trivy:0.68.0 is not a tag and never was;
    the scanners had only ever run on a machine that already had one pulled. 9ab5587cd.
  4. The scan reports went to the host. scan-image.sh passed -v "$OUT:/out", and a
    bind mount is resolved by the daemon, against the host filesystem — so running inside a
    container with the socket mounted in, the scanner wrote where the caller could not read.
    Trivy writes to stdout now, which grype already did. 1f42054ae.

Two of those needed a file to be read, not run, so tests/test_shell_scripts.py now
does: every tracked script parses under bash -n, is recorded executable, and has a
shebang. Both faults were reintroduced to check the tests catch them. b3b8339e3.

What the gate actually found, and the decision it needs

Debian: 0. Every OS finding — sqlite, systemd, ncurses, util-linux, perl, zlib — is
fix_deferred, will_not_fix or has no fix, and --ignore-unfixed excluded all of them.
That is the design working: "six unfixable CRITICALs is the normal state of a Debian base
image"
.

Python: 2, and both are the same thing — msgpack 1.1.2, GHSA-6v7p-g79w-8964, HIGH,
fixed in 1.2.1.

It is not ours. msgpack is not in uv.lock, not in pyproject.toml, and not in the
virtual environment. It lives at:

/usr/local/lib/python3.14/site-packages/pip/_vendor/msgpack

It is the copy pip vendors for its own HTTP cache, arriving with
python:3.14-slim-bookworm. No dependency bump can move it.

So the choice is not "upgrade a package". It is one of:

  • Remove pip from the runtime image. The Dockerfile already keeps uv deliberately —
    "plugins/installing.py prefers it 'where the image put it', falling back to pip" — and
    uv is present and on PATH. If the fallback is genuinely never taken when uv is there,
    dropping pip removes the finding, shrinks the image and cuts attack surface. But
    installing.py still has that fallback
    , and removing pip turns a documented path into a
    crash on an image where uv somehow is not usable. That wants checking, not assuming.
  • Upgrade pip in the image, and hope the newer one vendors msgpack ≥ 1.2.1. Outside our
    control and it will drift back.
  • Record an exception for this CVE with the reason above: it is pip's vendored cache
    parser, and pip only runs when an administrator installs a plugin.

I have not chosen. Removing pip touches plugin installation and that is not a call to make
while proving a CI channel works.

So, where this issue stands

The channel is built, running, and correct. It publishes nothing yet because the first thing
it scanned had a fixable HIGH — which is the gate, not a fault. Whichever way the msgpack
question is answered, the next push publishes :dev and :<version>-dev.<sha>.

Still open from the original body: retention (nothing prunes old dev tags), and :latest
remains untouched by design.

## The channel works. Its first scan stopped the push, which is the channel working. `dev-image.yml` landed and runs on every push to `main`. The last run built the image, scanned it with both scanners, and then refused to publish: Fixable findings at HIGH,CRITICAL or above. That is the gate doing exactly what #156 built it for. Nothing is wrong with the workflow. ### Getting there took four fixes, none of them in the new workflow All four were in the release path, none had ever been exercised, and each was found only by running the thing: 1. **The registry address.** Both workflows computed `${GITHUB_SERVER_URL#https://}`. `GITHUB_SERVER_URL` here is `http://server:3000` — the *runner's* route to Forgejo — so the strip left the scheme on, and the pushing daemon runs on the host and cannot resolve a compose name regardless. Now a `REGISTRY_HOST` repository variable with a fallback. `image.yml` had the identical line and would have failed the same way on the first release anybody tried to publish. Fixed in `7fd3a6d4f`. 2. **The shell scripts were not executable.** `scan-image.sh` and `check-image.sh` were recorded `100644`, so `./scripts/scan-image.sh` — how `CONTRIBUTING.md` tells a person to run it and how both workflows call it — was *Permission denied* on any fresh clone. `e8cc7c2bd`. 3. **The pinned Trivy did not exist.** `aquasec/trivy:0.68.0` is not a tag and never was; the scanners had only ever run on a machine that already had one pulled. `9ab5587cd`. 4. **The scan reports went to the host.** `scan-image.sh` passed `-v "$OUT:/out"`, and a bind mount is resolved by the *daemon*, against the host filesystem — so running inside a container with the socket mounted in, the scanner wrote where the caller could not read. Trivy writes to stdout now, which grype already did. `1f42054ae`. Two of those needed a file to be **read**, not run, so `tests/test_shell_scripts.py` now does: every tracked script parses under `bash -n`, is recorded executable, and has a shebang. Both faults were reintroduced to check the tests catch them. `b3b8339e3`. ### What the gate actually found, and the decision it needs **Debian: 0.** Every OS finding — sqlite, systemd, ncurses, util-linux, perl, zlib — is `fix_deferred`, `will_not_fix` or has no fix, and `--ignore-unfixed` excluded all of them. That is the design working: *"six unfixable CRITICALs is the normal state of a Debian base image"*. **Python: 2, and both are the same thing** — `msgpack` 1.1.2, GHSA-6v7p-g79w-8964, HIGH, fixed in 1.2.1. **It is not ours.** `msgpack` is not in `uv.lock`, not in `pyproject.toml`, and not in the virtual environment. It lives at: /usr/local/lib/python3.14/site-packages/pip/_vendor/msgpack It is the copy **pip vendors for its own HTTP cache**, arriving with `python:3.14-slim-bookworm`. No dependency bump can move it. So the choice is not "upgrade a package". It is one of: - **Remove pip from the runtime image.** The Dockerfile already keeps `uv` deliberately — *"`plugins/installing.py` prefers it 'where the image put it', falling back to pip"* — and `uv` is present and on PATH. If the fallback is genuinely never taken when uv is there, dropping pip removes the finding, shrinks the image and cuts attack surface. **But `installing.py` still has that fallback**, and removing pip turns a documented path into a crash on an image where uv somehow is not usable. That wants checking, not assuming. - **Upgrade pip in the image**, and hope the newer one vendors msgpack ≥ 1.2.1. Outside our control and it will drift back. - **Record an exception** for this CVE with the reason above: it is pip's vendored cache parser, and pip only runs when an administrator installs a plugin. I have not chosen. Removing pip touches plugin installation and that is not a call to make while proving a CI channel works. ### So, where this issue stands The channel is built, running, and correct. It publishes nothing yet because the first thing it scanned had a fixable HIGH — which is the gate, not a fault. Whichever way the msgpack question is answered, the next push publishes `:dev` and `:<version>-dev.<sha>`. Still open from the original body: retention (nothing prunes old dev tags), and `:latest` remains untouched by design.
Author
Owner

Built, published, and running on ragnar

3ac8448c7 published :dev and :0.2.1-dev.3ac8448, for amd64 and arm64 —
digest sha256:3425488e…. The gate passed with nothing to report once pip left the image.

ragnar is deployed from it and healthy:

{status: ok, database: ok, version: 0.2.1}

on both http://100.66.30.95:8000/healthz and https://postulo.tiagoagueda.com/healthz.
compose.yml pulls now instead of building — that is roughly ten minutes of Pi build time
per deploy gone. Config and a verified database backup were taken first
(…bak-20260912-1614), and deploy.log records the tag, the pinned tag, the digest and
the commit, so a rollback has something to name.

Two more faults, which makes six

Both in the release path, both would have hit the first real release:

  1. The pip removal used the wrong interpreter. python in the runtime stage is
    /app/.venv/bin/python — PATH puts the environment first — and that one never had
    pip, so python -m pip uninstall failed before removing anything. Removed by path now,
    checked against /usr/local/bin/python. Worth keeping: this means installer()'s pip
    branch could never have worked in the container anyway, because sys.executable is the
    environment's python.
  2. The image name carried a capital. GITHUB_REPOSITORY is Postulo/postulo and a
    Docker repository name may not have one, so the tag was refused after a full
    two-architecture build. image.yml computes it identically.

Changed from what this issue proposed

The dev image is multi-arch after all. The workflow said native-only, arguing that
anybody needing another architecture wanted a release. That was wrong here: the instance
these images exist to be run on is a Raspberry Pi, so an amd64-only dev image is one nobody
can deploy. QEMU binfmt was already registered on the runner, so it cost time and nothing
else.

Still open

  • Retention. Nothing prunes old 0.x.y-dev.<sha> tags, and there will be one per push.
  • :latest remains untouched, as intended.
  • Deployment is by hand, as decided — no watcher pulls on ragnar.
## Built, published, and running on ragnar `3ac8448c7` published **`:dev`** and **`:0.2.1-dev.3ac8448`**, for **amd64 and arm64** — digest `sha256:3425488e…`. The gate passed with nothing to report once pip left the image. ragnar is deployed from it and healthy: {status: ok, database: ok, version: 0.2.1} on both `http://100.66.30.95:8000/healthz` and https://postulo.tiagoagueda.com/healthz. `compose.yml` pulls now instead of building — that is roughly ten minutes of Pi build time per deploy gone. Config and a verified database backup were taken first (`…bak-20260912-1614`), and `deploy.log` records the tag, the pinned tag, the digest and the commit, so a rollback has something to name. ### Two more faults, which makes six Both in the release path, both would have hit the first real release: 5. **The pip removal used the wrong interpreter.** `python` in the runtime stage is `/app/.venv/bin/python` — PATH puts the environment first — and *that* one never had pip, so `python -m pip uninstall` failed before removing anything. Removed by path now, checked against `/usr/local/bin/python`. Worth keeping: this means `installer()`'s pip branch could never have worked in the container anyway, because `sys.executable` is the environment's python. 6. **The image name carried a capital.** `GITHUB_REPOSITORY` is `Postulo/postulo` and a Docker repository name may not have one, so the tag was refused *after* a full two-architecture build. `image.yml` computes it identically. ### Changed from what this issue proposed **The dev image is multi-arch after all.** The workflow said native-only, arguing that anybody needing another architecture wanted a release. That was wrong here: the instance these images exist to be run on is a Raspberry Pi, so an amd64-only dev image is one nobody can deploy. QEMU binfmt was already registered on the runner, so it cost time and nothing else. ### Still open - **Retention.** Nothing prunes old `0.x.y-dev.<sha>` tags, and there will be one per push. - `:latest` remains untouched, as intended. - Deployment is by hand, as decided — no watcher pulls on ragnar.
Author
Owner

Retention is in: 1683088a8

dev-image.yml ends with scripts/prune-dev-images.py now. The newest five pinned dev
tags stay (DEV_IMAGES_KEPT); the rest go, and so do the per-architecture manifests only
they referenced -- a manifest nothing names still holds its layers, so deleting a tag alone
would free nothing. Releases, latest, dev, and anything untagged the run did not itself
orphan are never candidates; that last case is what a push in flight from image.yml looks
like, which is why the script does not sweep "everything unreferenced".

What was exercised where

  • The logic -- choosing, ordering, and what counts as orphaned -- against a fake
    registry in tests/test_prune_dev_images.py (9 tests).
  • The HTTP half, for real, with my own token on a disposable package
    (postulo/prune-test, two multi-architecture pushes from ouranos, since removed):
    the listing, the /v2/token exchange, reading an index, and DELETE of a tagged
    version and of sha256: versions -- the older tag went with its four manifests, the
    kept build's four survived. Then the package was deleted.
  • In CI, run 315, on the runner's Ubuntu Python with REGISTRY_TOKEN:
    nothing to prune: 3 tags, none past the newest 5 dev tags -- which is the listing
    working with the CI credential, and nothing doomed yet.

The one thing still unproven

Whether REGISTRY_TOKEN may delete a package version. It can push, and the listing
worked with it, but delete needs write:package and nothing has asked it to delete yet.
The first time it will is the sixth pinned tag -- four pushes to main from now. A
failing prune fails the run on purpose, so if the scope is missing it will be a red run
with prune failed: DELETE … -> 403 in the log, and the fix is the token's scope, not the
code. Worth glancing at that run.

Also corrected in the same commit: CONTRIBUTING.md still said the dev image was built
for the native architecture only.

## Retention is in: `1683088a8` `dev-image.yml` ends with `scripts/prune-dev-images.py` now. The newest **five** pinned dev tags stay (`DEV_IMAGES_KEPT`); the rest go, and so do the per-architecture manifests only they referenced -- a manifest nothing names still holds its layers, so deleting a tag alone would free nothing. Releases, `latest`, `dev`, and anything untagged the run did not itself orphan are never candidates; that last case is what a push in flight from `image.yml` looks like, which is why the script does not sweep "everything unreferenced". ### What was exercised where - **The logic** -- choosing, ordering, and what counts as orphaned -- against a fake registry in `tests/test_prune_dev_images.py` (9 tests). - **The HTTP half, for real**, with my own token on a disposable package (`postulo/prune-test`, two multi-architecture pushes from ouranos, since removed): the listing, the `/v2/token` exchange, reading an index, and `DELETE` of a tagged version and of `sha256:` versions -- the older tag went with its four manifests, the kept build's four survived. Then the package was deleted. - **In CI**, run 315, on the runner's Ubuntu Python with `REGISTRY_TOKEN`: `nothing to prune: 3 tags, none past the newest 5 dev tags` -- which is the listing working with the CI credential, and nothing doomed yet. ### The one thing still unproven Whether `REGISTRY_TOKEN` may **delete** a package version. It can push, and the listing worked with it, but delete needs `write:package` and nothing has asked it to delete yet. The first time it will is the **sixth** pinned tag -- four pushes to `main` from now. A failing prune fails the run on purpose, so if the scope is missing it will be a red run with `prune failed: DELETE … -> 403` in the log, and the fix is the token's scope, not the code. Worth glancing at that run. Also corrected in the same commit: `CONTRIBUTING.md` still said the dev image was built for the native architecture only.
Author
Owner

The one thing left unproven is proven

The sixth pinned tag arrived with 789f0d9d4 (run 327), and the prune step made its first
real deletion with REGISTRY_TOKEN:

deleted 0.2.1-dev.3ac8448
deleted sha256:12e8faab…  deleted sha256:2c638ae8…
deleted sha256:6cc84254…  deleted sha256:ec757934…
1 dev tags and 4 manifests; 5 dev tags kept

Those four are exactly the manifests only the first build referenced -- its two
architectures and two attestations -- and nothing that a remaining tag names was touched.
The registry holds dev, five pinned tags and 18 manifests. So the token may delete,
retention works unattended, and there is nothing left open on this issue.

One consequence worth knowing: ragnar's deploy.log names 0.2.1-dev.3ac8448, which is
the tag that just went. A rollback to it by tag is no longer possible from the registry;
ragnar's own daemon still holds that image. The next deploy moves the rollback window along.

## The one thing left unproven is proven The sixth pinned tag arrived with `789f0d9d4` (run 327), and the prune step made its first real deletion **with `REGISTRY_TOKEN`**: deleted 0.2.1-dev.3ac8448 deleted sha256:12e8faab… deleted sha256:2c638ae8… deleted sha256:6cc84254… deleted sha256:ec757934… 1 dev tags and 4 manifests; 5 dev tags kept Those four are exactly the manifests only the first build referenced -- its two architectures and two attestations -- and nothing that a remaining tag names was touched. The registry holds `dev`, five pinned tags and 18 manifests. So the token may delete, retention works unattended, and there is nothing left open on this issue. One consequence worth knowing: ragnar's `deploy.log` names `0.2.1-dev.3ac8448`, which is the tag that just went. A rollback to it by tag is no longer possible from the registry; ragnar's own daemon still holds that image. The next deploy moves the rollback window along.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
Postulo/postulo#190
No description provided.