The image scan writes its reports to the host, so the gate never runs and its failure reads as a finding #192

Closed
opened 2026-09-12 12:44:26 +00:00 by tiagoagueda · 0 comments
Owner

What happens

The first time image.yml runs, the Scan it, with both scanners step fails with

cat: /workspace/Postulo/postulo/.scan/trivy-fixable.txt: No such file or directory

and nothing is published. The image is fine. The scanners are fine. What is broken is where
their output goes.

The gate never runs. scan-image.sh dies at the unguarded cat on line 91, which is
before the two commands that actually decide anything — the --exit-code 1 trivy on line 92
and the --fail-on grype on line 95. The script exits 1 either way, so from the outside a
plumbing failure and a genuine finding are the same event: a red Scan it, with both scanners
step. That is the part worth fixing quickly, ahead of the lost files. A gate that cannot be
told apart from its own failure is not reporting anything.

Why

scan-image.sh runs the scanners as containers and passes the report directory in as a bind
mount (lines 60 and 68):

$DOCKER run --rm \
    -v /var/run/docker.sock:/var/run/docker.sock \
    -v "$OUT:/out" \
    "$TRIVY" "$@"

That is correct when a person runs the script on their own machine, which is the only way it
has ever been run — $OUT is a real path on the same filesystem as the daemon, and
--ignore-unfixed output lands where the script then reads it.

In CI it is not the same filesystem. The job runs in a container, and $DOCKER talks to
the host's daemon through the mounted socket. So -v "$OUT:/out" is resolved by the host,
against a path that only exists inside the job container. The host has nothing there, creates
an empty directory, and the scanners write their reports into it — on the host, where the job
cannot see them.

$GITHUB_WORKSPACE in a Forgejo job is a per-task named volume, not a bind from the host.
Measured on the runner rather than assumed:

4846 4605 259:2 /var/lib/docker/volumes/FORGEJO-ACTIONS-TASK-2245_WORKFLOW-probe_JOB-probe/_data
  /workspace/tiagoagueda/docker-recipes rw,relatime master:1 - ext4 /dev/nvme0n1p2 rw

So there is no host path that corresponds to $OUT at all, and no arrangement of
SCAN_OUTPUT_DIR inside the workspace will produce one.

What survives and what does not

Everything that goes through a bind mount is lost; everything the job's own shell writes is
fine.

trivy-full.txt, trivy-fixable.txt, sbom.cdx.json written via --output /out/... — lost to the host
grype-fixable.txt (line 95) written by a shell redirect in the job — fine, but unreachable, the script has already exited
The $GITHUB_STEP_SUMMARY block runs under if: always(), finds nothing, prints no report
upload-artifact for the SBOM runs under if: always(), matches nothing, warns

Line 82 also prints wrote $OUT/sbom.cdx.json for a file that is not there, which is worth
removing whatever else changes — it is the one line that actively says the wrong thing.

The fix

The smallest change that works in both places is to stop bind-mounting for output and let the
calling shell place the bytes. Both tools write to stdout when --output is omitted:

trivy image --scanners vuln,secret,misconfig --format table "$IMAGE" > "$OUT/trivy-full.txt"
trivy image --format cyclonedx "$IMAGE" > "$OUT/sbom.cdx.json"

The redirect is performed by whoever ran the script, so it lands in the job container in CI and
on the host for a person, with no branch between the two cases and no -v "$OUT:/out" at all.
That is the same shape line 95 already uses for grype, which is why grype is the one that works.

Two alternatives, both worse here and noted so they do not get rediscovered: --volumes-from
the job's own container inherits the workspace correctly but only works when the script is
already inside a container, so it needs a branch on something the script cannot reliably
detect; and mounting the task volume by name requires knowing a name that only the runner
knows.

Whatever the mechanism, line 91 should not be the thing that decides the job's exit code.
Guarding it like line 76 already is would at least mean a missing report no longer impersonates
a finding.

Why this has not been seen before

Nothing had ever built an image in CI — image.yml needed a runner advertising docker and
none existed, which is #81 and the Check this first section of #190. That runner now exists,
registered to this repository only, so this is reachable for the first time. Expect it to be
the first of the usual first-run problems rather than the last.

The scan itself is not in question: run by hand on a host, scan-image.sh does exactly what it
says, and it is what found #155 and #157.

## What happens The first time `image.yml` runs, the *Scan it, with both scanners* step fails with ``` cat: /workspace/Postulo/postulo/.scan/trivy-fixable.txt: No such file or directory ``` and nothing is published. The image is fine. The scanners are fine. What is broken is where their output goes. **The gate never runs.** `scan-image.sh` dies at the unguarded `cat` on line 91, which is *before* the two commands that actually decide anything — the `--exit-code 1` trivy on line 92 and the `--fail-on` grype on line 95. The script exits 1 either way, so from the outside a plumbing failure and a genuine finding are the same event: a red *Scan it, with both scanners* step. That is the part worth fixing quickly, ahead of the lost files. A gate that cannot be told apart from its own failure is not reporting anything. ## Why `scan-image.sh` runs the scanners as containers and passes the report directory in as a bind mount (lines 60 and 68): ```sh $DOCKER run --rm \ -v /var/run/docker.sock:/var/run/docker.sock \ -v "$OUT:/out" \ "$TRIVY" "$@" ``` That is correct when a person runs the script on their own machine, which is the only way it has ever been run — `$OUT` is a real path on the same filesystem as the daemon, and `--ignore-unfixed` output lands where the script then reads it. **In CI it is not the same filesystem.** The job runs in a container, and `$DOCKER` talks to the *host's* daemon through the mounted socket. So `-v "$OUT:/out"` is resolved by the host, against a path that only exists inside the job container. The host has nothing there, creates an empty directory, and the scanners write their reports into it — on the host, where the job cannot see them. `$GITHUB_WORKSPACE` in a Forgejo job is a **per-task named volume**, not a bind from the host. Measured on the runner rather than assumed: ``` 4846 4605 259:2 /var/lib/docker/volumes/FORGEJO-ACTIONS-TASK-2245_WORKFLOW-probe_JOB-probe/_data /workspace/tiagoagueda/docker-recipes rw,relatime master:1 - ext4 /dev/nvme0n1p2 rw ``` So there is no host path that corresponds to `$OUT` at all, and no arrangement of `SCAN_OUTPUT_DIR` inside the workspace will produce one. ## What survives and what does not Everything that goes through a bind mount is lost; everything the job's own shell writes is fine. | | | |---|---| | `trivy-full.txt`, `trivy-fixable.txt`, `sbom.cdx.json` | written via `--output /out/...` — **lost to the host** | | `grype-fixable.txt` (line 95) | written by a shell redirect in the job — **fine**, but unreachable, the script has already exited | | The `$GITHUB_STEP_SUMMARY` block | runs under `if: always()`, finds nothing, prints `no report` | | `upload-artifact` for the SBOM | runs under `if: always()`, matches nothing, warns | Line 82 also prints `wrote $OUT/sbom.cdx.json` for a file that is not there, which is worth removing whatever else changes — it is the one line that actively says the wrong thing. ## The fix The smallest change that works in both places is to stop bind-mounting for output and let the *calling* shell place the bytes. Both tools write to stdout when `--output` is omitted: ```sh trivy image --scanners vuln,secret,misconfig --format table "$IMAGE" > "$OUT/trivy-full.txt" trivy image --format cyclonedx "$IMAGE" > "$OUT/sbom.cdx.json" ``` The redirect is performed by whoever ran the script, so it lands in the job container in CI and on the host for a person, with no branch between the two cases and no `-v "$OUT:/out"` at all. That is the same shape line 95 already uses for grype, which is why grype is the one that works. Two alternatives, both worse here and noted so they do not get rediscovered: `--volumes-from` the job's own container inherits the workspace correctly but only works when the script is *already* inside a container, so it needs a branch on something the script cannot reliably detect; and mounting the task volume by name requires knowing a name that only the runner knows. Whatever the mechanism, **line 91 should not be the thing that decides the job's exit code.** Guarding it like line 76 already is would at least mean a missing report no longer impersonates a finding. ## Why this has not been seen before Nothing had ever built an image in CI — `image.yml` needed a runner advertising `docker` and none existed, which is #81 and the *Check this first* section of #190. That runner now exists, registered to this repository only, so this is reachable for the first time. Expect it to be the first of the usual first-run problems rather than the last. The scan itself is not in question: run by hand on a host, `scan-image.sh` does exactly what it says, and it is what found #155 and #157.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
Postulo/postulo#192
No description provided.