The release run never finishes: a job queued against a label no runner has #81

Closed
opened 2026-09-07 11:44:15 +00:00 by tiagoagueda · 3 comments
Owner

Observation

Found while releasing 0.2.0. The release itself worked — v0.2.0 is published with its sdist,
its wheel and the changelog section as notes. But the workflow run for the tag is stuck in
waiting and will stay there.

What is wrong

release.yml's second job:

  image:
    if: ${{ vars.BUILD_IMAGE == 'true' }}
    runs-on: docker

BUILD_IMAGE is set nowhere — not on the repository, not on the organisation — so the job
should be skipped. Forgejo schedules it anyway: it looks for a runner advertising the
docker label, finds none, and queues the job. Nothing ever picks it up, so the run never
finishes.

The comment at the top of ci.yml describes this exact behaviour, in advance:

A label no runner has does not fail the job: it queues it for ever, which looks exactly
like CI passing until somebody checks.

It was written about ci.yml and it came true in release.yml.

Why it matters more than it looks

Same shape as #75, #76 and #77: a signal that means something other than what it appears to
mean. A run sitting in waiting is not a failure and not a success — it is a release that
looks unfinished for ever, on the one workflow whose output a person checks precisely when
they want to know whether a release worked.

The fix

Schedule the job on a runner that exists, then let the if skip it:

  image:
    if: ${{ vars.BUILD_IMAGE == 'true' }}
    runs-on: ${{ vars.BUILD_IMAGE == 'true' && 'docker' || 'ubuntu-latest' }}

With the variable unset the job lands on the ordinary runner and is skipped immediately, so
the run completes. With it set to true and a docker-labelled runner registered, nothing
changes from today's intent.

Worth checking at the same time whether ci.yml has the same latent problem — every job
there says runs-on: ubuntu-latest, which the runner clearly does advertise, so probably
not, but the comment deserves to be true.

Not included

Actually building the image. That still needs a docker-capable runner, BUILD_IMAGE=true
and REGISTRY_USER/REGISTRY_TOKEN, which is infrastructure rather than code.

Classification

Bug.

## Observation Found while releasing 0.2.0. The release itself worked — v0.2.0 is published with its sdist, its wheel and the changelog section as notes. But **the workflow run for the tag is stuck in `waiting` and will stay there.** ## What is wrong `release.yml`'s second job: ```yaml image: if: ${{ vars.BUILD_IMAGE == 'true' }} runs-on: docker ``` `BUILD_IMAGE` is set nowhere — not on the repository, not on the organisation — so the job should be skipped. Forgejo schedules it anyway: it looks for a runner advertising the `docker` label, finds none, and queues the job. Nothing ever picks it up, so the run never finishes. The comment at the top of `ci.yml` describes this exact behaviour, in advance: > A label no runner has does not fail the job: it queues it for ever, which looks exactly > like CI passing until somebody checks. It was written about `ci.yml` and it came true in `release.yml`. ## Why it matters more than it looks Same shape as #75, #76 and #77: a signal that means something other than what it appears to mean. A run sitting in `waiting` is not a failure and not a success — it is a release that looks unfinished for ever, on the one workflow whose output a person checks precisely when they want to know whether a release worked. ## The fix Schedule the job on a runner that exists, then let the `if` skip it: ```yaml image: if: ${{ vars.BUILD_IMAGE == 'true' }} runs-on: ${{ vars.BUILD_IMAGE == 'true' && 'docker' || 'ubuntu-latest' }} ``` With the variable unset the job lands on the ordinary runner and is skipped immediately, so the run completes. With it set to `true` and a docker-labelled runner registered, nothing changes from today's intent. Worth checking at the same time whether `ci.yml` has the same latent problem — every job there says `runs-on: ubuntu-latest`, which the runner clearly does advertise, so probably not, but the comment deserves to be true. ## Not included Actually building the image. That still needs a docker-capable runner, `BUILD_IMAGE=true` and `REGISTRY_USER`/`REGISTRY_TOKEN`, which is infrastructure rather than code. ## Classification Bug.
tiagoagueda added this to the 0.3.0 milestone 2026-09-07 11:44:15 +00:00
Author
Owner

Half fixed, and reopening for the other half.

runs-on: ${{ vars.BUILD_IMAGE == 'true' && 'docker' || 'ubuntu-latest' }} landed in v0.2.1. The run for that tag no longer hangs — so the queue-against-a-label-nothing-advertises part is genuinely gone, and that was the part that made a finished release look unfinished for ever.

But run 750 now reports failure rather than success. The release itself is fine: v0.2.1 is published with both the wheel and the sdist attached and the changelog as its notes, so the release job did its work. Something about the image job still ends the run badly — either this Forgejo's runner does not evaluate an expression in runs-on, or it schedules the job and then fails it rather than skipping it.

A visible failure beats an invisible hang, so this is progress rather than a wash. But 'the release worked and the run says it failed' is still a signal that means the wrong thing, which is the whole complaint of this issue.

Blocked on the job log, which this Forgejo does not expose through the API (no actions/runs/{id}/jobs, no actions/jobs/{id}/logs, no artifacts endpoint, and the web UI's JSON returns 500: task ... does not exist even for successful runs). If the image job's step output is pasted in, the fix is likely one line. The certain fallback, if expressions are not supported: move the image build into its own workflow file so the release workflow has one job and always completes.

Half fixed, and reopening for the other half. `runs-on: ${{ vars.BUILD_IMAGE == 'true' && 'docker' || 'ubuntu-latest' }}` landed in v0.2.1. The run for that tag **no longer hangs** — so the queue-against-a-label-nothing-advertises part is genuinely gone, and that was the part that made a finished release look unfinished for ever. But run 750 now reports **failure** rather than success. The release itself is fine: v0.2.1 is published with both the wheel and the sdist attached and the changelog as its notes, so the `release` job did its work. Something about the `image` job still ends the run badly — either this Forgejo's runner does not evaluate an expression in `runs-on`, or it schedules the job and then fails it rather than skipping it. A visible failure beats an invisible hang, so this is progress rather than a wash. But 'the release worked and the run says it failed' is still a signal that means the wrong thing, which is the whole complaint of this issue. **Blocked on the job log**, which this Forgejo does not expose through the API (no `actions/runs/{id}/jobs`, no `actions/jobs/{id}/logs`, no artifacts endpoint, and the web UI's JSON returns `500: task ... does not exist` even for successful runs). If the `image` job's step output is pasted in, the fix is likely one line. The certain fallback, if expressions are not supported: move the image build into its own workflow file so the release workflow has one job and always completes.
Author
Owner

Closed in 4ea68c9, taking the fallback from the comment above — and the evidence that settles why turned out to be reachable after all.

The API does expose per-job detail, at /repos/{owner}/{repo}/actions/tasks. Not runs/{id}/jobs, which 404s, and not the web UI's JSON, which still returns 500 — but tasks gives a row per job with its name, status and timings. It says:

run job status duration
111 (v0.2.0) release success 10s
111 image never became a task — the run sat in waiting —
120 (v0.2.1) release success 11s
120 image failure 0s

Zero seconds, having run no steps. That is not a step failing, it is Forgejo declining to start the job — so the answer to the open question is the first branch: the expression in runs-on is not evaluated, the job gets an unusable label, and where it used to queue for ever it now errors immediately. The if never gets a say either way, because scheduling happens first. No log needed.

The fix. A workflow nobody starts cannot queue. The image build is image.yml now, workflow_dispatch with the tag to build as its input, and release.yml has one job and always completes. It checks out the tag rather than a branch — an image built from something that moves is one nobody can reproduce — and the tag reaches the shell through the environment rather than pasted into it.

BUILD_IMAGE went with the automatic trigger it gated. A switch on something that only happens when you press the button is a second way of saying no. Nothing that worked is lost: that job ran three times, failed three times, and never once built an image.

What is and is not proven. zizmor audits all three workflow files clean and release.yml parses to a single job; CI is green. The release path itself needs a tag, which is a deliberate act and not mine to take, so the first real proof is the next release. I also did not dispatch image.yml to try it: without a docker runner registered that would queue exactly the job this removes.

One loose end: run 741 is still waiting — v0.2.0's original stuck run, the artefact this issue was filed about. Nothing new will join it, but it will sit there until somebody cancels it. Say the word and I will.

Closed in 4ea68c9, taking the fallback from the comment above — and the evidence that settles *why* turned out to be reachable after all. **The API does expose per-job detail**, at `/repos/{owner}/{repo}/actions/tasks`. Not `runs/{id}/jobs`, which 404s, and not the web UI's JSON, which still returns 500 — but `tasks` gives a row per job with its name, status and timings. It says: | run | job | status | duration | |---|---|---|---| | 111 (v0.2.0) | `release` | success | 10s | | 111 | `image` | *never became a task* — the run sat in `waiting` | — | | 120 (v0.2.1) | `release` | success | 11s | | 120 | `image` | **failure** | **0s** | **Zero seconds, having run no steps.** That is not a step failing, it is Forgejo declining to start the job — so the answer to the open question is the first branch: the expression in `runs-on` is not evaluated, the job gets an unusable label, and where it used to queue for ever it now errors immediately. The `if` never gets a say either way, because scheduling happens first. No log needed. **The fix.** A workflow nobody starts cannot queue. The image build is `image.yml` now, `workflow_dispatch` with the tag to build as its input, and `release.yml` has one job and always completes. It checks out the tag rather than a branch — an image built from something that moves is one nobody can reproduce — and the tag reaches the shell through the environment rather than pasted into it. `BUILD_IMAGE` went with the automatic trigger it gated. A switch on something that only happens when you press the button is a second way of saying no. **Nothing that worked is lost**: that job ran three times, failed three times, and never once built an image. **What is and is not proven.** zizmor audits all three workflow files clean and `release.yml` parses to a single job; CI is green. The release path itself needs a tag, which is a deliberate act and not mine to take, so the first real proof is the next release. I also did not dispatch `image.yml` to try it: without a docker runner registered that would queue exactly the job this removes. One loose end: **run 741 is still `waiting`** — v0.2.0's original stuck run, the artefact this issue was filed about. Nothing new will join it, but it will sit there until somebody cancels it. Say the word and I will.
Author
Owner

A third consequence of no runner building the image, after #121: nothing has ever scanned it. Trivy and Grype were run against postulo:latest by hand on the test instance today and found #154 (the browser test tooling and uv's cache shipped in the runtime image, ~430 MB including a bundled Node.js runtime) and #155 (a Debian security update the image does not carry). Both had been there since the image was built.

#156 is the scan itself, blocked by this.

A third consequence of no runner building the image, after #121: nothing has ever scanned it. Trivy and Grype were run against `postulo:latest` by hand on the test instance today and found #154 (the browser test tooling and uv's cache shipped in the runtime image, ~430 MB including a bundled Node.js runtime) and #155 (a Debian security update the image does not carry). Both had been there since the image was built. #156 is the scan itself, blocked by this.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Reference
Postulo/postulo#81
No description provided.