The release run never finishes: a job queued against a label no runner has #81
Labels
No labels
accessibility
authentication
breaking change
bug
documentation
enhancement
interface
internationalisation
observability
security
tier
1
tier
2
tier
3
tier/4
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Blocks
Reference
Postulo/postulo#81
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Observation
Found while releasing 0.2.0. The release itself worked — v0.2.0 is published with its sdist,
its wheel and the changelog section as notes. But the workflow run for the tag is stuck in
waitingand will stay there.What is wrong
release.yml's second job:BUILD_IMAGEis set nowhere — not on the repository, not on the organisation — so the jobshould be skipped. Forgejo schedules it anyway: it looks for a runner advertising the
dockerlabel, finds none, and queues the job. Nothing ever picks it up, so the run neverfinishes.
The comment at the top of
ci.ymldescribes this exact behaviour, in advance:It was written about
ci.ymland it came true inrelease.yml.Why it matters more than it looks
Same shape as #75, #76 and #77: a signal that means something other than what it appears to
mean. A run sitting in
waitingis not a failure and not a success — it is a release thatlooks unfinished for ever, on the one workflow whose output a person checks precisely when
they want to know whether a release worked.
The fix
Schedule the job on a runner that exists, then let the
ifskip it:With the variable unset the job lands on the ordinary runner and is skipped immediately, so
the run completes. With it set to
trueand a docker-labelled runner registered, nothingchanges from today's intent.
Worth checking at the same time whether
ci.ymlhas the same latent problem — every jobthere says
runs-on: ubuntu-latest, which the runner clearly does advertise, so probablynot, but the comment deserves to be true.
Not included
Actually building the image. That still needs a docker-capable runner,
BUILD_IMAGE=trueand
REGISTRY_USER/REGISTRY_TOKEN, which is infrastructure rather than code.Classification
Bug.
Half fixed, and reopening for the other half.
runs-on: ${{ vars.BUILD_IMAGE == 'true' && 'docker' || 'ubuntu-latest' }}landed in v0.2.1. The run for that tag no longer hangs — so the queue-against-a-label-nothing-advertises part is genuinely gone, and that was the part that made a finished release look unfinished for ever.But run 750 now reports failure rather than success. The release itself is fine: v0.2.1 is published with both the wheel and the sdist attached and the changelog as its notes, so the
releasejob did its work. Something about theimagejob still ends the run badly — either this Forgejo's runner does not evaluate an expression inruns-on, or it schedules the job and then fails it rather than skipping it.A visible failure beats an invisible hang, so this is progress rather than a wash. But 'the release worked and the run says it failed' is still a signal that means the wrong thing, which is the whole complaint of this issue.
Blocked on the job log, which this Forgejo does not expose through the API (no
actions/runs/{id}/jobs, noactions/jobs/{id}/logs, no artifacts endpoint, and the web UI's JSON returns500: task ... does not existeven for successful runs). If theimagejob's step output is pasted in, the fix is likely one line. The certain fallback, if expressions are not supported: move the image build into its own workflow file so the release workflow has one job and always completes.Closed in
4ea68c9, taking the fallback from the comment above — and the evidence that settles why turned out to be reachable after all.The API does expose per-job detail, at
/repos/{owner}/{repo}/actions/tasks. Notruns/{id}/jobs, which 404s, and not the web UI's JSON, which still returns 500 — buttasksgives a row per job with its name, status and timings. It says:releaseimagewaitingreleaseimageZero seconds, having run no steps. That is not a step failing, it is Forgejo declining to start the job — so the answer to the open question is the first branch: the expression in
runs-onis not evaluated, the job gets an unusable label, and where it used to queue for ever it now errors immediately. Theifnever gets a say either way, because scheduling happens first. No log needed.The fix. A workflow nobody starts cannot queue. The image build is
image.ymlnow,workflow_dispatchwith the tag to build as its input, andrelease.ymlhas one job and always completes. It checks out the tag rather than a branch — an image built from something that moves is one nobody can reproduce — and the tag reaches the shell through the environment rather than pasted into it.BUILD_IMAGEwent with the automatic trigger it gated. A switch on something that only happens when you press the button is a second way of saying no. Nothing that worked is lost: that job ran three times, failed three times, and never once built an image.What is and is not proven. zizmor audits all three workflow files clean and
release.ymlparses to a single job; CI is green. The release path itself needs a tag, which is a deliberate act and not mine to take, so the first real proof is the next release. I also did not dispatchimage.ymlto try it: without a docker runner registered that would queue exactly the job this removes.One loose end: run 741 is still
waiting— v0.2.0's original stuck run, the artefact this issue was filed about. Nothing new will join it, but it will sit there until somebody cancels it. Say the word and I will.A third consequence of no runner building the image, after #121: nothing has ever scanned it. Trivy and Grype were run against
postulo:latestby hand on the test instance today and found #154 (the browser test tooling and uv's cache shipped in the runtime image, ~430 MB including a bundled Node.js runtime) and #155 (a Debian security update the image does not carry). Both had been there since the image was built.#156 is the scan itself, blocked by this.