Docker: provision.sh must install and configure the daemon #69

Closed
opened 2026-08-30 11:38:48 +00:00 by tiagoagueda · 2 comments
Owner

Split out of the Docker enablement work.

The kernel side is carried in the repo as patches/configs/docker.config and folded
into the tracked defconfig, so a kernel rebuilt from git is Docker-capable. The
userspace side is not carried anywhere yet.

As it stands, a rootfs rebuilt from provision.sh would come up with a Docker-capable
kernel and no Docker on it, and nobody would notice until they tried to start a
container. That is the same drift that produced #64 and the a15=off gap.

Needs to be in provision.sh

  • the download.docker.com apt repo and signing key, pinned to trixie (verified
    present for armhf: docker-ce 5:29.7.2, containerd.io 2.3.4,
    docker-compose-plugin 5.5.0, docker-buildx-plugin 0.36.1)
  • docker-ce docker-ce-cli containerd.io docker-buildx-plugin docker-compose-plugin
  • daemon.json - log caps and data-root, whatever gets decided
  • the firewall decision
  • a chk assertion in the style of the existing ones that fails loudly if the running
    kernel cannot support the daemon. /sys/fs/cgroup/cgroup.controllers being non-empty
    is a good single probe: it read back empty on the pre-Tier-0 kernel, and it is the
    first thing to disappear if someone rebuilds from a stale defconfig.

Done when

provision.sh produces a host that can run a container, and its verify pass fails
loudly if the kernel underneath cannot.

Split out of the Docker enablement work. The kernel side is carried in the repo as `patches/configs/docker.config` and folded into the tracked defconfig, so a kernel rebuilt from git is Docker-capable. **The userspace side is not carried anywhere yet.** As it stands, a rootfs rebuilt from `provision.sh` would come up with a Docker-capable kernel and no Docker on it, and nobody would notice until they tried to start a container. That is the same drift that produced #64 and the `a15=off` gap. ## Needs to be in `provision.sh` - the `download.docker.com` apt repo and signing key, pinned to `trixie` (verified present for armhf: `docker-ce 5:29.7.2`, `containerd.io 2.3.4`, `docker-compose-plugin 5.5.0`, `docker-buildx-plugin 0.36.1`) - `docker-ce docker-ce-cli containerd.io docker-buildx-plugin docker-compose-plugin` - `daemon.json` - log caps and data-root, whatever gets decided - the firewall decision - a `chk` assertion in the style of the existing ones that fails loudly if the running kernel cannot support the daemon. `/sys/fs/cgroup/cgroup.controllers` being non-empty is a good single probe: it read back **empty** on the pre-Tier-0 kernel, and it is the first thing to disappear if someone rebuilds from a stale defconfig. ## Done when `provision.sh` produces a host that can run a container, and its verify pass fails loudly if the kernel underneath cannot.
Author
Owner

One addition to the scope, found after this was filed.

The board's u-boot environment is now part of what makes this kernel bootable, not just the
kernel config. bootcmd sets kernel_addr_r=0x20007fc0 so the image clears the 8 MiB bootm
ceiling (#72). A rebuilt or restored board whose env does not carry that will load a
Docker-capable kernel and reset before Starting kernel is printed, looking exactly like an
instant kernel panic.

So the chk assertion suggested above should be two:

  • /sys/fs/cgroup/cgroup.controllers is non-empty (the kernel can host containers)
  • fw_printenv bootcmd contains 0x20007fc0, or u-boot has been fixed per #72 (the kernel can
    actually be started)

The second is the one that fails silently and expensively. A backup of the pre-change env is on
the board at /root/uboot-env.bak-pre-bootm-20260830-120132.txt.

One addition to the scope, found after this was filed. The board's **u-boot environment** is now part of what makes this kernel bootable, not just the kernel config. `bootcmd` sets `kernel_addr_r=0x20007fc0` so the image clears the 8 MiB bootm ceiling (#72). A rebuilt or restored board whose env does not carry that will load a Docker-capable kernel and **reset before `Starting kernel` is printed**, looking exactly like an instant kernel panic. So the `chk` assertion suggested above should be two: - `/sys/fs/cgroup/cgroup.controllers` is non-empty (the kernel can host containers) - `fw_printenv bootcmd` contains `0x20007fc0`, *or* u-boot has been fixed per #72 (the kernel can actually be started) The second is the one that fails silently and expensively. A backup of the pre-change env is on the board at `/root/uboot-env.bak-pre-bootm-20260830-120132.txt`.
Author
Owner

Done. provision.sh reports 0 failures on the board, across a reboot.

Installed and configured

  • repo + key: files/etc/apt/sources.list.d/docker.list pinned to trixie with
    arch=armhf stated explicitly, and the signing key shipped in
    files/etc/apt/keyrings/docker.asc rather than fetched, so provisioning is deterministic and
    does not trust-on-first-use whatever download.docker.com serves that day. Fingerprint
    9DC858229FC7DD38854AE2D88D81803C0EBFCD88, Docker Release (CE deb).
  • packages in packages.txt: docker-ce docker-ce-cli containerd.io docker-buildx-plugin docker-compose-plugin.
  • daemon.json with the log caps from #65.
  • units: docker.service, draco-storage-health.timer, docker-prune.timer.
    docker.socket is deliberately not enabled - socket activation would start the daemon on
    first client access, which hides "the daemon is dead" until whatever needed it also fails.

One ordering trap worth naming: the keyring and sources file are copied in section 1, before
apt-get update
. Section 2 is where files/ normally lands, and by then the install has
already failed with Unable to locate package docker-ce.

The verify pass, which is the point of this issue

20 new assertions. The file-based ones check the repo pin, the key, the log caps, the prune not
being -a, the scoped nftables teardown, and that the forward chain both permits egress and does
not blanket-accept into docker0.

Then the live ones, exactly as suggested:

ok   kernel has cgroup controllers
ok     controller cpu / cpuset / io / memory / pids
ok   overlayfs, not vfs
ok   u-boot can load this kernel size

cgroup.controllers is checked on the running kernel, not inferred from files, because that
file read back empty before the kernel work and it is the first thing to disappear if someone
rebuilds from a stale defconfig. A daemon on a controller-less kernel installs, starts, reports
healthy, accepts every --memory/--cpus flag and enforces none of them.

The u-boot assertion is the second half of #72 and was added because the failure mode is silent:
over SYS_BOOTM_LEN, bootm resets before printing Starting kernel, which is indistinguishable
from an instant panic. It passes if the running u-boot has the raised limit or if bootcmd
loads at the XIP address so the limit never applies - and it reads the version from the board
rather than the build tree, since the tree is wrong precisely in the window where a bootloader is
built but not yet flashed.

Two things fixed on the way

  • .gitattributes did not cover tools/provision/**. These files were on their way to being
    CRLF in a Windows working tree, and they are scp'd straight from it to the board. That is the
    same trap that left tools/flash-emmc-uboot.sh unrunnable until today.
  • the trailing note claiming deploy-kernel.sh never installs a module tree (#8) is false and
    now actively misleading - 49 modules are deployed, and with the Docker netfilter/bridge/
    overlayfs code built as modules, modprobe working is load-bearing rather than theoretical.
Done. `provision.sh` reports **0 failures** on the board, across a reboot. ## Installed and configured - **repo + key**: `files/etc/apt/sources.list.d/docker.list` pinned to `trixie` with `arch=armhf` stated explicitly, and the signing key **shipped** in `files/etc/apt/keyrings/docker.asc` rather than fetched, so provisioning is deterministic and does not trust-on-first-use whatever `download.docker.com` serves that day. Fingerprint `9DC858229FC7DD38854AE2D88D81803C0EBFCD88`, *Docker Release (CE deb)*. - **packages** in `packages.txt`: `docker-ce docker-ce-cli containerd.io docker-buildx-plugin docker-compose-plugin`. - **`daemon.json`** with the log caps from #65. - **units**: `docker.service`, `draco-storage-health.timer`, `docker-prune.timer`. `docker.socket` is deliberately **not** enabled - socket activation would start the daemon on first client access, which hides "the daemon is dead" until whatever needed it also fails. One ordering trap worth naming: the keyring and sources file are copied **in section 1, before `apt-get update`**. Section 2 is where `files/` normally lands, and by then the install has already failed with `Unable to locate package docker-ce`. ## The verify pass, which is the point of this issue 20 new assertions. The file-based ones check the repo pin, the key, the log caps, the prune not being `-a`, the scoped nftables teardown, and that the forward chain both permits egress and does **not** blanket-accept into `docker0`. Then the live ones, exactly as suggested: ``` ok kernel has cgroup controllers ok controller cpu / cpuset / io / memory / pids ok overlayfs, not vfs ok u-boot can load this kernel size ``` `cgroup.controllers` is checked on the **running kernel**, not inferred from files, because that file read back **empty** before the kernel work and it is the first thing to disappear if someone rebuilds from a stale defconfig. A daemon on a controller-less kernel installs, starts, reports healthy, accepts every `--memory`/`--cpus` flag and enforces none of them. The u-boot assertion is the second half of #72 and was added because the failure mode is silent: over `SYS_BOOTM_LEN`, bootm resets before printing `Starting kernel`, which is indistinguishable from an instant panic. It passes if the running u-boot has the raised limit **or** if `bootcmd` loads at the XIP address so the limit never applies - and it reads the version from the board rather than the build tree, since the tree is wrong precisely in the window where a bootloader is built but not yet flashed. ## Two things fixed on the way - **`.gitattributes` did not cover `tools/provision/**`.** These files were on their way to being CRLF in a Windows working tree, and they are scp'd straight from it to the board. That is the same trap that left `tools/flash-emmc-uboot.sh` unrunnable until today. - the trailing note claiming `deploy-kernel.sh` never installs a module tree (#8) is false and now actively misleading - 49 modules are deployed, and with the Docker netfilter/bridge/ overlayfs code built as modules, `modprobe` working is load-bearing rather than theoretical.
Sign in to join this conversation.
No description provided.