Docker: provision.sh must install and configure the daemon #69
Labels
No labels
blocked-physical
cleanup
hardware
infra
kernel
P1-critical
P2-high
P3-normal
P4-later
reliability
security
upstream
wontfix
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
tiagoagueda/a80#69
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Split out of the Docker enablement work.
The kernel side is carried in the repo as
patches/configs/docker.configand foldedinto the tracked defconfig, so a kernel rebuilt from git is Docker-capable. The
userspace side is not carried anywhere yet.
As it stands, a rootfs rebuilt from
provision.shwould come up with a Docker-capablekernel and no Docker on it, and nobody would notice until they tried to start a
container. That is the same drift that produced #64 and the
a15=offgap.Needs to be in
provision.shdownload.docker.comapt repo and signing key, pinned totrixie(verifiedpresent for armhf:
docker-ce 5:29.7.2,containerd.io 2.3.4,docker-compose-plugin 5.5.0,docker-buildx-plugin 0.36.1)docker-ce docker-ce-cli containerd.io docker-buildx-plugin docker-compose-plugindaemon.json- log caps and data-root, whatever gets decidedchkassertion in the style of the existing ones that fails loudly if the runningkernel cannot support the daemon.
/sys/fs/cgroup/cgroup.controllersbeing non-emptyis a good single probe: it read back empty on the pre-Tier-0 kernel, and it is the
first thing to disappear if someone rebuilds from a stale defconfig.
Done when
provision.shproduces a host that can run a container, and its verify pass failsloudly if the kernel underneath cannot.
One addition to the scope, found after this was filed.
The board's u-boot environment is now part of what makes this kernel bootable, not just the
kernel config.
bootcmdsetskernel_addr_r=0x20007fc0so the image clears the 8 MiB bootmceiling (#72). A rebuilt or restored board whose env does not carry that will load a
Docker-capable kernel and reset before
Starting kernelis printed, looking exactly like aninstant kernel panic.
So the
chkassertion suggested above should be two:/sys/fs/cgroup/cgroup.controllersis non-empty (the kernel can host containers)fw_printenv bootcmdcontains0x20007fc0, or u-boot has been fixed per #72 (the kernel canactually be started)
The second is the one that fails silently and expensively. A backup of the pre-change env is on
the board at
/root/uboot-env.bak-pre-bootm-20260830-120132.txt.Done.
provision.shreports 0 failures on the board, across a reboot.Installed and configured
files/etc/apt/sources.list.d/docker.listpinned totrixiewitharch=armhfstated explicitly, and the signing key shipped infiles/etc/apt/keyrings/docker.ascrather than fetched, so provisioning is deterministic anddoes not trust-on-first-use whatever
download.docker.comserves that day. Fingerprint9DC858229FC7DD38854AE2D88D81803C0EBFCD88, Docker Release (CE deb).packages.txt:docker-ce docker-ce-cli containerd.io docker-buildx-plugin docker-compose-plugin.daemon.jsonwith the log caps from #65.docker.service,draco-storage-health.timer,docker-prune.timer.docker.socketis deliberately not enabled - socket activation would start the daemon onfirst client access, which hides "the daemon is dead" until whatever needed it also fails.
One ordering trap worth naming: the keyring and sources file are copied in section 1, before
apt-get update. Section 2 is wherefiles/normally lands, and by then the install hasalready failed with
Unable to locate package docker-ce.The verify pass, which is the point of this issue
20 new assertions. The file-based ones check the repo pin, the key, the log caps, the prune not
being
-a, the scoped nftables teardown, and that the forward chain both permits egress and doesnot blanket-accept into
docker0.Then the live ones, exactly as suggested:
cgroup.controllersis checked on the running kernel, not inferred from files, because thatfile read back empty before the kernel work and it is the first thing to disappear if someone
rebuilds from a stale defconfig. A daemon on a controller-less kernel installs, starts, reports
healthy, accepts every
--memory/--cpusflag and enforces none of them.The u-boot assertion is the second half of #72 and was added because the failure mode is silent:
over
SYS_BOOTM_LEN, bootm resets before printingStarting kernel, which is indistinguishablefrom an instant panic. It passes if the running u-boot has the raised limit or if
bootcmdloads at the XIP address so the limit never applies - and it reads the version from the board
rather than the build tree, since the tree is wrong precisely in the window where a bootloader is
built but not yet flashed.
Two things fixed on the way
.gitattributesdid not covertools/provision/**. These files were on their way to beingCRLF in a Windows working tree, and they are scp'd straight from it to the board. That is the
same trap that left
tools/flash-emmc-uboot.shunrunnable until today.deploy-kernel.shnever installs a module tree (#8) is false andnow actively misleading - 49 modules are deployed, and with the Docker netfilter/bridge/
overlayfs code built as modules,
modprobeworking is load-bearing rather than theoretical.