Docker: /var/lib/docker needs a wear budget before any workload lands #65
Labels
No labels
blocked-physical
cleanup
hardware
infra
kernel
P1-critical
P2-high
P3-normal
P4-later
reliability
security
upstream
wontfix
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
tiagoagueda/a80#65
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Split out of the Docker enablement work. The kernel side is handled separately; this
is the storage consequence.
Measured
life_time0x02 0x02- 10-20 % of rated write endurance already consumedpre_eol_info0x01(normal)//dev/mmcblk1p1, 29 G, 25 G free, single partition -/varis not separateEvery container layer, every writable container filesystem and every log line lands on
the same eMMC that holds the bootloader, the kernel and the rootfs. There is no second
partition absorbing the churn, and no quota.
Why it matters here specifically
Two Docker defaults are hostile to flash:
json-filelogging is unbounded by default. A chatty container will write untilthe disk fills. On this board a full
/also means no/bootwrites, sodeploy-kernel.shand the known-good rollout stop working.overlay2only became available with the Tier 0 kernel change. WithoutCONFIG_OVERLAY_FSthe daemon falls back tovfs, which makes a full physical copyof every layer. Confirm the daemon actually selected
overlay2rather than quietlyfalling back.
Options
daemon.jsonwithlog-optsmax-size/max-file, or thelocallog driver.Cheapest; do this regardless.
data-rootto USB. USB3 works (32-USB3-WORKING.md), but it adds a boot-orderdependency: the daemon must not start before the mount.
docker system prune.life_time/pre_eol_infoas a metric - belongs with #10.Done when
daemon.jsoncaps log size, and the cap is inprovision.shdocker infoconfirmsoverlay2, notvfsPartly addressed while enabling Docker, so the scope here is smaller than when it was filed.
Done:
/etc/docker/daemon.jsoncaps logs at10mx 3 from the moment of install. Leaving acontainer host with unbounded
json-filelogging on an eMMC atlife_time 0x02was not worthdeferring even for one session.
Answered: the storage driver question.
docker inforeportsoverlayfswithdriver-type: io.containerd.snapshotter.v1- Docker 29 uses the containerd image store, so it isthe overlayfs snapshotter rather than the old
overlay2graph driver. Either way it is notvfs, which is what mattered:CONFIG_OVERLAY_FSis now built and the full-copy-per-layerfallback is not in play.
Still open, and the substantive part: data-root placement (eMMC vs USB), a prune policy, and
getting
life_time/pre_eol_infosomewhere that will be noticed (#10). Baseline for latercomparison:
/was 2.4 G used before Docker, 2.9 G after the install plus analpineand annginx:alpineimage, of 29 G.Done, with two corrections to the issue as I filed it.
Correction 1: the bytes are not where this issue said
With Docker 29's containerd image store,
/var/lib/dockerholds 240K. The images are in/var/lib/containerd(67M) - content store 22M plus overlayfs snapshots 45M. So adata-rootmove, which is what the issue proposed, relocates 240K of metadata and looks done.Anything that moves container storage has to move containerd's
rootas well.Correction 2:
overlay2vsoverlayfsdocker inforeportsStorage Driver: overlayfswithdriver-type: io.containerd.snapshotter.v1- the containerd snapshotter, not the old graph driver. Differentname, same thing that matters: it is not
vfs, so the full-copy-per-layer fallback is not inplay.
The data-root decision, recorded
It stays on the eMMC.
lsusbshows a hub and a keyboard - there is no USB mass storageattached to move it to. Adding some would buy a boot-order dependency and a new failure mode
(the daemon must not start before the mount) in exchange for relieving a device currently at
11% full with 25 G free. Revisit if a workload actually generates churn; the health check
below is what would tell you.
What now exists
daemon.jsoncaps logs at10mx 3, inprovision.sh(#69).draco-storage-health, daily. Reads eMMClife_time/pre_eol_infoand rootfs free spaceand fails the unit, so degradation lands in
systemctl --failed. That is deliberate: while#10 is open, nothing else on this board surfaces anything. A number in a log would not have
helped. Thresholds:
pre_eol_info != 0x01,life_time >= 0x05(~40% of rated endurance), orunder 12% free.
Tested in both directions - normal run exits 0; with the thresholds tightened it exits 1
and does appear in
systemctl --failed.It resolves the root block device from sysfs rather than from the mmc host name, because
deriving
mmc1->1and substring-matching also matches/dev/mmcblk0p1- with the rescuecard inserted it would have reported on the wrong device.
docker-prune.timer, weekly, deliberately not-aand never--volumes.-aremovesunused tagged images, and on 4 armv7 cores re-pulling and re-extracting one costs exactly the
eMMC writes this is meant to reduce.
Baseline for later comparison:
/at 2.9 G of 29 G, eMMClife_time 0x02 0x02,pre_eol_info 0x01.