Docker: /var/lib/docker needs a wear budget before any workload lands #65

Closed
opened 2026-08-30 11:38:48 +00:00 by tiagoagueda · 2 comments
Owner

Split out of the Docker enablement work. The kernel side is handled separately; this
is the storage consequence.

Measured

eMMC life_time 0x02 0x02 - 10-20 % of rated write endurance already consumed
eMMC pre_eol_info 0x01 (normal)
/ /dev/mmcblk1p1, 29 G, 25 G free, single partition - /var is not separate

Every container layer, every writable container filesystem and every log line lands on
the same eMMC that holds the bootloader, the kernel and the rootfs. There is no second
partition absorbing the churn, and no quota.

Why it matters here specifically

Two Docker defaults are hostile to flash:

  • json-file logging is unbounded by default. A chatty container will write until
    the disk fills. On this board a full / also means no /boot writes, so
    deploy-kernel.sh and the known-good rollout stop working.
  • overlay2 only became available with the Tier 0 kernel change. Without
    CONFIG_OVERLAY_FS the daemon falls back to vfs, which makes a full physical copy
    of every layer. Confirm the daemon actually selected overlay2 rather than quietly
    falling back.

Options

  1. daemon.json with log-opts max-size / max-file, or the local log driver.
    Cheapest; do this regardless.
  2. Move data-root to USB. USB3 works (32-USB3-WORKING.md), but it adds a boot-order
    dependency: the daemon must not start before the mount.
  3. Scheduled docker system prune.
  4. Sample life_time / pre_eol_info as a metric - belongs with #10.

Done when

  • daemon.json caps log size, and the cap is in provision.sh
  • docker info confirms overlay2, not vfs
  • a decision is recorded on data-root placement, with the reasoning
  • eMMC health is sampled somewhere that will actually be noticed
Split out of the Docker enablement work. The kernel side is handled separately; this is the storage consequence. ## Measured | | | |---|---| | eMMC `life_time` | `0x02 0x02` - 10-20 % of rated write endurance already consumed | | eMMC `pre_eol_info` | `0x01` (normal) | | `/` | `/dev/mmcblk1p1`, 29 G, 25 G free, **single partition - `/var` is not separate** | Every container layer, every writable container filesystem and every log line lands on the same eMMC that holds the bootloader, the kernel and the rootfs. There is no second partition absorbing the churn, and no quota. ## Why it matters here specifically Two Docker defaults are hostile to flash: - **`json-file` logging is unbounded by default.** A chatty container will write until the disk fills. On this board a full `/` also means no `/boot` writes, so `deploy-kernel.sh` and the known-good rollout stop working. - **`overlay2` only became available with the Tier 0 kernel change.** Without `CONFIG_OVERLAY_FS` the daemon falls back to `vfs`, which makes a full physical copy of every layer. Confirm the daemon actually selected `overlay2` rather than quietly falling back. ## Options 1. `daemon.json` with `log-opts` `max-size` / `max-file`, or the `local` log driver. Cheapest; do this regardless. 2. Move `data-root` to USB. USB3 works (`32-USB3-WORKING.md`), but it adds a boot-order dependency: the daemon must not start before the mount. 3. Scheduled `docker system prune`. 4. Sample `life_time` / `pre_eol_info` as a metric - belongs with #10. ## Done when - `daemon.json` caps log size, and the cap is in `provision.sh` - `docker info` confirms `overlay2`, not `vfs` - a decision is recorded on data-root placement, with the reasoning - eMMC health is sampled somewhere that will actually be noticed
Author
Owner

Partly addressed while enabling Docker, so the scope here is smaller than when it was filed.

Done: /etc/docker/daemon.json caps logs at 10m x 3 from the moment of install. Leaving a
container host with unbounded json-file logging on an eMMC at life_time 0x02 was not worth
deferring even for one session.

Answered: the storage driver question. docker info reports overlayfs with
driver-type: io.containerd.snapshotter.v1 - Docker 29 uses the containerd image store, so it is
the overlayfs snapshotter rather than the old overlay2 graph driver. Either way it is not
vfs
, which is what mattered: CONFIG_OVERLAY_FS is now built and the full-copy-per-layer
fallback is not in play.

Still open, and the substantive part: data-root placement (eMMC vs USB), a prune policy, and
getting life_time / pre_eol_info somewhere that will be noticed (#10). Baseline for later
comparison: / was 2.4 G used before Docker, 2.9 G after the install plus an alpine and an
nginx:alpine image, of 29 G.

Partly addressed while enabling Docker, so the scope here is smaller than when it was filed. **Done:** `/etc/docker/daemon.json` caps logs at `10m` x 3 from the moment of install. Leaving a container host with unbounded `json-file` logging on an eMMC at `life_time 0x02` was not worth deferring even for one session. **Answered:** the storage driver question. `docker info` reports **`overlayfs`** with `driver-type: io.containerd.snapshotter.v1` - Docker 29 uses the containerd image store, so it is the overlayfs snapshotter rather than the old `overlay2` graph driver. Either way it is **not `vfs`**, which is what mattered: `CONFIG_OVERLAY_FS` is now built and the full-copy-per-layer fallback is not in play. **Still open, and the substantive part:** data-root placement (eMMC vs USB), a prune policy, and getting `life_time` / `pre_eol_info` somewhere that will be noticed (#10). Baseline for later comparison: `/` was 2.4 G used before Docker, **2.9 G after** the install plus an `alpine` and an `nginx:alpine` image, of 29 G.
Author
Owner

Done, with two corrections to the issue as I filed it.

Correction 1: the bytes are not where this issue said

With Docker 29's containerd image store, /var/lib/docker holds 240K. The images are in
/var/lib/containerd (67M) - content store 22M plus overlayfs snapshots 45M. So a
data-root move, which is what the issue proposed, relocates 240K of metadata and looks done.
Anything that moves container storage has to move containerd's root as well.

Correction 2: overlay2 vs overlayfs

docker info reports Storage Driver: overlayfs with driver-type: io.containerd.snapshotter.v1 - the containerd snapshotter, not the old graph driver. Different
name, same thing that matters: it is not vfs, so the full-copy-per-layer fallback is not in
play.

The data-root decision, recorded

It stays on the eMMC. lsusb shows a hub and a keyboard - there is no USB mass storage
attached to move it to. Adding some would buy a boot-order dependency and a new failure mode
(the daemon must not start before the mount) in exchange for relieving a device currently at
11% full with 25 G free. Revisit if a workload actually generates churn; the health check
below is what would tell you.

What now exists

  • daemon.json caps logs at 10m x 3, in provision.sh (#69).
  • draco-storage-health, daily. Reads eMMC life_time/pre_eol_info and rootfs free space
    and fails the unit, so degradation lands in systemctl --failed. That is deliberate: while
    #10 is open, nothing else on this board surfaces anything. A number in a log would not have
    helped. Thresholds: pre_eol_info != 0x01, life_time >= 0x05 (~40% of rated endurance), or
    under 12% free.
    Tested in both directions - normal run exits 0; with the thresholds tightened it exits 1
    and does appear in systemctl --failed.
    It resolves the root block device from sysfs rather than from the mmc host name, because
    deriving mmc1 -> 1 and substring-matching also matches /dev/mmcblk0p1 - with the rescue
    card inserted it would have reported on the wrong device.
  • docker-prune.timer, weekly, deliberately not -a and never --volumes. -a removes
    unused tagged images, and on 4 armv7 cores re-pulling and re-extracting one costs exactly the
    eMMC writes this is meant to reduce.

Baseline for later comparison: / at 2.9 G of 29 G, eMMC life_time 0x02 0x02,
pre_eol_info 0x01.

Done, with two corrections to the issue as I filed it. ## Correction 1: the bytes are not where this issue said With Docker 29's containerd image store, `/var/lib/docker` holds **240K**. The images are in **`/var/lib/containerd` (67M)** - content store 22M plus overlayfs snapshots 45M. So a `data-root` move, which is what the issue proposed, relocates 240K of metadata and looks done. Anything that moves container storage has to move containerd's `root` as well. ## Correction 2: `overlay2` vs `overlayfs` `docker info` reports `Storage Driver: overlayfs` with `driver-type: io.containerd.snapshotter.v1` - the containerd snapshotter, not the old graph driver. Different name, same thing that matters: it is **not `vfs`**, so the full-copy-per-layer fallback is not in play. ## The data-root decision, recorded **It stays on the eMMC.** `lsusb` shows a hub and a keyboard - there is no USB mass storage attached to move it to. Adding some would buy a boot-order dependency and a new failure mode (the daemon must not start before the mount) in exchange for relieving a device currently at **11% full with 25 G free**. Revisit if a workload actually generates churn; the health check below is what would tell you. ## What now exists - **`daemon.json`** caps logs at `10m` x 3, in `provision.sh` (#69). - **`draco-storage-health`**, daily. Reads eMMC `life_time`/`pre_eol_info` and rootfs free space and **fails the unit**, so degradation lands in `systemctl --failed`. That is deliberate: while #10 is open, nothing else on this board surfaces anything. A number in a log would not have helped. Thresholds: `pre_eol_info != 0x01`, `life_time >= 0x05` (~40% of rated endurance), or under 12% free. Tested in **both** directions - normal run exits 0; with the thresholds tightened it exits 1 and does appear in `systemctl --failed`. It resolves the root block device from sysfs rather than from the mmc host name, because deriving `mmc1` -> `1` and substring-matching also matches `/dev/mmcblk0p1` - with the rescue card inserted it would have reported on the wrong device. - **`docker-prune.timer`**, weekly, deliberately **not** `-a` and never `--volumes`. `-a` removes unused *tagged* images, and on 4 armv7 cores re-pulling and re-extracting one costs exactly the eMMC writes this is meant to reduce. Baseline for later comparison: `/` at **2.9 G of 29 G**, eMMC `life_time 0x02 0x02`, `pre_eol_info 0x01`.
Sign in to join this conversation.
No description provided.