No fsck policy for unclean shutdowns #14

Closed
opened 2026-08-27 23:22:25 +00:00 by tiagoagueda · 2 comments
Owner

Mount count is already 82 and the board has now taken several unclean resets (two
deliberate watchdog tests, one SPL hang, one power cycle). Root is a single ext4 with
errors=remount-ro and no separate /var.

errors=remount-ro is good — it fails safe — but an automatic remount-ro on a headless box
with no monitoring is an invisible outage.

Done when

  • periodic/boot-time fsck behaviour is decided and configured
  • a remount-ro event raises an alert (depends on the monitoring issue)
Mount count is already 82 and the board has now taken several unclean resets (two deliberate watchdog tests, one SPL hang, one power cycle). Root is a single ext4 with `errors=remount-ro` and no separate `/var`. `errors=remount-ro` is good — it fails safe — but an automatic remount-ro on a headless box with no monitoring is an invisible outage. **Done when** - [ ] periodic/boot-time fsck behaviour is decided and configured - [ ] a remount-ro event raises an alert (depends on the monitoring issue)
Author
Owner

First checkbox is effectively done, 2026-08-28 - and it was quietly broken for part
of the day, which is worth recording.

Boot-time fsck happens in the initramfs and was verified running:

[/sbin/fsck.ext4 (1) -- /dev/mmcblk1p1] fsck.ext4 -a -C0 /dev/mmcblk1p1
debian-a80: clean, 27665/1908736 files, 510469/7633664 blocks

The trap: while making the A15 default safe I added a boot entry with no initrd, and that
silently removed the root fsck. Without an initramfs the kernel mounts root rw directly and
systemd-fsck-root then skips itself:

systemd-fsck-root.service ... skipped, unmet condition check ConditionPathIsReadWrite=!/

Fixed: both main entries load the initramfs again, and the two no-initrd entries
(known-good, prev) now use ro so systemd checks root for them. Anyone adding a boot
entry here needs to know that rw without an initramfs means the filesystem is never checked.

The second checkbox - a remount-ro event raising an alert - still depends on #10. And the
board took several more unclean resets today from #53, so this matters more than it did.

**First checkbox is effectively done, 2026-08-28** - and it was quietly broken for part of the day, which is worth recording. Boot-time fsck happens in the initramfs and was verified running: ``` [/sbin/fsck.ext4 (1) -- /dev/mmcblk1p1] fsck.ext4 -a -C0 /dev/mmcblk1p1 debian-a80: clean, 27665/1908736 files, 510469/7633664 blocks ``` **The trap:** while making the A15 default safe I added a boot entry with no initrd, and that silently removed the root fsck. Without an initramfs the kernel mounts root `rw` directly and `systemd-fsck-root` then skips itself: ``` systemd-fsck-root.service ... skipped, unmet condition check ConditionPathIsReadWrite=!/ ``` Fixed: both main entries load the initramfs again, and the two no-initrd entries (`known-good`, `prev`) now use `ro` so systemd checks root for them. Anyone adding a boot entry here needs to know that `rw` without an initramfs means the filesystem is never checked. The second checkbox - a remount-ro event raising an alert - still depends on #10. And the board took several more unclean resets today from #53, so this matters more than it did.
Author
Owner

Done — and the filesystem had never been checked, not once

The cause was not a missing policy but an unsatisfiable condition. systemd-fsck-root.service carries ConditionPathIsReadWrite=!/, and every boot entry passed rw, so the initramfs mounted root read-write and the unit skipped itself on every boot since the board was built. Debian's initramfs only checks root when the command line says ro — scripts/local calls checkfs() under [ "$readonly" = y ].

All four eMMC entries and all four rescue-card entries now pass ro; systemd remounts rw from /etc/fstab immediately after. Proven on the serial console rather than assumed:

Begin: Will now check root file system ... fsck from util-linux 2.41.5
[/sbin/fsck.ext4 (1) -- /dev/mmcblk1p1] fsck.ext4 -a -C0 /dev/mmcblk1p1
debian-a80: |========================================| 100.0%
debian-a80: 27841/1908736 files (0.3% non-contiguous), 511858/7633664 blocks

One trap worth recording

Last checked does not move when fsck -a finds nothing to change. It still read Aug 28 18:42 after a full verified check had just run, so it is not merely unhelpful — it will actively convince you no check happened. The mount counter is reset reliably, and that is the signal to trust.

That also decides the policy. A time-based interval (-i 2w) would be a trap: with lastcheck frozen, it would force a full check on every boot once the interval elapsed. So the policy is mount-count based:

tune2fs -c 30 -i 0   # on both /dev/mmcblk1p1 and /dev/mmcblk0p1

Unclean shutdowns are caught by the dirty flag on the next boot, which is the failure mode this board actually has; the 30-mount check catches slow drift.

## Done — and the filesystem had never been checked, not once The cause was not a missing policy but an unsatisfiable condition. `systemd-fsck-root.service` carries `ConditionPathIsReadWrite=!/`, and every boot entry passed **`rw`**, so the initramfs mounted root read-write and the unit skipped itself on every boot since the board was built. Debian's initramfs only checks root when the command line says `ro` — `scripts/local` calls `checkfs()` under `[ "$readonly" = y ]`. All four eMMC entries and all four rescue-card entries now pass `ro`; systemd remounts `rw` from `/etc/fstab` immediately after. Proven on the serial console rather than assumed: ``` Begin: Will now check root file system ... fsck from util-linux 2.41.5 [/sbin/fsck.ext4 (1) -- /dev/mmcblk1p1] fsck.ext4 -a -C0 /dev/mmcblk1p1 debian-a80: |========================================| 100.0% debian-a80: 27841/1908736 files (0.3% non-contiguous), 511858/7633664 blocks ``` ### One trap worth recording **`Last checked` does not move when `fsck -a` finds nothing to change.** It still read `Aug 28 18:42` after a full verified check had just run, so it is not merely unhelpful — it will actively convince you no check happened. The **mount counter** is reset reliably, and that is the signal to trust. That also decides the policy. A time-based interval (`-i 2w`) would be a trap: with `lastcheck` frozen, it would force a full check on *every* boot once the interval elapsed. So the policy is mount-count based: ``` tune2fs -c 30 -i 0 # on both /dev/mmcblk1p1 and /dev/mmcblk0p1 ``` Unclean shutdowns are caught by the dirty flag on the next boot, which is the failure mode this board actually has; the 30-mount check catches slow drift.
Sign in to join this conversation.
No description provided.