No fsck policy for unclean shutdowns #14
Labels
No labels
blocked-physical
cleanup
hardware
infra
kernel
P1-critical
P2-high
P3-normal
P4-later
reliability
security
upstream
wontfix
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
tiagoagueda/a80#14
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Mount count is already 82 and the board has now taken several unclean resets (two
deliberate watchdog tests, one SPL hang, one power cycle). Root is a single ext4 with
errors=remount-roand no separate/var.errors=remount-rois good — it fails safe — but an automatic remount-ro on a headless boxwith no monitoring is an invisible outage.
Done when
First checkbox is effectively done, 2026-08-28 - and it was quietly broken for part
of the day, which is worth recording.
Boot-time fsck happens in the initramfs and was verified running:
The trap: while making the A15 default safe I added a boot entry with no initrd, and that
silently removed the root fsck. Without an initramfs the kernel mounts root
rwdirectly andsystemd-fsck-rootthen skips itself:Fixed: both main entries load the initramfs again, and the two no-initrd entries
(
known-good,prev) now useroso systemd checks root for them. Anyone adding a bootentry here needs to know that
rwwithout an initramfs means the filesystem is never checked.The second checkbox - a remount-ro event raising an alert - still depends on #10. And the
board took several more unclean resets today from #53, so this matters more than it did.
Done — and the filesystem had never been checked, not once
The cause was not a missing policy but an unsatisfiable condition.
systemd-fsck-root.servicecarriesConditionPathIsReadWrite=!/, and every boot entry passedrw, so the initramfs mounted root read-write and the unit skipped itself on every boot since the board was built. Debian's initramfs only checks root when the command line saysro—scripts/localcallscheckfs()under[ "$readonly" = y ].All four eMMC entries and all four rescue-card entries now pass
ro; systemd remountsrwfrom/etc/fstabimmediately after. Proven on the serial console rather than assumed:One trap worth recording
Last checkeddoes not move whenfsck -afinds nothing to change. It still readAug 28 18:42after a full verified check had just run, so it is not merely unhelpful — it will actively convince you no check happened. The mount counter is reset reliably, and that is the signal to trust.That also decides the policy. A time-based interval (
-i 2w) would be a trap: withlastcheckfrozen, it would force a full check on every boot once the interval elapsed. So the policy is mount-count based:Unclean shutdowns are caught by the dirty flag on the next boot, which is the failure mode this board actually has; the 30-mount check catches slow drift.