U-Boot: automatic fallback to the known-good kernel (CONFIG_BOOTCOUNT_LIMIT) #36

Closed
opened 2026-08-28 05:59:30 +00:00 by tiagoagueda · 1 comment
Owner

uImage.known-good and dtb.known-good exist in /boot, but selecting them requires
a human on the serial console. On a headless board that means the fallback only works
when someone is already standing next to the failure.

Measured

# CONFIG_BOOTCOUNT_LIMIT is not set
CONFIG_ENV_IS_IN_MMC=y
CONFIG_ENV_OFFSET=0xC0000

Change

CONFIG_BOOTCOUNT_LIMIT=y
CONFIG_BOOTCOUNT_ENV=y

plus bootlimit and an altbootcmd pointing at the known-good extlinux entry.
Userspace clears the counter from a late systemd unit once the boot is judged good:

fw_setenv bootcount 0

/etc/fw_env.config already points at 0xC0000, so no new plumbing is needed.

Why it is possible now

BOOTCOUNT_ENV needs somewhere persistent to write. The move from ENV_IS_IN_FAT to
raw MMC at 0xC0000 in 38-userspace-hardening.md is what
makes this available — before that, saveenv could never work.

Cost

One 64 KiB env write per boot at 0xC0000. Against eMMC life_time 0x02 0x02
(10-20% used, per 41-production-plan.md) that is negligible.

Open questions

  • What counts as "boot is good"? Simplest useful definition: network reachable and no
    failed systemd units. Reset the counter from a unit ordered after network-online.target.
  • bootlimit value. 3 is the usual choice.

Acceptance

  • Deliberately install a kernel that does not boot; confirm the board falls back to
    known-good by itself and stays reachable.
  • Confirm the counter is cleared on a healthy boot and does not creep upward.
`uImage.known-good` and `dtb.known-good` exist in `/boot`, but selecting them requires a human on the serial console. On a headless board that means the fallback only works when someone is already standing next to the failure. ## Measured ``` # CONFIG_BOOTCOUNT_LIMIT is not set CONFIG_ENV_IS_IN_MMC=y CONFIG_ENV_OFFSET=0xC0000 ``` ## Change ``` CONFIG_BOOTCOUNT_LIMIT=y CONFIG_BOOTCOUNT_ENV=y ``` plus `bootlimit` and an `altbootcmd` pointing at the known-good extlinux entry. Userspace clears the counter from a late systemd unit once the boot is judged good: ```sh fw_setenv bootcount 0 ``` `/etc/fw_env.config` already points at `0xC0000`, so no new plumbing is needed. ## Why it is possible now `BOOTCOUNT_ENV` needs somewhere persistent to write. The move from `ENV_IS_IN_FAT` to raw MMC at `0xC0000` in [38-userspace-hardening.md](38-userspace-hardening.md) is what makes this available — before that, `saveenv` could never work. ## Cost One 64 KiB env write per boot at `0xC0000`. Against eMMC `life_time 0x02 0x02` (10-20% used, per [41-production-plan.md](41-production-plan.md)) that is negligible. ## Open questions - What counts as "boot is good"? Simplest useful definition: network reachable and no failed systemd units. Reset the counter from a unit ordered after `network-online.target`. - `bootlimit` value. 3 is the usual choice. ## Acceptance - Deliberately install a kernel that does not boot; confirm the board falls back to known-good by itself and stays reachable. - Confirm the counter is cleared on a healthy boot and does not creep upward.
Author
Owner

Done and verified 2026-08-28.

Two gates this issue did not mention, both of which would have made it look broken:

  • bootcount_env does nothing unless upgrade_available is non-zero. Both
    bootcount_store() and bootcount_load() open with that check and return early.
  • The trigger is bootcount > bootlimit, strictly greater - so bootlimit=3 falls back on
    the fourth unconfirmed boot, not the third.

And a dependency worth naming: BOOTCOUNT_ENV needs a working environment, and the
checked-in defconfig carried CONFIG_ENV_IS_NOWHERE=y, which silently disables it. Fixed as
part of this work - it would have failed here in a way that looked like a bootcount bug.

altbootcmd deliberately does not go through the extlinux menu. The menu is another thing
that can be wrong; the fallback should depend on as little as possible.

Verified by breaking it on purpose. bootcount forced to 5 against bootlimit=3 with the
marking unit disabled, then a reboot:

kernel: 7.2.0-gaef0384da620-dirty      <- uImage.known-good

rather than the ge2df6e79588f test kernel that was installed. Unattended, no console.

The open question here - what counts as a good boot - is answered narrowly, in
tools/boot-mark-good/: a default route plus sshd, and deliberately not "no failed
units". On a headless board "good" means somebody can get in to fix it, and an unrelated
failed unit must not condemn a kernel that is otherwise fine and reachable. It also skips the
write when the counter is already zero, halving what this costs the eMMC.

bootlimit=3 is still a guess; nobody has argued for it over 2 or 5. See
50-uboot-resilience.md.

**Done and verified 2026-08-28.** Two gates this issue did not mention, both of which would have made it look broken: - **`bootcount_env` does nothing unless `upgrade_available` is non-zero.** Both `bootcount_store()` and `bootcount_load()` open with that check and return early. - The trigger is `bootcount > bootlimit`, strictly greater - so `bootlimit=3` falls back on the fourth unconfirmed boot, not the third. And a dependency worth naming: `BOOTCOUNT_ENV` needs a working environment, and the checked-in defconfig carried `CONFIG_ENV_IS_NOWHERE=y`, which silently disables it. Fixed as part of this work - it would have failed here in a way that looked like a bootcount bug. `altbootcmd` deliberately does **not** go through the extlinux menu. The menu is another thing that can be wrong; the fallback should depend on as little as possible. **Verified by breaking it on purpose.** `bootcount` forced to 5 against `bootlimit=3` with the marking unit disabled, then a reboot: ``` kernel: 7.2.0-gaef0384da620-dirty <- uImage.known-good ``` rather than the `ge2df6e79588f` test kernel that was installed. Unattended, no console. The open question here - what counts as a good boot - is answered narrowly, in `tools/boot-mark-good/`: a default route plus sshd, and deliberately **not** "no failed units". On a headless board "good" means somebody can get in to fix it, and an unrelated failed unit must not condemn a kernel that is otherwise fine and reachable. It also skips the write when the counter is already zero, halving what this costs the eMMC. `bootlimit=3` is still a guess; nobody has argued for it over 2 or 5. See `50-uboot-resilience.md`.
Sign in to join this conversation.
No description provided.