U-Boot: arm a watchdog in SPL — the stage that actually hung #37

Closed
opened 2026-08-28 05:59:30 +00:00 by tiagoagueda · 1 comment
Owner

The one real outage on this board happened in SPL, and none of the other U-Boot
availability work covers it.

The event

2026-08-27, one occurrence in 68 boots that day, per gap 2 of
41-production-plan.md: DRAM initialised, then two garbage
characters instead of Trying to boot from MMC2, and the board stayed dead until it
was physically power-cycled.

Every recovery path that exists — hardware watchdog under Linux, extlinux fallback,
rescue SD — assumes the SoC got at least as far as U-Boot proper. None of them apply
here.

Why it is not just a config switch

# CONFIG_SPL_WATCHDOG is not set
# CONFIG_SPL_DM is not set          <- no WDT uclass available in SPL
CONFIG_SPL_GPIO=y
CONFIG_SPL_SERIAL=y

With no driver model in SPL, SPL_WATCHDOG has no uclass to drive. Realistic options:

  1. Enable SPL_DM + SPL_WDT and accept the SPL size growth
    (CONFIG_SPL_SIZE_LIMIT=0x0 today, and the boot area has ~544 KiB spare before
    p1, so there is room — see the layout in
    38-userspace-hardening.md).
  2. A small patch arming R_WDT (8001000.watchdog) with direct register writes early
    in sunxi_board_init(), before DRAM init. Both instances are already proven to
    reset the SoC.

Option 2 is smaller and does not perturb the SPL that currently works. Option 1 is
more likely to be acceptable upstream.

Interaction

If SPL arms the watchdog and U-Boot proper does not service it, U-Boot must either
service it (CONFIG_WATCHDOG_AUTOSTART, see that issue) or explicitly disarm it.
Ordering matters — do this issue after the U-Boot-proper watchdog is in.

Prerequisite

This is hard to validate without being able to power-cycle remotely and loop boots,
which is Phase 1 item 5. Characterising the SPL hang honestly (is it 1-in-70 or
1-in-1000?) needs the same thing.

Acceptance

  • SPL arms the watchdog before DRAM init.
  • An induced hang between SPL start and U-Boot proper resets the board.
  • 50 consecutive power cycles still boot normally — the watchdog must not fire on a
    healthy slow boot.
The one real outage on this board happened in **SPL**, and none of the other U-Boot availability work covers it. ## The event 2026-08-27, one occurrence in 68 boots that day, per gap 2 of [41-production-plan.md](41-production-plan.md): DRAM initialised, then two garbage characters instead of `Trying to boot from MMC2`, and the board stayed dead until it was physically power-cycled. Every recovery path that exists — hardware watchdog under Linux, extlinux fallback, rescue SD — assumes the SoC got at least as far as U-Boot proper. None of them apply here. ## Why it is not just a config switch ``` # CONFIG_SPL_WATCHDOG is not set # CONFIG_SPL_DM is not set <- no WDT uclass available in SPL CONFIG_SPL_GPIO=y CONFIG_SPL_SERIAL=y ``` With no driver model in SPL, `SPL_WATCHDOG` has no uclass to drive. Realistic options: 1. Enable `SPL_DM` + `SPL_WDT` and accept the SPL size growth (`CONFIG_SPL_SIZE_LIMIT=0x0` today, and the boot area has ~544 KiB spare before `p1`, so there is room — see the layout in [38-userspace-hardening.md](38-userspace-hardening.md)). 2. A small patch arming R_WDT (`8001000.watchdog`) with direct register writes early in `sunxi_board_init()`, before DRAM init. Both instances are already proven to reset the SoC. Option 2 is smaller and does not perturb the SPL that currently works. Option 1 is more likely to be acceptable upstream. ## Interaction If SPL arms the watchdog and U-Boot proper does not service it, U-Boot must either service it (`CONFIG_WATCHDOG_AUTOSTART`, see that issue) or explicitly disarm it. Ordering matters — do this issue after the U-Boot-proper watchdog is in. ## Prerequisite This is hard to validate without being able to power-cycle remotely and loop boots, which is Phase 1 item 5. Characterising the SPL hang honestly (is it 1-in-70 or 1-in-1000?) needs the same thing. ## Acceptance - SPL arms the watchdog before DRAM init. - An induced hang between SPL start and U-Boot proper resets the board. - 50 consecutive power cycles still boot normally — the watchdog must not fire on a healthy slow boot.
Author
Owner

Done 2026-08-28, with an honest caveat.

Took option 2, as this issue recommended - open-code the sun6i sequence from
drivers/watchdog/sunxi_wdt.c, three register writes, as the first thing
sunxi_board_init() does. That function is SPL-only (CONFIG_XPL_BUILD) and runs before
DRAM init, which is the stage that had already completed when the observed hang happened.

Guarded by a new SPL_SUNXI_EARLY_WATCHDOG, default off, so no other sunxi board is affected.
SPL grew 16 bytes, so option 1's size concern never came up.

The handoff was checked on the board rather than inferred:

stage what happens
SPL arms R_WDT at 0x08001000, 16 s
U-Boot proper adopts the same instance via the WDT uclass (#34)
Linux sunxi-wdt binds both 6000ca0 and 8001000, stops them at probe
systemd re-arms watchdog0 at 15 s
sunxi-wdt 6000ca0.watchdog: Watchdog enabled (timeout=16 sec, nowayout=0)
sunxi-wdt 8001000.watchdog: Watchdog enabled (timeout=16 sec, nowayout=0)

R_WDT was chosen over the main instance because it is in the always-on domain, clocked from
the 24 MHz oscillator, needing no gate or reset released first - which is what makes it usable
that early.

What is not proven

The disassembly shows the three writes landing ahead of sunxi_dram_init(), and the board
boots reliably across repeated cycles. Nothing here demonstrates that it recovers a real SPL
hang
, because there is no way to induce one - the observed event was one boot in 68. This is
a change whose value can only be reasoned about, and that is worth saying rather than filing
under "verified".

**Done 2026-08-28, with an honest caveat.** Took option 2, as this issue recommended - open-code the sun6i sequence from `drivers/watchdog/sunxi_wdt.c`, three register writes, as the first thing `sunxi_board_init()` does. That function is SPL-only (`CONFIG_XPL_BUILD`) and runs before DRAM init, which is the stage that had already completed when the observed hang happened. Guarded by a new `SPL_SUNXI_EARLY_WATCHDOG`, default off, so no other sunxi board is affected. **SPL grew 16 bytes**, so option 1's size concern never came up. The handoff was checked on the board rather than inferred: | stage | what happens | |---|---| | SPL | arms R_WDT at `0x08001000`, 16 s | | U-Boot proper | adopts the same instance via the WDT uclass (#34) | | Linux | `sunxi-wdt` binds **both** `6000ca0` and `8001000`, stops them at probe | | systemd | re-arms `watchdog0` at 15 s | ``` sunxi-wdt 6000ca0.watchdog: Watchdog enabled (timeout=16 sec, nowayout=0) sunxi-wdt 8001000.watchdog: Watchdog enabled (timeout=16 sec, nowayout=0) ``` R_WDT was chosen over the main instance because it is in the always-on domain, clocked from the 24 MHz oscillator, needing no gate or reset released first - which is what makes it usable that early. ### What is not proven The disassembly shows the three writes landing ahead of `sunxi_dram_init()`, and the board boots reliably across repeated cycles. **Nothing here demonstrates that it recovers a real SPL hang**, because there is no way to induce one - the observed event was one boot in 68. This is a change whose value can only be reasoned about, and that is worth saying rather than filing under "verified".
Sign in to join this conversation.
No description provided.