U-Boot: arm a watchdog in SPL — the stage that actually hung #37
Labels
No labels
blocked-physical
cleanup
hardware
infra
kernel
P1-critical
P2-high
P3-normal
P4-later
reliability
security
upstream
wontfix
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
tiagoagueda/a80#37
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
The one real outage on this board happened in SPL, and none of the other U-Boot
availability work covers it.
The event
2026-08-27, one occurrence in 68 boots that day, per gap 2 of
41-production-plan.md: DRAM initialised, then two garbage
characters instead of
Trying to boot from MMC2, and the board stayed dead until itwas physically power-cycled.
Every recovery path that exists — hardware watchdog under Linux, extlinux fallback,
rescue SD — assumes the SoC got at least as far as U-Boot proper. None of them apply
here.
Why it is not just a config switch
With no driver model in SPL,
SPL_WATCHDOGhas no uclass to drive. Realistic options:SPL_DM+SPL_WDTand accept the SPL size growth(
CONFIG_SPL_SIZE_LIMIT=0x0today, and the boot area has ~544 KiB spare beforep1, so there is room — see the layout in38-userspace-hardening.md).
8001000.watchdog) with direct register writes earlyin
sunxi_board_init(), before DRAM init. Both instances are already proven toreset the SoC.
Option 2 is smaller and does not perturb the SPL that currently works. Option 1 is
more likely to be acceptable upstream.
Interaction
If SPL arms the watchdog and U-Boot proper does not service it, U-Boot must either
service it (
CONFIG_WATCHDOG_AUTOSTART, see that issue) or explicitly disarm it.Ordering matters — do this issue after the U-Boot-proper watchdog is in.
Prerequisite
This is hard to validate without being able to power-cycle remotely and loop boots,
which is Phase 1 item 5. Characterising the SPL hang honestly (is it 1-in-70 or
1-in-1000?) needs the same thing.
Acceptance
healthy slow boot.
Done 2026-08-28, with an honest caveat.
Took option 2, as this issue recommended - open-code the sun6i sequence from
drivers/watchdog/sunxi_wdt.c, three register writes, as the first thingsunxi_board_init()does. That function is SPL-only (CONFIG_XPL_BUILD) and runs beforeDRAM init, which is the stage that had already completed when the observed hang happened.
Guarded by a new
SPL_SUNXI_EARLY_WATCHDOG, default off, so no other sunxi board is affected.SPL grew 16 bytes, so option 1's size concern never came up.
The handoff was checked on the board rather than inferred:
0x08001000, 16 ssunxi-wdtbinds both6000ca0and8001000, stops them at probewatchdog0at 15 sR_WDT was chosen over the main instance because it is in the always-on domain, clocked from
the 24 MHz oscillator, needing no gate or reset released first - which is what makes it usable
that early.
What is not proven
The disassembly shows the three writes landing ahead of
sunxi_dram_init(), and the boardboots reliably across repeated cycles. Nothing here demonstrates that it recovers a real SPL
hang, because there is no way to induce one - the observed event was one boot in 68. This is
a change whose value can only be reasoned about, and that is worth saying rather than filing
under "verified".