Board sometimes hangs in SPL and needs a physical power cycle #3
Labels
No labels
blocked-physical
cleanup
hardware
infra
kernel
P1-critical
P2-high
P3-normal
P4-later
reliability
security
upstream
wontfix
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
tiagoagueda/a80#3
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
On 2026-08-27 the board failed to come back from an ordinary
systemctl reboot. Theserial console shows SPL starting, initialising DRAM, then emitting two garbage characters
where the next message should be:
SSis not a prefix of anything SPL prints, so the SoC glitched rather than failing down acode path. It stayed dead — no network, ARP
INCOMPLETEon both interfaces — until it wasphysically power-cycled, after which it booted cleanly and has not repeated.
This board has prior history of non-deterministic boot failures (see the regulator comments
in the board DTS).
Why this is the issue that decides "production". A host that occasionally does not return
from a reboot is not a production host, and with no remote power control a single occurrence
ends the service until someone walks over.
Done when
and count (needs remote power control first)
Partly mitigated 2026-08-28, not closed. #37 arms a watchdog in SPL, which is the
stage this issue is about and which nothing else covered - the hardware watchdog under Linux,
the extlinux fallback and the rescue card all assume the SoC reached U-Boot proper.
Three register writes on R_WDT as the first thing
sunxi_board_init()does, before DRAM init,which is the point the observed hang had already passed. Handoff verified: U-Boot proper
adopts the same instance through the WDT uclass, and Linux stops it in
sunxi_wdt_probe().It stays open because the mitigation is unproven. There is no way to induce an SPL hang on
demand - the observed event was one boot in 68 - so the disassembly and repeated clean boots
are all the evidence there is. See
50-uboot-resilience.md.