Board sometimes hangs in SPL and needs a physical power cycle #3

Open
opened 2026-08-27 23:22:23 +00:00 by tiagoagueda · 1 comment
Owner

On 2026-08-27 the board failed to come back from an ordinary systemctl reboot. The
serial console shows SPL starting, initialising DRAM, then emitting two garbage characters
where the next message should be:

DRAM: 3584 MiB
Trying to boot from MMC2      <- all 68 other boots that day
DRAM: 3584 MiB
SS                            <- this one, then silence

SS is not a prefix of anything SPL prints, so the SoC glitched rather than failing down a
code path. It stayed dead — no network, ARP INCOMPLETE on both interfaces — until it was
physically power-cycled, after which it booted cleanly and has not repeated.

This board has prior history of non-deterministic boot failures (see the regulator comments
in the board DTS).

Why this is the issue that decides "production". A host that occasionally does not return
from a reboot is not a production host, and with no remote power control a single occurrence
ends the service until someone walks over.

Done when

  • the failure rate is measured, not guessed — loop power-cycles against the serial capture
    and count (needs remote power control first)
  • a cause is identified, or the rate is shown low enough to accept and is written down
On 2026-08-27 the board failed to come back from an ordinary `systemctl reboot`. The serial console shows SPL starting, initialising DRAM, then emitting two garbage characters where the next message should be: ``` DRAM: 3584 MiB Trying to boot from MMC2 <- all 68 other boots that day DRAM: 3584 MiB SS <- this one, then silence ``` `SS` is not a prefix of anything SPL prints, so the SoC glitched rather than failing down a code path. It stayed dead — no network, ARP `INCOMPLETE` on both interfaces — until it was physically power-cycled, after which it booted cleanly and has not repeated. This board has prior history of non-deterministic boot failures (see the regulator comments in the board DTS). **Why this is the issue that decides "production".** A host that occasionally does not return from a reboot is not a production host, and with no remote power control a single occurrence ends the service until someone walks over. **Done when** - [ ] the failure rate is measured, not guessed — loop power-cycles against the serial capture and count (needs remote power control first) - [ ] a cause is identified, or the rate is shown low enough to accept and is written down
Author
Owner

Partly mitigated 2026-08-28, not closed. #37 arms a watchdog in SPL, which is the
stage this issue is about and which nothing else covered - the hardware watchdog under Linux,
the extlinux fallback and the rescue card all assume the SoC reached U-Boot proper.

Three register writes on R_WDT as the first thing sunxi_board_init() does, before DRAM init,
which is the point the observed hang had already passed. Handoff verified: U-Boot proper
adopts the same instance through the WDT uclass, and Linux stops it in sunxi_wdt_probe().

It stays open because the mitigation is unproven. There is no way to induce an SPL hang on
demand - the observed event was one boot in 68 - so the disassembly and repeated clean boots
are all the evidence there is. See 50-uboot-resilience.md.

**Partly mitigated 2026-08-28, not closed.** #37 arms a watchdog in SPL, which is the stage this issue is about and which nothing else covered - the hardware watchdog under Linux, the extlinux fallback and the rescue card all assume the SoC reached U-Boot proper. Three register writes on R_WDT as the first thing `sunxi_board_init()` does, before DRAM init, which is the point the observed hang had already passed. Handoff verified: U-Boot proper adopts the same instance through the WDT uclass, and Linux stops it in `sunxi_wdt_probe()`. **It stays open because the mitigation is unproven.** There is no way to induce an SPL hang on demand - the observed event was one boot in 68 - so the disassembly and repeated clean boots are all the evidence there is. See `50-uboot-resilience.md`.
Sign in to join this conversation.
No description provided.