Watchdog recovery from a genuine hard lockup is unproven #12
Labels
No labels
blocked-physical
cleanup
hardware
infra
kernel
P1-critical
P2-high
P3-normal
P4-later
reliability
security
upstream
wontfix
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
tiagoagueda/a80#12
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Both
sunxi-wdtinstances were proven to reset the SoC (watchdog0at uptime 184 s,watchdog1at 447 s), and systemd re-armswatchdog0automatically on boot.But both tests stopped the pings deliberately, using magic-close semantics, rather than by
wedging the kernel. The watchdog is hardware and should still fire with the CPU stuck — but a
dead-bus hang like the CCI incident could plausibly take its clock with it, and the SPL hang
was not recovered by the watchdog at all.
Done when
exists, since a failure then costs nothing)
A real data point 2026-08-28, and it is a negative one.
This issue notes the watchdog tests stopped the pings deliberately rather than by wedging the
kernel. Today the board wedged for real, and the watchdog did not fire.
Booting after a reset, it Oopsed in ext4 inode handling at 10.8 s (memory corruption from
#53),
systemd-udevdfailed after a 7.5 minute timeout, and the board sat unreachable onnetwork and serial for roughly ten minutes until it was recovered by hand with the rescue
card.
/dev/watchdog0was open and systemd was still petting it the whole time.So the gap is now concrete and worth adding to the "done when" list: the hardware watchdog
does not cover "PID 1 alive, system unusable". That needs either a systemd watchdog on a
service that actually proves the system works, or an external check - which is #10 and #5.
On the other side of the ledger, two things did improve: #37 arms a watchdog in SPL, and #35
gives U-Boot a 60 s boot-retry that was verified by reproducing the fault.
The concrete gap is closed; the original done-when is not, and still needs #5
What this issue's comment identified
That is now covered.
sunxi-wdtregisters two instances and systemd claims onlywatchdog0.draco-healthdogownswatchdog1— R_WDT, always-on domain, clocked from the 24 MHz oscillator — and pets it only while the board is actually usable:/proc/net/tcprather than forkingss, since this runs forever)That last check is the one the 2026-08-28 incident would have tripped: ext4 wedged while everything else looked alive.
A check that hangs is a failure too, and that falls out for free. If the rootfs blocks, the
touchnever returns, the pet never happens, and the watchdog fires. No timeout logic to get wrong.The definition of healthy is deliberately narrow and matches
boot-mark-good: not "no failed units", because an unrelated failed unit must not reset a board that is otherwise fine and reachable.Proven end to end
Removing both default routes on a running board (it is dual-homed, so removing one proves nothing — my first attempt made exactly that mistake):
Thirty-one seconds from unusable to back and healthy, against ten minutes and a manual rescue-card trip.
Verified across a reboot: unit
active/enabled, auto-starts, holds/dev/watchdog1on fd 3, 0 failed units.Two bugs the testing found, both worth recording
The first version reset the board when you stopped it. The magic close wrote to the device path, but a watchdog is exclusive-open and the descriptor was already held, so the reopen failed with
Device or resource busy, theVwas never written, and stopping the service armed a reset that fired 16 s later. It writes to fd 3 now. Had I not tested the stop path, this would have shipped as "stopping the health daemon reboots the board".A single failed check would have reset the board.
TOLERATE=5was added after realising a DHCP renewal that takes a moment is indistinguishable from a wedge — and on a boot where DHCP is slow, that is a reset loop, far worse than the fault being guarded against. Five consecutive failures at 4 s is 20 s of continuous unhealth before petting stops.An accidental data point in this issue's favour
The magic-close bug above meant the watchdog fired on a genuinely unpetted condition rather than a deliberate magic close — and
watchdog1reset the SoC correctly. Not the hard-lockup test, but one more piece of evidence that the hardware path works.What is still NOT done
The original done-when stands:
I have not wedged the kernel. This issue gates that on remote power control existing, and #5 is still open — if the watchdog failed to fire, the board would need a physical power cycle. That reasoning is still correct and I have not overridden it.
Worth noting the risk is probably lower than it reads: a
local_irq_disable(); while(1);lockup does not stop the watchdog's clock, which is the concern the issue raises about "a dead-bus hang like the CCI incident". Those are different failures. But "probably" is not the standard for a test whose failure mode is a dead board.