Watchdog recovery from a genuine hard lockup is unproven #12

Open
opened 2026-08-27 23:22:25 +00:00 by tiagoagueda · 2 comments
Owner

Both sunxi-wdt instances were proven to reset the SoC (watchdog0 at uptime 184 s,
watchdog1 at 447 s), and systemd re-arms watchdog0 automatically on boot.

But both tests stopped the pings deliberately, using magic-close semantics, rather than by
wedging the kernel. The watchdog is hardware and should still fire with the CPU stuck — but a
dead-bus hang like the CCI incident could plausibly take its clock with it, and the SPL hang
was not recovered by the watchdog at all.

Done when

  • recovery is tested against a real lockup (safe to attempt once remote power control
    exists, since a failure then costs nothing)
Both `sunxi-wdt` instances were proven to reset the SoC (`watchdog0` at uptime 184 s, `watchdog1` at 447 s), and systemd re-arms `watchdog0` automatically on boot. But both tests stopped the pings *deliberately*, using magic-close semantics, rather than by wedging the kernel. The watchdog is hardware and should still fire with the CPU stuck — but a dead-bus hang like the CCI incident could plausibly take its clock with it, and the SPL hang was not recovered by the watchdog at all. **Done when** - [ ] recovery is tested against a real lockup (safe to attempt once remote power control exists, since a failure then costs nothing)
Author
Owner

A real data point 2026-08-28, and it is a negative one.

This issue notes the watchdog tests stopped the pings deliberately rather than by wedging the
kernel. Today the board wedged for real, and the watchdog did not fire.

Booting after a reset, it Oopsed in ext4 inode handling at 10.8 s (memory corruption from
#53), systemd-udevd failed after a 7.5 minute timeout, and the board sat unreachable on
network and serial for roughly ten minutes until it was recovered by hand with the rescue
card. /dev/watchdog0 was open and systemd was still petting it the whole time.

So the gap is now concrete and worth adding to the "done when" list: the hardware watchdog
does not cover "PID 1 alive, system unusable".
That needs either a systemd watchdog on a
service that actually proves the system works, or an external check - which is #10 and #5.

On the other side of the ledger, two things did improve: #37 arms a watchdog in SPL, and #35
gives U-Boot a 60 s boot-retry that was verified by reproducing the fault.

**A real data point 2026-08-28, and it is a negative one.** This issue notes the watchdog tests stopped the pings deliberately rather than by wedging the kernel. Today the board wedged for real, and **the watchdog did not fire.** Booting after a reset, it Oopsed in ext4 inode handling at 10.8 s (memory corruption from #53), `systemd-udevd` failed after a 7.5 minute timeout, and the board sat unreachable on network *and* serial for roughly ten minutes until it was recovered by hand with the rescue card. `/dev/watchdog0` was open and systemd was still petting it the whole time. So the gap is now concrete and worth adding to the "done when" list: **the hardware watchdog does not cover "PID 1 alive, system unusable".** That needs either a systemd watchdog on a service that actually proves the system works, or an external check - which is #10 and #5. On the other side of the ledger, two things did improve: #37 arms a watchdog in SPL, and #35 gives U-Boot a 60 s boot-retry that was verified by reproducing the fault.
Author
Owner

The concrete gap is closed; the original done-when is not, and still needs #5

What this issue's comment identified

the hardware watchdog does not cover "PID 1 alive, system unusable"

That is now covered. sunxi-wdt registers two instances and systemd claims only watchdog0. draco-healthdog owns watchdog1 — R_WDT, always-on domain, clocked from the 24 MHz oscillator — and pets it only while the board is actually usable:

  • a default route exists
  • sshd is listening on 22 (read from /proc/net/tcp rather than forking ss, since this runs forever)
  • the root filesystem still accepts a write

That last check is the one the 2026-08-28 incident would have tripped: ext4 wedged while everything else looked alive.

A check that hangs is a failure too, and that falls out for free. If the rootfs blocks, the touch never returns, the pet never happens, and the watchdog fires. No timeout logic to get wrong.

The definition of healthy is deliberately narrow and matches boot-mark-good: not "no failed units", because an unrelated failed unit must not reset a board that is otherwise fine and reachable.

Proven end to end

Removing both default routes on a running board (it is dual-homed, so removing one proves nothing — my first attempt made exactly that mistake):

t+4s   health check failed (1/5) - still petting
t+8s   health check failed (2/5) - still petting
...
t+20s  health check FAILED 5 times - NOT petting /dev/watchdog1; reset expected
       board down ~16s later - watchdog fired
       back after ~15s, uptime 24 s, routes restored by DHCP

Thirty-one seconds from unusable to back and healthy, against ten minutes and a manual rescue-card trip.

Verified across a reboot: unit active/enabled, auto-starts, holds /dev/watchdog1 on fd 3, 0 failed units.

Two bugs the testing found, both worth recording

The first version reset the board when you stopped it. The magic close wrote to the device path, but a watchdog is exclusive-open and the descriptor was already held, so the reopen failed with Device or resource busy, the V was never written, and stopping the service armed a reset that fired 16 s later. It writes to fd 3 now. Had I not tested the stop path, this would have shipped as "stopping the health daemon reboots the board".

A single failed check would have reset the board. TOLERATE=5 was added after realising a DHCP renewal that takes a moment is indistinguishable from a wedge — and on a boot where DHCP is slow, that is a reset loop, far worse than the fault being guarded against. Five consecutive failures at 4 s is 20 s of continuous unhealth before petting stops.

An accidental data point in this issue's favour

The magic-close bug above meant the watchdog fired on a genuinely unpetted condition rather than a deliberate magic close — and watchdog1 reset the SoC correctly. Not the hard-lockup test, but one more piece of evidence that the hardware path works.

What is still NOT done

The original done-when stands:

  • recovery is tested against a real lockup

I have not wedged the kernel. This issue gates that on remote power control existing, and #5 is still open — if the watchdog failed to fire, the board would need a physical power cycle. That reasoning is still correct and I have not overridden it.

Worth noting the risk is probably lower than it reads: a local_irq_disable(); while(1); lockup does not stop the watchdog's clock, which is the concern the issue raises about "a dead-bus hang like the CCI incident". Those are different failures. But "probably" is not the standard for a test whose failure mode is a dead board.

## The concrete gap is closed; the original done-when is not, and still needs #5 ### What this issue's comment identified > the hardware watchdog does not cover "PID 1 alive, system unusable" That is now covered. `sunxi-wdt` registers **two** instances and systemd claims only `watchdog0`. `draco-healthdog` owns `watchdog1` — R_WDT, always-on domain, clocked from the 24 MHz oscillator — and pets it **only while the board is actually usable**: - a default route exists - sshd is listening on 22 (read from `/proc/net/tcp` rather than forking `ss`, since this runs forever) - the root filesystem still accepts a write That last check is the one the 2026-08-28 incident would have tripped: ext4 wedged while everything else looked alive. **A check that hangs is a failure too, and that falls out for free.** If the rootfs blocks, the `touch` never returns, the pet never happens, and the watchdog fires. No timeout logic to get wrong. The definition of healthy is deliberately narrow and matches `boot-mark-good`: *not* "no failed units", because an unrelated failed unit must not reset a board that is otherwise fine and reachable. ### Proven end to end Removing both default routes on a running board (it is dual-homed, so removing one proves nothing — my first attempt made exactly that mistake): ``` t+4s health check failed (1/5) - still petting t+8s health check failed (2/5) - still petting ... t+20s health check FAILED 5 times - NOT petting /dev/watchdog1; reset expected board down ~16s later - watchdog fired back after ~15s, uptime 24 s, routes restored by DHCP ``` **Thirty-one seconds from unusable to back and healthy**, against ten minutes and a manual rescue-card trip. Verified across a reboot: unit `active`/`enabled`, auto-starts, holds `/dev/watchdog1` on fd 3, 0 failed units. ### Two bugs the testing found, both worth recording **The first version reset the board when you stopped it.** The magic close wrote to the device *path*, but a watchdog is exclusive-open and the descriptor was already held, so the reopen failed with `Device or resource busy`, the `V` was never written, and stopping the service armed a reset that fired 16 s later. It writes to fd 3 now. Had I not tested the stop path, this would have shipped as "stopping the health daemon reboots the board". **A single failed check would have reset the board.** `TOLERATE=5` was added after realising a DHCP renewal that takes a moment is indistinguishable from a wedge — and on a boot where DHCP is slow, that is a *reset loop*, far worse than the fault being guarded against. Five consecutive failures at 4 s is 20 s of continuous unhealth before petting stops. ### An accidental data point in this issue's favour The magic-close bug above meant the watchdog fired on a genuinely unpetted condition rather than a deliberate magic close — and `watchdog1` reset the SoC correctly. Not the hard-lockup test, but one more piece of evidence that the hardware path works. ### What is still NOT done The original done-when stands: - [ ] recovery is tested against a real lockup I have **not** wedged the kernel. This issue gates that on remote power control existing, and **#5 is still open** — if the watchdog failed to fire, the board would need a physical power cycle. That reasoning is still correct and I have not overridden it. Worth noting the risk is probably lower than it reads: a `local_irq_disable(); while(1);` lockup does not stop the watchdog's clock, which is the concern the issue raises about "a dead-bus hang like the CCI incident". Those are different failures. But "probably" is not the standard for a test whose failure mode is a dead board.
Sign in to join this conversation.
No description provided.