AR100 voltage service dies if the coprocessor is started too early #47

Closed
opened 2026-08-28 11:12:26 +00:00 by tiagoagueda · 1 comment
Owner

sunxi_arisc_probe() waits 3 s before its first request_firmware(), and that delay is
load-bearing for a reason nobody understands.

Started ~600 ms earlier — which is exactly what happens when the blob is reachable from an
initramfs instead of the rootfs — the AR100 completes its startup handshake, answers the
first voltage query correctly, and from a second or so later reports 0 mV for every rail
type for the rest of the session
. It never recovers.

That is not cosmetic. With the voltage service dead the coprocessor still accepts the
cluster power-on request, does not raise PLL_C1, and brings the A15 up at whatever
c1cpux was left at:

cpu0 (A7)  openssl sha256   35149k
cpu4 (A15) openssl sha256   33483k     <- the big core, slower than the little one

Every return code said success. It was caught by benchmarking, not by reading the log.

What is measured

probe/voltwatch.ko, alternating the firmware's location across reboots with nothing else
changed:

AR100 starts type 29 (A15 DC-DC) type 14 (AXP806 DCDCA)
firmware in the initramfs 3.8 s 0 mV 0 mV
firmware on the rootfs 4.4 s 1200 mV 600 mV

Ruled out, by measurement

  • our own message traffic — reads 0 mV at 70 s on a boot where nothing touched the
    coprocessor after probe
  • settling time — 40 one-second retries, every one 0 mV
  • a corrupt copy — the blob in the initramfs md5s identical to the one on disk
  • the initramfs itself — an initramfs with the firmware removed works fine, so the
    variable is when the AR100 starts, not whether an initramfs exists
  • RSB / AXP806 ordering — sunxi-rsb and axp20x-rsb probe before the beacon either
    way

So what its own init races with in that window is still unknown.

Why the current state is not good enough

A timer is a fragile thing to depend on, and the failure is silent by nature. The driver
should verify its voltage service a second after the handshake and restart the firmware
if it is dead
— it can, since it already reads the CPU-B rail at probe. Restarting means
asserting the AR100 reset, rewriting SRAM A2 and redoing the beacon handshake, all of which
the driver already does once.

Until then the readback guard in mc_smp (4a71fdf3b) is what stops a crippled cluster
coming up: an implausible rail reading returns -EAGAIN and the board boots on four cores
rather than eight slow ones.

⚠️ If the 3 s delay is ever shortened, the A15 will come up at 408 MHz and nothing will
say so.

Done when

  • the driver detects a dead voltage service and recovers without a fixed delay, or the
    race is identified and avoided properly

See 45-a15-cluster-solved.md.

`sunxi_arisc_probe()` waits 3 s before its first `request_firmware()`, and that delay is load-bearing for a reason nobody understands. Started ~600 ms earlier — which is exactly what happens when the blob is reachable from an initramfs instead of the rootfs — the AR100 completes its startup handshake, answers the first voltage query correctly, and from a second or so later **reports 0 mV for every rail type for the rest of the session**. It never recovers. That is not cosmetic. With the voltage service dead the coprocessor still **accepts** the cluster power-on request, does not raise `PLL_C1`, and brings the A15 up at whatever `c1cpux` was left at: ``` cpu0 (A7) openssl sha256 35149k cpu4 (A15) openssl sha256 33483k <- the big core, slower than the little one ``` Every return code said success. It was caught by benchmarking, not by reading the log. ## What is measured `probe/voltwatch.ko`, alternating the firmware's location across reboots with nothing else changed: | | AR100 starts | type 29 (A15 DC-DC) | type 14 (AXP806 DCDCA) | |---|---|---|---| | firmware in the initramfs | **3.8 s** | 0 mV | 0 mV | | firmware on the rootfs | **4.4 s** | 1200 mV | 600 mV | ## Ruled out, by measurement - **our own message traffic** — reads 0 mV at 70 s on a boot where nothing touched the coprocessor after probe - **settling time** — 40 one-second retries, every one 0 mV - **a corrupt copy** — the blob in the initramfs md5s identical to the one on disk - **the initramfs itself** — an initramfs with the firmware *removed* works fine, so the variable is when the AR100 starts, not whether an initramfs exists - **RSB / AXP806 ordering** — `sunxi-rsb` and `axp20x-rsb` probe before the beacon either way So what its own init races with in that window is still unknown. ## Why the current state is not good enough A timer is a fragile thing to depend on, and the failure is silent by nature. The driver should **verify its voltage service a second after the handshake and restart the firmware if it is dead** — it can, since it already reads the CPU-B rail at probe. Restarting means asserting the AR100 reset, rewriting SRAM A2 and redoing the beacon handshake, all of which the driver already does once. Until then the readback guard in `mc_smp` (`4a71fdf3b`) is what stops a crippled cluster coming up: an implausible rail reading returns `-EAGAIN` and the board boots on four cores rather than eight slow ones. ⚠️ **If the 3 s delay is ever shortened, the A15 will come up at 408 MHz and nothing will say so.** ## Done when - [x] the driver detects a dead voltage service and recovers without a fixed delay, or the race is identified and avoided properly See [45-a15-cluster-solved.md](../src/branch/main/45-a15-cluster-solved.md).
Author
Owner

Done — the driver now detects the dead voltage service and restarts the firmware

sunxi_arisc_start_verified() re-reads the CPU-B rail a second after the handshake and, if it is implausible, restarts the firmware — the same reset/load/release/handshake used to start it — up to three times. If none works the coprocessor is deliberately left not ready, so sunxi_arisc_get_volt() returns -EAGAIN and mc_smp refuses the cluster rather than bringing it up at 408 MHz with every return code reporting success. Four good cores beat eight bad ones.

Kernel 2e487726b2ba, exported to patches/linux-sunxi-arisc/.

The race is reproducible on demand now, which it never was

A start_delay_ms module parameter is what made this testable. At 2500 the beacon lands at 3.76 s and it fails every time — and is then recovered:

[3.757441] startup beacon on ch2 (handle 0x27000)
[3.761152] coprocessor up, loopback ok; CPUB rail reads 1200 mV
[4.829143] voltage service dead after start 1 (CPUB reads 0 mV); restarting firmware
[5.119272] startup beacon on ch2 (handle 0x27000)
[6.179819] voltage service healthy after 2 starts (CPUB 1200 mV)

Every element of the original description reproduced: beacon in the hazard window, first query correct, 0 mV a second later. One restart was enough, and voltwatch reads type29 = 1200 mV, type14 = 600 mV ninety seconds later, so the recovery is durable rather than momentary.

It no longer reproduces by itself — and that is luck, not a fix

The trigger recorded here was the firmware being reachable from an initramfs, which put the beacon at 3.8 s. Today's initramfs work did exactly that, so this should have been failing. Measured across three boots:

beacon at probe at +25 s
boot 1 4.327559 s 1200 mV 1200 mV
boot 2 4.327580 s 1200 mV 1200 mV
boot 3 4.327556 s 1200 mV 1200 mV

Healthy every time, with sub-millisecond variance. Today's kernel is simply bigger — zram, netfilter, IPv6 — so the start drifted past the hazard on its own. The margin is now set by something as arbitrary as kernel size, which is exactly why the driver should not depend on it.

A worse failure mode this issue did not know about

Building the reproduction found something more serious. At start_delay_ms=500 the beacon lands at 1.76 s and the board does not boot at all:

1c0f000.mmc:   sunxi-mmc: Error applying setting, reverse things back
830000.ethernet: sun7i-dwmac: Error applying setting, reverse things back
usb1-vbus:     reg-fixed-voltage: can't get GPIO
wifi-pwrseq:   pwrseq_simple: reset GPIOs not ready

Every driver needing a GPIO fails, the MMC controllers never probe, there is no root device, and it drops to an initramfs shell. Notably the voltage service itself read 1200 mV on that boot and the recovery never fired — this is a different failure, not a worse version of the same one.

So the hazard is graded, not binary:

beacon outcome
~1.76 s pinctrl/GPIO subsystem dies; board does not boot
~3.8 s voltage service dies silently; A15 would come up at 408 MHz
~4.33 s healthy

Starting the coprocessor early enough takes the pinctrl subsystem with it, not just its own voltage service. That widens the question this issue was originally asking.

Recovery, and the first real use of the serial console

Getting back from the 1.76 s hang used the path that has existed all project but never been needed in anger: interrupt U-Boot's autoboot over /dev/ttyUSB0, setenv bootretry -1, sysboot mmc 1:1 any ${scriptaddr} /boot/extlinux/extlinux.conf, pick the good entry. No physical access, no rescue card swap. Worth knowing it works, given #3 and #5.

Also fixed on the way

deploy-kernel.sh deleted the module tree it had just installed. A uImage stores its name in a 32-byte field, so with the Linux- prefix only the first 26 characters of the version survive — a -dirty suffix is invisible to strings, and the prune could not match it. It now compares that prefix and never prunes the version just installed.

Done when

  • the driver detects a dead voltage service and recovers without a fixed delay

The 3 s delay remains as a preference rather than a dependency: crossing the race is now recoverable instead of silent. The underlying race is still unidentified, and the 1.76 s finding above suggests it is worth reopening as a pinctrl-ordering question rather than an AR100 one.

## Done — the driver now detects the dead voltage service and restarts the firmware `sunxi_arisc_start_verified()` re-reads the CPU-B rail a second after the handshake and, if it is implausible, restarts the firmware — the same reset/load/release/handshake used to start it — up to three times. If none works the coprocessor is deliberately left **not ready**, so `sunxi_arisc_get_volt()` returns `-EAGAIN` and `mc_smp` refuses the cluster rather than bringing it up at 408 MHz with every return code reporting success. Four good cores beat eight bad ones. Kernel `2e487726b2ba`, exported to `patches/linux-sunxi-arisc/`. ### The race is reproducible on demand now, which it never was A `start_delay_ms` module parameter is what made this testable. At 2500 the beacon lands at 3.76 s and it fails **every time** — and is then recovered: ``` [3.757441] startup beacon on ch2 (handle 0x27000) [3.761152] coprocessor up, loopback ok; CPUB rail reads 1200 mV [4.829143] voltage service dead after start 1 (CPUB reads 0 mV); restarting firmware [5.119272] startup beacon on ch2 (handle 0x27000) [6.179819] voltage service healthy after 2 starts (CPUB 1200 mV) ``` Every element of the original description reproduced: beacon in the hazard window, first query correct, 0 mV a second later. One restart was enough, and `voltwatch` reads `type29 = 1200 mV, type14 = 600 mV` ninety seconds later, so the recovery is durable rather than momentary. ### It no longer reproduces by itself — and that is luck, not a fix The trigger recorded here was the firmware being reachable from an initramfs, which put the beacon at 3.8 s. **Today's initramfs work did exactly that**, so this should have been failing. Measured across three boots: | | beacon | at probe | at +25 s | |---|---|---|---| | boot 1 | 4.327559 s | 1200 mV | 1200 mV | | boot 2 | 4.327580 s | 1200 mV | 1200 mV | | boot 3 | 4.327556 s | 1200 mV | 1200 mV | Healthy every time, with sub-millisecond variance. Today's kernel is simply bigger — zram, netfilter, IPv6 — so the start drifted past the hazard on its own. The margin is now set by something as arbitrary as kernel size, which is exactly why the driver should not depend on it. ### A worse failure mode this issue did not know about Building the reproduction found something more serious. At `start_delay_ms=500` the beacon lands at **1.76 s and the board does not boot at all**: ``` 1c0f000.mmc: sunxi-mmc: Error applying setting, reverse things back 830000.ethernet: sun7i-dwmac: Error applying setting, reverse things back usb1-vbus: reg-fixed-voltage: can't get GPIO wifi-pwrseq: pwrseq_simple: reset GPIOs not ready ``` Every driver needing a GPIO fails, the MMC controllers never probe, there is no root device, and it drops to an initramfs shell. Notably the voltage service itself read **1200 mV** on that boot and the recovery never fired — this is a different failure, not a worse version of the same one. So the hazard is graded, not binary: | beacon | outcome | |---|---| | ~1.76 s | pinctrl/GPIO subsystem dies; board does not boot | | ~3.8 s | voltage service dies silently; A15 would come up at 408 MHz | | ~4.33 s | healthy | **Starting the coprocessor early enough takes the pinctrl subsystem with it, not just its own voltage service.** That widens the question this issue was originally asking. ### Recovery, and the first real use of the serial console Getting back from the 1.76 s hang used the path that has existed all project but never been needed in anger: interrupt U-Boot's autoboot over `/dev/ttyUSB0`, `setenv bootretry -1`, `sysboot mmc 1:1 any ${scriptaddr} /boot/extlinux/extlinux.conf`, pick the good entry. No physical access, no rescue card swap. Worth knowing it works, given #3 and #5. ### Also fixed on the way `deploy-kernel.sh` deleted the module tree it had just installed. A uImage stores its name in a 32-byte field, so with the `Linux-` prefix only the first 26 characters of the version survive — a `-dirty` suffix is invisible to `strings`, and the prune could not match it. It now compares that prefix and never prunes the version just installed. ### Done when - [x] the driver detects a dead voltage service and recovers without a fixed delay The 3 s delay remains as a preference rather than a dependency: crossing the race is now recoverable instead of silent. The underlying race is still unidentified, and the 1.76 s finding above suggests it is worth reopening as a pinctrl-ordering question rather than an AR100 one.
Sign in to join this conversation.
No description provided.