AR100 voltage service dies if the coprocessor is started too early #47
Labels
No labels
blocked-physical
cleanup
hardware
infra
kernel
P1-critical
P2-high
P3-normal
P4-later
reliability
security
upstream
wontfix
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
tiagoagueda/a80#47
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
sunxi_arisc_probe()waits 3 s before its firstrequest_firmware(), and that delay isload-bearing for a reason nobody understands.
Started ~600 ms earlier — which is exactly what happens when the blob is reachable from an
initramfs instead of the rootfs — the AR100 completes its startup handshake, answers the
first voltage query correctly, and from a second or so later reports 0 mV for every rail
type for the rest of the session. It never recovers.
That is not cosmetic. With the voltage service dead the coprocessor still accepts the
cluster power-on request, does not raise
PLL_C1, and brings the A15 up at whateverc1cpuxwas left at:Every return code said success. It was caught by benchmarking, not by reading the log.
What is measured
probe/voltwatch.ko, alternating the firmware's location across reboots with nothing elsechanged:
Ruled out, by measurement
coprocessor after probe
variable is when the AR100 starts, not whether an initramfs exists
sunxi-rsbandaxp20x-rsbprobe before the beacon eitherway
So what its own init races with in that window is still unknown.
Why the current state is not good enough
A timer is a fragile thing to depend on, and the failure is silent by nature. The driver
should verify its voltage service a second after the handshake and restart the firmware
if it is dead — it can, since it already reads the CPU-B rail at probe. Restarting means
asserting the AR100 reset, rewriting SRAM A2 and redoing the beacon handshake, all of which
the driver already does once.
Until then the readback guard in
mc_smp(4a71fdf3b) is what stops a crippled clustercoming up: an implausible rail reading returns
-EAGAINand the board boots on four coresrather than eight slow ones.
⚠️ If the 3 s delay is ever shortened, the A15 will come up at 408 MHz and nothing will
say so.
Done when
race is identified and avoided properly
See 45-a15-cluster-solved.md.
Done — the driver now detects the dead voltage service and restarts the firmware
sunxi_arisc_start_verified()re-reads the CPU-B rail a second after the handshake and, if it is implausible, restarts the firmware — the same reset/load/release/handshake used to start it — up to three times. If none works the coprocessor is deliberately left not ready, sosunxi_arisc_get_volt()returns-EAGAINandmc_smprefuses the cluster rather than bringing it up at 408 MHz with every return code reporting success. Four good cores beat eight bad ones.Kernel
2e487726b2ba, exported topatches/linux-sunxi-arisc/.The race is reproducible on demand now, which it never was
A
start_delay_msmodule parameter is what made this testable. At 2500 the beacon lands at 3.76 s and it fails every time — and is then recovered:Every element of the original description reproduced: beacon in the hazard window, first query correct, 0 mV a second later. One restart was enough, and
voltwatchreadstype29 = 1200 mV, type14 = 600 mVninety seconds later, so the recovery is durable rather than momentary.It no longer reproduces by itself — and that is luck, not a fix
The trigger recorded here was the firmware being reachable from an initramfs, which put the beacon at 3.8 s. Today's initramfs work did exactly that, so this should have been failing. Measured across three boots:
Healthy every time, with sub-millisecond variance. Today's kernel is simply bigger — zram, netfilter, IPv6 — so the start drifted past the hazard on its own. The margin is now set by something as arbitrary as kernel size, which is exactly why the driver should not depend on it.
A worse failure mode this issue did not know about
Building the reproduction found something more serious. At
start_delay_ms=500the beacon lands at 1.76 s and the board does not boot at all:Every driver needing a GPIO fails, the MMC controllers never probe, there is no root device, and it drops to an initramfs shell. Notably the voltage service itself read 1200 mV on that boot and the recovery never fired — this is a different failure, not a worse version of the same one.
So the hazard is graded, not binary:
Starting the coprocessor early enough takes the pinctrl subsystem with it, not just its own voltage service. That widens the question this issue was originally asking.
Recovery, and the first real use of the serial console
Getting back from the 1.76 s hang used the path that has existed all project but never been needed in anger: interrupt U-Boot's autoboot over
/dev/ttyUSB0,setenv bootretry -1,sysboot mmc 1:1 any ${scriptaddr} /boot/extlinux/extlinux.conf, pick the good entry. No physical access, no rescue card swap. Worth knowing it works, given #3 and #5.Also fixed on the way
deploy-kernel.shdeleted the module tree it had just installed. A uImage stores its name in a 32-byte field, so with theLinux-prefix only the first 26 characters of the version survive — a-dirtysuffix is invisible tostrings, and the prune could not match it. It now compares that prefix and never prunes the version just installed.Done when
The 3 s delay remains as a preference rather than a dependency: crossing the race is now recoverable instead of silent. The underlying race is still unidentified, and the 1.76 s finding above suggests it is worth reopening as a pinctrl-ordering question rather than an AR100 one.