A15 cluster does not boot; cpufreq blocked behind it #11
Labels
No labels
blocked-physical
cleanup
hardware
infra
kernel
P1-critical
P2-high
P3-normal
P4-later
reliability
security
upstream
wontfix
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
tiagoagueda/a80#11
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Four of eight cores are unusable.
dmesgreportsCPU4: failed to boot: -11throughCPU7, and the CPU runs at a fixed clock with no cpufreq at all.Verified correct at the moment of failure:
PLL_C1=0x82003200, clamp0xff, gating0x0e,PO_RST=0x01,RST_CTRL=0x010111f1, ACINACTM deasserted, L2RSTDISABLE low,SOFT_ENTRYat0x164, rail asserted, thermals tracking voltage. Excluded: voltage, voltage identity, theDVFS request, the soft-entry register, NEON reset ordering, CCI.
Remaining hypotheses: something boot0 does for cluster 1 that mainline does not, or a hardware
/ strap difference.
Two tools now exist that did not last time this was attempted:
/dev/mem, whichhung the SoC twice (see
38-userspace-hardening.md)⚠️ Do not write to the CCI slave port. It hung the SoC once already; see
37-arisc-driver.md.Done when
Solved 2026-08-28. All eight cores online, and they come up at boot.
The remaining blocker was our own patch. #26
concluded from the vendor BSP that the A15 inverts the power-clamp encoding —
0xffmeanson — and
mc_smpwas changed to match. That is the rev A branch of a function thatopens with
and the excerpt that was read started below the
if. This board is rev B(
R_PRCM + 0x190bit 3 = 1, read at0x08001590), where both clusters use the A7convention — which is what mainline already implemented. So the "fix" clamped each A15 core
off at the instant it was released from reset. That is why power, voltage, clocks, resets,
gating and the boot vector all measured correct and the core never fetched, and why the
register dump in note 35 reads
clamp0 = 0x000000ff closedright above the claim that thecluster was powered.
Verified, not merely enumerated: 20 s of load pinned to each A15 core accumulates 2000
jiffies each while the A7s stay idle, thermals rise 11–18 °C, per-core and whole-cluster
hotplug both survive, and
They come up at boot from an initramfs that carries the AR100 firmware and onlines
cpu4-7 in
init-bottom.smp_init()runs before any driver probes, so the fourCPU4: failed to boot: -11lines at 0.18 s are permanent and expected.That work exposed a second fault worth more than the first — the coprocessor silently
bringing the cluster up at 408 MHz — tracked separately as #47.
Kernel commits on
draco-aw80:d8ecb3974,4a71fdf3b,aef0384da.Written up in 45-a15-cluster-solved.md.
Reopened in spirit as #53, not here.
This issue's scope - "A15 cluster does not boot" - is genuinely resolved: the cores come up,
execute, and survive hotplug. But an hour after it closed, the cluster turned out to corrupt
memory silently under load, 14 of 15 tmpfs buffer copies against 0 of 15 with
cpu4-7offline, and to hard-reset the board under heavier load.
That is a different fault from the one tracked here, with its own evidence, so it has its own
issue: #53. Keeping this one closed so the bring-up result stays legible; the new issue is
the one to watch.
Practical note for anyone reading this issue as "eight cores work": they do, and you should
keep
cpu4-7offline until #53 is resolved.