A15 cluster does not boot; cpufreq blocked behind it #11

Closed
opened 2026-08-27 23:22:25 +00:00 by tiagoagueda · 2 comments
Owner

Four of eight cores are unusable. dmesg reports CPU4: failed to boot: -11 through
CPU7, and the CPU runs at a fixed clock with no cpufreq at all.

Verified correct at the moment of failure: PLL_C1=0x82003200, clamp 0xff, gating 0x0e,
PO_RST=0x01, RST_CTRL=0x010111f1, ACINACTM deasserted, L2RSTDISABLE low, SOFT_ENTRY at
0x164, rail asserted, thermals tracking voltage. Excluded: voltage, voltage identity, the
DVFS request, the soft-entry register, NEON reset ordering, CCI.

Remaining hypotheses: something boot0 does for cluster 1 that mainline does not, or a hardware
/ strap difference.

Two tools now exist that did not last time this was attempted:

  • loadable test modules — register pokes in kernel context instead of /dev/mem, which
    hung the SoC twice (see 38-userspace-hardening.md)
  • ramoops — the oops survives the reset instead of being lost

⚠️ Do not write to the CCI slave port. It hung the SoC once already; see 37-arisc-driver.md.

Done when

  • A15 cores come online, or the blocker is identified and documented as hardware
Four of eight cores are unusable. `dmesg` reports `CPU4: failed to boot: -11` through `CPU7`, and the CPU runs at a fixed clock with no cpufreq at all. Verified correct at the moment of failure: `PLL_C1=0x82003200`, clamp `0xff`, gating `0x0e`, `PO_RST=0x01`, `RST_CTRL=0x010111f1`, ACINACTM deasserted, L2RSTDISABLE low, `SOFT_ENTRY` at `0x164`, rail asserted, thermals tracking voltage. Excluded: voltage, voltage identity, the DVFS request, the soft-entry register, NEON reset ordering, CCI. Remaining hypotheses: something boot0 does for cluster 1 that mainline does not, or a hardware / strap difference. Two tools now exist that did not last time this was attempted: - **loadable test modules** — register pokes in kernel context instead of `/dev/mem`, which hung the SoC twice (see `38-userspace-hardening.md`) - **ramoops** — the oops survives the reset instead of being lost ⚠️ Do not write to the CCI slave port. It hung the SoC once already; see `37-arisc-driver.md`. **Done when** - [ ] A15 cores come online, or the blocker is identified and documented as hardware
Author
Owner

Solved 2026-08-28. All eight cores online, and they come up at boot.

processor : 4    CPU part : 0xc0f
processor : 5    CPU part : 0xc0f
processor : 6    CPU part : 0xc0f
processor : 7    CPU part : 0xc0f

The remaining blocker was our own patch. #26
concluded from the vendor BSP that the A15 inverts the power-clamp encoding — 0xff means
on — and mc_smp was changed to match. That is the rev A branch of a function that
opens with

if (sunxi_get_soc_ver() >= SUN9IW1P1_REV_B)

and the excerpt that was read started below the if. This board is rev B
(R_PRCM + 0x190 bit 3 = 1, read at 0x08001590), where both clusters use the A7
convention — which is what mainline already implemented. So the "fix" clamped each A15 core
off at the instant it was released from reset. That is why power, voltage, clocks, resets,
gating and the boot vector all measured correct and the core never fetched, and why the
register dump in note 35 reads clamp0 = 0x000000ff closed right above the claim that the
cluster was powered.

Verified, not merely enumerated: 20 s of load pinned to each A15 core accumulates 2000
jiffies each while the A7s stay idle, thermals rise 11–18 °C, per-core and whole-cluster
hotplug both survive, and

cpu0 (A7)  openssl sha256    35149k
cpu4 (A15) openssl sha256   148127k     4.2x

They come up at boot from an initramfs that carries the AR100 firmware and onlines
cpu4-7 in init-bottom. smp_init() runs before any driver probes, so the four
CPU4: failed to boot: -11 lines at 0.18 s are permanent and expected.

That work exposed a second fault worth more than the first — the coprocessor silently
bringing the cluster up at 408 MHz — tracked separately as #47.

Kernel commits on draco-aw80: d8ecb3974, 4a71fdf3b, aef0384da.
Written up in 45-a15-cluster-solved.md.

**Solved 2026-08-28.** All eight cores online, and they come up at boot. ``` processor : 4 CPU part : 0xc0f processor : 5 CPU part : 0xc0f processor : 6 CPU part : 0xc0f processor : 7 CPU part : 0xc0f ``` The remaining blocker was **our own patch**. [#26](../src/branch/main/26-a15-cluster.md) concluded from the vendor BSP that the A15 inverts the power-clamp encoding — `0xff` means on — and `mc_smp` was changed to match. That is the **rev A** branch of a function that opens with ```c if (sunxi_get_soc_ver() >= SUN9IW1P1_REV_B) ``` and the excerpt that was read started below the `if`. This board is **rev B** (`R_PRCM + 0x190` bit 3 = 1, read at `0x08001590`), where both clusters use the A7 convention — which is what mainline already implemented. So the "fix" clamped each A15 core off at the instant it was released from reset. That is why power, voltage, clocks, resets, gating and the boot vector all measured correct and the core never fetched, and why the register dump in note 35 reads `clamp0 = 0x000000ff closed` right above the claim that the cluster was powered. **Verified, not merely enumerated:** 20 s of load pinned to each A15 core accumulates 2000 jiffies each while the A7s stay idle, thermals rise 11–18 °C, per-core and whole-cluster hotplug both survive, and ``` cpu0 (A7) openssl sha256 35149k cpu4 (A15) openssl sha256 148127k 4.2x ``` **They come up at boot** from an initramfs that carries the AR100 firmware and onlines cpu4-7 in `init-bottom`. `smp_init()` runs before any driver probes, so the four `CPU4: failed to boot: -11` lines at 0.18 s are permanent and expected. That work exposed a second fault worth more than the first — the coprocessor silently bringing the cluster up at 408 MHz — tracked separately as #47. Kernel commits on `draco-aw80`: `d8ecb3974`, `4a71fdf3b`, `aef0384da`. Written up in [45-a15-cluster-solved.md](../src/branch/main/45-a15-cluster-solved.md).
Author
Owner

Reopened in spirit as #53, not here.

This issue's scope - "A15 cluster does not boot" - is genuinely resolved: the cores come up,
execute, and survive hotplug. But an hour after it closed, the cluster turned out to corrupt
memory silently under load
, 14 of 15 tmpfs buffer copies against 0 of 15 with cpu4-7
offline, and to hard-reset the board under heavier load.

That is a different fault from the one tracked here, with its own evidence, so it has its own
issue: #53. Keeping this one closed so the bring-up result stays legible; the new issue is
the one to watch.

Practical note for anyone reading this issue as "eight cores work": they do, and you should
keep cpu4-7 offline
until #53 is resolved.

**Reopened in spirit as #53, not here.** This issue's scope - "A15 cluster does not boot" - is genuinely resolved: the cores come up, execute, and survive hotplug. But an hour after it closed, the cluster turned out to **corrupt memory silently under load**, 14 of 15 tmpfs buffer copies against 0 of 15 with `cpu4-7` offline, and to hard-reset the board under heavier load. That is a different fault from the one tracked here, with its own evidence, so it has its own issue: **#53**. Keeping this one closed so the bring-up result stays legible; the new issue is the one to watch. Practical note for anyone reading this issue as "eight cores work": they do, and **you should keep `cpu4-7` offline** until #53 is resolved.
Sign in to join this conversation.
No description provided.