A15 cluster: board hard-resets under any sustained A15 load, at ~80 C #53

Open
opened 2026-08-28 16:38:22 +00:00 by tiagoagueda · 5 comments
Owner

With cpu4-7 online the board silently corrupts memory under load, and hard-resets
under more of it.
The A15 cores come up and execute - #11 is genuinely resolved - but the
system is not usable with them running.

Found 2026-08-28, by accident, an hour after #11 was closed.

Measured

Network-free: a 32 MiB buffer copied inside tmpfs and checked, under two busy loops. No disk,
no network.

A15 cluster online    14 of 15 copies corrupt
A15 cluster offline    0 of 15

The same split shows in transfers - 10/10 clean with cpu4-7 down, 6/10 with them up.

Under heavier CPU load the board hard-resets, twice during that session, with
/sys/fs/pstore empty both times. It is not an oops; the machine goes away.

It is one bit

Every difference, in every sample:

$ cmp -l good bad
2650526 311 313
7787687 135 137

Twelve flips across five corrupted 8 MiB transfers, and all twelve are bit 1, always 0 to 1,
never any other bit and never the other direction.
Software does not have a favourite bit.

How it surfaced, and how close it came to not surfacing

A routine kernel deploy failed its checksum:

>> uImage f0378ab482012fa38f953780dea2eb0c
uImage md5 MISMATCH

deploy-kernel.sh md5-checks both ends. This is the first time that check has ever caught
anything, and without it the corrupt kernel would have been installed and booted.

Ruled out

thought actually
truncated transfer sizes matched to the byte
eMMC write corruption destination was tmpfs; 40 MiB written to eMMC locally, clean
network corruption SSH is integrity-checked; corruption in flight aborts the connection
A15 cores compute wrong 8 cores x 20 x 8 MiB of taskset-pinned md5sum: 160/160 correct

It needs load - an idle board copies 40 MiB without a flip. Two busy loops make it
near-certain.

Cause not established

The obvious suspect is power: the rail work in 35-a15-rail-enable.md brought up a cluster
drawing far more than this board had run before, and a sagging supply makes DRAM marginal. A
single data bit failing first is what a marginal DRAM interface looks like. That is a story
that fits the evidence, not a measurement.

Worth settling, cheapest first:

  1. Read the PMIC rails under A15 load - is anything drooping?
  2. Compare against the vendor kernel, which runs eight cores on this same board. Clean there
    would put the fault in our configuration, not the silicon.
  3. Check whether U-Boot's DRAM training ran with the A15 cluster powered down (#28,
    22-dram-4gb.md).
  4. Separate "cluster powered" from "cluster executing" - both tests here let the scheduler use
    the A15 cores.

Consequences until it is understood

  • Keep cpu4-7 offline for development. Transfers, builds and deploys are reliable with
    them down and are not with them up.
  • The A15 half of the mc_smp series must not be offered upstream - see #27 and
    patches/linux-mc-smp-a15/. The cpu_table fix in that series is an independent mainline
    bug and is unaffected.
  • Any measurement taken with eight cores online since #11 closed is suspect.

The part worth keeping

Closing #11 on "nproc returns 8 and the cores execute" was true and insufficient. A cluster
that corrupts memory under load is worse than one that does not come up
, because the second
failure is loud. The first time anything actually leaned on those cores, they corrupted data,
and the only thing standing between that and a week of unexplainable results was a checksum in
a deploy script.

See 48-a15-memory-corruption.md.

**With `cpu4-7` online the board silently corrupts memory under load, and hard-resets under more of it.** The A15 cores come up and execute - #11 is genuinely resolved - but the system is not usable with them running. Found 2026-08-28, by accident, an hour after #11 was closed. ## Measured Network-free: a 32 MiB buffer copied inside tmpfs and checked, under two busy loops. No disk, no network. ``` A15 cluster online 14 of 15 copies corrupt A15 cluster offline 0 of 15 ``` The same split shows in transfers - 10/10 clean with `cpu4-7` down, 6/10 with them up. Under heavier CPU load the board **hard-resets**, twice during that session, with `/sys/fs/pstore` empty both times. It is not an oops; the machine goes away. ## It is one bit Every difference, in every sample: ``` $ cmp -l good bad 2650526 311 313 7787687 135 137 ``` Twelve flips across five corrupted 8 MiB transfers, and **all twelve are bit 1, always 0 to 1, never any other bit and never the other direction.** Software does not have a favourite bit. ## How it surfaced, and how close it came to not surfacing A routine kernel deploy failed its checksum: ``` >> uImage f0378ab482012fa38f953780dea2eb0c uImage md5 MISMATCH ``` `deploy-kernel.sh` md5-checks both ends. This is the first time that check has ever caught anything, and **without it the corrupt kernel would have been installed and booted.** ## Ruled out | thought | actually | |---|---| | truncated transfer | sizes matched to the byte | | eMMC write corruption | destination was tmpfs; 40 MiB written to eMMC locally, clean | | network corruption | SSH is integrity-checked; corruption in flight aborts the connection | | A15 cores compute wrong | 8 cores x 20 x 8 MiB of `taskset`-pinned `md5sum`: 160/160 correct | It needs load - an idle board copies 40 MiB without a flip. Two busy loops make it near-certain. ## Cause not established The obvious suspect is power: the rail work in `35-a15-rail-enable.md` brought up a cluster drawing far more than this board had run before, and a sagging supply makes DRAM marginal. A single data bit failing first is what a marginal DRAM interface looks like. **That is a story that fits the evidence, not a measurement.** Worth settling, cheapest first: 1. Read the PMIC rails under A15 load - is anything drooping? 2. Compare against the vendor kernel, which runs eight cores on this same board. Clean there would put the fault in our configuration, not the silicon. 3. Check whether U-Boot's DRAM training ran with the A15 cluster powered down (#28, `22-dram-4gb.md`). 4. Separate "cluster powered" from "cluster executing" - both tests here let the scheduler use the A15 cores. ## Consequences until it is understood - **Keep `cpu4-7` offline for development.** Transfers, builds and deploys are reliable with them down and are not with them up. - **The A15 half of the mc_smp series must not be offered upstream** - see #27 and `patches/linux-mc-smp-a15/`. The `cpu_table` fix in that series is an independent mainline bug and is unaffected. - Any measurement taken with eight cores online since #11 closed is suspect. ## The part worth keeping Closing #11 on "`nproc` returns 8 and the cores execute" was true and insufficient. **A cluster that corrupts memory under load is worse than one that does not come up**, because the second failure is loud. The first time anything actually leaned on those cores, they corrupted data, and the only thing standing between that and a week of unexplainable results was a checksum in a deploy script. See `48-a15-memory-corruption.md`.
Author
Owner

Investigated at length 2026-08-28. Root cause narrowed, not found. Mitigated, not fixed.
Full write-up in 52-a15-not-dram.md.

The DRAM hypothesis in this issue is wrong

Every piece of prior art pointed at DRAM - the sunxi community's whole reliability-testing
apparatus exists for load-dependent single-bit DRAM errors - and it is not that.

Same test throughout: a 32 MiB buffer copied inside tmpfs under CPU load and checked, 15
rounds. Baseline 14/15 corrupt.

hypothesis change result
DRAM impedance CONFIG_DRAM_ZQ 0x3f3fdd -> 0x3b3bbb (this board's own) 14/15
DRAM clock CONFIG_DRAM_CLK 672 -> 600 (this board's own) 14/15
Thermal measured during the failing run peak 69 C, trip at 95 C
Cluster in the coherency domain online, all work pinned to A7 0/15
A15 executing at all A15 spinning, A7 copying 0/15
A15 doing the memory work A7 spinning, A15 copying 0/15
Under-volt at 1800 MHz rail forced to 1200 mV, its maximum 14/15

A 12% DRAM clock change produced exactly 14/15 three times. Marginal DRAM timing is
rate-sensitive; this is not. That invariance was more useful than any of the failures.

This also answers the "cluster powered vs cluster executing" question left open here: neither.
The cluster can be online, and its cores can execute, and it stays clean.

What it does track

1200 MHz /  900 mV	 0/15
1608 MHz / 1000 mV	 0/15
1800 MHz / 1100 mV	15/15
1800 MHz / 1200 mV	14/15	<- rail at maximum

and the number of A15 cores busy at once - 1 or 2 clean, 3 hard reset, 4 corrupt. That is a
current-draw curve.

The frequency is genuine, not a misprogrammed PLL: a fixed loop pinned to an A15 core runs
4024 ms at 1200, 2973 at 1608, 2644 at 1800 - ratios 1.35 and 1.52 against an expected 1.34
and 1.50. The cores reach 1800 MHz and compute correctly there, and still corrupt memory
when several work at once.

Mitigation, and why it is not a fix

mc_smp.c now caps the cluster at 1608 MHz / 1000 mV. Two results from after that:

  • 1608 MHz with 4 busy threads passed 0/10 in one run and reset the board in another.
  • 1200 MHz with 8 busy threads reset the board.

So the cap moves the odds and nothing more. The eMMC now boots A7-only by default, via an
entry that omits the initramfs - maxcpus=4 alone does not work, because the initramfs onlines
the cores through sysfs afterwards.

The severity in the description is understated

It is not only silent flips and clean resets. The board came up from a reset and Oopsed in
ext4 inode handling at 10.8 s into an ordinary boot
- no synthetic load, nothing but udev -
and was unreachable on network and serial. Recovered with the rescue card; fsck.ext4 -f -y
on the unmounted eMMC then found the filesystem completely clean, so that was in-memory
corruption of the ext4_inode_cache slab, not disk damage.

The board is not reliably bootable on eight cores.

What is left, and it needs hardware

Power delivery to the external A15 DC-DC. It cannot be confirmed from software: the only
ADCs this board exposes are its four thermal sensors, no iio devices at all, so the regulator
framework reports the rail's setpoint and nothing else. Step 1 of this issue -"read the PMIC
rails under load" - can only ever answer "it is set to what we asked for".

Next step: a multimeter or scope on the A15 rail with several cores loaded.

Then, in order: compare against the vendor kernel on this same board (it has working cpufreq
and thermal throttling, so it rarely sustains a high OPP on four cores, which may be the whole
reason 1800 MHz is an exposed operating point); and check the board's 5 V input, since if the
DC-DC is fine and the supply is not, no software change will help.

**Investigated at length 2026-08-28. Root cause narrowed, not found. Mitigated, not fixed.** Full write-up in `52-a15-not-dram.md`. ## The DRAM hypothesis in this issue is wrong Every piece of prior art pointed at DRAM - the sunxi community's whole reliability-testing apparatus exists for load-dependent single-bit DRAM errors - and it is not that. Same test throughout: a 32 MiB buffer copied inside tmpfs under CPU load and checked, 15 rounds. Baseline **14/15 corrupt**. | hypothesis | change | result | |---|---|---| | DRAM impedance | `CONFIG_DRAM_ZQ` `0x3f3fdd` -> `0x3b3bbb` (this board's own) | 14/15 | | DRAM clock | `CONFIG_DRAM_CLK` 672 -> 600 (this board's own) | 14/15 | | Thermal | measured during the failing run | peak **69 C**, trip at 95 C | | Cluster in the coherency domain | online, all work pinned to A7 | **0/15** | | A15 executing at all | A15 spinning, A7 copying | **0/15** | | A15 doing the memory work | A7 spinning, A15 copying | **0/15** | | Under-volt at 1800 MHz | rail forced to 1200 mV, its maximum | 14/15 | A 12% DRAM clock change produced **exactly 14/15 three times**. Marginal DRAM timing is rate-sensitive; this is not. That invariance was more useful than any of the failures. This also answers the "cluster powered vs cluster executing" question left open here: neither. The cluster can be online, and its cores can execute, and it stays clean. ## What it does track ``` 1200 MHz / 900 mV 0/15 1608 MHz / 1000 mV 0/15 1800 MHz / 1100 mV 15/15 1800 MHz / 1200 mV 14/15 <- rail at maximum ``` and the number of A15 cores busy at once - 1 or 2 clean, 3 hard reset, 4 corrupt. That is a current-draw curve. The frequency is genuine, not a misprogrammed PLL: a fixed loop pinned to an A15 core runs 4024 ms at 1200, 2973 at 1608, 2644 at 1800 - ratios 1.35 and 1.52 against an expected 1.34 and 1.50. **The cores reach 1800 MHz and compute correctly there, and still corrupt memory when several work at once.** ## Mitigation, and why it is not a fix `mc_smp.c` now caps the cluster at **1608 MHz / 1000 mV**. Two results from after that: - **1608 MHz with 4 busy threads passed 0/10 in one run and reset the board in another.** - **1200 MHz with 8 busy threads reset the board.** So the cap moves the odds and nothing more. **The eMMC now boots A7-only by default**, via an entry that omits the initramfs - `maxcpus=4` alone does not work, because the initramfs onlines the cores through sysfs afterwards. ## The severity in the description is understated It is not only silent flips and clean resets. The board came up from a reset and **Oopsed in ext4 inode handling at 10.8 s into an ordinary boot** - no synthetic load, nothing but udev - and was unreachable on network *and* serial. Recovered with the rescue card; `fsck.ext4 -f -y` on the unmounted eMMC then found the filesystem **completely clean**, so that was in-memory corruption of the `ext4_inode_cache` slab, not disk damage. **The board is not reliably bootable on eight cores.** ## What is left, and it needs hardware Power delivery to the external A15 DC-DC. It **cannot be confirmed from software**: the only ADCs this board exposes are its four thermal sensors, no iio devices at all, so the regulator framework reports the rail's *setpoint* and nothing else. Step 1 of this issue -"read the PMIC rails under load" - can only ever answer "it is set to what we asked for". **Next step: a multimeter or scope on the A15 rail with several cores loaded.** Then, in order: compare against the vendor kernel on this same board (it has working cpufreq and thermal throttling, so it rarely sustains a high OPP on four cores, which may be the whole reason 1800 MHz is an exposed operating point); and check the board's 5 V input, since if the DC-DC is fine and the supply is not, no software change will help.
Author
Owner

Cross-checked against the offline linux-sunxi mirror. The Sunchip CX-A99 page documents an
A80 board that is close to a sibling of ours - same RTL8211E, AP6335, AC100, S/PDIF, IR receiver

  • and it independently documents the A15 supply.

It corroborates our pinmap exactly

Their power tree and pin table give:

  • OZ80120 switching regulator, default 0.9 V, supplying 4 x ARM Cortex-A15 cores
  • PL2 - PL5: "OZ80120 voltage regulator output control"

which is precisely what 33-REAL-PINMAP.md derived from this board's own script.bin
(a15_pwr_en = PL02, a15_vset1..3 = PL03/PL04/PL05). Two independent sources, same answer, so
the pin mapping and the "external DC-DC, not the AXP" conclusion in 31-power-topology.md can be
treated as settled rather than inferred.

It does not support an undervolt theory, and I am not proposing one - 52-a15-not-dram.md
already forced the rail to its 1200 mV maximum and still saw 14/15 corrupt runs. The setpoint is
not the problem.

What is actually new

The exact part number: OZ80120. We had it as "an OZ8012x". That is worth having because it
makes the next step a datasheet lookup rather than an experiment:

What is the OZ80120's maximum continuous output current, and what do 4 x Cortex-A15 at
1.8 GHz actually draw?

That question is worth answering before any more software work, because 52-a15-not-dram.md
already concluded the evidence points at power delivery rather than setpoint, and severity
scaling with the number of simultaneously busy cores is a current-draw signature. If the part is
simply undersized for four A15 cores at the top operating point, then #53 is a board design
limit, not a kernel bug
, and the honest resolution is to cap the cluster (as mc_smp.c already
does) and say so - not to keep hunting for a software cure that cannot exist.

That would also explain the one result that has never fitted a software explanation: raising the
setpoint to the rail's maximum made no difference, because the limit is how much current the
regulator can deliver, not what voltage it is asked for.

Two smaller notes

  • The wiki lists the mainline binding for this regulator as regulator-gpio
    (regulator/gpio-regulator.c), which is a cleaner way to describe the rail in DT than the
    ad-hoc pokes in probe/a15opp.c, if the cluster is ever made usable.
  • It also warns that the gpio-regulator driver glitches the output during initialisation,
    causing a crash if the system boots on a Cortex-A15 core
    . Not something we hit - we boot
    A7-only - but it would matter the moment a regulator-gpio node is added.

From a deep pass over the offline linux-sunxi.org mirror (a80/linux-sunxi.org/wiki-backup, 4095 pages, snapshot 2026-08-29). linux-sunxi.org content is CC BY-SA; the pages named above are the attribution.

Cross-checked against the offline linux-sunxi mirror. The **Sunchip CX-A99** page documents an A80 board that is close to a sibling of ours - same RTL8211E, AP6335, AC100, S/PDIF, IR receiver - and it independently documents the A15 supply. ## It corroborates our pinmap exactly Their power tree and pin table give: - `OZ80120` switching regulator, **default 0.9 V**, supplying *4 x ARM Cortex-A15 cores* - **PL2 - PL5**: "OZ80120 voltage regulator output control" which is precisely what `33-REAL-PINMAP.md` derived from this board's own `script.bin` (`a15_pwr_en = PL02`, `a15_vset1..3 = PL03/PL04/PL05`). Two independent sources, same answer, so the pin mapping and the "external DC-DC, not the AXP" conclusion in `31-power-topology.md` can be treated as settled rather than inferred. **It does not support an undervolt theory, and I am not proposing one** - `52-a15-not-dram.md` already forced the rail to its 1200 mV maximum and still saw 14/15 corrupt runs. The setpoint is not the problem. ## What is actually new **The exact part number: `OZ80120`.** We had it as "an OZ8012x". That is worth having because it makes the next step a datasheet lookup rather than an experiment: > **What is the OZ80120's maximum continuous output current, and what do 4 x Cortex-A15 at > 1.8 GHz actually draw?** That question is worth answering before any more software work, because `52-a15-not-dram.md` already concluded the evidence points at *power delivery* rather than setpoint, and severity scaling with the number of simultaneously busy cores is a current-draw signature. If the part is simply undersized for four A15 cores at the top operating point, then **#53 is a board design limit, not a kernel bug**, and the honest resolution is to cap the cluster (as `mc_smp.c` already does) and say so - not to keep hunting for a software cure that cannot exist. That would also explain the one result that has never fitted a software explanation: raising the setpoint to the rail's maximum made no difference, because the limit is how much current the regulator can deliver, not what voltage it is asked for. ## Two smaller notes - The wiki lists the mainline binding for this regulator as **`regulator-gpio`** (`regulator/gpio-regulator.c`), which is a cleaner way to describe the rail in DT than the ad-hoc pokes in `probe/a15opp.c`, if the cluster is ever made usable. - It also warns that the gpio-regulator driver **glitches the output during initialisation, causing a crash if the system boots on a Cortex-A15 core**. Not something we hit - we boot A7-only - but it would matter the moment a `regulator-gpio` node is added. --- *From a deep pass over the offline linux-sunxi.org mirror (`a80/linux-sunxi.org/wiki-backup`, 4095 pages, snapshot 2026-08-29). linux-sunxi.org content is CC BY-SA; the pages named above are the attribution.*
Author
Owner

The datasheet question is answered — by a different datasheet

Full write-up in 60-a15-power-budget.md.

The OZ80120 datasheet is not public, and would not have helped

Searched O2Micro's own site (EN and the ap. mirror), their power and mobile-comm catalogues,
datasheetarchive, electronicsdatasheets, the Chinese aggregators (icspec, semiee, 114ic), and
all 166 PDFs in our linux-sunxi mirror. Nothing. O2Micro data is customer-NDA.

But the more useful point is that O2Micro files this part under "CPU Core DC/DC Controller" —
their notebook CPU-core line. A controller switches external MOSFETs through an external
inductor. Its datasheet gives phase count, input range and the VID interface; the amps come from
the power stage on our PCB.
So the answer would have been "whatever the board's FETs and
inductor are rated for".

The step in my last comment - "make the next step a datasheet lookup rather than an
experiment" - was wrong.
The current capability is a property of this board, and only this
board can be asked.

What the A80 datasheet says instead

A80 Datasheet rev 1.3, which we have locally (and rev 1.1 / 1.0 are in the wiki mirror):

§5.1 Absolute Maximum Ratings MIN MAX
VDD-CPUB - Power Supply for CPUB -0.3 V 1.1 V
§5.2 Recommended Operating Conditions MIN TYP MAX
VDD-CPUB - Power Supply for Cluster B 0.8 V 0.9 V 1.1 V

🔴 The 1800 MHz / 1200 mV row in 52-a15-not-dram.md was an out-of-spec test. 1200 mV is
100 mV above the SoC's absolute maximum rating - not above a recommendation, above the rating
beyond which Allwinner makes no claim that the part survives. It ran twice. Do not repeat it.
Upstream agrees: sun9i-a80-cubieboard4.dts caps vdd-cpua and vdd-cpub at 1100000 µV.

And 1800 MHz has zero voltage headroom by design. The vendor's vf_table0 asks 1100 mV at
1800 MHz
- exactly the datasheet ceiling. Its 2016 MHz @ 1200 mV entry is never selected
because B_max_freq = 1800000000, so the vendor never runs this rail above 1100 mV either.
That closes the voltage line of enquiry on documentary grounds: the rail cannot legally go higher
than it already does, and "raise the setpoint" was never going to work.

VDD-CPUB is the largest supply rail on the package

Counted from the §6.1 ball map (via kicad/gen/a80-ballmap.csv, parsed from the datasheet and
asserting the 636-ball total):

rail balls
GND 112
VDD-CPUB (A15) 24
VCC-DRAM 20
VDD-GPU 11
VDD-CPUA (A7) 10
VDD-SYS 10
VDD-VPU 5
VDD-CPUS 3

Plus dedicated remote-sense pins VDDFB-CPUB / GNDFB-CPUB (J14 / K14): the package expects
the CPUB regulator to close its loop at the die, which is what you do for a high-current rail
with real IR drop - and independent confirmation of the external-DC-DC topology.

24 balls against the A7 cluster's 10. At the ~0.5 A/ball that 0.65 mm-pitch FCBGA power delivery
is normally designed around (a rule of thumb, not a datasheet figure) that is a rail provisioned
for roughly 12 A, ~13 W at 1.1 V. The per-ball figure is arguable; the 2.4:1 ratio and the
absolute size are not.

Also worth recording: VDD-CPUB appears in neither the power-on nor the power-down sequence
(§5.5), while CPUA, SYS, GPU and CPUS all do. The A15 rail is designed to be off at power-on and
raised later under software control - the topology we found the hard way in
35-a15-rail-enable.md - so we are violating no documented sequencing constraint.

The budget

The Draco ships a 12 V / 2 A brick - 24 W for the whole box.

A15 cluster, provisioned    ~13 W at the die
  through the DC-DC (~90%)  ~14.5 W from the 12 V input
leaving for everything else ~9.5 W of 24 W

and "everything else" is 4 GiB of DDR3 across four chips at 672 MHz, the A7 cluster, GPU, VPU,
eMMC, a gigabit PHY, AP6335, a 4-port USB hub and its downstream ports, USB 3.0 and a SATA bridge.
A fully loaded A15 cluster and a normally loaded board do not both fit in 24 W.

That rests on one rule of thumb, so the margin is indicative, not exact. But it points the same
way as every measurement in this issue, and it fits the one symptom that has never had a software
explanation: hard reset at three busy A15 cores, /sys/fs/pstore empty, nothing on serial.
That is not what a kernel bug looks like. It is what an over-current or under-voltage trip looks
like.

The half I could not answer

What 4 x A15 at 1.8 GHz on 28 nm actually draw. The comparable part is the Exynos 5422 (same
node, same quad-A15 cluster, per-rail INA231 sensors on the ODROID-XU3) and the figures usually
quoted put a saturated A15 cluster in the 8-13 W range - but I did not land a primary source, so
that is hearsay and does no work in the argument above.
The ball count and the 24 W budget are
the load-bearing parts.

Next steps, cheapest first — and step 1 is new

  1. 🔌 Swap the power supply. A bench supply at 12 V with a current readout, or any 12 V / 4-5 A
    adapter, then re-run the corruption test with 4 A15 cores busy. Never tried, needs no probe
    on the board, and if the failure moves, the answer is the input supply.
    A bench supply also
    gives the total-draw curve against busy-core count for free - the measurement this issue has
    been asking for, taken at the barrel jack instead of at a 0.65 mm ball.
  2. 📷 Photograph the DC-DC beside the SoC. Count the phases, read the inductor marking. That
    is the current rating, and it is exactly what the missing datasheet would not have told us.
  3. 🔬 Only then, a scope on the rail, for ripple and droop across a core-count step.

⚠️ And a caveat that should travel with every voltage in this issue: all of them come from the
AR100's own table, never from a meter.
840 / 900 / 1000 / 1100 / 1200 mV are what the blob
says the vset codes mean. Given the 1.1 V ceiling, confirming that mapping is worth doing - and a
step-and-measure sweep of PL03/04/05 needs only the same multimeter as step 1.

If the cluster is ever made usable

The regulator-gpio binding suggested earlier stands, with two constraints: declare
regulator-max-microvolt = <1100000> and omit the 1200 mV state from states entirely (the
vendor driver offers it and the vendor's DVFS never uses it); and not regulator-always-on,
which would enable the rail at probe into a domain U-Boot left clamped at 0xff and reproduce the
attempt-1 hang.

## The datasheet question is answered — by a different datasheet Full write-up in `60-a15-power-budget.md`. ### The OZ80120 datasheet is not public, and would not have helped Searched O2Micro's own site (EN and the `ap.` mirror), their power and mobile-comm catalogues, datasheetarchive, electronicsdatasheets, the Chinese aggregators (icspec, semiee, 114ic), and **all 166 PDFs in our linux-sunxi mirror**. Nothing. O2Micro data is customer-NDA. But the more useful point is that O2Micro files this part under **"CPU Core DC/DC Controller"** — their notebook CPU-core line. A *controller* switches external MOSFETs through an external inductor. Its datasheet gives phase count, input range and the VID interface; **the amps come from the power stage on our PCB.** So the answer would have been "whatever the board's FETs and inductor are rated for". > **The step in my last comment - "make the next step a datasheet lookup rather than an > experiment" - was wrong.** The current capability is a property of this board, and only this > board can be asked. ### What the A80 datasheet says instead **A80 Datasheet rev 1.3**, which we have locally (and rev 1.1 / 1.0 are in the wiki mirror): | §5.1 Absolute Maximum Ratings | MIN | MAX | | --- | --- | --- | | VDD-CPUB - Power Supply for CPUB | -0.3 V | **1.1 V** | | §5.2 Recommended Operating Conditions | MIN | TYP | MAX | | --- | --- | --- | --- | | VDD-CPUB - Power Supply for Cluster B | 0.8 V | 0.9 V | **1.1 V** | 🔴 **The `1800 MHz / 1200 mV` row in `52-a15-not-dram.md` was an out-of-spec test.** 1200 mV is 100 mV above the SoC's *absolute maximum rating* - not above a recommendation, above the rating beyond which Allwinner makes no claim that the part survives. It ran twice. **Do not repeat it.** Upstream agrees: `sun9i-a80-cubieboard4.dts` caps `vdd-cpua` and `vdd-cpub` at `1100000` µV. **And 1800 MHz has zero voltage headroom by design.** The vendor's `vf_table0` asks **1100 mV at 1800 MHz** - exactly the datasheet ceiling. Its `2016 MHz @ 1200 mV` entry is never selected because `B_max_freq = 1800000000`, so **the vendor never runs this rail above 1100 mV either.** That closes the voltage line of enquiry on documentary grounds: the rail cannot legally go higher than it already does, and "raise the setpoint" was never going to work. ### VDD-CPUB is the largest supply rail on the package Counted from the §6.1 ball map (via `kicad/gen/a80-ballmap.csv`, parsed from the datasheet and asserting the 636-ball total): | rail | balls | | --- | --- | | GND | 112 | | **VDD-CPUB** (A15) | **24** | | VCC-DRAM | 20 | | VDD-GPU | 11 | | VDD-CPUA (A7) | 10 | | VDD-SYS | 10 | | VDD-VPU | 5 | | VDD-CPUS | 3 | Plus dedicated remote-sense pins **`VDDFB-CPUB` / `GNDFB-CPUB`** (J14 / K14): the package expects the CPUB regulator to close its loop *at the die*, which is what you do for a high-current rail with real IR drop - and independent confirmation of the external-DC-DC topology. 24 balls against the A7 cluster's 10. At the ~0.5 A/ball that 0.65 mm-pitch FCBGA power delivery is normally designed around (a rule of thumb, not a datasheet figure) that is a rail provisioned for roughly **12 A, ~13 W at 1.1 V**. The per-ball figure is arguable; the 2.4:1 ratio and the absolute size are not. Also worth recording: **VDD-CPUB appears in neither the power-on nor the power-down sequence** (§5.5), while CPUA, SYS, GPU and CPUS all do. The A15 rail is *designed* to be off at power-on and raised later under software control - the topology we found the hard way in `35-a15-rail-enable.md` - so we are violating no documented sequencing constraint. ### The budget The Draco ships a **12 V / 2 A** brick - **24 W for the whole box**. ``` A15 cluster, provisioned ~13 W at the die through the DC-DC (~90%) ~14.5 W from the 12 V input leaving for everything else ~9.5 W of 24 W ``` and "everything else" is 4 GiB of DDR3 across four chips at 672 MHz, the A7 cluster, GPU, VPU, eMMC, a gigabit PHY, AP6335, a 4-port USB hub and its downstream ports, USB 3.0 and a SATA bridge. **A fully loaded A15 cluster and a normally loaded board do not both fit in 24 W.** That rests on one rule of thumb, so the margin is indicative, not exact. But it points the same way as every measurement in this issue, and it fits the one symptom that has never had a software explanation: **hard reset at three busy A15 cores, `/sys/fs/pstore` empty, nothing on serial.** That is not what a kernel bug looks like. It is what an over-current or under-voltage trip looks like. ### The half I could not answer *What 4 x A15 at 1.8 GHz on 28 nm actually draw.* The comparable part is the Exynos 5422 (same node, same quad-A15 cluster, per-rail INA231 sensors on the ODROID-XU3) and the figures usually quoted put a saturated A15 cluster in the 8-13 W range - **but I did not land a primary source, so that is hearsay and does no work in the argument above.** The ball count and the 24 W budget are the load-bearing parts. ### Next steps, cheapest first — and step 1 is new 1. 🔌 **Swap the power supply.** A bench supply at 12 V with a current readout, or any 12 V / 4-5 A adapter, then re-run the corruption test with 4 A15 cores busy. **Never tried, needs no probe on the board, and if the failure moves, the answer is the input supply.** A bench supply also gives the total-draw curve against busy-core count for free - the measurement this issue has been asking for, taken at the barrel jack instead of at a 0.65 mm ball. 2. 📷 **Photograph the DC-DC beside the SoC.** Count the phases, read the inductor marking. That *is* the current rating, and it is exactly what the missing datasheet would not have told us. 3. 🔬 Only then, a scope on the rail, for ripple and droop across a core-count step. ⚠️ And a caveat that should travel with every voltage in this issue: **all of them come from the AR100's own table, never from a meter.** `840 / 900 / 1000 / 1100 / 1200 mV` are what the blob says the vset codes mean. Given the 1.1 V ceiling, confirming that mapping is worth doing - and a step-and-measure sweep of PL03/04/05 needs only the same multimeter as step 1. ### If the cluster is ever made usable The `regulator-gpio` binding suggested earlier stands, with two constraints: declare `regulator-max-microvolt = <1100000>` and **omit the 1200 mV state from `states` entirely** (the vendor driver offers it and the vendor's DVFS never uses it); and **not** `regulator-always-on`, which would enable the rail at probe into a domain U-Boot left clamped at `0xff` and reproduce the attempt-1 hang.
Author
Owner

Tested on a second power supply: no change at all

A different, more stable 12 V / 2 A brick was fitted 2026-08-30. The failure is identical, and
the threshold is identical to the core.

The threshold is exactly where it was

Escalating the number of simultaneously busy A15 cores, 45 s per step, with continuous 32 MiB
tmpfs copy+md5 traffic throughout, cluster at 1608 MHz / 1000 mV, logged to disk with sync
after every line so the record survives the reset:

20:16:45  A15 online: 0-7
20:16:45  --> starting 1 busy A15 thread(s)
20:17:30      SURVIVED 1 thread(s)   temps=65 65 67 68
20:17:35  --> starting 2 busy A15 thread(s)
20:18:21      SURVIVED 2 thread(s)   temps=70 69 74 77
20:18:26  --> starting 3 busy A15 thread(s)
                                       <- log ends here; board reset

1 and 2 busy A15 cores are clean; 3 hard-resets the board. That is the same
"1-2 clean, 3 hard reset" boundary recorded on the old supply, reproduced on the new one.

A separate run beforehand, 8 tight ALU loops pinned one per core plus the same memory churn, reset
the board within ~12-40 s of the load starting:

stress launched at uptime 407s
t+10 / t+20 / t+30   UNREACHABLE
t+40                 uptime 33s     <- already rebooted

Two independent resets on the new supply. Every signature matches what this issue already
records: /sys/fs/pstore empty, no oops, no panic, the previous boot's journal ending
mid-ssh-session with nothing after it. The machine simply goes away. Temperatures peaked at
77 C against a 95 C trip, so it is not thermal, and the filesystem needed no journal recovery
on either reboot.

What that does and does not settle

✅ Supply quality is exonerated. Ripple, sag and tired electrolytics in the original brick are
no longer candidates - a second, better-behaved unit behaves identically, to the same core count.

❌ The power budget is NOT exonerated, because both bricks are rated 12 V / 2 A - the same 24 W.
This swap changed supply quality, not headroom, so the central claim in 60-a15-power-budget.md
is untouched by it. (The other supply to hand is 19 V, which is not usable on a 12 V input.)

So the two power hypotheses have now been separated, and only one of them has been tested.

Also measured on the way

  • The A15 cluster really is at 1608 MHz. clk_summary reports c1cpux at 408 MHz, which
    is wrong - the AR100 owns that PLL and Linux's cached rate is stale. A tight subs/bne loop
    retires 1 iteration per cycle and gives 1605 Miter/s on an A15 core against 1004 on an A7
    whose clock is independently known to be 1008 MHz. ⚠️ Do not read c1cpux out of
    clk_summary and believe it
    ; time a loop instead.
  • The rail comes up reading 1200 mV before mc_smp sets it down to 1000
    (sunxi mc_smp: A15 rail at 1200 mV, target 1000 mV) - i.e. the power-on default of the vset
    code is, per the AR100's table, above the SoC's 1.1 V absolute maximum. See
    60-a15-power-budget.md.
  • a15=off is not a kernel parameter. There is no such string anywhere in
    arch/arm/mach-sunxi or drivers/soc/sunxi. maxcpus=4 limits boot-time bring-up only, so
    cpu4-7 are present and can be onlined live from sysfs with no reboot - which is how these
    runs were done. The comment in extlinux.conf claiming a15=off "is what actually does the
    work" is only true of the initramfs script that reads it; nothing in the kernel honours it.
  • The 15-round corruption test returned 0 of 15 at 1608 MHz, matching the pre-swap baseline.
    Expected and uninformative: only the 1800 MHz operating points ever corrupted. At the current
    cap the hard reset is the only failure that reproduces
    , so it is the only useful probe.

Next

The one experiment that would settle the budget question is still outstanding, and it is now the
only cheap one left:

  1. 🔌 A 12 V supply rated well above 2 A - a 12 V / 5 A CCTV or LED-strip brick, same
    5.5/2.5 mm centre-positive barrel. If 3 busy cores still resets on that, then power delivery
    upstream of the board is fully exonerated and the suspect narrows to the on-board power stage,
    or to something that is not power at all.
  2. 📷 A photo of the DC-DC beside the SoC - phase count and the inductor marking. That is the
    current rating, and the thing no datasheet was going to tell us.

Worth noting how sharp the boundary is: 2 cores indefinitely fine, 3 cores gone in seconds. A step
that abrupt looks like a protection trip - over-current or under-voltage - rather than gradual
marginality. That is consistent with a budget limit, but equally consistent with an on-board OCP,
and the two are distinguished by test 1.

## Tested on a second power supply: no change at all **A different, more stable 12 V / 2 A brick was fitted 2026-08-30.** The failure is identical, and the threshold is identical to the core. ### The threshold is exactly where it was Escalating the number of simultaneously busy A15 cores, 45 s per step, with continuous 32 MiB tmpfs copy+md5 traffic throughout, cluster at **1608 MHz / 1000 mV**, logged to disk with `sync` after every line so the record survives the reset: ``` 20:16:45 A15 online: 0-7 20:16:45 --> starting 1 busy A15 thread(s) 20:17:30 SURVIVED 1 thread(s) temps=65 65 67 68 20:17:35 --> starting 2 busy A15 thread(s) 20:18:21 SURVIVED 2 thread(s) temps=70 69 74 77 20:18:26 --> starting 3 busy A15 thread(s) <- log ends here; board reset ``` **1 and 2 busy A15 cores are clean; 3 hard-resets the board.** That is the same "1-2 clean, 3 hard reset" boundary recorded on the old supply, reproduced on the new one. A separate run beforehand, 8 tight ALU loops pinned one per core plus the same memory churn, reset the board within ~12-40 s of the load starting: ``` stress launched at uptime 407s t+10 / t+20 / t+30 UNREACHABLE t+40 uptime 33s <- already rebooted ``` Two independent resets on the new supply. Every signature matches what this issue already records: `/sys/fs/pstore` **empty**, no oops, no panic, the previous boot's journal ending mid-ssh-session with nothing after it. The machine simply goes away. Temperatures peaked at **77 C** against a 95 C trip, so it is not thermal, and the filesystem needed no journal recovery on either reboot. ### What that does and does not settle ✅ **Supply quality is exonerated.** Ripple, sag and tired electrolytics in the original brick are no longer candidates - a second, better-behaved unit behaves identically, to the same core count. ❌ **The power budget is NOT exonerated, because both bricks are rated 12 V / 2 A - the same 24 W.** This swap changed supply *quality*, not *headroom*, so the central claim in `60-a15-power-budget.md` is untouched by it. (The other supply to hand is 19 V, which is not usable on a 12 V input.) So the two power hypotheses have now been separated, and only one of them has been tested. ### Also measured on the way - **The A15 cluster really is at 1608 MHz.** `clk_summary` reports `c1cpux` at **408 MHz**, which is wrong - the AR100 owns that PLL and Linux's cached rate is stale. A tight `subs`/`bne` loop retires 1 iteration per cycle and gives **1605 Miter/s** on an A15 core against **1004** on an A7 whose clock is independently known to be 1008 MHz. ⚠️ **Do not read `c1cpux` out of `clk_summary` and believe it**; time a loop instead. - **The rail comes up reading 1200 mV** before `mc_smp` sets it down to 1000 (`sunxi mc_smp: A15 rail at 1200 mV, target 1000 mV`) - i.e. the power-on default of the vset code is, per the AR100's table, **above the SoC's 1.1 V absolute maximum**. See `60-a15-power-budget.md`. - **`a15=off` is not a kernel parameter.** There is no such string anywhere in `arch/arm/mach-sunxi` or `drivers/soc/sunxi`. `maxcpus=4` limits boot-time bring-up only, so `cpu4-7` are *present* and can be onlined live from sysfs with no reboot - which is how these runs were done. The comment in `extlinux.conf` claiming `a15=off` "is what actually does the work" is only true of the initramfs script that reads it; nothing in the kernel honours it. - The 15-round corruption test returned **0 of 15** at 1608 MHz, matching the pre-swap baseline. Expected and uninformative: only the 1800 MHz operating points ever corrupted. **At the current cap the hard reset is the only failure that reproduces**, so it is the only useful probe. ### Next The one experiment that would settle the budget question is still outstanding, and it is now the *only* cheap one left: 1. 🔌 **A 12 V supply rated well above 2 A** - a 12 V / 5 A CCTV or LED-strip brick, same 5.5/2.5 mm centre-positive barrel. If 3 busy cores still resets on that, then power delivery *upstream of the board* is fully exonerated and the suspect narrows to the on-board power stage, or to something that is not power at all. 2. 📷 **A photo of the DC-DC beside the SoC** - phase count and the inductor marking. That is the current rating, and the thing no datasheet was going to tell us. Worth noting how sharp the boundary is: 2 cores indefinitely fine, 3 cores gone in seconds. A step that abrupt looks like a protection trip - over-current or under-voltage - rather than gradual marginality. That is consistent with a budget limit, but equally consistent with an on-board OCP, and the two are distinguished by test 1.
Author
Owner

⚠️ Correction: there is no core-count threshold. There is ~80 °C.

My comment earlier today said "1-2 busy A15 cores clean, 3 hard-resets... 2 cores indefinitely
fine". That is wrong.
So is the same claim in the issue body and in 52-a15-not-dram.md. Full
write-up in 61-a15-eighty-degrees.md.

The reason is boring: every test step was 45 seconds long, and 2 cores take about 45 seconds to
kill the board.
The escalation caught 2 cores at 46 s with zone3 at 77 °C and logged
"SURVIVED" — it passed by a couple of seconds, and I read a threshold into it.

Re-measured with 5 s heartbeats written to disk with sync

The last line before the log stops is the survival time.

busy A15 cores A7 load time to hard reset thermal_zone3 at last heartbeat
1 idle ~85 s 81 °C
2 idle ~40 s 80 °C
2 4 cores busy ~60 s 81 °C
3 idle seconds (warm board) —
4 4 cores busy 12-40 s —
0 4 cores busy, 7 min never 59-60 °C, flat

Load does not decide whether the board dies, only how fast. One busy A15 core is enough, given
ninety seconds.

What the failures have in common

Every reset lands at 77-81 °C on thermal_zone3; every survival is at 77 °C or below.
Four independent failures, one number.

Not the kernel's thermal management: the SoC trip is 95 °C, pstore stays empty, no oops, no
panic, no emergency-shutdown line. 52-a15-not-dram.md ruled thermal out on a run that peaked at
69 °C — below this line, which is why it never appeared. Correct for that run, wrong as a
general conclusion; that note now carries a banner saying so.

The A7 control is the sharpest number

Four busy A7 cores at 1008 MHz do not raise the board's temperature at all — seven minutes,
59-60 °C, unmoved. One busy A15 core at 1608 MHz takes it 62 → 81 °C in 85 seconds. The A15
cluster and its regulator are in a completely different regime from the rest of the SoC, and that
is why an A7-only board is rock solid.

🌀 The next test is now a fan, not a power supply

Two readings survive, and they make opposite predictions:

  1. An over-temperature protection is tripping — most plausibly the external DC-DC's, since the
    regulator and its inductor heat with the current they pass and nothing in software watches
    them. Airflow fixes it.
  2. It is a current limit and SoC temperature is only a proxy for dissipated power. Airflow
    changes nothing.

Point a fan at the board and re-run one busy A15 core. Costs nothing, needs no instrument, and
is now cheaper and more decisive than the larger 12 V supply.

Stated plainly: ~80 °C is a correlation across four failures, not an established cause.
thermal_zone3 is an on-die SoC sensor, not the regulator's, so it may be reporting the symptom.
The fan test is what would turn it into a cause.

Consequence for the "limit the cores" workaround

There is no N for which "at most N busy A15 cores" is safe, because N=1 fails. For the
record, since it came up:

  • No cpufreq governor gates core count — governors scale frequency; the cpufreq layer does not
    model how many cores are simultaneously busy.
  • Onlining only cpu4/cpu5 is a genuine structural guarantee, but of the wrong thing.
    Measured: reset in ~40 s idle, ~60 s with the A7s loaded.
  • Idle injection (cpuidle_cooling, mainline's nearest analogue to the vendor's
    sunxi-budget-cooling) idles a cooling domain in sync — during each active window every core
    runs at once, preserving exactly the peak this board cannot take while lowering the average.
    Wrong tool in principle.
  • The vendor's own [cooler_table] does carry "1200000 4 1440000 2" states, but it is a 3.4 BSP
    driver, thermally driven, and on this evidence 2 A15 cores would not have saved it either.

A7-only is not a conservative default; on current evidence it is the only working one.

Method note

What made this measurable was logging heartbeats to disk with sync, so the record survives
the reset that is the failure. And the mistake worth not repeating: a fixed-duration test step
measures your timeout, not the hardware
— vary the duration before believing any threshold.

## ⚠️ Correction: there is no core-count threshold. There is ~80 °C. **My comment earlier today said "1-2 busy A15 cores clean, 3 hard-resets... 2 cores indefinitely fine". That is wrong.** So is the same claim in the issue body and in `52-a15-not-dram.md`. Full write-up in `61-a15-eighty-degrees.md`. The reason is boring: **every test step was 45 seconds long, and 2 cores take about 45 seconds to kill the board.** The escalation caught 2 cores at 46 s with `zone3` at 77 °C and logged "SURVIVED" — it passed by a couple of seconds, and I read a threshold into it. ### Re-measured with 5 s heartbeats written to disk with `sync` The last line before the log stops is the survival time. | busy A15 cores | A7 load | time to hard reset | `thermal_zone3` at last heartbeat | | --- | --- | --- | --- | | 1 | idle | **~85 s** | **81 °C** | | 2 | idle | **~40 s** | **80 °C** | | 2 | 4 cores busy | **~60 s** | **81 °C** | | 3 | idle | seconds (warm board) | — | | 4 | 4 cores busy | 12-40 s | — | | **0** | **4 cores busy, 7 min** | **never** | **59-60 °C, flat** | **Load does not decide whether the board dies, only how fast.** One busy A15 core is enough, given ninety seconds. ### What the failures have in common Every reset lands at **77-81 °C on `thermal_zone3`**; every survival is at **77 °C or below**. Four independent failures, one number. Not the kernel's thermal management: **the SoC trip is 95 °C**, pstore stays empty, no oops, no panic, no emergency-shutdown line. `52-a15-not-dram.md` ruled thermal out on a run that peaked at **69 °C** — below this line, which is why it never appeared. Correct for that run, wrong as a general conclusion; that note now carries a banner saying so. ### The A7 control is the sharpest number **Four busy A7 cores at 1008 MHz do not raise the board's temperature at all** — seven minutes, 59-60 °C, unmoved. One busy A15 core at 1608 MHz takes it 62 → 81 °C in 85 seconds. The A15 cluster and its regulator are in a completely different regime from the rest of the SoC, and that is why an A7-only board is rock solid. ### 🌀 The next test is now a fan, not a power supply Two readings survive, and they make opposite predictions: 1. **An over-temperature protection is tripping** — most plausibly the external DC-DC's, since the regulator and its inductor heat with the current they pass and nothing in software watches them. **Airflow fixes it.** 2. **It is a current limit** and SoC temperature is only a proxy for dissipated power. **Airflow changes nothing.** **Point a fan at the board and re-run one busy A15 core.** Costs nothing, needs no instrument, and is now cheaper and more decisive than the larger 12 V supply. Stated plainly: ~80 °C is a **correlation across four failures, not an established cause**. `thermal_zone3` is an on-die SoC sensor, not the regulator's, so it may be reporting the symptom. The fan test is what would turn it into a cause. ### Consequence for the "limit the cores" workaround There is **no N** for which "at most N busy A15 cores" is safe, because **N=1 fails**. For the record, since it came up: - **No cpufreq governor gates core count** — governors scale frequency; the cpufreq layer does not model how many cores are simultaneously busy. - **Onlining only `cpu4`/`cpu5`** is a genuine structural guarantee, but of the wrong thing. Measured: reset in ~40 s idle, ~60 s with the A7s loaded. - **Idle injection** (`cpuidle_cooling`, mainline's nearest analogue to the vendor's `sunxi-budget-cooling`) idles a cooling domain **in sync** — during each active window every core runs at once, preserving exactly the peak this board cannot take while lowering the average. Wrong tool in principle. - The vendor's own `[cooler_table]` does carry `"1200000 4 1440000 2"` states, but it is a 3.4 BSP driver, thermally driven, and on this evidence 2 A15 cores would not have saved it either. **A7-only is not a conservative default; on current evidence it is the only working one.** ### Method note What made this measurable was **logging heartbeats to disk with `sync`**, so the record survives the reset that *is* the failure. And the mistake worth not repeating: **a fixed-duration test step measures your timeout, not the hardware** — vary the duration before believing any threshold.
tiagoagueda changed title from A15 cluster corrupts memory under load, and hard-resets under more to A15 cluster: board hard-resets under any sustained A15 load, at ~80 C 2026-08-30 20:49:24 +00:00
Sign in to join this conversation.
No description provided.