A15 cluster: board hard-resets under any sustained A15 load, at ~80 C #53
Labels
No labels
blocked-physical
cleanup
hardware
infra
kernel
P1-critical
P2-high
P3-normal
P4-later
reliability
security
upstream
wontfix
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
tiagoagueda/a80#53
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
With
cpu4-7online the board silently corrupts memory under load, and hard-resetsunder more of it. The A15 cores come up and execute - #11 is genuinely resolved - but the
system is not usable with them running.
Found 2026-08-28, by accident, an hour after #11 was closed.
Measured
Network-free: a 32 MiB buffer copied inside tmpfs and checked, under two busy loops. No disk,
no network.
The same split shows in transfers - 10/10 clean with
cpu4-7down, 6/10 with them up.Under heavier CPU load the board hard-resets, twice during that session, with
/sys/fs/pstoreempty both times. It is not an oops; the machine goes away.It is one bit
Every difference, in every sample:
Twelve flips across five corrupted 8 MiB transfers, and all twelve are bit 1, always 0 to 1,
never any other bit and never the other direction. Software does not have a favourite bit.
How it surfaced, and how close it came to not surfacing
A routine kernel deploy failed its checksum:
deploy-kernel.shmd5-checks both ends. This is the first time that check has ever caughtanything, and without it the corrupt kernel would have been installed and booted.
Ruled out
taskset-pinnedmd5sum: 160/160 correctIt needs load - an idle board copies 40 MiB without a flip. Two busy loops make it
near-certain.
Cause not established
The obvious suspect is power: the rail work in
35-a15-rail-enable.mdbrought up a clusterdrawing far more than this board had run before, and a sagging supply makes DRAM marginal. A
single data bit failing first is what a marginal DRAM interface looks like. That is a story
that fits the evidence, not a measurement.
Worth settling, cheapest first:
would put the fault in our configuration, not the silicon.
22-dram-4gb.md).the A15 cores.
Consequences until it is understood
cpu4-7offline for development. Transfers, builds and deploys are reliable withthem down and are not with them up.
patches/linux-mc-smp-a15/. Thecpu_tablefix in that series is an independent mainlinebug and is unaffected.
The part worth keeping
Closing #11 on "
nprocreturns 8 and the cores execute" was true and insufficient. A clusterthat corrupts memory under load is worse than one that does not come up, because the second
failure is loud. The first time anything actually leaned on those cores, they corrupted data,
and the only thing standing between that and a week of unexplainable results was a checksum in
a deploy script.
See
48-a15-memory-corruption.md.Investigated at length 2026-08-28. Root cause narrowed, not found. Mitigated, not fixed.
Full write-up in
52-a15-not-dram.md.The DRAM hypothesis in this issue is wrong
Every piece of prior art pointed at DRAM - the sunxi community's whole reliability-testing
apparatus exists for load-dependent single-bit DRAM errors - and it is not that.
Same test throughout: a 32 MiB buffer copied inside tmpfs under CPU load and checked, 15
rounds. Baseline 14/15 corrupt.
CONFIG_DRAM_ZQ0x3f3fdd->0x3b3bbb(this board's own)CONFIG_DRAM_CLK672 -> 600 (this board's own)A 12% DRAM clock change produced exactly 14/15 three times. Marginal DRAM timing is
rate-sensitive; this is not. That invariance was more useful than any of the failures.
This also answers the "cluster powered vs cluster executing" question left open here: neither.
The cluster can be online, and its cores can execute, and it stays clean.
What it does track
and the number of A15 cores busy at once - 1 or 2 clean, 3 hard reset, 4 corrupt. That is a
current-draw curve.
The frequency is genuine, not a misprogrammed PLL: a fixed loop pinned to an A15 core runs
4024 ms at 1200, 2973 at 1608, 2644 at 1800 - ratios 1.35 and 1.52 against an expected 1.34
and 1.50. The cores reach 1800 MHz and compute correctly there, and still corrupt memory
when several work at once.
Mitigation, and why it is not a fix
mc_smp.cnow caps the cluster at 1608 MHz / 1000 mV. Two results from after that:So the cap moves the odds and nothing more. The eMMC now boots A7-only by default, via an
entry that omits the initramfs -
maxcpus=4alone does not work, because the initramfs onlinesthe cores through sysfs afterwards.
The severity in the description is understated
It is not only silent flips and clean resets. The board came up from a reset and Oopsed in
ext4 inode handling at 10.8 s into an ordinary boot - no synthetic load, nothing but udev -
and was unreachable on network and serial. Recovered with the rescue card;
fsck.ext4 -f -yon the unmounted eMMC then found the filesystem completely clean, so that was in-memory
corruption of the
ext4_inode_cacheslab, not disk damage.The board is not reliably bootable on eight cores.
What is left, and it needs hardware
Power delivery to the external A15 DC-DC. It cannot be confirmed from software: the only
ADCs this board exposes are its four thermal sensors, no iio devices at all, so the regulator
framework reports the rail's setpoint and nothing else. Step 1 of this issue -"read the PMIC
rails under load" - can only ever answer "it is set to what we asked for".
Next step: a multimeter or scope on the A15 rail with several cores loaded.
Then, in order: compare against the vendor kernel on this same board (it has working cpufreq
and thermal throttling, so it rarely sustains a high OPP on four cores, which may be the whole
reason 1800 MHz is an exposed operating point); and check the board's 5 V input, since if the
DC-DC is fine and the supply is not, no software change will help.
Cross-checked against the offline linux-sunxi mirror. The Sunchip CX-A99 page documents an
A80 board that is close to a sibling of ours - same RTL8211E, AP6335, AC100, S/PDIF, IR receiver
It corroborates our pinmap exactly
Their power tree and pin table give:
OZ80120switching regulator, default 0.9 V, supplying 4 x ARM Cortex-A15 coreswhich is precisely what
33-REAL-PINMAP.mdderived from this board's ownscript.bin(
a15_pwr_en = PL02,a15_vset1..3 = PL03/PL04/PL05). Two independent sources, same answer, sothe pin mapping and the "external DC-DC, not the AXP" conclusion in
31-power-topology.mdcan betreated as settled rather than inferred.
It does not support an undervolt theory, and I am not proposing one -
52-a15-not-dram.mdalready forced the rail to its 1200 mV maximum and still saw 14/15 corrupt runs. The setpoint is
not the problem.
What is actually new
The exact part number:
OZ80120. We had it as "an OZ8012x". That is worth having because itmakes the next step a datasheet lookup rather than an experiment:
That question is worth answering before any more software work, because
52-a15-not-dram.mdalready concluded the evidence points at power delivery rather than setpoint, and severity
scaling with the number of simultaneously busy cores is a current-draw signature. If the part is
simply undersized for four A15 cores at the top operating point, then #53 is a board design
limit, not a kernel bug, and the honest resolution is to cap the cluster (as
mc_smp.calreadydoes) and say so - not to keep hunting for a software cure that cannot exist.
That would also explain the one result that has never fitted a software explanation: raising the
setpoint to the rail's maximum made no difference, because the limit is how much current the
regulator can deliver, not what voltage it is asked for.
Two smaller notes
regulator-gpio(
regulator/gpio-regulator.c), which is a cleaner way to describe the rail in DT than thead-hoc pokes in
probe/a15opp.c, if the cluster is ever made usable.causing a crash if the system boots on a Cortex-A15 core. Not something we hit - we boot
A7-only - but it would matter the moment a
regulator-gpionode is added.From a deep pass over the offline linux-sunxi.org mirror (
a80/linux-sunxi.org/wiki-backup, 4095 pages, snapshot 2026-08-29). linux-sunxi.org content is CC BY-SA; the pages named above are the attribution.The datasheet question is answered — by a different datasheet
Full write-up in
60-a15-power-budget.md.The OZ80120 datasheet is not public, and would not have helped
Searched O2Micro's own site (EN and the
ap.mirror), their power and mobile-comm catalogues,datasheetarchive, electronicsdatasheets, the Chinese aggregators (icspec, semiee, 114ic), and
all 166 PDFs in our linux-sunxi mirror. Nothing. O2Micro data is customer-NDA.
But the more useful point is that O2Micro files this part under "CPU Core DC/DC Controller" —
their notebook CPU-core line. A controller switches external MOSFETs through an external
inductor. Its datasheet gives phase count, input range and the VID interface; the amps come from
the power stage on our PCB. So the answer would have been "whatever the board's FETs and
inductor are rated for".
What the A80 datasheet says instead
A80 Datasheet rev 1.3, which we have locally (and rev 1.1 / 1.0 are in the wiki mirror):
🔴 The
1800 MHz / 1200 mVrow in52-a15-not-dram.mdwas an out-of-spec test. 1200 mV is100 mV above the SoC's absolute maximum rating - not above a recommendation, above the rating
beyond which Allwinner makes no claim that the part survives. It ran twice. Do not repeat it.
Upstream agrees:
sun9i-a80-cubieboard4.dtscapsvdd-cpuaandvdd-cpubat1100000µV.And 1800 MHz has zero voltage headroom by design. The vendor's
vf_table0asks 1100 mV at1800 MHz - exactly the datasheet ceiling. Its
2016 MHz @ 1200 mVentry is never selectedbecause
B_max_freq = 1800000000, so the vendor never runs this rail above 1100 mV either.That closes the voltage line of enquiry on documentary grounds: the rail cannot legally go higher
than it already does, and "raise the setpoint" was never going to work.
VDD-CPUB is the largest supply rail on the package
Counted from the §6.1 ball map (via
kicad/gen/a80-ballmap.csv, parsed from the datasheet andasserting the 636-ball total):
Plus dedicated remote-sense pins
VDDFB-CPUB/GNDFB-CPUB(J14 / K14): the package expectsthe CPUB regulator to close its loop at the die, which is what you do for a high-current rail
with real IR drop - and independent confirmation of the external-DC-DC topology.
24 balls against the A7 cluster's 10. At the ~0.5 A/ball that 0.65 mm-pitch FCBGA power delivery
is normally designed around (a rule of thumb, not a datasheet figure) that is a rail provisioned
for roughly 12 A, ~13 W at 1.1 V. The per-ball figure is arguable; the 2.4:1 ratio and the
absolute size are not.
Also worth recording: VDD-CPUB appears in neither the power-on nor the power-down sequence
(§5.5), while CPUA, SYS, GPU and CPUS all do. The A15 rail is designed to be off at power-on and
raised later under software control - the topology we found the hard way in
35-a15-rail-enable.md- so we are violating no documented sequencing constraint.The budget
The Draco ships a 12 V / 2 A brick - 24 W for the whole box.
and "everything else" is 4 GiB of DDR3 across four chips at 672 MHz, the A7 cluster, GPU, VPU,
eMMC, a gigabit PHY, AP6335, a 4-port USB hub and its downstream ports, USB 3.0 and a SATA bridge.
A fully loaded A15 cluster and a normally loaded board do not both fit in 24 W.
That rests on one rule of thumb, so the margin is indicative, not exact. But it points the same
way as every measurement in this issue, and it fits the one symptom that has never had a software
explanation: hard reset at three busy A15 cores,
/sys/fs/pstoreempty, nothing on serial.That is not what a kernel bug looks like. It is what an over-current or under-voltage trip looks
like.
The half I could not answer
What 4 x A15 at 1.8 GHz on 28 nm actually draw. The comparable part is the Exynos 5422 (same
node, same quad-A15 cluster, per-rail INA231 sensors on the ODROID-XU3) and the figures usually
quoted put a saturated A15 cluster in the 8-13 W range - but I did not land a primary source, so
that is hearsay and does no work in the argument above. The ball count and the 24 W budget are
the load-bearing parts.
Next steps, cheapest first — and step 1 is new
adapter, then re-run the corruption test with 4 A15 cores busy. Never tried, needs no probe
on the board, and if the failure moves, the answer is the input supply. A bench supply also
gives the total-draw curve against busy-core count for free - the measurement this issue has
been asking for, taken at the barrel jack instead of at a 0.65 mm ball.
is the current rating, and it is exactly what the missing datasheet would not have told us.
⚠️ And a caveat that should travel with every voltage in this issue: all of them come from the
AR100's own table, never from a meter.
840 / 900 / 1000 / 1100 / 1200 mVare what the blobsays the vset codes mean. Given the 1.1 V ceiling, confirming that mapping is worth doing - and a
step-and-measure sweep of PL03/04/05 needs only the same multimeter as step 1.
If the cluster is ever made usable
The
regulator-gpiobinding suggested earlier stands, with two constraints: declareregulator-max-microvolt = <1100000>and omit the 1200 mV state fromstatesentirely (thevendor driver offers it and the vendor's DVFS never uses it); and not
regulator-always-on,which would enable the rail at probe into a domain U-Boot left clamped at
0xffand reproduce theattempt-1 hang.
Tested on a second power supply: no change at all
A different, more stable 12 V / 2 A brick was fitted 2026-08-30. The failure is identical, and
the threshold is identical to the core.
The threshold is exactly where it was
Escalating the number of simultaneously busy A15 cores, 45 s per step, with continuous 32 MiB
tmpfs copy+md5 traffic throughout, cluster at 1608 MHz / 1000 mV, logged to disk with
syncafter every line so the record survives the reset:
1 and 2 busy A15 cores are clean; 3 hard-resets the board. That is the same
"1-2 clean, 3 hard reset" boundary recorded on the old supply, reproduced on the new one.
A separate run beforehand, 8 tight ALU loops pinned one per core plus the same memory churn, reset
the board within ~12-40 s of the load starting:
Two independent resets on the new supply. Every signature matches what this issue already
records:
/sys/fs/pstoreempty, no oops, no panic, the previous boot's journal endingmid-ssh-session with nothing after it. The machine simply goes away. Temperatures peaked at
77 C against a 95 C trip, so it is not thermal, and the filesystem needed no journal recovery
on either reboot.
What that does and does not settle
✅ Supply quality is exonerated. Ripple, sag and tired electrolytics in the original brick are
no longer candidates - a second, better-behaved unit behaves identically, to the same core count.
❌ The power budget is NOT exonerated, because both bricks are rated 12 V / 2 A - the same 24 W.
This swap changed supply quality, not headroom, so the central claim in
60-a15-power-budget.mdis untouched by it. (The other supply to hand is 19 V, which is not usable on a 12 V input.)
So the two power hypotheses have now been separated, and only one of them has been tested.
Also measured on the way
clk_summaryreportsc1cpuxat 408 MHz, whichis wrong - the AR100 owns that PLL and Linux's cached rate is stale. A tight
subs/bneloopretires 1 iteration per cycle and gives 1605 Miter/s on an A15 core against 1004 on an A7
whose clock is independently known to be 1008 MHz. ⚠️ Do not read
c1cpuxout ofclk_summaryand believe it; time a loop instead.mc_smpsets it down to 1000(
sunxi mc_smp: A15 rail at 1200 mV, target 1000 mV) - i.e. the power-on default of the vsetcode is, per the AR100's table, above the SoC's 1.1 V absolute maximum. See
60-a15-power-budget.md.a15=offis not a kernel parameter. There is no such string anywhere inarch/arm/mach-sunxiordrivers/soc/sunxi.maxcpus=4limits boot-time bring-up only, socpu4-7are present and can be onlined live from sysfs with no reboot - which is how theseruns were done. The comment in
extlinux.confclaiminga15=off"is what actually does thework" is only true of the initramfs script that reads it; nothing in the kernel honours it.
Expected and uninformative: only the 1800 MHz operating points ever corrupted. At the current
cap the hard reset is the only failure that reproduces, so it is the only useful probe.
Next
The one experiment that would settle the budget question is still outstanding, and it is now the
only cheap one left:
5.5/2.5 mm centre-positive barrel. If 3 busy cores still resets on that, then power delivery
upstream of the board is fully exonerated and the suspect narrows to the on-board power stage,
or to something that is not power at all.
current rating, and the thing no datasheet was going to tell us.
Worth noting how sharp the boundary is: 2 cores indefinitely fine, 3 cores gone in seconds. A step
that abrupt looks like a protection trip - over-current or under-voltage - rather than gradual
marginality. That is consistent with a budget limit, but equally consistent with an on-board OCP,
and the two are distinguished by test 1.
⚠️ Correction: there is no core-count threshold. There is ~80 °C.
My comment earlier today said "1-2 busy A15 cores clean, 3 hard-resets... 2 cores indefinitely
fine". That is wrong. So is the same claim in the issue body and in
52-a15-not-dram.md. Fullwrite-up in
61-a15-eighty-degrees.md.The reason is boring: every test step was 45 seconds long, and 2 cores take about 45 seconds to
kill the board. The escalation caught 2 cores at 46 s with
zone3at 77 °C and logged"SURVIVED" — it passed by a couple of seconds, and I read a threshold into it.
Re-measured with 5 s heartbeats written to disk with
syncThe last line before the log stops is the survival time.
thermal_zone3at last heartbeatLoad does not decide whether the board dies, only how fast. One busy A15 core is enough, given
ninety seconds.
What the failures have in common
Every reset lands at 77-81 °C on
thermal_zone3; every survival is at 77 °C or below.Four independent failures, one number.
Not the kernel's thermal management: the SoC trip is 95 °C, pstore stays empty, no oops, no
panic, no emergency-shutdown line.
52-a15-not-dram.mdruled thermal out on a run that peaked at69 °C — below this line, which is why it never appeared. Correct for that run, wrong as a
general conclusion; that note now carries a banner saying so.
The A7 control is the sharpest number
Four busy A7 cores at 1008 MHz do not raise the board's temperature at all — seven minutes,
59-60 °C, unmoved. One busy A15 core at 1608 MHz takes it 62 → 81 °C in 85 seconds. The A15
cluster and its regulator are in a completely different regime from the rest of the SoC, and that
is why an A7-only board is rock solid.
🌀 The next test is now a fan, not a power supply
Two readings survive, and they make opposite predictions:
regulator and its inductor heat with the current they pass and nothing in software watches
them. Airflow fixes it.
changes nothing.
Point a fan at the board and re-run one busy A15 core. Costs nothing, needs no instrument, and
is now cheaper and more decisive than the larger 12 V supply.
Stated plainly: ~80 °C is a correlation across four failures, not an established cause.
thermal_zone3is an on-die SoC sensor, not the regulator's, so it may be reporting the symptom.The fan test is what would turn it into a cause.
Consequence for the "limit the cores" workaround
There is no N for which "at most N busy A15 cores" is safe, because N=1 fails. For the
record, since it came up:
model how many cores are simultaneously busy.
cpu4/cpu5is a genuine structural guarantee, but of the wrong thing.Measured: reset in ~40 s idle, ~60 s with the A7s loaded.
cpuidle_cooling, mainline's nearest analogue to the vendor'ssunxi-budget-cooling) idles a cooling domain in sync — during each active window every coreruns at once, preserving exactly the peak this board cannot take while lowering the average.
Wrong tool in principle.
[cooler_table]does carry"1200000 4 1440000 2"states, but it is a 3.4 BSPdriver, thermally driven, and on this evidence 2 A15 cores would not have saved it either.
A7-only is not a conservative default; on current evidence it is the only working one.
Method note
What made this measurable was logging heartbeats to disk with
sync, so the record survivesthe reset that is the failure. And the mistake worth not repeating: a fixed-duration test step
measures your timeout, not the hardware — vary the duration before believing any threshold.
A15 cluster corrupts memory under load, and hard-resets under moreto A15 cluster: board hard-resets under any sustained A15 load, at ~80 C