Recover the last 512 MiB: 4 GiB fitted, 3.5 GiB reachable #61
Labels
No labels
blocked-physical
cleanup
hardware
infra
kernel
P1-critical
P2-high
P3-normal
P4-later
reliability
security
upstream
wontfix
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
tiagoagueda/a80#61
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
The board has 4 GiB fitted and Linux sees 3.5 GiB. sun9i DRAM is based at
0x20000000, so0x100000000 - 0x20000000= 3584 MiB is all a 32-bit physical address can describe, andsunxi_dram_init()clamps to exactly that. Background in 22-dram-4gb.md;that note calls recovering the rest "LPAE work well beyond a constant change", which turns
out to be half right — see below.
Current state on the board:
The address decode exists — the manual says so
A80 User Manual rev 1.1, memory map (p. 79). DRAM is two windows, not one:
0x2000_0000 - 0xFFFF_FFFF0x1_0000_0000 - 0x2_1FFF_FFFF8 G total, matching "Support 8GB address space" in §2.1.3.2. So the missing 512 MiB is not
lost to an MMIO hole or an unpopulated rank: it is decoded at
0x1_0000_0000 - 0x1_1FFF_FFFF, and nothing but 32-bit addressing stands in front of it.Upstream's
SUNXI_DRAM_MAX_SIZEKconfig used to carry# TODO: try out A80's 8GiB DRAM space— this issue is that TODO.The kernel side is already done
~/a80/linux/.configon the build host:The running kernel can already address above 4 GiB. There is no LPAE port to do — the
kernel is simply never told the memory is there.
The blocker is entirely U-Boot
Work:
CONFIG_PHYS_64BITforMACH_SUN9I. Without itphys_addr_tandgd->ram_sizeare32-bit and the size arithmetic wraps at exactly 4096 MiB — bug 2 in note 22, which was
papered over with a clamp rather than fixed.
arch/arm/mach-sunxi/dram_sun9i.c:unsigned long sunxi_dram_init(void)needs toreturn
phys_size_t, and theif (size_mb > 3584) size_mb = 3584;clamp comes out.The comment above that clamp already spells out why it is there.
SUNXI_DRAM_MAX_SIZEto0x100000000for sun9i — the value H616/A133/A523already use — and rewrite the sun9i comment beneath it.
memorynode must express a range crossing the boundary:#address-cells = <2>,or two banks. Today it is one bank,
0x20000000+0xe0000000.CONFIG_PHYS_64BITon 32-bit sunxi is not a well-travelled path. Expect fallout in codethat stores or prints addresses as
ulong.LOADADDR=0x20008000is unaffected and still mandatory.Then verify it properly — this is the failure mode that bites silently
rows=16reported the right size, booted Linux, and passed every spot check whilecorrupting memory under sustained load. Same trap here, same standard of proof:
mtestacross the whole range in one go, including above0x1_0000_0000. Chunked runs cannot see aliasing between chunks (note 22).wdt stopon both watchdogs first, andsetenv bootretry -1, or the run resets theboard around 45 minutes in.
step to
/dev/consolebefore writing it.memtesteris still notinstalled. Fold that in rather than deferring it a second time.
configuration already known to corrupt memory proves nothing either way.
Driver audit
dma_addr_tis already 64-bit, but a device whosedma_maskis 32-bit cannot reach the newpages — it will bounce through SWIOTLB, or quietly truncate. The hardware side is fine: the
manual gives the DMAC "script memory and device space support 34-bit address". Whether each
driver agrees is the question. Worth checking
sun4i-drm— see #48, where a display layeraddress already needs a physical/DMA fixup nobody has explained, which is exactly the shape
of an address-truncation bug — plus
dwc3,sun6i-dma,sun7i-dwmac, and the AR100mailbox pool.
Is it worth doing
+512 MiB on 3584 is +14%. Highmem:lowmem goes from 4.3:1 to 5.1:1, still comfortable at
the current 3G/1G split, so no
VMSPLITchange and no lowmem pressure. Per-process addressspace does not change and never can on ARMv7 — nothing here helps a single process that
wants more than ~3 GiB.
Two non-strategies, recorded so nobody proposes them later:
Not available: the A80 is Cortex-A7/A15, ARMv7 only, no AArch64 state.
reserved region. There is no partial win available. Without 64-bit physical addresses out
of the bootloader the kernel is never told that memory exists at all.
Suggest this sits behind #53. It is a bounded U-Boot change onto a kernel that is already
configured for it, but it reflashes the eMMC bootloader for a 14% gain, and the validation
run is the expensive part rather than the patch.
Done 2026-08-30 — 4 GiB reachable, and U-Boot was never the blocker
MemTotalHighTotal/proc/iomem20000000-ffffffff20000000-11fffffffFull write-up in 57-4gb-reclaimed.md.
There is no 4 GB / 8 GB mode
Searching all 1008 pages of the user manual, the high window
0x1 0000 0000 - 0x2 1FFF FFFFappears exactly once - the memory map on p.80. No mode register, no second map. "Up to 8G"
is the decode capacity, not a selectable mode. Nothing to switch on: the decode was always
there, and 32-bit physical addressing was the only thing in front of it.
The U-Boot work in this issue is not required
arm_add_memory()only truncates a bank above 4 GiB inside#ifndef CONFIG_PHYS_ADDR_T_64BIT,and this kernel already sets it. ARM's
mem=takes a base as well as a size, and the firstmem=wipes the bootloader's banks - so whatever U-Boot writes into the DTmemorynode isdiscarded. The entire change is a kernel command line:
No
CONFIG_PHYS_64BIT, nophys_size_treturn, noSUNXI_DRAM_MAX_SIZEchange, no two-cellmemory node, and the eMMC boot area was never rewritten. The U-Boot work is now an
improvement rather than a prerequisite, and should be judged on its own merits.
CMA was the real blocker - the driver risk this issue predicted, arriving through CMA
First attempt booted a long way then died:
CMA took the top of memory, which is exactly what a 32-bit
dma_maskcannot reach; SWIOTLB wasplaced low but CMA bypasses it.
cma=size@base-limitwithbase + size != limitis a caprather than a fixed placement, so the pool lands low:
After that: eth0, wlan0, eMMC, 8 USB devices with zero "HC died", Bluetooth - all up.
Verified to the standard this issue asked for
3680 MiB resident is 208 MiB more than the old
MemTotalof 3472 MiB, so it could not havebeen satisfied before at any price, and with zero swap it was genuinely in RAM. Zones agree:
196608 + 851968 = 1048576 pages = exactly 4 GiB.
⚠️ One boot and a bounded integrity test, not a soak. Per this issue's own warning about the
rows=16trap, treat 4 GiB as demonstrated, not proven stable.The
mem8gentry is NOT the defaulta7-onlystill is, so an ordinary reboot returns to 3.5 GiB. Promoting it is a decision.Cost, and a lesson worth keeping
The first attempt stranded the board and needed a physical power cycle. The entry had a safe
default but not a safe failure: with no root the initramfs dropped to a shell that never
executed anything,
CONFIG_MAGIC_SYSRQis unset so there was no serial escape, and bothwatchdogs are stopped by
sunxi-wdtat probe. Network, serial and watchdog gone at once -the #5 case exactly.
panic=10fixes it and is now on the entry: Debian'sscripts/functionsreboots instead ofspawning a shell, and if
reboot -ffails it forces a kernel panic which also reboots. Anyexperimental boot entry on this board should carry it.
Still open
dma_maskperdriver;
sun4i-drmmost of all, given #48's unexplainedPHYS_OFFSETfixupmem=hardcodes the split in the boot menu, so the clean fix is still the U-Boot work aboveSoaked clean, and promoted to the default
Soak, 2026-08-30
Four processes holding 3600 MiB between them — more than the old
MemTotalof 3472 MiB, sothe reclaimed region is in use throughout — each looping write-then-verify with a fresh seed per
pass, twenty minutes at load ~4, cpu4-7 offline:
For contrast, the A15 corruption in #53 needed only a 32 MiB buffer and two busy loops to produce
14 bad copies out of 15. Nothing of that kind appears here.
Promoted
default mem8gin the eMMCextlinux.conf. Verified with an unattended boot — no serialinteraction — coming up in 24 s with
MemTotal 4075420 kBandbootcountback to 0.There is no brick risk, and that was checked rather than assumed:
bootlimit = 3+altbootcmduImage.known-good, and the altbootcmd bootargs are hardcoded with nomem=, so the fallback cannot inherit the faultboot-mark-good.servicebootcountonce userspace is healthytimeout 30bootdelay 2Reverting is one line —
default a7-only— and the previous file is backed up as/root/extlinux.conf.bak-promote-*.⚠️ Twenty minutes is a soak, not a burn-in. The honest claim is that 4 GiB survives sustained
load as far as it has been pushed, not that it is proven. If anything starts corrupting data
later, revert this first.
Remaining
than checking
dma_maskper driver;sun4i-drmmost of all, given #48mem=hardcodes the split in the boot menu, so the clean fix is still the U-Boot workoriginally described here — now an improvement rather than a prerequisite
Closing: the 512 MiB is recovered and in production
MemTotal4075420 kB against 3555228,/proc/iomemreading20000000-11fffffff,mem8gpromoted to the default in the eMMC
extlinux.conf, and an unattended boot verified. Soaked fortwenty minutes at load ~4 with 3600 MiB resident: 235 passes, 206.5 GiB written and read back,
0 blocks mismatched, 0 kB swap, no OOM, bad page or ext4 errors. Write-up in
57-4gb-reclaimed.md.
Three things this issue got wrong, worth recording
were needed — not
CONFIG_PHYS_64BIT, not thephys_size_treturn, notSUNXI_DRAM_MAX_SIZE, not a two-cell memory node.arm_add_memory()only truncates above4 GiB inside
#ifndef CONFIG_PHYS_ADDR_T_64BIT, which this kernel already sets, so twomem=arguments were enough and the bootloader was never rewritten.where 32-bit masters cannot reach.
cpu4-7 offline throughout, and #53 is untouched by it.
Also settled: there is no 4 GB / 8 GB mode. The high window appears exactly once in all 1008
pages of the user manual — the memory map on p.80. "Up to 8G" is the decode capacity, not
something selectable, so there was nothing to switch on.
What is not closed, and where it went
It was a precondition here and is now a live production concern, since the board boots 4 GiB
by default and any driver can be handed a page above
0x1_0000_0000. Tracking it inside a"recover the RAM" issue would hide it.
mtestacross the whole range asked for here is not achievable: U-Boot'smtestis 32-bit and cannot address above0x1_0000_0000. The 206.5 GiB verified under Linuxis the stronger equivalent. The top 81 MiB left untested by #38 remains untested for the same
reason — U-Boot lives there.
prerequisite:
mem=hardcodes the split in the boot menu, so a board with different DRAMfitted would need it edited. It should be judged on its own merits now, not as the price of
the extra 512 MiB.
⚠️ Twenty minutes is a soak, not a burn-in. If anything starts corrupting data later, revert
default a7-onlyin the eMMCextlinux.confas the first suspect.