Recover the last 512 MiB: 4 GiB fitted, 3.5 GiB reachable #61

Closed
opened 2026-08-29 16:36:21 +00:00 by tiagoagueda · 3 comments
Owner

The board has 4 GiB fitted and Linux sees 3.5 GiB. sun9i DRAM is based at 0x20000000, so
0x100000000 - 0x20000000 = 3584 MiB is all a 32-bit physical address can describe, and
sunxi_dram_init() clamps to exactly that. Background in 22-dram-4gb.md;
that note calls recovering the rest "LPAE work well beyond a constant change", which turns
out to be half right — see below.

Current state on the board:

MemTotal:        3555228 kB
HighTotal:       2883584 kB
LowTotal:         671644 kB

The address decode exists — the manual says so

A80 User Manual rev 1.1, memory map (p. 79). DRAM is two windows, not one:

Range Size
0x2000_0000 - 0xFFFF_FFFF 3.5 G
0x1_0000_0000 - 0x2_1FFF_FFFF 4.5 G

8 G total, matching "Support 8GB address space" in §2.1.3.2. So the missing 512 MiB is not
lost to an MMIO hole or an unpopulated rank: it is decoded at
0x1_0000_0000 - 0x1_1FFF_FFFF, and nothing but 32-bit addressing stands in front of it.
Upstream's SUNXI_DRAM_MAX_SIZE Kconfig used to carry # TODO: try out A80's 8GiB DRAM space — this issue is that TODO.

The kernel side is already done

~/a80/linux/.config on the build host:

CONFIG_ARM_LPAE=y
CONFIG_HIGHMEM=y
CONFIG_HIGHPTE=y
CONFIG_PHYS_ADDR_T_64BIT=y
CONFIG_ARCH_DMA_ADDR_T_64BIT=y
CONFIG_SWIOTLB=y

The running kernel can already address above 4 GiB. There is no LPAE port to do — the
kernel is simply never told the memory is there.

The blocker is entirely U-Boot

CONFIG_SUNXI_DRAM_MAX_SIZE=0xE0000000
# CONFIG_PHYS_64BIT is not set

Work:

  1. CONFIG_PHYS_64BIT for MACH_SUN9I. Without it phys_addr_t and gd->ram_size are
    32-bit and the size arithmetic wraps at exactly 4096 MiB — bug 2 in note 22, which was
    papered over with a clamp rather than fixed.
  2. arch/arm/mach-sunxi/dram_sun9i.c: unsigned long sunxi_dram_init(void) needs to
    return phys_size_t, and the if (size_mb > 3584) size_mb = 3584; clamp comes out.
    The comment above that clamp already spells out why it is there.
  3. Raise SUNXI_DRAM_MAX_SIZE to 0x100000000 for sun9i — the value H616/A133/A523
    already use — and rewrite the sun9i comment beneath it.
  4. The DT memory node must express a range crossing the boundary: #address-cells = <2>,
    or two banks. Today it is one bank, 0x20000000 + 0xe0000000.
  5. CONFIG_PHYS_64BIT on 32-bit sunxi is not a well-travelled path. Expect fallout in code
    that stores or prints addresses as ulong.

LOADADDR=0x20008000 is unaffected and still mandatory.

Then verify it properly — this is the failure mode that bites silently

rows=16 reported the right size, booted Linux, and passed every spot check while
corrupting memory under sustained load. Same trap here, same standard of proof:

  • A single-run mtest across the whole range in one go, including above
    0x1_0000_0000. Chunked runs cannot see aliasing between chunks (note 22).
  • wdt stop on both watchdogs first, and setenv bootretry -1, or the run resets the
    board around 45 minutes in.
  • A sustained-load ramp under Linux that actually touches the new pages, reporting each
    step to /dev/console before writing it.
  • The top 81 MiB that #38 left untested is still untested and memtester is still not
    installed. Fold that in rather than deferring it a second time.
  • With cpu4-7 offline. #53 is open and unexplained; validating new memory on top of a
    configuration already known to corrupt memory proves nothing either way.

Driver audit

dma_addr_t is already 64-bit, but a device whose dma_mask is 32-bit cannot reach the new
pages — it will bounce through SWIOTLB, or quietly truncate. The hardware side is fine: the
manual gives the DMAC "script memory and device space support 34-bit address". Whether each
driver agrees is the question. Worth checking sun4i-drm — see #48, where a display layer
address already needs a physical/DMA fixup nobody has explained, which is exactly the shape
of an address-truncation bug — plus dwc3, sun6i-dma, sun7i-dwmac, and the AR100
mailbox pool.

Is it worth doing

+512 MiB on 3584 is +14%. Highmem:lowmem goes from 4.3:1 to 5.1:1, still comfortable at
the current 3G/1G split, so no VMSPLIT change and no lowmem pressure. Per-process address
space does not change and never can on ARMv7 — nothing here helps a single process that
wants more than ~3 GiB.

Two non-strategies, recorded so nobody proposes them later:

  • An arm64 kernel with an armhf userspace, the usual escape from 32-bit memory limits.
    Not available: the A80 is Cortex-A7/A15, ARMv7 only, no AArch64 state.
  • Spending the top 512 MiB on something outside the kernel linear map — CMA, ramoops, a
    reserved region. There is no partial win available. Without 64-bit physical addresses out
    of the bootloader the kernel is never told that memory exists at all.

Suggest this sits behind #53. It is a bounded U-Boot change onto a kernel that is already
configured for it, but it reflashes the eMMC bootloader for a 14% gain, and the validation
run is the expensive part rather than the patch.

The board has 4 GiB fitted and Linux sees 3.5 GiB. sun9i DRAM is based at `0x20000000`, so `0x100000000 - 0x20000000` = 3584 MiB is all a 32-bit physical address can describe, and `sunxi_dram_init()` clamps to exactly that. Background in [22-dram-4gb.md](22-dram-4gb.md); that note calls recovering the rest "LPAE work well beyond a constant change", which turns out to be half right — see below. Current state on the board: ``` MemTotal: 3555228 kB HighTotal: 2883584 kB LowTotal: 671644 kB ``` ## The address decode exists — the manual says so A80 User Manual rev 1.1, memory map (p. 79). DRAM is **two** windows, not one: | Range | Size | |---|---| | `0x2000_0000 - 0xFFFF_FFFF` | 3.5 G | | `0x1_0000_0000 - 0x2_1FFF_FFFF` | 4.5 G | 8 G total, matching "Support 8GB address space" in §2.1.3.2. So the missing 512 MiB is not lost to an MMIO hole or an unpopulated rank: it is decoded at `0x1_0000_0000 - 0x1_1FFF_FFFF`, and nothing but 32-bit addressing stands in front of it. Upstream's `SUNXI_DRAM_MAX_SIZE` Kconfig used to carry `# TODO: try out A80's 8GiB DRAM space` — this issue is that TODO. ## The kernel side is already done `~/a80/linux/.config` on the build host: ``` CONFIG_ARM_LPAE=y CONFIG_HIGHMEM=y CONFIG_HIGHPTE=y CONFIG_PHYS_ADDR_T_64BIT=y CONFIG_ARCH_DMA_ADDR_T_64BIT=y CONFIG_SWIOTLB=y ``` The running kernel can already address above 4 GiB. There is no LPAE port to do — the kernel is simply never told the memory is there. ## The blocker is entirely U-Boot ``` CONFIG_SUNXI_DRAM_MAX_SIZE=0xE0000000 # CONFIG_PHYS_64BIT is not set ``` Work: 1. `CONFIG_PHYS_64BIT` for `MACH_SUN9I`. Without it `phys_addr_t` and `gd->ram_size` are 32-bit and the size arithmetic wraps at exactly 4096 MiB — bug 2 in note 22, which was papered over with a clamp rather than fixed. 2. `arch/arm/mach-sunxi/dram_sun9i.c`: `unsigned long sunxi_dram_init(void)` needs to return `phys_size_t`, and the `if (size_mb > 3584) size_mb = 3584;` clamp comes out. The comment above that clamp already spells out why it is there. 3. Raise `SUNXI_DRAM_MAX_SIZE` to `0x100000000` for sun9i — the value H616/A133/A523 already use — and rewrite the sun9i comment beneath it. 4. The DT `memory` node must express a range crossing the boundary: `#address-cells = <2>`, or two banks. Today it is one bank, `0x20000000` + `0xe0000000`. 5. `CONFIG_PHYS_64BIT` on 32-bit sunxi is not a well-travelled path. Expect fallout in code that stores or prints addresses as `ulong`. `LOADADDR=0x20008000` is unaffected and still mandatory. ## Then verify it properly — this is the failure mode that bites silently `rows=16` reported the right size, booted Linux, and passed every spot check while corrupting memory under sustained load. Same trap here, same standard of proof: - A **single-run** `mtest` across the whole range in one go, including above `0x1_0000_0000`. Chunked runs cannot see aliasing between chunks (note 22). - `wdt stop` on both watchdogs first, and `setenv bootretry -1`, or the run resets the board around 45 minutes in. - A sustained-load ramp under Linux that actually touches the new pages, reporting each step to `/dev/console` before writing it. - The top 81 MiB that #38 left untested is still untested and `memtester` is still not installed. Fold that in rather than deferring it a second time. - **With cpu4-7 offline.** #53 is open and unexplained; validating new memory on top of a configuration already known to corrupt memory proves nothing either way. ## Driver audit `dma_addr_t` is already 64-bit, but a device whose `dma_mask` is 32-bit cannot reach the new pages — it will bounce through SWIOTLB, or quietly truncate. The hardware side is fine: the manual gives the DMAC "script memory and device space support 34-bit address". Whether each driver agrees is the question. Worth checking `sun4i-drm` — see #48, where a display layer address already needs a physical/DMA fixup nobody has explained, which is exactly the shape of an address-truncation bug — plus `dwc3`, `sun6i-dma`, `sun7i-dwmac`, and the AR100 mailbox pool. ## Is it worth doing +512 MiB on 3584 is **+14%**. Highmem:lowmem goes from 4.3:1 to 5.1:1, still comfortable at the current 3G/1G split, so no `VMSPLIT` change and no lowmem pressure. Per-process address space does not change and never can on ARMv7 — nothing here helps a single process that wants more than ~3 GiB. Two non-strategies, recorded so nobody proposes them later: - **An arm64 kernel with an armhf userspace**, the usual escape from 32-bit memory limits. Not available: the A80 is Cortex-A7/A15, ARMv7 only, no AArch64 state. - **Spending the top 512 MiB on something outside the kernel linear map** — CMA, ramoops, a reserved region. There is no partial win available. Without 64-bit physical addresses out of the bootloader the kernel is never told that memory exists at all. Suggest this sits behind #53. It is a bounded U-Boot change onto a kernel that is already configured for it, but it reflashes the eMMC bootloader for a 14% gain, and the validation run is the expensive part rather than the patch.
Author
Owner

Done 2026-08-30 — 4 GiB reachable, and U-Boot was never the blocker

before after
MemTotal 3555228 kB 4075420 kB
HighTotal 2883584 kB 3407872 kB
/proc/iomem 20000000-ffffffff 20000000-11fffffff

Full write-up in 57-4gb-reclaimed.md.

There is no 4 GB / 8 GB mode

Searching all 1008 pages of the user manual, the high window 0x1 0000 0000 - 0x2 1FFF FFFF
appears exactly once - the memory map on p.80. No mode register, no second map. "Up to 8G"
is the decode capacity, not a selectable mode. Nothing to switch on: the decode was always
there, and 32-bit physical addressing was the only thing in front of it.

The U-Boot work in this issue is not required

arm_add_memory() only truncates a bank above 4 GiB inside #ifndef CONFIG_PHYS_ADDR_T_64BIT,
and this kernel already sets it. ARM's mem= takes a base as well as a size, and the first
mem= wipes the bootloader's banks - so whatever U-Boot writes into the DT memory node is
discarded. The entire change is a kernel command line:

mem=3584M@0x20000000 mem=512M@0x100000000 cma=64M@0x20000000-0x100000000

No CONFIG_PHYS_64BIT, no phys_size_t return, no SUNXI_DRAM_MAX_SIZE change, no two-cell
memory node, and the eMMC boot area was never rewritten. The U-Boot work is now an
improvement rather than a prerequisite, and should be judged on its own merits.

CMA was the real blocker - the driver risk this issue predicted, arriving through CMA

First attempt booted a long way then died:

cma: Reserved 64 MiB at 0x0000000100000000      <- the whole pool above 4 GiB
ohci-platform a00400.usb: HC died; cleaning up
ALERT!  /dev/mmcblk1p1 does not exist.  Dropping to a shell!

CMA took the top of memory, which is exactly what a 32-bit dma_mask cannot reach; SWIOTLB was
placed low but CMA bypasses it. cma=size@base-limit with base + size != limit is a cap
rather than a fixed placement, so the pool lands low:

cma: Reserved 64 MiB at 0x00000000fc000000

After that: eth0, wlan0, eMMC, 8 USB devices with zero "HC died", Bluetooth - all up.

Verified to the standard this issue asked for

Test Result
4 x 850 MiB, distinct pattern per 1 MiB block, write then verify 0 mismatched blocks
4 x 920 MiB = 3680 MiB resident at once, same check 0 mismatched blocks
Swap used during that run 0 kB
OOM / bad page / corruption in dmesg none

3680 MiB resident is 208 MiB more than the old MemTotal of 3472 MiB, so it could not have
been satisfied before at any price, and with zero swap it was genuinely in RAM. Zones agree:
196608 + 851968 = 1048576 pages = exactly 4 GiB.

⚠️ One boot and a bounded integrity test, not a soak. Per this issue's own warning about the
rows=16 trap, treat 4 GiB as demonstrated, not proven stable.

The mem8g entry is NOT the default

a7-only still is, so an ordinary reboot returns to 3.5 GiB. Promoting it is a decision.

Cost, and a lesson worth keeping

The first attempt stranded the board and needed a physical power cycle. The entry had a safe
default but not a safe failure: with no root the initramfs dropped to a shell that never
executed anything, CONFIG_MAGIC_SYSRQ is unset so there was no serial escape, and both
watchdogs are stopped by sunxi-wdt at probe. Network, serial and watchdog gone at once -
the #5 case exactly.

panic=10 fixes it and is now on the entry: Debian's scripts/functions reboots instead of
spawning a shell, and if reboot -f fails it forces a kernel panic which also reboots. Any
experimental boot entry on this board should carry it.

Still open

  • a sustained-load soak, with cpu4-7 offline so #53 is not confounding
  • the driver audit - everything probes, but that is weaker than checking dma_mask per
    driver; sun4i-drm most of all, given #48's unexplained PHYS_OFFSET fixup
  • mem= hardcodes the split in the boot menu, so the clean fix is still the U-Boot work above
## Done 2026-08-30 — 4 GiB reachable, and U-Boot was never the blocker | | before | after | |---|---|---| | `MemTotal` | 3555228 kB | **4075420 kB** | | `HighTotal` | 2883584 kB | 3407872 kB | | `/proc/iomem` | `20000000-ffffffff` | **`20000000-11fffffff`** | Full write-up in [57-4gb-reclaimed.md](57-4gb-reclaimed.md). ### There is no 4 GB / 8 GB mode Searching all **1008 pages** of the user manual, the high window `0x1 0000 0000 - 0x2 1FFF FFFF` appears **exactly once** - the memory map on p.80. No mode register, no second map. "Up to 8G" is the decode capacity, not a selectable mode. Nothing to switch on: the decode was always there, and 32-bit physical addressing was the only thing in front of it. ### The U-Boot work in this issue is not required `arm_add_memory()` only truncates a bank above 4 GiB inside `#ifndef CONFIG_PHYS_ADDR_T_64BIT`, and this kernel already sets it. ARM's `mem=` takes a base as well as a size, and the first `mem=` wipes the bootloader's banks - so whatever U-Boot writes into the DT `memory` node is discarded. The entire change is a kernel command line: ``` mem=3584M@0x20000000 mem=512M@0x100000000 cma=64M@0x20000000-0x100000000 ``` No `CONFIG_PHYS_64BIT`, no `phys_size_t` return, no `SUNXI_DRAM_MAX_SIZE` change, no two-cell memory node, and **the eMMC boot area was never rewritten**. The U-Boot work is now an improvement rather than a prerequisite, and should be judged on its own merits. ### CMA was the real blocker - the driver risk this issue predicted, arriving through CMA First attempt booted a long way then died: ``` cma: Reserved 64 MiB at 0x0000000100000000 <- the whole pool above 4 GiB ohci-platform a00400.usb: HC died; cleaning up ALERT! /dev/mmcblk1p1 does not exist. Dropping to a shell! ``` CMA took the top of memory, which is exactly what a 32-bit `dma_mask` cannot reach; SWIOTLB was placed low but CMA bypasses it. `cma=size@base-limit` with `base + size != limit` is a **cap** rather than a fixed placement, so the pool lands low: ``` cma: Reserved 64 MiB at 0x00000000fc000000 ``` After that: eth0, wlan0, eMMC, **8 USB devices with zero "HC died"**, Bluetooth - all up. ### Verified to the standard this issue asked for | Test | Result | |---|---| | 4 x 850 MiB, distinct pattern per 1 MiB block, write then verify | 0 mismatched blocks | | 4 x 920 MiB = **3680 MiB resident at once**, same check | **0 mismatched blocks** | | Swap used during that run | **0 kB** | | OOM / bad page / corruption in dmesg | none | 3680 MiB resident is **208 MiB more than the old `MemTotal` of 3472 MiB**, so it could not have been satisfied before at any price, and with zero swap it was genuinely in RAM. Zones agree: 196608 + 851968 = 1048576 pages = exactly 4 GiB. ⚠️ **One boot and a bounded integrity test, not a soak.** Per this issue's own warning about the `rows=16` trap, treat 4 GiB as demonstrated, not proven stable. ### The `mem8g` entry is NOT the default `a7-only` still is, so an ordinary reboot returns to 3.5 GiB. Promoting it is a decision. ### Cost, and a lesson worth keeping The first attempt stranded the board and needed a physical power cycle. The entry had a safe *default* but not a safe *failure*: with no root the initramfs dropped to a shell that never executed anything, `CONFIG_MAGIC_SYSRQ` is unset so there was no serial escape, and both watchdogs are stopped by `sunxi-wdt` at probe. Network, serial and watchdog gone at once - the #5 case exactly. **`panic=10` fixes it** and is now on the entry: Debian's `scripts/functions` reboots instead of spawning a shell, and if `reboot -f` fails it forces a kernel panic which also reboots. Any experimental boot entry on this board should carry it. ### Still open - a **sustained-load soak**, with cpu4-7 offline so #53 is not confounding - the **driver audit** - everything probes, but that is weaker than checking `dma_mask` per driver; `sun4i-drm` most of all, given #48's unexplained `PHYS_OFFSET` fixup - `mem=` hardcodes the split in the boot menu, so the clean fix is still the U-Boot work above
Author
Owner

Soaked clean, and promoted to the default

Soak, 2026-08-30

Four processes holding 3600 MiB between them — more than the old MemTotal of 3472 MiB, so
the reclaimed region is in use throughout — each looping write-then-verify with a fresh seed per
pass, twenty minutes at load ~4, cpu4-7 offline:

Passes 235 across four processes (58, 58, 59, 60)
Data written and read back 206.5 GiB
Blocks mismatched 0
Swap consumed 0 kB — working set stayed resident
OOM / bad page / corruption / hardware error none
ext4 errors none

For contrast, the A15 corruption in #53 needed only a 32 MiB buffer and two busy loops to produce
14 bad copies out of 15. Nothing of that kind appears here.

Promoted

default mem8g in the eMMC extlinux.conf. Verified with an unattended boot — no serial
interaction — coming up in 24 s with MemTotal 4075420 kB and bootcount back to 0.

There is no brick risk, and that was checked rather than assumed:

Net State
bootlimit = 3 + altbootcmd three failed boots fall back unattended to uImage.known-good, and the altbootcmd bootargs are hardcoded with no mem=, so the fallback cannot inherit the fault
boot-mark-good.service enabled and active, clears bootcount once userspace is healthy
extlinux timeout 30 3 s serial window to pick another entry
bootdelay 2 serial escape to the U-Boot prompt
rescue SD last resort

Reverting is one line — default a7-only — and the previous file is backed up as
/root/extlinux.conf.bak-promote-*.

⚠️ Twenty minutes is a soak, not a burn-in. The honest claim is that 4 GiB survives sustained
load as far as it has been pushed, not that it is proven. If anything starts corrupting data
later, revert this first.

Remaining

  • the driver audit — everything probes and ran through the soak, but that is still weaker
    than checking dma_mask per driver; sun4i-drm most of all, given #48
  • mem= hardcodes the split in the boot menu, so the clean fix is still the U-Boot work
    originally described here — now an improvement rather than a prerequisite
## Soaked clean, and promoted to the default ### Soak, 2026-08-30 Four processes holding **3600 MiB** between them — more than the old `MemTotal` of 3472 MiB, so the reclaimed region is in use throughout — each looping write-then-verify with a fresh seed per pass, twenty minutes at load ~4, cpu4-7 offline: | | | |---|---| | Passes | 235 across four processes (58, 58, 59, 60) | | Data written **and read back** | **206.5 GiB** | | Blocks mismatched | **0** | | Swap consumed | **0 kB** — working set stayed resident | | OOM / bad page / corruption / hardware error | none | | ext4 errors | none | For contrast, the A15 corruption in #53 needed only a 32 MiB buffer and two busy loops to produce 14 bad copies out of 15. Nothing of that kind appears here. ### Promoted `default mem8g` in the eMMC `extlinux.conf`. Verified with an unattended boot — no serial interaction — coming up in 24 s with `MemTotal 4075420 kB` and `bootcount` back to 0. **There is no brick risk, and that was checked rather than assumed:** | Net | State | |---|---| | `bootlimit = 3` + `altbootcmd` | three failed boots fall back **unattended** to `uImage.known-good`, and the altbootcmd bootargs are hardcoded with **no `mem=`**, so the fallback cannot inherit the fault | | `boot-mark-good.service` | enabled and active, clears `bootcount` once userspace is healthy | | extlinux `timeout 30` | 3 s serial window to pick another entry | | `bootdelay 2` | serial escape to the U-Boot prompt | | rescue SD | last resort | Reverting is one line — `default a7-only` — and the previous file is backed up as `/root/extlinux.conf.bak-promote-*`. ⚠️ **Twenty minutes is a soak, not a burn-in.** The honest claim is that 4 GiB survives sustained load as far as it has been pushed, not that it is proven. If anything starts corrupting data later, revert this first. ### Remaining - the **driver audit** — everything probes and ran through the soak, but that is still weaker than checking `dma_mask` per driver; `sun4i-drm` most of all, given #48 - `mem=` hardcodes the split in the boot menu, so the clean fix is still the U-Boot work originally described here — now an improvement rather than a prerequisite
Author
Owner

Closing: the 512 MiB is recovered and in production

MemTotal 4075420 kB against 3555228, /proc/iomem reading 20000000-11fffffff, mem8g
promoted to the default in the eMMC extlinux.conf, and an unattended boot verified. Soaked for
twenty minutes at load ~4 with 3600 MiB resident: 235 passes, 206.5 GiB written and read back,
0 blocks mismatched, 0 kB swap, no OOM, bad page or ext4 errors. Write-up in
57-4gb-reclaimed.md.

Three things this issue got wrong, worth recording

  • "The blocker is entirely U-Boot." It was not. None of the four U-Boot changes listed here
    were needed — not CONFIG_PHYS_64BIT, not the phys_size_t return, not
    SUNXI_DRAM_MAX_SIZE, not a two-cell memory node. arm_add_memory() only truncates above
    4 GiB inside #ifndef CONFIG_PHYS_ADDR_T_64BIT, which this kernel already sets, so two
    mem= arguments were enough and the bootloader was never rewritten.
  • The real blocker was CMA, which took the top of memory by default and put every DMA buffer
    where 32-bit masters cannot reach.
  • "Suggest this sits behind #53." It did not need to. The work was done ahead of #53 with
    cpu4-7 offline throughout, and #53 is untouched by it.

Also settled: there is no 4 GB / 8 GB mode. The high window appears exactly once in all 1008
pages of the user manual — the memory map on p.80. "Up to 8G" is the decode capacity, not
something selectable, so there was nothing to switch on.

What is not closed, and where it went

  • The driver audit is now #63.
    It was a precondition here and is now a live production concern, since the board boots 4 GiB
    by default and any driver can be handed a page above 0x1_0000_0000. Tracking it inside a
    "recover the RAM" issue would hide it.
  • The single-run mtest across the whole range asked for here is not achievable: U-Boot's
    mtest is 32-bit and cannot address above 0x1_0000_0000. The 206.5 GiB verified under Linux
    is the stronger equivalent. The top 81 MiB left untested by #38 remains untested for the same
    reason — U-Boot lives there.
  • The U-Boot work described here is still worth doing, but as an improvement rather than a
    prerequisite: mem= hardcodes the split in the boot menu, so a board with different DRAM
    fitted would need it edited. It should be judged on its own merits now, not as the price of
    the extra 512 MiB.

⚠️ Twenty minutes is a soak, not a burn-in. If anything starts corrupting data later, revert
default a7-only in the eMMC extlinux.conf as the first suspect.

## Closing: the 512 MiB is recovered and in production `MemTotal` 4075420 kB against 3555228, `/proc/iomem` reading `20000000-11fffffff`, `mem8g` promoted to the default in the eMMC `extlinux.conf`, and an unattended boot verified. Soaked for twenty minutes at load ~4 with 3600 MiB resident: 235 passes, 206.5 GiB written and read back, **0 blocks mismatched**, 0 kB swap, no OOM, bad page or ext4 errors. Write-up in [57-4gb-reclaimed.md](57-4gb-reclaimed.md). ### Three things this issue got wrong, worth recording - **"The blocker is entirely U-Boot."** It was not. None of the four U-Boot changes listed here were needed — not `CONFIG_PHYS_64BIT`, not the `phys_size_t` return, not `SUNXI_DRAM_MAX_SIZE`, not a two-cell memory node. `arm_add_memory()` only truncates above 4 GiB inside `#ifndef CONFIG_PHYS_ADDR_T_64BIT`, which this kernel already sets, so two `mem=` arguments were enough and the bootloader was never rewritten. - **The real blocker was CMA**, which took the top of memory by default and put every DMA buffer where 32-bit masters cannot reach. - **"Suggest this sits behind #53."** It did not need to. The work was done ahead of #53 with cpu4-7 offline throughout, and #53 is untouched by it. Also settled: there is **no 4 GB / 8 GB mode**. The high window appears exactly once in all 1008 pages of the user manual — the memory map on p.80. "Up to 8G" is the decode capacity, not something selectable, so there was nothing to switch on. ### What is not closed, and where it went - **The driver audit** is now [#63](https://source.tiagoagueda.com/tiagoagueda/a80/issues/63). It was a precondition here and is now a live production concern, since the board boots 4 GiB by default and any driver can be handed a page above `0x1_0000_0000`. Tracking it inside a "recover the RAM" issue would hide it. - **The single-run `mtest` across the whole range** asked for here is not achievable: U-Boot's `mtest` is 32-bit and cannot address above `0x1_0000_0000`. The 206.5 GiB verified under Linux is the stronger equivalent. The top 81 MiB left untested by #38 remains untested for the same reason — U-Boot lives there. - **The U-Boot work** described here is still worth doing, but as an improvement rather than a prerequisite: `mem=` hardcodes the split in the boot menu, so a board with different DRAM fitted would need it edited. It should be judged on its own merits now, not as the price of the extra 512 MiB. ⚠️ Twenty minutes is a soak, not a burn-in. If anything starts corrupting data later, revert `default a7-only` in the eMMC `extlinux.conf` as the first suspect.
Sign in to join this conversation.
No description provided.