Audit dma_mask on every DMA-capable driver, now that RAM extends above 4 GiB #63

Closed
opened 2026-08-30 10:13:49 +00:00 by tiagoagueda · 1 comment
Owner

Since 2026-08-30 the board boots with 4 GiB reachable by default — see
#61 and
57-4gb-reclaimed.md. Physical memory now runs to 0x1_1FFF_FFFF, so
any driver can be handed a page whose address does not fit in 32 bits.

This was listed inside #61 as a precondition and never done. It has changed character: it is no
longer a question about a proposed configuration, it guards one that is live.

Why it matters

A device with a 32-bit dma_mask is fine if the mask is declared honestly — the DMA API
bounces the buffer through SWIOTLB. The failure mode is a driver that

  • sets a 64-bit mask it cannot actually honour, or
  • stores a dma_addr_t in a u32, or writes only the low word into a descriptor

Either truncates silently. There is no error, no warning, and the write lands somewhere else in
DRAM. That is the worst class of bug this board can have, and
#53 is a standing reminder of how
long silent memory corruption can go unexplained.

What already protects us

  • cma=64M@0x20000000-0x100000000 caps the CMA pool below 4 GiB. Without it the pool landed
    at 0x1_0000_0000 and OHCI died and the eMMC never appeared — that is what made this concrete
    rather than theoretical.
  • SWIOTLB is reserved low (0x49eb7000) and bounces correctly declared 32-bit masters.
  • The A80 DMAC itself is fine: the manual says "Script memory and device space support 34-bit
    address" (p.218).

Evidence so far, which is encouraging but not an audit

Everything probed and survived a twenty-minute soak at load ~4 with 3600 MiB resident: eth0,
wlan0, the eMMC, 8 USB devices with zero "HC died", Bluetooth, 206.5 GiB written and read
back with 0 mismatches. But "it worked for twenty minutes" is not the same as "the mask is
right", because a truncating driver only misbehaves once it is actually given a high page.

Worth checking, roughly in order of risk

Driver Why
sun4i-drm highest. #48 records a display layer address needing an unexplained PHYS_OFFSET fixup, and the layer address is split across two registers as ((h4add << 32) | l32add) >> 3. That is exactly the shape of an address-truncation bug, and it now has memory above 4 GiB to truncate
dwc3 / xhci large descriptor rings, and USB was the first thing to break when CMA went high
sun6i-dma descriptor addresses; also the subject of #40
sun7i-dwmac descriptor rings in DRAM
sunxi-mmc the boot path — a truncation here corrupts the filesystem
sunxi_arisc mailbox pool is in SRAM rather than DRAM, so probably not affected; confirm rather than assume

How

  • grep for dma_set_mask, dma_set_coherent_mask, dma_coerce_mask_and_coherent in each and
    check the declared width against what the hardware can address
  • look for dma_addr_t narrowed to u32, and for descriptor writes that take only the low word
  • check whether any node needs a dma-ranges — there is none anywhere in this SoC's tree today,
    which #48 already flags as unexplained
  • a targeted test is possible: force an allocation into the high region and hand it to each
    driver, rather than waiting for the page allocator to do it by chance

Done when

  • every DMA-capable driver on this board has had its mask confirmed correct, or had a
    constraint added
  • sun4i-drm specifically is understood well enough to say whether #48's fixup is related
  • the finding is written down, so the 4 GiB default is either justified or qualified
Since 2026-08-30 the board boots with **4 GiB reachable** by default — see [#61](https://source.tiagoagueda.com/tiagoagueda/a80/issues/61) and [57-4gb-reclaimed.md](57-4gb-reclaimed.md). Physical memory now runs to `0x1_1FFF_FFFF`, so **any driver can be handed a page whose address does not fit in 32 bits.** This was listed inside #61 as a precondition and never done. It has changed character: it is no longer a question about a proposed configuration, it guards one that is **live**. ## Why it matters A device with a 32-bit `dma_mask` is fine *if the mask is declared honestly* — the DMA API bounces the buffer through SWIOTLB. The failure mode is a driver that - sets a 64-bit mask it cannot actually honour, or - stores a `dma_addr_t` in a `u32`, or writes only the low word into a descriptor Either truncates silently. There is no error, no warning, and the write lands somewhere else in DRAM. That is the worst class of bug this board can have, and [#53](https://source.tiagoagueda.com/tiagoagueda/a80/issues/53) is a standing reminder of how long silent memory corruption can go unexplained. ## What already protects us - `cma=64M@0x20000000-0x100000000` caps the CMA pool **below 4 GiB**. Without it the pool landed at `0x1_0000_0000` and OHCI died and the eMMC never appeared — that is what made this concrete rather than theoretical. - SWIOTLB is reserved low (`0x49eb7000`) and bounces correctly declared 32-bit masters. - The A80 DMAC itself is fine: the manual says "Script memory and device space support 34-bit address" (p.218). ## Evidence so far, which is encouraging but not an audit Everything probed and survived a twenty-minute soak at load ~4 with 3600 MiB resident: eth0, wlan0, the eMMC, **8 USB devices with zero "HC died"**, Bluetooth, 206.5 GiB written and read back with 0 mismatches. But "it worked for twenty minutes" is not the same as "the mask is right", because a truncating driver only misbehaves once it is actually given a high page. ## Worth checking, roughly in order of risk | Driver | Why | |---|---| | `sun4i-drm` | **highest.** [#48](https://source.tiagoagueda.com/tiagoagueda/a80/issues/48) records a display layer address needing an unexplained `PHYS_OFFSET` fixup, and the layer address is split across two registers as `((h4add << 32) \| l32add) >> 3`. That is exactly the shape of an address-truncation bug, and it now has memory above 4 GiB to truncate | | `dwc3` / `xhci` | large descriptor rings, and USB was the first thing to break when CMA went high | | `sun6i-dma` | descriptor addresses; also the subject of #40 | | `sun7i-dwmac` | descriptor rings in DRAM | | `sunxi-mmc` | the boot path — a truncation here corrupts the filesystem | | `sunxi_arisc` mailbox | pool is in SRAM rather than DRAM, so probably not affected; confirm rather than assume | ## How - `grep` for `dma_set_mask`, `dma_set_coherent_mask`, `dma_coerce_mask_and_coherent` in each and check the declared width against what the hardware can address - look for `dma_addr_t` narrowed to `u32`, and for descriptor writes that take only the low word - check whether any node needs a `dma-ranges` — there is none anywhere in this SoC's tree today, which #48 already flags as unexplained - a targeted test is possible: force an allocation into the high region and hand it to each driver, rather than waiting for the page allocator to do it by chance ## Done when - [ ] every DMA-capable driver on this board has had its mask confirmed correct, or had a constraint added - [ ] `sun4i-drm` specifically is understood well enough to say whether #48's fixup is related - [ ] the finding is written down, so the 4 GiB default is either justified or qualified
Author
Owner

Audit done — clean. Closing.

Masks read from kernel context with a throwaway module over
bus_for_each_dev(&platform_bus_type), rather than inferred from source. Every DMA-capable
device on this board reports 0x00000000ffffffff — 32 bits.

The three that looked dangerous in source all turn out to be honest:

Device Why it looked risky Actual
900000.usb dwc3 dma_set_mask_and_coherent(..., DMA_BIT_MASK(64)) 32 — gated on DWC3_GHWPARAMS0_AWIDTH == 64; this core reports 32
xhci-hcd.3.auto xhci-plat.c sets 64 unconditionally 32 — xhci_gen_setup() narrows it back when HCC_64BIT_ADDR is clear
830000.ethernet sun7i-dwmac mask from dma_cap.host_dma_width 32 — consistent with "No HW DMA feature register supported" at probe

sunxi-mmc x3, sun6i-dma, sun8i-ss, ohci-platform, ehci-platform, sun4i-backend and
the rest never widen at all. There is no dma-ranges anywhere, so of_dma_configure() falls
to end = dev->coherent_dma_mask — the platform default DMA_BIT_MASK(32) — and narrows with
&=. It leaves bus_dma_limit at 0, so nothing would stop a driver widening; none does.

This also answers something for #48

display-engine / sun4i-drm sets 32 bits explicitly, so it can never be handed an address
above 4 GiB. Whatever #48's
PHYS_OFFSET fixup is, it is not high-address truncation, and reclaiming the 512 MiB cannot
have made it worse. That removes one hypothesis from #48.

The bounce path, forced and verified

With every mask at 32 bits, correctness above 4 GiB rests on SWIOTLB — which had never been
used
. io_tlb_used_hiwater was still 0 after the twenty-minute soak and after 2.5 GiB of disk
reads, because the allocator hands out low pages first and the reclaimed region is effectively
the last memory touched; 512 single allocations never landed in it.

So it was forced — walk the allocator down until a page above 4 GiB appears, then map it against
a real 32-bit device:

dmabounce: 1c0f000.mmc dma_mask=0xffffffff
dmabounce: held 128830 pages (503 MiB) before reaching the high region
dmabounce: phys 0x114400000 -> dma 0x4a6b7000 : bounced below 4 GiB - CORRECT
dmabounce: data after round trip: 0xa5c31729 (intact)

0x4a6b7000 is inside the SWIOTLB reservation (0x49eb7000-0x4deb7000), so the bounce
happened, the device got an address it can reach, and the data survived.

That the region is used last is worth keeping: the reclaimed 512 MiB only comes into play
under memory pressure — which is exactly when bouncing begins. The path that sat untested for
the first hour of uptime is the one that matters under load, and it works.

Done when

  • every DMA-capable driver's mask confirmed correct — all 32-bit, measured not inferred
  • sun4i-drm specifically understood — explicitly 32-bit, so #48 is unrelated to this
  • written down — 57-4gb-reclaimed.md; modules kept as
    probe/dmamask.c and probe/dmabounce.c

The 4 GiB default is justified rather than merely working.

## Audit done — clean. Closing. Masks read from **kernel context** with a throwaway module over `bus_for_each_dev(&platform_bus_type)`, rather than inferred from source. **Every DMA-capable device on this board reports `0x00000000ffffffff` — 32 bits.** The three that looked dangerous in source all turn out to be honest: | Device | Why it looked risky | Actual | |---|---|---| | `900000.usb` `dwc3` | `dma_set_mask_and_coherent(..., DMA_BIT_MASK(64))` | **32** — gated on `DWC3_GHWPARAMS0_AWIDTH == 64`; this core reports 32 | | `xhci-hcd.3.auto` | `xhci-plat.c` sets 64 **unconditionally** | **32** — `xhci_gen_setup()` narrows it back when `HCC_64BIT_ADDR` is clear | | `830000.ethernet` `sun7i-dwmac` | mask from `dma_cap.host_dma_width` | **32** — consistent with "No HW DMA feature register supported" at probe | `sunxi-mmc` x3, `sun6i-dma`, `sun8i-ss`, `ohci-platform`, `ehci-platform`, `sun4i-backend` and the rest never widen at all. There is **no `dma-ranges` anywhere**, so `of_dma_configure()` falls to `end = dev->coherent_dma_mask` — the platform default `DMA_BIT_MASK(32)` — and narrows with `&=`. It leaves `bus_dma_limit` at 0, so nothing would *stop* a driver widening; none does. ### This also answers something for #48 `display-engine` / `sun4i-drm` sets **32 bits explicitly**, so it can never be handed an address above 4 GiB. Whatever [#48](https://source.tiagoagueda.com/tiagoagueda/a80/issues/48)'s `PHYS_OFFSET` fixup is, **it is not high-address truncation**, and reclaiming the 512 MiB cannot have made it worse. That removes one hypothesis from #48. ### The bounce path, forced and verified With every mask at 32 bits, correctness above 4 GiB rests on SWIOTLB — which had **never been used**. `io_tlb_used_hiwater` was still 0 after the twenty-minute soak and after 2.5 GiB of disk reads, because the allocator hands out low pages first and the reclaimed region is effectively the *last* memory touched; 512 single allocations never landed in it. So it was forced — walk the allocator down until a page above 4 GiB appears, then map it against a real 32-bit device: ``` dmabounce: 1c0f000.mmc dma_mask=0xffffffff dmabounce: held 128830 pages (503 MiB) before reaching the high region dmabounce: phys 0x114400000 -> dma 0x4a6b7000 : bounced below 4 GiB - CORRECT dmabounce: data after round trip: 0xa5c31729 (intact) ``` `0x4a6b7000` is inside the SWIOTLB reservation (`0x49eb7000-0x4deb7000`), so the bounce happened, the device got an address it can reach, and the data survived. **That the region is used last is worth keeping**: the reclaimed 512 MiB only comes into play under memory pressure — which is exactly when bouncing begins. The path that sat untested for the first hour of uptime is the one that matters under load, and it works. ### Done when - [x] every DMA-capable driver's mask confirmed correct — all 32-bit, measured not inferred - [x] `sun4i-drm` specifically understood — explicitly 32-bit, so #48 is unrelated to this - [x] written down — [57-4gb-reclaimed.md](57-4gb-reclaimed.md); modules kept as `probe/dmamask.c` and `probe/dmabounce.c` The 4 GiB default is justified rather than merely working.
Sign in to join this conversation.
No description provided.