Audit dma_mask on every DMA-capable driver, now that RAM extends above 4 GiB #63
Labels
No labels
blocked-physical
cleanup
hardware
infra
kernel
P1-critical
P2-high
P3-normal
P4-later
reliability
security
upstream
wontfix
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
tiagoagueda/a80#63
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Since 2026-08-30 the board boots with 4 GiB reachable by default — see
#61 and
57-4gb-reclaimed.md. Physical memory now runs to
0x1_1FFF_FFFF, soany driver can be handed a page whose address does not fit in 32 bits.
This was listed inside #61 as a precondition and never done. It has changed character: it is no
longer a question about a proposed configuration, it guards one that is live.
Why it matters
A device with a 32-bit
dma_maskis fine if the mask is declared honestly — the DMA APIbounces the buffer through SWIOTLB. The failure mode is a driver that
dma_addr_tin au32, or writes only the low word into a descriptorEither truncates silently. There is no error, no warning, and the write lands somewhere else in
DRAM. That is the worst class of bug this board can have, and
#53 is a standing reminder of how
long silent memory corruption can go unexplained.
What already protects us
cma=64M@0x20000000-0x100000000caps the CMA pool below 4 GiB. Without it the pool landedat
0x1_0000_0000and OHCI died and the eMMC never appeared — that is what made this concreterather than theoretical.
0x49eb7000) and bounces correctly declared 32-bit masters.address" (p.218).
Evidence so far, which is encouraging but not an audit
Everything probed and survived a twenty-minute soak at load ~4 with 3600 MiB resident: eth0,
wlan0, the eMMC, 8 USB devices with zero "HC died", Bluetooth, 206.5 GiB written and read
back with 0 mismatches. But "it worked for twenty minutes" is not the same as "the mask is
right", because a truncating driver only misbehaves once it is actually given a high page.
Worth checking, roughly in order of risk
sun4i-drmPHYS_OFFSETfixup, and the layer address is split across two registers as((h4add << 32) | l32add) >> 3. That is exactly the shape of an address-truncation bug, and it now has memory above 4 GiB to truncatedwc3/xhcisun6i-dmasun7i-dwmacsunxi-mmcsunxi_ariscmailboxHow
grepfordma_set_mask,dma_set_coherent_mask,dma_coerce_mask_and_coherentin each andcheck the declared width against what the hardware can address
dma_addr_tnarrowed tou32, and for descriptor writes that take only the low worddma-ranges— there is none anywhere in this SoC's tree today,which #48 already flags as unexplained
driver, rather than waiting for the page allocator to do it by chance
Done when
constraint added
sun4i-drmspecifically is understood well enough to say whether #48's fixup is relatedAudit done — clean. Closing.
Masks read from kernel context with a throwaway module over
bus_for_each_dev(&platform_bus_type), rather than inferred from source. Every DMA-capabledevice on this board reports
0x00000000ffffffff— 32 bits.The three that looked dangerous in source all turn out to be honest:
900000.usbdwc3dma_set_mask_and_coherent(..., DMA_BIT_MASK(64))DWC3_GHWPARAMS0_AWIDTH == 64; this core reports 32xhci-hcd.3.autoxhci-plat.csets 64 unconditionallyxhci_gen_setup()narrows it back whenHCC_64BIT_ADDRis clear830000.ethernetsun7i-dwmacdma_cap.host_dma_widthsunxi-mmcx3,sun6i-dma,sun8i-ss,ohci-platform,ehci-platform,sun4i-backendandthe rest never widen at all. There is no
dma-rangesanywhere, soof_dma_configure()fallsto
end = dev->coherent_dma_mask— the platform defaultDMA_BIT_MASK(32)— and narrows with&=. It leavesbus_dma_limitat 0, so nothing would stop a driver widening; none does.This also answers something for #48
display-engine/sun4i-drmsets 32 bits explicitly, so it can never be handed an addressabove 4 GiB. Whatever #48's
PHYS_OFFSETfixup is, it is not high-address truncation, and reclaiming the 512 MiB cannothave made it worse. That removes one hypothesis from #48.
The bounce path, forced and verified
With every mask at 32 bits, correctness above 4 GiB rests on SWIOTLB — which had never been
used.
io_tlb_used_hiwaterwas still 0 after the twenty-minute soak and after 2.5 GiB of diskreads, because the allocator hands out low pages first and the reclaimed region is effectively
the last memory touched; 512 single allocations never landed in it.
So it was forced — walk the allocator down until a page above 4 GiB appears, then map it against
a real 32-bit device:
0x4a6b7000is inside the SWIOTLB reservation (0x49eb7000-0x4deb7000), so the bouncehappened, the device got an address it can reach, and the data survived.
That the region is used last is worth keeping: the reclaimed 512 MiB only comes into play
under memory pressure — which is exactly when bouncing begins. The path that sat untested for
the first hour of uptime is the one that matters under load, and it works.
Done when
sun4i-drmspecifically understood — explicitly 32-bit, so #48 is unrelated to thisprobe/dmamask.candprobe/dmabounce.cThe 4 GiB default is justified rather than merely working.