u-boot: SYS_BOOTM_LEN is 8 MiB and the kernel was 170 KiB from it #72
Labels
No labels
blocked-physical
cleanup
hardware
infra
kernel
P1-critical
P2-high
P3-normal
P4-later
reliability
security
upstream
wontfix
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
tiagoagueda/a80#72
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
The board has been running within 170 KiB of an undiagnosable boot failure for its
entire life, and nothing anywhere said so. Found on 2026-08-30 while enabling Docker; see
58-docker.md.What happens
CONFIG_SYS_BOOTM_LEN=0x800000inconfigs/draco-extlinux_defconfig- exactly 8 MiB. It isnot a considered value: it is the generic fallback default in
boot/Kconfigfor anything that isnot PPC/ARM64/RISCV/X86. arm64 boards get
0x8000000(128 MiB).The pre-Docker kernel was 7.83 MiB. Adding the Docker options took it to 8.72 MiB, and:
Starting kernel ...is never printed. The kernel is never entered. From userspace afterwardsthis is indistinguishable from a kernel that panics on its first instruction - and three separate
pieces of evidence on the running board (
/proc/cmdline,bootcount,journalctl --list-bootstimestamps) all actively pointed away from the real cause. Only the serial capture identified
it, and only because one happened to be running from an unrelated session.
Current state: worked around, not fixed
bootcmdnow setskernel_addr_r=0x20007fc0instead of0x22000000. A legacy uImage header is64 bytes, so the payload lands exactly on the kernel load address
0x20008000, andbootm_load_os()takes this branch inboot/image.c:skipping the ceiling entirely. Confirmed on the console:
XIP Kernel Image to 20008000.This is an environment change only, which is why it was the right thing to try first - no
bootloader flash, and
altbootcmdkept its own hardcoded addresses throughout, so the recoverypath was never at risk while testing the mechanism that does the recovering.
Why it still needs a real fix
1. It is a trick, and it is silent when it stops working. It depends on
kernel_addr_rbeingexactly 64 bytes below the load address. Anyone who changes
LOADADDR, the image type, orkernel_addr_rre-arms the original trap, and the symptom will again look like a kernel panic.2.
altbootcmddid not get the dodge. It hardcodes0x22000000and takes the copy path, soa known-good kernel over 8 MiB is one the recovery path cannot load. A fallback that resets in
a loop is worse than no fallback, because
bootlimithas already been spent by the time it runs.tools/deploy-kernel.shnow refuses--known-goodfor an image that size unlessaltbootcmdhasbeen updated - but that is a guard rail, not a fix.
3. 🔴 The rescue SD card is exposed and the card is the last resort. Its environment is blank,
so it boots from its compiled-in
bootcmd, which loads at a copy address. It carries a smallkernel today, so it is fine today. A sync that copies the current 8.42 MiB kernel onto the
card would make the card unbootable, and discovering that requires physical access to the one
thing that exists to avoid needing physical access. This must be checked at the next sync -
see #64, which is already about that card's compiled-in bootcmd.
The fix
CONFIG_SYS_BOOTM_LEN=0x4000000(64 MiB) inconfigs/draco-extlinux_defconfig. There is noreason for 8 MiB on a board with 4 GiB of RAM;
bootm_sizeis already0xa000000.Then, in the same pass:
kernel_addr_rto0x22000000so the boot path is ordinary again, or keep the XIPaddress deliberately and write down why
altbootcmd, or confirm it no longer needs itWhy it is not done yet
Flashing u-boot to the eMMC is the one operation on this board that can strand it, and it is only
safe to attempt with the rescue card to hand. It should be done with the card in the room, not
remotely. Everything above works without it in the meantime.
Done when
SYS_BOOTM_LENis large on both mediaaltbootcmdis verified to load an oversized known-good, by actually exceedingbootlimitonce and watching it recover
deploy-kernel.shis updated or removed with a note saying whyDone, on both media, verified rather than reasoned about.
What was changed
2026.10-rc3-00011-g103b8bsetsCONFIG_SYS_BOOTM_LEN=0x2000000(32 MiB), flashed to the eMMCand the rescue card. 32 MiB, not the 64 MiB proposed above: the copy destination is the kernel
load address
0x20008000andfdt_addr_ris0x23000000, only ~47.9 MiB above it, so a largerbound would permit an image that silently overwrites the loaded device tree. 32 MiB is the clean
power of two under that, and ~3.8x the current kernel.
Sequencing, which is the part worth keeping
The flash does not touch the environment at 768 KiB, so the XIP workaround survived it. The board
booted on the new bootloader before anything depended on the new limit. Only then was
kernel_addr_rreverted to0x22000000- and that step was covered bybootlimit/altbootcmd.So the only step with no software fallback was "does the SPL run at all", and it was reduced to a
one-line Kconfig delta over a binary already proven on this board, with an md5-verified off-box
backup taken first. That backup turned out to matter: neither existing off-box backup matched
the running bootloader - the newest was
-00005while the board was running-00009. Thebootloader keeping this board alive had never been copied off it.
Done-when, item by item
SYS_BOOTM_LENlarge on both media - eMMC and card both-00011-g103b8b, md5906cc04bfrom a cold boot rather than a warm reboot:
altbootcmdverified to load an oversized known-good by actually exceedingbootlimitonce, with the 8.42 MiB kernel promoted to known-good:
An hour earlier that would have reset in a loop.
deploy-kernel.shreworked rather than removed. It no longer hardcodes 8 MiB:it reads
SYS_BOOTM_LENfrom the u-boot defconfig at the commit the board reports running,and falls back to the conservative default for a dirty or unreadable version - because the
tree's HEAD is wrong precisely in the window where a new bootloader is built but not yet
flashed, which is when someone promotes a known-good. Confirmed it resolves
0x2000000for therunning commit and
0x800000for the one before it.rather than assumed - came up as
a80-rescueon/dev/mmcblk0p1viaTrying to boot from MMC1. Handed back using the documentedemmcmaintenance entry so its rootfs was never pulledlive.
sync-rescue-sd.shnow identifies the card's bootloader from its own boot area andaccepts
>= -00011or an XIP env, failing closed on anything it cannot identify.Two things recorded rather than fixed
⚠️ The bootloader vintage split is gone. The card deliberately ran an older bootloader so a
bad eMMC flash still had a known-different fallback. Both media now run the identical binary, so
a latent bug in
-00011takes out the eMMC and its rescue together. Chosen knowingly over theenv-only card fix that would have kept the split; noted in
58-docker.mdso it is not latermistaken for an oversight.
⚠️ One caveat on the card: it booted its own 7.8 MiB kernel, which does not itself exercise
the raised limit. That rests on the eMMC having proven the byte-identical binary loads 8.42 MiB
through the copy path. Sound, but inference rather than a direct test - it will be settled the
first time the card is synced with a current kernel.
Commits
ee9a0dcand07bf715. Write-up in58-docker.md.