u-boot: SYS_BOOTM_LEN is 8 MiB and the kernel was 170 KiB from it #72

Closed
opened 2026-08-30 12:10:39 +00:00 by tiagoagueda · 1 comment
Owner

The board has been running within 170 KiB of an undiagnosable boot failure for its
entire life, and nothing anywhere said so. Found on 2026-08-30 while enabling Docker; see
58-docker.md.

What happens

CONFIG_SYS_BOOTM_LEN=0x800000 in configs/draco-extlinux_defconfig - exactly 8 MiB. It is
not a considered value: it is the generic fallback default in boot/Kconfig for anything that is
not PPC/ARM64/RISCV/X86. arm64 boards get 0x8000000 (128 MiB).

The pre-Docker kernel was 7.83 MiB. Adding the Docker options took it to 8.72 MiB, and:

## Booting kernel from Legacy Image at 22000000 ...
   Data Size:    9138504 Bytes = 8.7 MiB
   Load Address: 20008000
   Verifying Checksum ... OK
   Loading Kernel Image to 20008000
Image too large: increase CONFIG_SYS_BOOTM_LEN
Must RESET board to recover

Starting kernel ... is never printed. The kernel is never entered. From userspace afterwards
this is indistinguishable from a kernel that panics on its first instruction - and three separate
pieces of evidence on the running board (/proc/cmdline, bootcount, journalctl --list-boots
timestamps) all actively pointed away from the real cause. Only the serial capture identified
it, and only because one happened to be running from an unrelated session.

Current state: worked around, not fixed

bootcmd now sets kernel_addr_r=0x20007fc0 instead of 0x22000000. A legacy uImage header is
64 bytes, so the payload lands exactly on the kernel load address 0x20008000, and
bootm_load_os() takes this branch in boot/image.c:

case IH_COMP_NONE:
        ret = 0;
        if (load == image_start)
                break;            /* no copy - and no length check */

skipping the ceiling entirely. Confirmed on the console: XIP Kernel Image to 20008000.

This is an environment change only, which is why it was the right thing to try first - no
bootloader flash, and altbootcmd kept its own hardcoded addresses throughout, so the recovery
path was never at risk while testing the mechanism that does the recovering.

Why it still needs a real fix

1. It is a trick, and it is silent when it stops working. It depends on kernel_addr_r being
exactly 64 bytes below the load address. Anyone who changes LOADADDR, the image type, or
kernel_addr_r re-arms the original trap, and the symptom will again look like a kernel panic.

2. altbootcmd did not get the dodge. It hardcodes 0x22000000 and takes the copy path, so
a known-good kernel over 8 MiB is one the recovery path cannot load. A fallback that resets in
a loop is worse than no fallback, because bootlimit has already been spent by the time it runs.
tools/deploy-kernel.sh now refuses --known-good for an image that size unless altbootcmd has
been updated - but that is a guard rail, not a fix.

3. 🔴 The rescue SD card is exposed and the card is the last resort. Its environment is blank,
so it boots from its compiled-in bootcmd, which loads at a copy address. It carries a small
kernel today, so it is fine today. A sync that copies the current 8.42 MiB kernel onto the
card would make the card unbootable
, and discovering that requires physical access to the one
thing that exists to avoid needing physical access. This must be checked at the next sync -
see #64, which is already about that card's compiled-in bootcmd.

The fix

CONFIG_SYS_BOOTM_LEN=0x4000000 (64 MiB) in configs/draco-extlinux_defconfig. There is no
reason for 8 MiB on a board with 4 GiB of RAM; bootm_size is already 0xa000000.

Then, in the same pass:

  • rebuild and flash u-boot to the eMMC
  • revert kernel_addr_r to 0x22000000 so the boot path is ordinary again, or keep the XIP
    address deliberately and write down why
  • update altbootcmd, or confirm it no longer needs it
  • reflash or re-env the rescue card so it is not the only medium still carrying the trap

Why it is not done yet

Flashing u-boot to the eMMC is the one operation on this board that can strand it, and it is only
safe to attempt with the rescue card to hand. It should be done with the card in the room, not
remotely. Everything above works without it in the meantime.

Done when

  • SYS_BOOTM_LEN is large on both media
  • a kernel well over 8 MiB boots through the ordinary copy path, verified on the console
  • altbootcmd is verified to load an oversized known-good, by actually exceeding bootlimit
    once and watching it recover
  • the guard in deploy-kernel.sh is updated or removed with a note saying why
The board has been running within **170 KiB** of an undiagnosable boot failure for its entire life, and nothing anywhere said so. Found on 2026-08-30 while enabling Docker; see `58-docker.md`. ## What happens `CONFIG_SYS_BOOTM_LEN=0x800000` in `configs/draco-extlinux_defconfig` - **exactly 8 MiB**. It is not a considered value: it is the generic fallback default in `boot/Kconfig` for anything that is not PPC/ARM64/RISCV/X86. arm64 boards get `0x8000000` (128 MiB). The pre-Docker kernel was 7.83 MiB. Adding the Docker options took it to 8.72 MiB, and: ``` ## Booting kernel from Legacy Image at 22000000 ... Data Size: 9138504 Bytes = 8.7 MiB Load Address: 20008000 Verifying Checksum ... OK Loading Kernel Image to 20008000 Image too large: increase CONFIG_SYS_BOOTM_LEN Must RESET board to recover ``` **`Starting kernel ...` is never printed.** The kernel is never entered. From userspace afterwards this is indistinguishable from a kernel that panics on its first instruction - and three separate pieces of evidence on the running board (`/proc/cmdline`, `bootcount`, `journalctl --list-boots` timestamps) all actively pointed *away* from the real cause. Only the serial capture identified it, and only because one happened to be running from an unrelated session. ## Current state: worked around, not fixed `bootcmd` now sets `kernel_addr_r=0x20007fc0` instead of `0x22000000`. A legacy uImage header is 64 bytes, so the payload lands exactly on the kernel load address `0x20008000`, and `bootm_load_os()` takes this branch in `boot/image.c`: ```c case IH_COMP_NONE: ret = 0; if (load == image_start) break; /* no copy - and no length check */ ``` skipping the ceiling entirely. Confirmed on the console: `XIP Kernel Image to 20008000`. This is an **environment change only**, which is why it was the right thing to try first - no bootloader flash, and `altbootcmd` kept its own hardcoded addresses throughout, so the recovery path was never at risk while testing the mechanism that does the recovering. ## Why it still needs a real fix **1. It is a trick, and it is silent when it stops working.** It depends on `kernel_addr_r` being exactly 64 bytes below the load address. Anyone who changes `LOADADDR`, the image type, or `kernel_addr_r` re-arms the original trap, and the symptom will again look like a kernel panic. **2. `altbootcmd` did not get the dodge.** It hardcodes `0x22000000` and takes the copy path, so **a known-good kernel over 8 MiB is one the recovery path cannot load**. A fallback that resets in a loop is worse than no fallback, because `bootlimit` has already been spent by the time it runs. `tools/deploy-kernel.sh` now refuses `--known-good` for an image that size unless `altbootcmd` has been updated - but that is a guard rail, not a fix. **3. 🔴 The rescue SD card is exposed and the card is the last resort.** Its environment is blank, so it boots from its compiled-in `bootcmd`, which loads at a copy address. It carries a small kernel today, so it is fine *today*. **A sync that copies the current 8.42 MiB kernel onto the card would make the card unbootable**, and discovering that requires physical access to the one thing that exists to avoid needing physical access. This must be checked at the next sync - see #64, which is already about that card's compiled-in bootcmd. ## The fix `CONFIG_SYS_BOOTM_LEN=0x4000000` (64 MiB) in `configs/draco-extlinux_defconfig`. There is no reason for 8 MiB on a board with 4 GiB of RAM; `bootm_size` is already `0xa000000`. Then, in the same pass: - rebuild and flash u-boot to the eMMC - revert `kernel_addr_r` to `0x22000000` so the boot path is ordinary again, or keep the XIP address deliberately and write down why - update `altbootcmd`, or confirm it no longer needs it - reflash or re-env the rescue card so it is not the only medium still carrying the trap ## Why it is not done yet Flashing u-boot to the eMMC is the one operation on this board that can strand it, and it is only safe to attempt with the rescue card to hand. It should be done **with the card in the room**, not remotely. Everything above works without it in the meantime. ## Done when - `SYS_BOOTM_LEN` is large on both media - a kernel well over 8 MiB boots through the *ordinary* copy path, verified on the console - `altbootcmd` is verified to load an oversized known-good, by actually exceeding `bootlimit` once and watching it recover - the guard in `deploy-kernel.sh` is updated or removed with a note saying why
Author
Owner

Done, on both media, verified rather than reasoned about.

What was changed

2026.10-rc3-00011-g103b8b sets CONFIG_SYS_BOOTM_LEN=0x2000000 (32 MiB), flashed to the eMMC
and the rescue card. 32 MiB, not the 64 MiB proposed above: the copy destination is the kernel
load address 0x20008000 and fdt_addr_r is 0x23000000, only ~47.9 MiB above it, so a larger
bound would permit an image that silently overwrites the loaded device tree. 32 MiB is the clean
power of two under that, and ~3.8x the current kernel.

Sequencing, which is the part worth keeping

The flash does not touch the environment at 768 KiB, so the XIP workaround survived it. The board
booted on the new bootloader before anything depended on the new limit. Only then was
kernel_addr_r reverted to 0x22000000 - and that step was covered by bootlimit/altbootcmd.

So the only step with no software fallback was "does the SPL run at all", and it was reduced to a
one-line Kconfig delta over a binary already proven on this board, with an md5-verified off-box
backup taken first. That backup turned out to matter: neither existing off-box backup matched
the running bootloader
- the newest was -00005 while the board was running -00009. The
bootloader keeping this board alive had never been copied off it.

Done-when, item by item

  • ✅ SYS_BOOTM_LEN large on both media - eMMC and card both -00011-g103b8b, md5 906cc04b
  • ✅ a kernel well over 8 MiB boots through the ordinary copy path, verified on console, and
    from a cold boot rather than a warm reboot:
Trying to boot from MMC2
U-Boot 2026.10-rc3-00011-g103b8b4f43f7
## Booting kernel from Legacy Image at 22000000 ...
   Data Size:    8828792 Bytes = 8.4 MiB
   Loading Kernel Image to 20008000
Starting kernel ...
  • ✅ altbootcmd verified to load an oversized known-good by actually exceeding bootlimit
    once
    , with the 8.42 MiB kernel promoted to known-good:
*** bootlimit exceeded - booting the known-good kernel ***
   Data Size:    8828792 Bytes = 8.4 MiB
Starting kernel ...

An hour earlier that would have reset in a loop.

  • ✅ the guard in deploy-kernel.sh reworked rather than removed. It no longer hardcodes 8 MiB:
    it reads SYS_BOOTM_LEN from the u-boot defconfig at the commit the board reports running,
    and falls back to the conservative default for a dirty or unreadable version - because the
    tree's HEAD is wrong precisely in the window where a new bootloader is built but not yet
    flashed, which is when someone promotes a known-good. Confirmed it resolves 0x2000000 for the
    running commit and 0x800000 for the one before it.
  • ✅ the card, which was the sharp end of this. Flashed, then booted to confirm it still works
    rather than assumed - came up as a80-rescue on /dev/mmcblk0p1 via Trying to boot from MMC1. Handed back using the documented emmc maintenance entry so its rootfs was never pulled
    live. sync-rescue-sd.sh now identifies the card's bootloader from its own boot area and
    accepts >= -00011 or an XIP env, failing closed on anything it cannot identify.

Two things recorded rather than fixed

⚠️ The bootloader vintage split is gone. The card deliberately ran an older bootloader so a
bad eMMC flash still had a known-different fallback. Both media now run the identical binary, so
a latent bug in -00011 takes out the eMMC and its rescue together. Chosen knowingly over the
env-only card fix that would have kept the split; noted in 58-docker.md so it is not later
mistaken for an oversight.

⚠️ One caveat on the card: it booted its own 7.8 MiB kernel, which does not itself exercise
the raised limit. That rests on the eMMC having proven the byte-identical binary loads 8.42 MiB
through the copy path. Sound, but inference rather than a direct test - it will be settled the
first time the card is synced with a current kernel.

Commits ee9a0dc and 07bf715. Write-up in 58-docker.md.

Done, on both media, verified rather than reasoned about. ## What was changed `2026.10-rc3-00011-g103b8b` sets `CONFIG_SYS_BOOTM_LEN=0x2000000` (32 MiB), flashed to the eMMC and the rescue card. **32 MiB, not the 64 MiB proposed above:** the copy destination is the kernel load address `0x20008000` and `fdt_addr_r` is `0x23000000`, only ~47.9 MiB above it, so a larger bound would permit an image that silently overwrites the loaded device tree. 32 MiB is the clean power of two under that, and ~3.8x the current kernel. ## Sequencing, which is the part worth keeping The flash does not touch the environment at 768 KiB, so the XIP workaround survived it. The board booted on the new bootloader **before anything depended on the new limit**. Only then was `kernel_addr_r` reverted to `0x22000000` - and that step was covered by `bootlimit`/`altbootcmd`. So the only step with no software fallback was "does the SPL run at all", and it was reduced to a one-line Kconfig delta over a binary already proven on this board, with an md5-verified off-box backup taken first. That backup turned out to matter: **neither existing off-box backup matched the running bootloader** - the newest was `-00005` while the board was running `-00009`. The bootloader keeping this board alive had never been copied off it. ## Done-when, item by item - ✅ `SYS_BOOTM_LEN` large on both media - eMMC and card both `-00011-g103b8b`, md5 `906cc04b` - ✅ a kernel well over 8 MiB boots through the **ordinary copy path**, verified on console, and from a **cold boot** rather than a warm reboot: ``` Trying to boot from MMC2 U-Boot 2026.10-rc3-00011-g103b8b4f43f7 ## Booting kernel from Legacy Image at 22000000 ... Data Size: 8828792 Bytes = 8.4 MiB Loading Kernel Image to 20008000 Starting kernel ... ``` - ✅ `altbootcmd` verified to load an oversized known-good **by actually exceeding `bootlimit` once**, with the 8.42 MiB kernel promoted to known-good: ``` *** bootlimit exceeded - booting the known-good kernel *** Data Size: 8828792 Bytes = 8.4 MiB Starting kernel ... ``` An hour earlier that would have reset in a loop. - ✅ the guard in `deploy-kernel.sh` reworked rather than removed. It no longer hardcodes 8 MiB: it reads `SYS_BOOTM_LEN` from the u-boot defconfig **at the commit the board reports running**, and falls back to the conservative default for a dirty or unreadable version - because the tree's HEAD is wrong precisely in the window where a new bootloader is built but not yet flashed, which is when someone promotes a known-good. Confirmed it resolves `0x2000000` for the running commit and `0x800000` for the one before it. - ✅ the card, which was the sharp end of this. Flashed, then **booted to confirm it still works** rather than assumed - came up as `a80-rescue` on `/dev/mmcblk0p1` via `Trying to boot from MMC1`. Handed back using the documented `emmc` maintenance entry so its rootfs was never pulled live. `sync-rescue-sd.sh` now identifies the card's bootloader from its own boot area and accepts `>= -00011` or an XIP env, failing closed on anything it cannot identify. ## Two things recorded rather than fixed ⚠️ **The bootloader vintage split is gone.** The card deliberately ran an older bootloader so a bad eMMC flash still had a known-different fallback. Both media now run the identical binary, so a latent bug in `-00011` takes out the eMMC and its rescue together. Chosen knowingly over the env-only card fix that would have kept the split; noted in `58-docker.md` so it is not later mistaken for an oversight. ⚠️ **One caveat on the card:** it booted its own 7.8 MiB kernel, which does not itself exercise the raised limit. That rests on the eMMC having proven the byte-identical binary loads 8.42 MiB through the copy path. Sound, but inference rather than a direct test - it will be settled the first time the card is synced with a current kernel. Commits `ee9a0dc` and `07bf715`. Write-up in `58-docker.md`.
Sign in to join this conversation.
No description provided.