Rescue card: close the A15 guard gap in its compiled-in bootcmd, at the next sync #64

Closed
opened 2026-08-30 11:16:43 +00:00 by tiagoagueda · 2 comments
Owner

Deferred deliberately on 2026-08-30, to be picked up the next time
tools/sync-rescue-sd.sh is run
— it needs the card inserted and the board booted from the
eMMC, which is exactly the state a sync already requires.

The gap

The card's boot menu is fully correct: all five entries carry ro and maxcpus=4 a15=off, and
sync-rescue-sd.sh now verifies every entry rather than just finding one.

What is not covered is the layer beneath it. The card runs U-Boot
2026.10-rc3-00002-g6a5618, its env area at 768 KiB is blank, so it boots from its
compiled-in bootcmd — and that build predates the guard fix. Measured on the card:

card env area at 768 KiB : BLANK -> runs the compiled-in bootcmd
compiled-in fallback     : [FAIL] mmcblk0p1   (no a15=off)
                           [FAIL] mmcblk1p1   (no a15=off)

The eMMC had the same gap and it is now fixed there, in both the stored bootcmd and
altbootcmd.

Severity is genuinely third-order. That fallback only runs if sysboot cannot read
/boot/extlinux/extlinux.conf on either medium. If both are unreadable the card is already
in deep trouble. It is filed because it is the same class of defect that was just found on the
eMMC, not because it is likely.

Why it was not fixed on the spot

Both available fixes add risk to the medium whose entire job is to be reliable:

  • Reflash the card's U-Boot with the eMMC's -00009, whose compiled-in bootcmd is guarded.
    But the two vintages are deliberately different: the card is the last resort if an eMMC
    bootloader flash goes wrong, and making them identical removes that. See
    [a80-lab-infrastructure] and tools/sync-rescue-sd.sh, which has never touched the
    bootloader on purpose.
  • Write a stored env onto the card with a guarded bootcmd. Preserves the vintage split,
    but the card's blank env is itself a simplicity property — it runs on built-in defaults with
    nothing to corrupt — and a malformed env write would break the rescue medium.

Options, to decide at the sync

  1. Do nothing, keep it recorded. Defensible: the exposure needs both extlinux files
    unreadable.
  2. Write a minimal env setting only a guarded bootcmd. fw_env.config on the card already
    correctly names /dev/mmcblk0, so fw_setenv from the card would land in the right place.
    Verify by reading the env back before rebooting.
  3. Reflash the card's bootloader and accept a single vintage, ideally only once the eMMC's
    -00009 has more running time behind it.

Not a bug, so nobody "fixes" it

The eMMC's test entry has no a15=off by design — it is the labelled all-8-cores entry
and exists to bring the A15 up for #53 work. Both provision.sh and sync-rescue-sd.sh
exclude it from their checks by name.

Done when

  • the card's compiled-in fallback either carries the guard, or the decision not to close it
    is written down with its reasoning
  • whichever route is taken, sync-rescue-sd.sh checks it, so it cannot drift back silently
Deferred deliberately on 2026-08-30, to be picked up **the next time `tools/sync-rescue-sd.sh` is run** — it needs the card inserted and the board booted from the eMMC, which is exactly the state a sync already requires. ## The gap The card's boot menu is fully correct: all five entries carry `ro` and `maxcpus=4 a15=off`, and `sync-rescue-sd.sh` now verifies **every** entry rather than just finding one. What is not covered is the layer beneath it. The card runs U-Boot `2026.10-rc3-00002-g6a5618`, its env area at 768 KiB is **blank**, so it boots from its **compiled-in** `bootcmd` — and that build predates the guard fix. Measured on the card: ``` card env area at 768 KiB : BLANK -> runs the compiled-in bootcmd compiled-in fallback : [FAIL] mmcblk0p1 (no a15=off) [FAIL] mmcblk1p1 (no a15=off) ``` The eMMC had the same gap and it is now fixed there, in both the stored `bootcmd` and `altbootcmd`. **Severity is genuinely third-order.** That fallback only runs if `sysboot` cannot read `/boot/extlinux/extlinux.conf` on *either* medium. If both are unreadable the card is already in deep trouble. It is filed because it is the same class of defect that was just found on the eMMC, not because it is likely. ## Why it was not fixed on the spot Both available fixes add risk to the medium whose entire job is to be reliable: - **Reflash the card's U-Boot** with the eMMC's `-00009`, whose compiled-in bootcmd is guarded. But the two vintages are deliberately different: the card is the last resort if an eMMC bootloader flash goes wrong, and making them identical removes that. See [a80-lab-infrastructure] and `tools/sync-rescue-sd.sh`, which has never touched the bootloader on purpose. - **Write a stored env onto the card** with a guarded `bootcmd`. Preserves the vintage split, but the card's blank env is itself a simplicity property — it runs on built-in defaults with nothing to corrupt — and a malformed env write would break the rescue medium. ## Options, to decide at the sync 1. **Do nothing, keep it recorded.** Defensible: the exposure needs both extlinux files unreadable. 2. **Write a minimal env** setting only a guarded `bootcmd`. `fw_env.config` on the card already correctly names `/dev/mmcblk0`, so `fw_setenv` from the card would land in the right place. Verify by reading the env back before rebooting. 3. **Reflash the card's bootloader** and accept a single vintage, ideally only once the eMMC's `-00009` has more running time behind it. ## Not a bug, so nobody "fixes" it The eMMC's `test` entry has no `a15=off` **by design** — it is the labelled all-8-cores entry and exists to bring the A15 up for #53 work. Both `provision.sh` and `sync-rescue-sd.sh` exclude it from their checks by name. ## Done when - [ ] the card's compiled-in fallback either carries the guard, or the decision not to close it is written down with its reasoning - [ ] whichever route is taken, `sync-rescue-sd.sh` checks it, so it cannot drift back silently
Author
Owner

There is now a second, worse reason to deal with this card's compiled-in bootcmd, found
while enabling Docker: #72.

The card's env is blank, so it boots from the compiled-in bootcmd - which is also why that
bootcmd loads the kernel at a copy address. u-boot's CONFIG_SYS_BOOTM_LEN is exactly 8 MiB
and the current eMMC kernel is 8.42 MiB, so the eMMC now dodges the ceiling by loading at
0x20007fc0 (payload lands on the load address, bootm takes its XIP branch). The card gets no
such dodge.

So the exposure is no longer only the third-order a15=off gap this issue was opened for.
tools/sync-rescue-sd.sh line 103 copies /boot/uImage onto the card, and doing that today
would hand the card a kernel its own bootloader cannot start - discovered only by needing the
card, which is precisely the trip it exists to prevent.

That is now blocked in the tool: the sync refuses to run if the kernel is over the ceiling and
the card's boot area does not advertise the XIP address. Tested both ways against a real boot
area.

It also collapses the three options in the original post into one. "Leave it, recorded" is no
longer defensible, because the card is now on a divergent boot path from the eMMC, not merely an
older one. Whatever is done here should be done together with #72.

There is now a **second, worse reason** to deal with this card's compiled-in bootcmd, found while enabling Docker: **#72**. The card's env is blank, so it boots from the compiled-in bootcmd - which is also why that bootcmd loads the kernel at a *copy* address. u-boot's `CONFIG_SYS_BOOTM_LEN` is exactly 8 MiB and the current eMMC kernel is 8.42 MiB, so the eMMC now dodges the ceiling by loading at `0x20007fc0` (payload lands on the load address, bootm takes its XIP branch). **The card gets no such dodge.** So the exposure is no longer only the third-order `a15=off` gap this issue was opened for. `tools/sync-rescue-sd.sh` line 103 copies `/boot/uImage` onto the card, and doing that today would hand the card a kernel **its own bootloader cannot start** - discovered only by needing the card, which is precisely the trip it exists to prevent. That is now blocked in the tool: the sync **refuses to run** if the kernel is over the ceiling and the card's boot area does not advertise the XIP address. Tested both ways against a real boot area. It also collapses the three options in the original post into one. "Leave it, recorded" is no longer defensible, because the card is now on a divergent boot path from the eMMC, not merely an older one. Whatever is done here should be done together with #72.
Author
Owner

Closed by the #72 work, and closed properly rather than worked around.

The card was flashed with 2026.10-rc3-00011-g103b8b, and that build's compiled-in bootcmd
already carries the guard on both branches:

$ dd if=/dev/mmcblk0 bs=1K skip=8 count=600 | strings | grep -o 'maxcpus=4 a15=off'
maxcpus=4 a15=off
maxcpus=4 a15=off

one for the mmc 0:1 branch and one for mmc 1:1. So the third-order path this issue was about

  • extlinux unreadable on both media, falling through to the fixed-image fallback - now brings the
    A15 cluster up disabled, exactly like every other route.

That resolves it without needing any of the three options in the original post: no minimal env to
maintain, and the card did not have to keep a bootcmd that disagreed with the eMMC's.

Verified afterwards by booting the card: it came up as a80-rescue on /dev/mmcblk0p1.

The eMMC's test entry remains unguarded by design - it is the labelled all-8-cores entry for
#53 work, and both provision.sh and sync-rescue-sd.sh exclude it by name.

Note the tradeoff this took, recorded in #72 and 58-docker.md: the card no longer runs a
different bootloader vintage from the eMMC.

Closed by the #72 work, and closed properly rather than worked around. The card was flashed with `2026.10-rc3-00011-g103b8b`, and that build's **compiled-in** bootcmd already carries the guard on both branches: ``` $ dd if=/dev/mmcblk0 bs=1K skip=8 count=600 | strings | grep -o 'maxcpus=4 a15=off' maxcpus=4 a15=off maxcpus=4 a15=off ``` one for the `mmc 0:1` branch and one for `mmc 1:1`. So the third-order path this issue was about - extlinux unreadable on both media, falling through to the fixed-image fallback - now brings the A15 cluster up disabled, exactly like every other route. That resolves it without needing any of the three options in the original post: no minimal env to maintain, and the card did not have to keep a bootcmd that disagreed with the eMMC's. Verified afterwards by booting the card: it came up as `a80-rescue` on `/dev/mmcblk0p1`. The eMMC's `test` entry remains unguarded **by design** - it is the labelled all-8-cores entry for #53 work, and both `provision.sh` and `sync-rescue-sd.sh` exclude it by name. Note the tradeoff this took, recorded in #72 and `58-docker.md`: the card no longer runs a different bootloader vintage from the eMMC.
Sign in to join this conversation.
No description provided.