Rescue SD has drifted and backups are stale #9

Closed
opened 2026-08-27 23:22:24 +00:00 by tiagoagueda · 0 comments
Owner

Resolved 2026-08-28 — see the new section at the end of
36-rescue-sd.md.

Re-synced from the eMMC while the board was running from the card, so the eMMC stayed available
as a fallback throughout.

before after
U-Boot 27 Aug 09:43 — hardcoded image load, no menu 27 Aug 20:37, extlinux menu
Kernel 27 Aug 17:07 current, with uImage.known-good beside it
/boot 9 kernels + 9 dtbs of leftovers tidied into attic/
Config items 1 of 8 8 of 8
BT firmware absent BCM4335C0.hcd + nvram
gmac-rebind.service enabled disabled — in-kernel MDIO settle fix replaced it
/etc/network/interfaces full lo only, as on the eMMC

The card had no fallback, which mattered more than the drift

Its U-Boot predated the extlinux bootcmd, so it loaded the kernel with a hardcoded legacy
image load and ignored extlinux.conf entirely. One bad write to /boot/uImage and the
recovery path would not boot at all.
The eMMC's bootloader was copied over (8 KiB, 512 KiB,
verified by readback; the partition starts at 1 MiB so the write cannot reach the filesystem),
and the card now presents its own menu:

U-Boot 2026.10-rc3-dirty (Aug 27 2026 - 20:37:35 +0000)
Retrieving file: /boot/extlinux/extlinux.conf
Draco A80 - RESCUE CARD
Enter choice: 1:	Rescue kernel  (uImage)

⚠️ Copying units is not copying their dependencies

The first pass left the card with two failed units on boot — ir-protocols with
203/EXEC, bt-public-addr with a timeout — because ir-keytable, ir-ctl, btmgmt and
hciconfig were not installed there. The hardware was fine and every md5 matched; the result
was still broken. Fixed with bluez, v4l-utils and ir-keytable (a separate package on
Debian 13).

A rescue system that boots with failures is worse than one without those services at all.

Verified across two cold boots

hostname        a80-rescue        root  /dev/mmcblk0p1
failed units    none
ir-protocols    active            bt-public-addr  active
BD address      04:E6:76:D0:67:B2 — same derived address as the eMMC
network         eth0 + wlan0 up, single DHCP client
watchdog        15 s

Backups kept: old kernel as uImage.rescue-20260828-0942 on the card, old boot area as
sd-bootarea-20260828-0954.img.gz on ouranos (d9e1c8591653e4d832544a5371e74428) and recorded
in ARCHIVE.md — now the only copy of that bootloader build.

Tools committed: tools/sync-rescue-sd.sh, tools/flash-sd-uboot.sh.

⚠️ Still open from this issue's original scope: backups are stale. The rootfs tarballs in
ouranos:~/a80/backup/ predate two days of changes. That is worth its own issue.

**Resolved 2026-08-28** — see the new section at the end of [36-rescue-sd.md](36-rescue-sd.md). Re-synced from the eMMC while the board was running from the card, so the eMMC stayed available as a fallback throughout. | | before | after | |---|---|---| | U-Boot | 27 Aug **09:43** — hardcoded image load, no menu | 27 Aug 20:37, **extlinux menu** | | Kernel | 27 Aug 17:07 | current, with `uImage.known-good` beside it | | `/boot` | 9 kernels + 9 dtbs of leftovers | tidied into `attic/` | | Config items | 1 of 8 | **8 of 8** | | BT firmware | absent | `BCM4335C0.hcd` + nvram | | `gmac-rebind.service` | enabled | **disabled** — in-kernel MDIO settle fix replaced it | | `/etc/network/interfaces` | full | `lo` only, as on the eMMC | ### The card had no fallback, which mattered more than the drift Its U-Boot predated the extlinux `bootcmd`, so it loaded the kernel with a hardcoded legacy image load and ignored `extlinux.conf` entirely. **One bad write to `/boot/uImage` and the recovery path would not boot at all.** The eMMC's bootloader was copied over (8 KiB, 512 KiB, verified by readback; the partition starts at 1 MiB so the write cannot reach the filesystem), and the card now presents its own menu: ``` U-Boot 2026.10-rc3-dirty (Aug 27 2026 - 20:37:35 +0000) Retrieving file: /boot/extlinux/extlinux.conf Draco A80 - RESCUE CARD Enter choice: 1: Rescue kernel (uImage) ``` ### ⚠️ Copying units is not copying their dependencies The first pass left the card with **two failed units on boot** — `ir-protocols` with `203/EXEC`, `bt-public-addr` with a timeout — because `ir-keytable`, `ir-ctl`, `btmgmt` and `hciconfig` were not installed there. The hardware was fine and every md5 matched; the result was still broken. Fixed with `bluez`, `v4l-utils` and `ir-keytable` (a separate package on Debian 13). A rescue system that boots with failures is worse than one without those services at all. ### Verified across two cold boots ``` hostname a80-rescue root /dev/mmcblk0p1 failed units none ir-protocols active bt-public-addr active BD address 04:E6:76:D0:67:B2 — same derived address as the eMMC network eth0 + wlan0 up, single DHCP client watchdog 15 s ``` Backups kept: old kernel as `uImage.rescue-20260828-0942` on the card, old boot area as `sd-bootarea-20260828-0954.img.gz` on ouranos (`d9e1c8591653e4d832544a5371e74428`) and recorded in [ARCHIVE.md](ARCHIVE.md) — now the only copy of that bootloader build. Tools committed: `tools/sync-rescue-sd.sh`, `tools/flash-sd-uboot.sh`. ⚠️ Still open from this issue's original scope: **backups are stale**. The rootfs tarballs in `ouranos:~/a80/backup/` predate two days of changes. That is worth its own issue.
tiagoagueda 2026-08-27 23:22:24 +00:00
Sign in to join this conversation.
No description provided.