U-Boot: no MAC driver is built — enable the GMAC to unlock TFTP boot #39

Closed
opened 2026-08-28 05:59:30 +00:00 by tiagoagueda · 8 comments
Owner

U-Boot has the entire network stack compiled in and no ethernet driver, which is why
there is no Net: line in the banner.

Measured

CONFIG_NET=y
CONFIG_CMD_NET=y
CONFIG_CMD_DHCP=y
CONFIG_CMD_TFTPBOOT=y
CONFIG_DM_ETH=y
# CONFIG_ETH_DESIGNWARE is not set     <- no MAC driver
# CONFIG_SUN7I_GMAC is not set

Consistent with the banner in 13-first-mainline-boot.md,
which goes straight from MMC: to the prompt, and with the older
Net: No ethernet found. in 16-ethernet-broken.md.

The hard half

CONFIG_SUNXI_NO_PMIC=y

This is the config-level statement of the finding in
34-ethernet-cold-boot.md: U-Boot omits AXP*_POWER, so
axp_init() never runs and every AXP rail sits at its OTP default until Linux binds
axp20x at ~5.24 s. The PHY's two supplies are AXP rails (axp15_sw0,
axp15_aldo3). The PHY is unpowered for U-Boot's entire lifetime.

So enabling the driver symbols is necessary and not sufficient. PMIC support in U-Boot
is the real work, and the A80's AXP806 + AXP809 pair over RSB is not covered by
mainline sunxi's existing AXP support.

Why it is worth it

  1. TFTP boot. Kernel and dtb served from ouranos means iteration with zero eMMC
    writes and zero unbootable-kernel risk — a bad kernel is one reset away rather
    than a rescue-SD trip. That is what makes the A15 cluster work (Phase 4) cheap.
  2. Fixes the cold-boot PHY problem at the source. If the rails are up from early
    boot, the PHY is ready long before the kernel's one-shot MDIO scan, and the settle
    loop in linux-stmmac-mdio-settle becomes unnecessary rather than needing to be
    reshaped for upstream (see issue #19).
  3. Netconsole becomes available as a by-product — though it is a poor substitute for
    network-attached serial, since it cannot see SPL and is unauthenticated UDP.

Suggested order

  1. ETH_DESIGNWARE + SUN7I_GMAC + PHY_REALTEK, confirm the failure mode is
    "PHY not responding" rather than "no driver".
  2. PMIC bring-up, enough to raise axp15_sw0 and axp15_aldo3.
  3. dhcp / tftpboot from the U-Boot prompt.

Acceptance

  • Net: line present in the banner with a MAC address.
  • dhcp gets a lease from 192.168.27.1.
  • A kernel + dtb boot end to end over TFTP from ouranos.
U-Boot has the entire network stack compiled in and no ethernet driver, which is why there is no `Net:` line in the banner. ## Measured ``` CONFIG_NET=y CONFIG_CMD_NET=y CONFIG_CMD_DHCP=y CONFIG_CMD_TFTPBOOT=y CONFIG_DM_ETH=y # CONFIG_ETH_DESIGNWARE is not set <- no MAC driver # CONFIG_SUN7I_GMAC is not set ``` Consistent with the banner in [13-first-mainline-boot.md](13-first-mainline-boot.md), which goes straight from `MMC:` to the prompt, and with the older `Net: No ethernet found.` in [16-ethernet-broken.md](16-ethernet-broken.md). ## The hard half ``` CONFIG_SUNXI_NO_PMIC=y ``` This is the config-level statement of the finding in [34-ethernet-cold-boot.md](34-ethernet-cold-boot.md): U-Boot omits `AXP*_POWER`, so `axp_init()` never runs and every AXP rail sits at its OTP default until Linux binds `axp20x` at ~5.24 s. The PHY's two supplies are AXP rails (`axp15_sw0`, `axp15_aldo3`). The PHY is unpowered for U-Boot's entire lifetime. So enabling the driver symbols is necessary and not sufficient. PMIC support in U-Boot is the real work, and the A80's AXP806 + AXP809 pair over RSB is not covered by mainline sunxi's existing AXP support. ## Why it is worth it 1. **TFTP boot.** Kernel and dtb served from ouranos means iteration with zero eMMC writes and zero unbootable-kernel risk — a bad kernel is one `reset` away rather than a rescue-SD trip. That is what makes the A15 cluster work (Phase 4) cheap. 2. **Fixes the cold-boot PHY problem at the source.** If the rails are up from early boot, the PHY is ready long before the kernel's one-shot MDIO scan, and the settle loop in `linux-stmmac-mdio-settle` becomes unnecessary rather than needing to be reshaped for upstream (see issue #19). 3. Netconsole becomes available as a by-product — though it is a poor substitute for network-attached serial, since it cannot see SPL and is unauthenticated UDP. ## Suggested order 1. `ETH_DESIGNWARE` + `SUN7I_GMAC` + `PHY_REALTEK`, confirm the failure mode is "PHY not responding" rather than "no driver". 2. PMIC bring-up, enough to raise `axp15_sw0` and `axp15_aldo3`. 3. `dhcp` / `tftpboot` from the U-Boot prompt. ## Acceptance - `Net:` line present in the banner with a MAC address. - `dhcp` gets a lease from 192.168.27.1. - A kernel + dtb boot end to end over TFTP from ouranos.
Author
Owner

Progress 2026-08-29: the GMAC half is done; the PMIC half has a different blocker

Full write-up in 56-uboot-gmac-and-pmic.md. Three commits on draco-aw80. The board is back on its known-good bootloaders.

Two corrections to this issue

There is no AXP809 on this board. 31-power-topology.md established that electrically — RSB slave NACK at 0x3a3. It is an AXP806 alone at 0x745, so "the AXP806 + AXP809 pair" overstates the problem.

Mainline U-Boot already covers the AXP806. pmic/axp.c matches x-powers,axp806, axp_regulator.c carries a full axp806_regulators[] including aldo3 and sw, and sun8i_rsb.c is a UCLASS_I2C driver so the PMIC binds as an ordinary I2C child. DM_I2C was already on. There is no PMIC driver to write — which also supersedes note 09's "main U-Boot porting task", since that reasoned only about the legacy AXP*_POWER path.

⚠️ Step 1 of the suggested order is wrong: do not enable CONFIG_SUN7I_GMAC. board/sunxi/gmac.c writes ccm->gmac_clk_cfg at CCM + 0x164, the sun4i register layout. On sun9i that register is standalone at 0x00800030.

Done and committed

  • CLK_BUS_GMAC = GATE(0x584, BIT(17)) and RST_BUS_GMAC = RESET(0x5a4, BIT(17)) in clk_a80.c — without these the designware probe aborts in its clk_enable() loop
  • an ethernet0 alias, without which setup_environment() never generates the SID-derived MAC
  • the gmac node refreshed — it still carried a fixed-link and a "no PHY answers" comment from before the MDIO settle fix
  • ETH_DESIGNWARE, PHY_REALTEK
  • new: drivers/clk/sunxi/clk_a80_r.c, the sun9i CPUS-domain gates and resets

No work is needed on the MDIO divider: designware.c hardcodes MII_CLKRANGE_150_250M, the same range the kernel picks with snps,clk-csr = <4>.

The hang, and the fix, validated on hardware

The first attempt hung between DRAM: 3.5 GiB and Core: NN devices — i.e. inside dm_init_and_scan(), where always-on regulators are force-probed and drag up the RSB.

r_rsb takes clocks = <&apbs_gates 3> / resets = <&apbs_rst 3> — providers at 0x08001428 and 0x080014b0 that U-Boot had for no sun9i SoC. sun8i_rsb_probe() treats both lookups as best-effort and calls sun8i_rsb_init() regardless, so it wrote to a block still gated and in reset, hanging the SoC. CONFIG_SYS_I2C_SUN8I_RSB=y had been harmless for months only because nothing ever probed the bus.

With clk_a80_r.c in place, confirmed at the U-Boot prompt:

=> md 0x08001428 1
08001428: 00100009        RSB gate, bit 3, on
=> md 0x080014b0 1
080014b0: 00000008        RSB reset, bit 3, deasserted
=> md 0x08002c48 1
08002c48: 00000033        PN0/PN1 muxed to s_rsb

and Core: 79 devices, 19 uclasses against the previous build's 55/16.

🔴 The actual blocker

Every AXP806 transfer NACKs. The controller is healthy — ccr = 0x103 is exactly what sun8i_rsb_set_clk() computes, dmcr = 0x007c3e00 with START cleared means the mode switch completed, devaddr = 0x003a0000 shows the right runtime address. But driven by hand from the prompt:

set runtime address (cmd 0xe8, devaddr 0x003a0745)  ->  stat 0x01   TOVER, success
byte read of reg 0x03 at runtime 0x3a (cmd 0x8b)    ->  stat 0x103  TOVER|TERR, data 0xff

The chip answers on its hardware address and NACKs on its runtime address.

The consequence is what stops the boot: regulators exist but cannot be driven, so mmc fails on Error enabling VMMC supply : -5, U-Boot can read no kernel from either medium, and Net: reports Error enabling phy supply. So DM_PMIC/PMIC_AXP/REGULATOR_AXP are deliberately not enabled in the defconfig — the tree stays bootable.

The lesson worth carrying: the current bootloader works because it never touches the PMIC. Enabling the regulators makes previously-unconditional subsystems depend on RSB transfers, so this is not separable from making those transfers work.

Two leads: sun8i_rsb_init() writes RSB_CTRL_SOFT_RST and proceeds without waiting for it to clear, where Linux's sunxi_rsb_hw_init() polls; and nothing in U-Boot knows about x-powers,master-mode, for which Linux writes AXP806_REG_ADDR_EXT.

Revised remaining work

  1. RSB transfers to the AXP806 — everything else is behind this
  2. The RGMII TX clock at 0x00800030 — needed before a frame can leave the board
  3. Then dhcp / tftpboot

A safe intermediate step is already committed: ETH_DESIGNWARE without the PMIC cannot break booting, since mmc touches no regulator in that configuration. It should give the Net: line with a real MAC — acceptance criterion 1 — even if the PHY stays silent without aldo3 at 2500 mV.

## Progress 2026-08-29: the GMAC half is done; the PMIC half has a different blocker Full write-up in [56-uboot-gmac-and-pmic.md](56-uboot-gmac-and-pmic.md). Three commits on `draco-aw80`. The board is back on its known-good bootloaders. ### Two corrections to this issue **There is no AXP809 on this board.** `31-power-topology.md` established that electrically — RSB slave NACK at `0x3a3`. It is an AXP806 alone at `0x745`, so "the AXP806 + AXP809 pair" overstates the problem. **Mainline U-Boot already covers the AXP806.** `pmic/axp.c` matches `x-powers,axp806`, `axp_regulator.c` carries a full `axp806_regulators[]` including `aldo3` and `sw`, and `sun8i_rsb.c` is a `UCLASS_I2C` driver so the PMIC binds as an ordinary I2C child. `DM_I2C` was already on. **There is no PMIC driver to write** — which also supersedes note 09's "main U-Boot porting task", since that reasoned only about the legacy `AXP*_POWER` path. ⚠️ **Step 1 of the suggested order is wrong: do not enable `CONFIG_SUN7I_GMAC`.** `board/sunxi/gmac.c` writes `ccm->gmac_clk_cfg` at `CCM + 0x164`, the sun4i register layout. On sun9i that register is standalone at `0x00800030`. ### Done and committed - `CLK_BUS_GMAC` = `GATE(0x584, BIT(17))` and `RST_BUS_GMAC` = `RESET(0x5a4, BIT(17))` in `clk_a80.c` — without these the designware probe aborts in its `clk_enable()` loop - an `ethernet0` alias, without which `setup_environment()` never generates the SID-derived MAC - the `gmac` node refreshed — it still carried a `fixed-link` and a "no PHY answers" comment from before the MDIO settle fix - `ETH_DESIGNWARE`, `PHY_REALTEK` - **new: `drivers/clk/sunxi/clk_a80_r.c`**, the sun9i CPUS-domain gates and resets No work is needed on the MDIO divider: `designware.c` hardcodes `MII_CLKRANGE_150_250M`, the same range the kernel picks with `snps,clk-csr = <4>`. ### The hang, and the fix, validated on hardware The first attempt hung between `DRAM: 3.5 GiB` and `Core: NN devices` — i.e. inside `dm_init_and_scan()`, where `always-on` regulators are force-probed and drag up the RSB. `r_rsb` takes `clocks = <&apbs_gates 3>` / `resets = <&apbs_rst 3>` — providers at `0x08001428` and `0x080014b0` that **U-Boot had for no sun9i SoC**. `sun8i_rsb_probe()` treats both lookups as best-effort and calls `sun8i_rsb_init()` regardless, so it wrote to a block still gated and in reset, hanging the SoC. `CONFIG_SYS_I2C_SUN8I_RSB=y` had been harmless for months only because nothing ever probed the bus. With `clk_a80_r.c` in place, confirmed at the U-Boot prompt: ``` => md 0x08001428 1 08001428: 00100009 RSB gate, bit 3, on => md 0x080014b0 1 080014b0: 00000008 RSB reset, bit 3, deasserted => md 0x08002c48 1 08002c48: 00000033 PN0/PN1 muxed to s_rsb ``` and `Core: 79 devices, 19 uclasses` against the previous build's 55/16. ### 🔴 The actual blocker Every AXP806 transfer NACKs. The controller is healthy — `ccr = 0x103` is exactly what `sun8i_rsb_set_clk()` computes, `dmcr = 0x007c3e00` with START cleared means the mode switch completed, `devaddr = 0x003a0000` shows the right runtime address. But driven by hand from the prompt: ``` set runtime address (cmd 0xe8, devaddr 0x003a0745) -> stat 0x01 TOVER, success byte read of reg 0x03 at runtime 0x3a (cmd 0x8b) -> stat 0x103 TOVER|TERR, data 0xff ``` **The chip answers on its hardware address and NACKs on its runtime address.** The consequence is what stops the boot: regulators exist but cannot be driven, so `mmc` fails on `Error enabling VMMC supply : -5`, U-Boot can read no kernel from either medium, and `Net:` reports `Error enabling phy supply`. So `DM_PMIC`/`PMIC_AXP`/`REGULATOR_AXP` are **deliberately not enabled** in the defconfig — the tree stays bootable. The lesson worth carrying: the current bootloader works *because* it never touches the PMIC. Enabling the regulators makes previously-unconditional subsystems depend on RSB transfers, so this is not separable from making those transfers work. Two leads: `sun8i_rsb_init()` writes `RSB_CTRL_SOFT_RST` and proceeds without waiting for it to clear, where Linux's `sunxi_rsb_hw_init()` polls; and nothing in U-Boot knows about `x-powers,master-mode`, for which Linux writes `AXP806_REG_ADDR_EXT`. ### Revised remaining work 1. **RSB transfers to the AXP806** — everything else is behind this 2. The RGMII TX clock at `0x00800030` — needed before a frame can leave the board 3. Then `dhcp` / `tftpboot` A safe intermediate step is already committed: `ETH_DESIGNWARE` **without** the PMIC cannot break booting, since `mmc` touches no regulator in that configuration. It should give the `Net:` line with a real MAC — acceptance criterion 1 — even if the PHY stays silent without `aldo3` at 2500 mV.
Author
Owner

Progress 2026-08-30: RGMII transmit clock done; the RSB diagnosis was wrong and is now sharper

The transmit clock is programmed

The previous comment said not to enable CONFIG_SUN7I_GMAC because board/sunxi/gmac.c
writes ccm->gmac_clk_cfg at CCM + 0x164, the sun4i layout. That is now fixed properly:
sun9i keeps the same control in the AHB1 region at 0x00800030 - the dtsi's gmac_tx_clk,
carrying the A20's own compatible string, so the fields are identical and only the address
differs. gmac.c now selects the right one and the symbol is enabled.

The value was read off a running kernel rather than guessed, on a board where the link
negotiates at 1 Gbit:

0x00800030 = 0x00000006     tx clk src 2 (internal RGMII), GPIT 1 (RGMII), rx/tx delay 0

which is exactly CCM_GMAC_CTRL_TX_CLK_SRC_INT_RGMII | CCM_GMAC_CTRL_GPIT_RGMII with
CONFIG_GMAC_TX_DELAY at its default 0, agreeing with the board DTS carrying no
allwinner,tx-delay-ps. Verified in the object code:

mov.w  r3, #0x800000
ldr    r2, [r3, #0x30]
orr.w  r2, r2, #6
str    r2, [r3, #0x30]

Commit 1183c9b on draco-aw80. DM_PMIC remains off, so this build cannot break booting -
mmc touches no regulator without it.

Correction: the RSB failure is not a NACK

I previously described the blocker as "the chip acknowledges on its hardware address and NACKs
on its runtime address". That misread the status word:

#define RSB_INTS_TRANS_ERR_ACK   BIT(16)        /* NACK */
#define RSB_INTS_TRANS_ERR_DATA  GENMASK(11, 8) /* data error, value = bit position */

The observed 0x103 has bit 16 clear and TRANS_ERR_DATA reading 1. The AXP806 does
acknowledge; the transfer fails in its data phase, at bit 1. Addressing and authorisation
are not involved.

Clocking is ruled out too, from the tree alone

sun8i_rsb_set_clk() assumes a 24 MHz parent and computes ccr = 0x103, which is what the
register holds. That assumption is correct here:

  • apbs_rsb <- apbs <- ahbs <- cpus_clk, and ahbs is a fixed-factor 1:1 clock
  • arch/arm/mach-sunxi/clock_sun9i.c never references the CPUS domain, so U-Boot does not
    reprogram it
  • no assigned-clocks anywhere, so Linux does not force it either - it reports what it finds
  • and Linux reports the whole chain at 24 MHz

So U-Boot inherits the same 24 MHz and the divider is right.

Remaining suspects, in order

  1. The un-awaited soft reset. sun8i_rsb_init() writes RSB_CTRL_SOFT_RST then goes
    straight to set_clk and set_device_mode without waiting for the bit to clear; Linux's
    sunxi_rsb_hw_init() polls for it.
  2. x-powers,master-mode - nothing in U-Boot's AXP driver knows about it; Linux writes
    AXP806_REG_ADDR_EXT.
  3. The AC100 sharing the bus, at hardware address 0xe89 / runtime 0x4e under Linux.
    U-Boot's sun8i_rsb_get_runtime_address() knows only 0x3a3 and 0x745 so it never
    addresses it - but the device-mode command is a broadcast, and after a warm reboot the
    AC100 still holds the runtime address Linux gave it. This had not been considered before.

Ready to test, and low risk this time

The current build should produce the Net: line with a SID-derived MAC - acceptance criterion
1 of this issue - and cannot leave the board unbootable, because without the PMIC nothing in
the boot path depends on a regulator. Whether the PHY answers MDIO is the open question, since
aldo3 is not raised to 2500 mV without the regulator framework.

## Progress 2026-08-30: RGMII transmit clock done; the RSB diagnosis was wrong and is now sharper ### The transmit clock is programmed The previous comment said not to enable `CONFIG_SUN7I_GMAC` because `board/sunxi/gmac.c` writes `ccm->gmac_clk_cfg` at `CCM + 0x164`, the sun4i layout. That is now fixed properly: sun9i keeps the same control in the AHB1 region at `0x00800030` - the dtsi's `gmac_tx_clk`, carrying the A20's own compatible string, so the fields are identical and only the address differs. `gmac.c` now selects the right one and the symbol is enabled. The value was read off a running kernel rather than guessed, on a board where the link negotiates at 1 Gbit: ``` 0x00800030 = 0x00000006 tx clk src 2 (internal RGMII), GPIT 1 (RGMII), rx/tx delay 0 ``` which is exactly `CCM_GMAC_CTRL_TX_CLK_SRC_INT_RGMII | CCM_GMAC_CTRL_GPIT_RGMII` with `CONFIG_GMAC_TX_DELAY` at its default 0, agreeing with the board DTS carrying no `allwinner,tx-delay-ps`. Verified in the object code: ``` mov.w r3, #0x800000 ldr r2, [r3, #0x30] orr.w r2, r2, #6 str r2, [r3, #0x30] ``` Commit `1183c9b` on `draco-aw80`. `DM_PMIC` remains off, so this build cannot break booting - `mmc` touches no regulator without it. ### Correction: the RSB failure is not a NACK I previously described the blocker as "the chip acknowledges on its hardware address and NACKs on its runtime address". That misread the status word: ```c #define RSB_INTS_TRANS_ERR_ACK BIT(16) /* NACK */ #define RSB_INTS_TRANS_ERR_DATA GENMASK(11, 8) /* data error, value = bit position */ ``` The observed `0x103` has **bit 16 clear** and `TRANS_ERR_DATA` reading **1**. The AXP806 *does* acknowledge; the transfer fails in its **data phase, at bit 1**. Addressing and authorisation are not involved. ### Clocking is ruled out too, from the tree alone `sun8i_rsb_set_clk()` assumes a 24 MHz parent and computes `ccr = 0x103`, which is what the register holds. That assumption is correct here: - `apbs_rsb` <- `apbs` <- `ahbs` <- `cpus_clk`, and `ahbs` is a **fixed-factor 1:1** clock - `arch/arm/mach-sunxi/clock_sun9i.c` never references the CPUS domain, so U-Boot does not reprogram it - no `assigned-clocks` anywhere, so Linux does not force it either - it reports what it finds - and Linux reports the whole chain at 24 MHz So U-Boot inherits the same 24 MHz and the divider is right. ### Remaining suspects, in order 1. **The un-awaited soft reset.** `sun8i_rsb_init()` writes `RSB_CTRL_SOFT_RST` then goes straight to `set_clk` and `set_device_mode` without waiting for the bit to clear; Linux's `sunxi_rsb_hw_init()` polls for it. 2. **`x-powers,master-mode`** - nothing in U-Boot's AXP driver knows about it; Linux writes `AXP806_REG_ADDR_EXT`. 3. **The AC100 sharing the bus**, at hardware address `0xe89` / runtime `0x4e` under Linux. U-Boot's `sun8i_rsb_get_runtime_address()` knows only `0x3a3` and `0x745` so it never addresses it - but the device-mode command is a *broadcast*, and after a warm reboot the AC100 still holds the runtime address Linux gave it. This had not been considered before. ### Ready to test, and low risk this time The current build should produce the `Net:` line with a SID-derived MAC - acceptance criterion 1 of this issue - and cannot leave the board unbootable, because without the PMIC nothing in the boot path depends on a regulator. Whether the PHY answers MDIO is the open question, since `aldo3` is not raised to 2500 mV without the regulator framework.
Author
Owner

Tested on hardware 2026-08-30: Net: line achieved; the MAC transmits nothing

Flashed to the eMMC with the card out. The board booted normally all the way to Linux -
with DM_PMIC off nothing in the boot path depends on a regulator, as intended.

Core:  61 devices, 17 uclasses, devicetree: separate
MMC:   mmc@1c0f000: 0, mmc@1c11000: 1
Loading Environment from MMC...Reading from MMC(1)... OK
Net:   eth0: ethernet@830000

Net: eth0 is acceptance criterion 1 of this issue, replacing No ethernet found.

Working, measured at the U-Boot prompt

MAC ethaddr=02:2c:b6:3a:4d:0c - the SID-derived form, so the ethernet0 alias works
MDIO mii info -> OUI = 0x0732, Model = 0x11, Rev = 0x05 = RTL8211E (0x001cc915)
Autoneg Waiting for PHY auto negotiation to complete....... done
Link Speed: 1000, full duplex
PHY delays page 7 / ext 0xa4 / reg 0x1c = 0xBD75 - bits 13/12/11, i.e. CTRL_DELAY | TX_DELAY | RX_DELAY, what rtl8211e_config() writes for rgmii-id
PHY basics BMCR 0x1140, BMSR 0x796D - link up, autoneg complete

MDIO works without the PMIC. I predicted the PHY would stay silent until aldo3 reached
2500 mV. It does not - it answers and links at gigabit on the power-on defaults. That is a
useful result on its own: the PHY rail is not a prerequisite for MDIO here.

Not working: transmit

ping and dhcp both fail - ARP Retry count exceeded, BOOTP broadcast 1..7. Captured on
the build host with tcpdump -i eno1 arp while U-Boot pinged:

board ARP frames during the U-Boot ping:  0

with a control - the same capture on the same interface, with the board running Linux,
caught the same MAC:

02:2c:b6:3a:4d:0c > ff:ff:ff:ff:ff:ff, ARP, Request who-has 192.168.27.46 tell 192.168.27.44

Same board, same MAC, same switch. Linux's frames arrive, U-Boot's never leave. The fault is
transmit; receive is untested because nothing gets that far.

Ruled out, with evidence

  • CCU gate and reset - MDIO goes through the MAC's own registers, so the block is clocked
    and out of reset.
  • Transmit clock mux and gate - 0x00800030 = 0x6, byte-identical to a running Linux
    carrying traffic. Linux's clk-a20-gmac.c confirms the layout: mask 0x3 selects the source
    in bits [1:0], and SUN7I_A20_GMAC_GPIT 2 is the gate. U-Boot's ..._GPIT_RGMII name is
    misleading, but the bit is right and it is set.
  • Both clock parents - mii_phy_tx_clk and gmac_int_tx_clk are plain fixed-clock
    nodes; the dtsi states the actual TX rate is not controlled by this clock.
  • The PHY - configured, delayed, linked.
  • MBUS - sunxi_mbus.c lists only display devices.
  • CONFIG_PHY_GIGE - unset, but used only by common/miiphyutil.c, so it explains why
    mii info prints 10baseT, HDX while phylib reports 1000/full. Cosmetic; worth enabling.

Next

The remaining surface is the designware transmit DMA path - descriptor ring, the addresses
written into it, and cache maintenance. Worth holding #48 in mind, which records a sun9i block
needing an unexplained PHYS_OFFSET fixup, though U-Boot has no virtual mapping to confuse.

Test-rig note

Spamming a key at the serial port to interrupt autoboot is unreliable and wasted several
boots. The reliable method is fw_setenv bootdelay 25, reboot, send one newline, then
restore bootdelay 2.

Board left healthy: a80-debian on the eMMC running the new build, nproc 4, card out.

## Tested on hardware 2026-08-30: `Net:` line achieved; the MAC transmits nothing Flashed to the eMMC with the card out. **The board booted normally all the way to Linux** - with `DM_PMIC` off nothing in the boot path depends on a regulator, as intended. ``` Core: 61 devices, 17 uclasses, devicetree: separate MMC: mmc@1c0f000: 0, mmc@1c11000: 1 Loading Environment from MMC...Reading from MMC(1)... OK Net: eth0: ethernet@830000 ``` **`Net: eth0` is acceptance criterion 1 of this issue**, replacing `No ethernet found.` ### Working, measured at the U-Boot prompt | | | |---|---| | MAC | `ethaddr=02:2c:b6:3a:4d:0c` - the SID-derived form, so the `ethernet0` alias works | | MDIO | `mii info` -> `OUI = 0x0732, Model = 0x11, Rev = 0x05` = RTL8211E (`0x001cc915`) | | Autoneg | `Waiting for PHY auto negotiation to complete....... done` | | Link | `Speed: 1000, full duplex` | | PHY delays | page 7 / ext `0xa4` / reg `0x1c` = **`0xBD75`** - bits 13/12/11, i.e. `CTRL_DELAY \| TX_DELAY \| RX_DELAY`, what `rtl8211e_config()` writes for `rgmii-id` | | PHY basics | BMCR `0x1140`, BMSR `0x796D` - link up, autoneg complete | **MDIO works without the PMIC.** I predicted the PHY would stay silent until `aldo3` reached 2500 mV. It does not - it answers and links at gigabit on the power-on defaults. That is a useful result on its own: the PHY rail is not a prerequisite for MDIO here. ### Not working: transmit `ping` and `dhcp` both fail - `ARP Retry count exceeded`, `BOOTP broadcast 1..7`. Captured on the build host with `tcpdump -i eno1 arp` while U-Boot pinged: ``` board ARP frames during the U-Boot ping: 0 ``` with a control - the same capture on the same interface, with the board running **Linux**, caught the same MAC: ``` 02:2c:b6:3a:4d:0c > ff:ff:ff:ff:ff:ff, ARP, Request who-has 192.168.27.46 tell 192.168.27.44 ``` Same board, same MAC, same switch. Linux's frames arrive, U-Boot's never leave. **The fault is transmit; receive is untested because nothing gets that far.** ### Ruled out, with evidence - **CCU gate and reset** - MDIO goes through the MAC's own registers, so the block is clocked and out of reset. - **Transmit clock mux and gate** - `0x00800030 = 0x6`, byte-identical to a running Linux carrying traffic. Linux's `clk-a20-gmac.c` confirms the layout: mask `0x3` selects the source in bits [1:0], and `SUN7I_A20_GMAC_GPIT 2` is the **gate**. U-Boot's `..._GPIT_RGMII` name is misleading, but the bit is right and it is set. - **Both clock parents** - `mii_phy_tx_clk` and `gmac_int_tx_clk` are plain `fixed-clock` nodes; the dtsi states the actual TX rate is not controlled by this clock. - **The PHY** - configured, delayed, linked. - **MBUS** - `sunxi_mbus.c` lists only display devices. - **`CONFIG_PHY_GIGE`** - unset, but used only by `common/miiphyutil.c`, so it explains why `mii info` prints `10baseT, HDX` while phylib reports 1000/full. Cosmetic; worth enabling. ### Next The remaining surface is the designware **transmit DMA path** - descriptor ring, the addresses written into it, and cache maintenance. Worth holding #48 in mind, which records a sun9i block needing an unexplained `PHYS_OFFSET` fixup, though U-Boot has no virtual mapping to confuse. ### Test-rig note Spamming a key at the serial port to interrupt autoboot is unreliable and wasted several boots. The reliable method is `fw_setenv bootdelay 25`, reboot, send **one** newline, then restore `bootdelay 2`. Board left healthy: `a80-debian` on the eMMC running the new build, `nproc` 4, card out.
Author
Owner

Narrowed further: the MAC completes the frames and reports success

Read the transmit descriptor ring at the prompt after a failed ping. Ring at
0x00831010 = 0xfbf82200, 64-byte stride:

fbf82200: 00000000 6100003c fbf82a00 fbf82240
fbf82240: 00000000 6100003c fbf83200 fbf82280

des1 = 0x6100003c is FS | LS | TCH with TBS1 = 0x3c = 60 bytes - a real, padded ARP
frame. des0 = 0 means the OWN bit is cleared with zero error status. Two of them. DMA
status 0x44000546 also has ETI set, so transmission started.

So U-Boot queues complete frames, the DMA consumes them, reports no error, and nothing reaches
the wire. The fault is downstream of the transmit FIFO, in the RGMII output.

Register-by-register against a Linux that carries traffic

What U-Boot Linux
gmac_tx_clk 0x00800030 0x00000006 0x00000006 same
PA CFG0 0x27222222 0x27222222 same
PA CFG1 0x27227222 0x27227222 same
PA CFG2 0x00000022 0x00000022 same
PA DRV0 0xdf7fdfff 0xdf7fdfff same, drive level 3
PA DRV1 0x0000000f 0x0000000f same
Descriptor format normal (DW_ALTDESCRIPTOR unset) "Normal descriptors" per dmesg same

Also ruled out this round: CONFIG_PINCONF=y and PINCTRL_FULL=y, so drive-strength = <40>
really is applied - worth checking because the dtsi warns RGMII DDR needs it and MDIO at
2.35 MHz would tolerate a weak setting that 125 MHz would not. And MAC conf = 0x00202800
confirms full duplex with MII_PORTSELECT clear, i.e. GMII/1000.

Where to look next

Everything software can see matches a working Linux, so more register comparison is not the
way forward. In order:

  1. A scope on TXCK (PA12) during a ping - settles in one measurement whether the 125 MHz
    transmit clock is present at the pin. Everything above says it should be.
  2. clk_set_rate ordering. Linux reaches 0x6 as two operations - clk_set_rate picks the
    mux, clk_prepare_enable sets the gate. eth_init_board() ORs both in a single write. Same
    end state, different sequence.
  3. eth_init_board() runs from board_init(), long before the designware driver enables
    CLK_BUS_GMAC and deasserts RST_BUS_GMAC. The write demonstrably sticks, but whether the
    transmit clock tree latches it while the block is still in reset is not established. Moving
    the call into the driver probe would settle it, and is the cheapest of the three to try.

Note 56 has the detail. Board left healthy on the eMMC, bootdelay restored to 2.

## Narrowed further: the MAC completes the frames and reports success Read the transmit descriptor ring at the prompt after a failed `ping`. Ring at `0x00831010 = 0xfbf82200`, 64-byte stride: ``` fbf82200: 00000000 6100003c fbf82a00 fbf82240 fbf82240: 00000000 6100003c fbf83200 fbf82280 ``` `des1 = 0x6100003c` is FS | LS | TCH with **TBS1 = 0x3c = 60 bytes** - a real, padded ARP frame. `des0 = 0` means the **OWN bit is cleared with zero error status**. Two of them. DMA status `0x44000546` also has **ETI** set, so transmission started. So U-Boot queues complete frames, the DMA consumes them, reports no error, and nothing reaches the wire. **The fault is downstream of the transmit FIFO, in the RGMII output.** ## Register-by-register against a Linux that carries traffic | What | U-Boot | Linux | | |---|---|---|---| | `gmac_tx_clk` `0x00800030` | `0x00000006` | `0x00000006` | same | | PA `CFG0` | `0x27222222` | `0x27222222` | same | | PA `CFG1` | `0x27227222` | `0x27227222` | same | | PA `CFG2` | `0x00000022` | `0x00000022` | same | | PA `DRV0` | `0xdf7fdfff` | `0xdf7fdfff` | same, drive level 3 | | PA `DRV1` | `0x0000000f` | `0x0000000f` | same | | Descriptor format | normal (`DW_ALTDESCRIPTOR` unset) | "Normal descriptors" per dmesg | same | Also ruled out this round: `CONFIG_PINCONF=y` and `PINCTRL_FULL=y`, so `drive-strength = <40>` really is applied - worth checking because the dtsi warns RGMII DDR needs it and MDIO at 2.35 MHz would tolerate a weak setting that 125 MHz would not. And MAC `conf = 0x00202800` confirms full duplex with `MII_PORTSELECT` clear, i.e. GMII/1000. ## Where to look next Everything software can see matches a working Linux, so more register comparison is not the way forward. In order: 1. **A scope on TXCK (PA12) during a `ping`** - settles in one measurement whether the 125 MHz transmit clock is present at the pin. Everything above says it should be. 2. **`clk_set_rate` ordering.** Linux reaches `0x6` as two operations - `clk_set_rate` picks the mux, `clk_prepare_enable` sets the gate. `eth_init_board()` ORs both in a single write. Same end state, different sequence. 3. **`eth_init_board()` runs from `board_init()`**, long before the designware driver enables `CLK_BUS_GMAC` and deasserts `RST_BUS_GMAC`. The write demonstrably sticks, but whether the transmit clock tree latches it while the block is still in reset is not established. Moving the call into the driver probe would settle it, and is the cheapest of the three to try. Note 56 has the detail. Board left healthy on the eMMC, `bootdelay` restored to 2.
Author
Owner

Both code-side transmit hypotheses tested and dead

Checked at the U-Boot prompt with mw, after a first ping had brought the MAC up and out of
reset - no rebuild and no reflash, which is what made them worth trying before writing code.

Experiment Sequence Result
Re-latch the register with the block live mw 0x00800030 0 then 6, md confirms 6, ping ARP retry exceeded - no change
Separate mux from gate, Linux's order mw 0x00800030 2 (mux, = clk_set_rate) then 6 (gate, = clk_prepare_enable), ping ARP retry exceeded - no change

So neither the ordering of eth_init_board() against the driver's clock and reset handling,
nor the single-write-versus-two difference from Linux, is the cause. Moving the call into the
driver probe would not have helped and does not need trying
- which is the useful part of a
negative result.

#39 now parks for the bench

Everything reachable from software is done and matches a working Linux register for register.
The MAC accepts two complete 60-byte frames, clears OWN with a zero error status, and nothing
arrives. The next useful information is physical:

  • A scope on TXCK (PA12) during a ping - settles whether the 125 MHz transmit clock is
    present at the pin.
  • While the scope is attached, worth distinguishing "no frames" from "frames with bad timing".
    A switch drops a frame with a bad FCS silently, which looks identical from both ends: the MAC
    reporting success and tcpdump seeing nothing. That would point at RGMII skew despite the
    PHY's internal delays reading correct.

This puts #39 alongside #53 and #42 as bench work rather than desk work.

What was achieved

  • Net: eth0: ethernet@830000 with a SID-derived MAC - acceptance criterion 1 of this issue
  • RGMII transmit clock programmed for sun9i (commit 1183c9b)
  • MDIO, PHY detection, autonegotiation and a 1000/full link, all without the PMIC
  • The board boots normally throughout, because DM_PMIC stays off

dhcp and tftpboot remain out of reach until transmit works.

## Both code-side transmit hypotheses tested and dead Checked at the U-Boot prompt with `mw`, after a first `ping` had brought the MAC up and out of reset - no rebuild and no reflash, which is what made them worth trying before writing code. | Experiment | Sequence | Result | |---|---|---| | Re-latch the register with the block live | `mw 0x00800030 0` then `6`, `md` confirms `6`, `ping` | ARP retry exceeded - **no change** | | Separate mux from gate, Linux's order | `mw 0x00800030 2` (mux, = `clk_set_rate`) then `6` (gate, = `clk_prepare_enable`), `ping` | ARP retry exceeded - **no change** | So neither the ordering of `eth_init_board()` against the driver's clock and reset handling, nor the single-write-versus-two difference from Linux, is the cause. **Moving the call into the driver probe would not have helped and does not need trying** - which is the useful part of a negative result. ## #39 now parks for the bench Everything reachable from software is done and matches a working Linux register for register. The MAC accepts two complete 60-byte frames, clears OWN with a zero error status, and nothing arrives. The next useful information is physical: - **A scope on TXCK (PA12) during a `ping`** - settles whether the 125 MHz transmit clock is present at the pin. - While the scope is attached, worth distinguishing "no frames" from "frames with bad timing". A switch drops a frame with a bad FCS silently, which looks identical from both ends: the MAC reporting success and `tcpdump` seeing nothing. That would point at RGMII skew despite the PHY's internal delays reading correct. This puts #39 alongside #53 and #42 as bench work rather than desk work. ## What was achieved - `Net: eth0: ethernet@830000` with a SID-derived MAC - **acceptance criterion 1 of this issue** - RGMII transmit clock programmed for sun9i (commit `1183c9b`) - MDIO, PHY detection, autonegotiation and a 1000/full link, **all without the PMIC** - The board boots normally throughout, because `DM_PMIC` stays off `dhcp` and `tftpboot` remain out of reach until transmit works.
Author
Owner

Transmit is not dead. Gigabit transmit is dead.

This is the finding, and it reframes the issue. Forcing the PHY to 100/full at the u-boot
prompt and pinging:

=> mii write 1 9 0          (stop advertising 1000)
=> mii write 1 4 0x0101     (advertise 100/full only)
=> mii write 1 0 0x1200     (restart autoneg)
=> ping 192.168.27.46
Speed: 100, full duplex
host 192.168.27.46 is alive          <-- TRANSMIT WORKS

Reproduced twice in the same session. So the MAC, the descriptor rings, the DMA, the pinmux,
the drive strength, the MDIO path and the PHY are all fine, and every negative result recorded
above stands - they were just looking in the wrong place. The fault is confined to the
1000 Mbps RGMII clock domain
, i.e. the 125 MHz DDR transmit path, and nothing else.

That also retires the parked next step. A scope on TXCK would have shown a clock present,
because at 100 Mbps the same pin carries 25 MHz and works.

This unblocks the issue's actual goal

The acceptance criteria here are dhcp and a kernel booted over TFTP. Both are reachable
right now
by forcing 100 Mbps before the network is used - a preboot or a line in bootcmd
doing the three mii writes above. 100 Mbps is ~12 MB/s, which for an 8.6 MiB uImage plus a dtb
is roughly 1 s of transfer. TFTP boot does not need gigabit.

I would treat that as the fix for #39 and split the gigabit fault into its own issue, rather
than keeping TFTP boot blocked behind a hardware timing problem.

What was eliminated this round

  • CCU gates and resets. Diffed the whole CCU gate/reset block between a Linux carrying
    gigabit traffic and u-boot. Linux has several bits u-boot does not (0x584, 0x588, 0x590,
    0x594 and the matching reset words). Applying Linux's exact values with mw before the ping
    changed nothing at gigabit.
  • SYS_CTRL 0x00800030 reads 0x00000006 in both, confirmed again from kernel context.
  • The PHY's page-0 registers match between Linux and u-boot, register for register, with the
    sole exception of ANAR (r04: Linux 0x0de1, u-boot 0x01e1 - pause advertisement, which the
    driver rewrites at autoneg and which cannot affect transmit).
  • The RGMII delay register (ext-page 0xa4, 0x1c) reads 0xBD75 in both.

The remaining candidate is RGMII transmit timing at 125 MHz - skew, or the internal delay being
inappropriate for this board at gigabit only. Note the dtsi asks for rgmii-id, so both delays
are on; a board that needs rgmii-txid or external skew would look exactly like this.

⚠️ I damaged something, and it needs a power cycle

Being explicit because it is my fault and it is visible in the logs.

While testing PHY registers I wrote to the RTL8211E's delay register with the reserved field
(bits 10:0, Realtek test/debug settings) cleared, and at least one write landed on page 0
register 0x1c rather than the ext-page one, because the page-select did not take. Since then
gigabit autonegotiation flaps - the link comes up at 1 Gbps, drops, and after two or three
attempts settles at 100 Mbps.

The serial log dates it precisely:

boots  1-21   ups=[1Gbps/Full]  downs=0     <- stable, every boot
boot  22+     ups=[1Gbps 1Gbps 1Gbps 100Mbps]  downs=3

Boot 22 is the first boot after my first PHY write. Before that, gigabit was rock solid for
21 consecutive boots
, which also disproves a theory I briefly held - that Linux's gigabit was
marginal and u-boot was failing for the same reason. It was not marginal.

Restoring the register to its original 0xBD75 did not clear it, and neither did a BMCR
soft reset (0x8000): a soft reset does not restore the vendor's reserved bits, and there is no
PHY reset GPIO in the device tree - only phy-supply.

A power cycle should clear it, since that reloads the PHY straps. Until then the board is
fully usable at 100 Mbps: reachable on 192.168.27.44, 0 failed units, containers running.

Lesson worth keeping: on this PHY, read-modify-write the delay bits and never write the whole
register, and verify the page-select took before trusting any paged access.

## Transmit is not dead. Gigabit transmit is dead. This is the finding, and it reframes the issue. Forcing the PHY to 100/full at the u-boot prompt and pinging: ``` => mii write 1 9 0 (stop advertising 1000) => mii write 1 4 0x0101 (advertise 100/full only) => mii write 1 0 0x1200 (restart autoneg) => ping 192.168.27.46 Speed: 100, full duplex host 192.168.27.46 is alive <-- TRANSMIT WORKS ``` Reproduced twice in the same session. So the MAC, the descriptor rings, the DMA, the pinmux, the drive strength, the MDIO path and the PHY are all fine, and every negative result recorded above stands - they were just looking in the wrong place. **The fault is confined to the 1000 Mbps RGMII clock domain**, i.e. the 125 MHz DDR transmit path, and nothing else. That also retires the parked next step. A scope on TXCK would have shown a clock present, because at 100 Mbps the same pin carries 25 MHz and works. ## This unblocks the issue's actual goal The acceptance criteria here are `dhcp` and a kernel booted over TFTP. **Both are reachable right now** by forcing 100 Mbps before the network is used - a `preboot` or a line in `bootcmd` doing the three `mii write`s above. 100 Mbps is ~12 MB/s, which for an 8.6 MiB uImage plus a dtb is roughly 1 s of transfer. TFTP boot does not need gigabit. I would treat that as the fix for #39 and split the gigabit fault into its own issue, rather than keeping TFTP boot blocked behind a hardware timing problem. ## What was eliminated this round - **CCU gates and resets.** Diffed the whole CCU gate/reset block between a Linux carrying gigabit traffic and u-boot. Linux has several bits u-boot does not (`0x584`, `0x588`, `0x590`, `0x594` and the matching reset words). Applying Linux's exact values with `mw` before the ping changed nothing at gigabit. - **`SYS_CTRL 0x00800030`** reads `0x00000006` in both, confirmed again from kernel context. - **The PHY's page-0 registers** match between Linux and u-boot, register for register, with the sole exception of ANAR (`r04`: Linux `0x0de1`, u-boot `0x01e1` - pause advertisement, which the driver rewrites at autoneg and which cannot affect transmit). - **The RGMII delay register** (ext-page `0xa4`, `0x1c`) reads `0xBD75` in both. The remaining candidate is RGMII transmit timing at 125 MHz - skew, or the internal delay being inappropriate for this board at gigabit only. Note the dtsi asks for `rgmii-id`, so both delays are on; a board that needs `rgmii-txid` or external skew would look exactly like this. ## ⚠️ I damaged something, and it needs a power cycle Being explicit because it is my fault and it is visible in the logs. While testing PHY registers I wrote to the RTL8211E's delay register with the reserved field (bits 10:0, Realtek test/debug settings) cleared, and at least one write landed on **page 0** register `0x1c` rather than the ext-page one, because the page-select did not take. Since then gigabit autonegotiation **flaps** - the link comes up at 1 Gbps, drops, and after two or three attempts settles at 100 Mbps. The serial log dates it precisely: ``` boots 1-21 ups=[1Gbps/Full] downs=0 <- stable, every boot boot 22+ ups=[1Gbps 1Gbps 1Gbps 100Mbps] downs=3 ``` Boot 22 is the first boot after my first PHY write. **Before that, gigabit was rock solid for 21 consecutive boots**, which also disproves a theory I briefly held - that Linux's gigabit was marginal and u-boot was failing for the same reason. It was not marginal. Restoring the register to its original `0xBD75` did **not** clear it, and neither did a BMCR soft reset (`0x8000`): a soft reset does not restore the vendor's reserved bits, and there is no PHY reset GPIO in the device tree - only `phy-supply`. **A power cycle should clear it**, since that reloads the PHY straps. Until then the board is fully usable at 100 Mbps: reachable on `192.168.27.44`, 0 failed units, containers running. Lesson worth keeping: on this PHY, read-modify-write the delay bits and never write the whole register, and verify the page-select took before trusting any paged access.
Author
Owner

Recovered, and it confirms the mechanism

Power cycled. Gigabit is clean again on the first attempt:

[   16.753220] sun7i-dwmac 830000.ethernet eth0: Link is Up - 1Gbps/Full - flow control rx/tx
downs=0   final_speed=1000

One link event, no flap, versus three drops and a fallback to 100 Mbps immediately before.

This confirms the diagnosis rather than just clearing the symptom: the RTL8211E's reserved
configuration bits survive a warm reboot and a BMCR soft reset, and are only reloaded from
the straps at power-on. That is why restoring 0xBD75 over MDIO did not help - the damage was in
state that no software write could reach.

Board otherwise healthy: 0 failed units, docker and healthdog active, wlan0 back, 4 GiB, 4 cores,
bootcount=0, and all scratch u-boot variables cleared with bootdelay back to 2.

Practical rule for anyone touching this PHY again: read-modify-write the delay bits only,
never write the whole register, verify the page-select took before trusting a paged access, and
assume any mistake needs physical access to undo. That last part is what makes it worth doing
these experiments while someone can reach the board - or after #5.

Nothing above changes the substantive finding in the previous comment: transmit works at
100 Mbps
, the fault is confined to the gigabit RGMII clock domain, and TFTP boot is reachable
now without solving it.

## Recovered, and it confirms the mechanism Power cycled. Gigabit is clean again on the first attempt: ``` [ 16.753220] sun7i-dwmac 830000.ethernet eth0: Link is Up - 1Gbps/Full - flow control rx/tx downs=0 final_speed=1000 ``` One link event, no flap, versus three drops and a fallback to 100 Mbps immediately before. This confirms the diagnosis rather than just clearing the symptom: the RTL8211E's reserved configuration bits **survive a warm reboot and a BMCR soft reset**, and are only reloaded from the straps at power-on. That is why restoring `0xBD75` over MDIO did not help - the damage was in state that no software write could reach. Board otherwise healthy: 0 failed units, docker and healthdog active, wlan0 back, 4 GiB, 4 cores, `bootcount=0`, and all scratch u-boot variables cleared with `bootdelay` back to 2. **Practical rule for anyone touching this PHY again:** read-modify-write the delay bits only, never write the whole register, verify the page-select took before trusting a paged access, and assume any mistake needs physical access to undo. That last part is what makes it worth doing these experiments while someone can reach the board - or after #5. Nothing above changes the substantive finding in the previous comment: **transmit works at 100 Mbps**, the fault is confined to the gigabit RGMII clock domain, and TFTP boot is reachable now without solving it.
Author
Owner

All three acceptance criteria met

Net:   eth0: ethernet@830000                              <- 1, with a SID-derived MAC
DHCP client bound to address 192.168.27.44 (37 ms)        <- 2
Filename 'uImage'.      Bytes transferred = 9002048   702.1 KiB/s
Filename 'draco.dtb'.   Bytes transferred = 28469     1 MiB/s
Filename 'initrd.img'.  Bytes transferred = 9407610   1.1 MiB/s
   Verifying Checksum ... OK
   Loading Kernel Image to 20008000
   Loading Ramdisk to 29707000, end 29fffc7a ... OK
   Loading Device Tree to 296fd000, end 29706f34 ... OK
Starting kernel ...                                       <- 3

Linux came up on it: 7.2.0-14858, the TFTP bootargs in /proc/cmdline, eth0 back at
1000 Mbps, 0 failed units, docker active. iminfo's checksum pass is what makes this a
byte-exact transfer rather than a hopeful one.

The 100 Mbps forcing costs nothing

This is the part that makes it a fix rather than a hack. PHY registers do not survive a power
cycle, and Linux re-runs autonegotiation at boot regardless of what u-boot left behind - eth0
reads 1000 after every TFTP boot above. So forcing 100 Mbps for the duration of u-boot has no
runtime cost, and TFTP does not want gigabit anyway: 18.4 MB of kernel, dtb and initrd move in
about 20 seconds.

The gigabit fault is now #76, where it can be worked on without blocking anything.

Two things that cost time, recorded so they do not again

  • autoload=no. Without it, dhcp invents a bootfile name from the IP in hex
    (C0A81B2C.img) and spends ten timeouts trying to TFTP it from the DHCP server, which has no
    TFTP. It looks exactly like the network failing.
  • tftpwindowsize. CONFIG_TFTP_BLOCKSIZE is already 1468 but CONFIG_TFTP_WINDOWSIZE is
    1 - lockstep, one ACK per block, which is precisely the 490 KiB/s first seen. Setting it to 8
    roughly doubled throughput.

The environment, with the reasoning, is in tools/uboot-netboot-env.txt. Server side is
tftpd-hpa on ouranos serving /srv/tftp; refresh those files after any deploy-kernel.sh run.

What this unlocks, from the original post

TFTP boot. Kernel and dtb served from ouranos means iteration with zero eMMC writes and
zero unbootable-kernel risk - a bad kernel is one reset away rather than a rescue-SD trip.

That now holds. It also means a kernel can be tested without deploy-kernel.sh writing to the
eMMC at all, which matters for #65's wear budget.

Note the original post's prediction that PMIC support would be required turned out to be wrong -
MDIO, autonegotiation, DHCP and TFTP all work with CONFIG_SUNXI_NO_PMIC=y. The PHY's rails at
their OTP defaults are evidently sufficient.

Netconsole: built, not yet flashed

CONFIG_NETCONSOLE=y is committed on the u-boot branch (f6a7cf60921, version -00012) and
builds clean at 494896 bytes. It is compiled in but not enabled: stdin/stdout/stderr stay on
serial, so a board with no network still prints where it always did, and switching to netconsole
is a deliberate act - which is also the point at which someone accepts that it is unauthenticated
UDP, readable and writable by anyone on the LAN, at the u-boot prompt.

Flashing it needs approval, since writing the eMMC bootloader is the one operation with no
software fallback. Closing this issue on its stated criteria; netconsole is a separate,
optional step.

## All three acceptance criteria met ``` Net: eth0: ethernet@830000 <- 1, with a SID-derived MAC DHCP client bound to address 192.168.27.44 (37 ms) <- 2 Filename 'uImage'. Bytes transferred = 9002048 702.1 KiB/s Filename 'draco.dtb'. Bytes transferred = 28469 1 MiB/s Filename 'initrd.img'. Bytes transferred = 9407610 1.1 MiB/s Verifying Checksum ... OK Loading Kernel Image to 20008000 Loading Ramdisk to 29707000, end 29fffc7a ... OK Loading Device Tree to 296fd000, end 29706f34 ... OK Starting kernel ... <- 3 ``` Linux came up on it: `7.2.0-14858`, the TFTP `bootargs` in `/proc/cmdline`, eth0 back at **1000 Mbps**, 0 failed units, docker active. `iminfo`'s checksum pass is what makes this a byte-exact transfer rather than a hopeful one. ## The 100 Mbps forcing costs nothing This is the part that makes it a fix rather than a hack. PHY registers do not survive a power cycle, and Linux re-runs autonegotiation at boot regardless of what u-boot left behind - eth0 reads **1000** after every TFTP boot above. So forcing 100 Mbps for the duration of u-boot has no runtime cost, and TFTP does not want gigabit anyway: 18.4 MB of kernel, dtb and initrd move in about 20 seconds. The gigabit fault is now **#76**, where it can be worked on without blocking anything. ## Two things that cost time, recorded so they do not again - **`autoload=no`.** Without it, `dhcp` invents a bootfile name from the IP in hex (`C0A81B2C.img`) and spends ten timeouts trying to TFTP it from the DHCP server, which has no TFTP. It looks exactly like the network failing. - **`tftpwindowsize`.** `CONFIG_TFTP_BLOCKSIZE` is already 1468 but `CONFIG_TFTP_WINDOWSIZE` is 1 - lockstep, one ACK per block, which is precisely the 490 KiB/s first seen. Setting it to 8 roughly doubled throughput. The environment, with the reasoning, is in `tools/uboot-netboot-env.txt`. Server side is `tftpd-hpa` on ouranos serving `/srv/tftp`; refresh those files after any `deploy-kernel.sh` run. ## What this unlocks, from the original post > **TFTP boot.** Kernel and dtb served from ouranos means iteration with zero eMMC writes and > zero unbootable-kernel risk - a bad kernel is one `reset` away rather than a rescue-SD trip. That now holds. It also means a kernel can be tested without `deploy-kernel.sh` writing to the eMMC at all, which matters for #65's wear budget. Note the original post's prediction that PMIC support would be required turned out to be wrong - MDIO, autonegotiation, DHCP and TFTP all work with `CONFIG_SUNXI_NO_PMIC=y`. The PHY's rails at their OTP defaults are evidently sufficient. ## Netconsole: built, not yet flashed `CONFIG_NETCONSOLE=y` is committed on the u-boot branch (`f6a7cf60921`, version `-00012`) and builds clean at 494896 bytes. It is compiled in but **not enabled**: stdin/stdout/stderr stay on serial, so a board with no network still prints where it always did, and switching to netconsole is a deliberate act - which is also the point at which someone accepts that it is unauthenticated UDP, readable and writable by anyone on the LAN, at the u-boot prompt. Flashing it needs approval, since writing the eMMC bootloader is the one operation with no software fallback. Closing this issue on its stated criteria; netconsole is a separate, optional step.
Sign in to join this conversation.
No description provided.