eth0 does not come up after a cold boot: networkd gives up permanently #7

Open
opened 2026-08-27 23:22:24 +00:00 by tiagoagueda · 3 comments
Owner

Distinct from the MDIO settle fix in 34-ethernet-cold-boot.md, which works correctly
(PHY answered PHYSID1=001c after 0 ms of settle, PHY binds as RTL8211E).

After a real power-cycle:

22:20:07 eth0: Configuring with /etc/systemd/network/10-any-eth.network
22:20:08 eth0: Could not bring up interface, ignoring: Connection timed out

Note ignoring — networkd never retries, so the link stays administratively DOWN forever. A
manual ip link set eth0 up succeeds instantly and DHCP returns .44 in four seconds.

Cause: the interface open blocks ~5.4 s on a cold boot — driver probes at 1.58 s, PHY
attaches at 6.94 s — which is far longer than networkd's netlink timeout of about a second.
On a warm reboot the PHY is already awake and the open returns immediately, which is why a
full day of warm reboots never showed this.

Mitigated with ActivationPolicy=always-up in the .network file. Unverified until the
next cold boot.

Done when

  • the mitigation is confirmed across a real power-cycle
  • the kernel-side cause is fixed: stmmac_open / phylink_of_phy_connect should not block
    five seconds after power-on
Distinct from the MDIO settle fix in `34-ethernet-cold-boot.md`, which works correctly (`PHY answered PHYSID1=001c after 0 ms of settle`, PHY binds as RTL8211E). After a real power-cycle: ``` 22:20:07 eth0: Configuring with /etc/systemd/network/10-any-eth.network 22:20:08 eth0: Could not bring up interface, ignoring: Connection timed out ``` Note `ignoring` — networkd never retries, so the link stays administratively DOWN forever. A manual `ip link set eth0 up` succeeds instantly and DHCP returns `.44` in four seconds. **Cause:** the interface open blocks ~5.4 s on a cold boot — driver probes at `1.58 s`, PHY attaches at `6.94 s` — which is far longer than networkd's netlink timeout of about a second. On a warm reboot the PHY is already awake and the open returns immediately, which is why a full day of warm reboots never showed this. **Mitigated** with `ActivationPolicy=always-up` in the `.network` file. **Unverified until the next cold boot.** **Done when** - [ ] the mitigation is confirmed across a real power-cycle - [ ] the kernel-side cause is fixed: `stmmac_open` / `phylink_of_phy_connect` should not block five seconds after power-on
Author
Owner

First genuine cold-boot evidence since the ActivationPolicy=always-up mitigation went in.
This issue has been sitting on "unverified until the next power-cycle" - the board was fully
powered down and powered back on today (2026-08-30), which is that power-cycle.

The mitigation held.

[  12.061883] sun7i-dwmac 830000.ethernet eth0: configuring for phy/rgmii-id link mode
[  16.263541] sun7i-dwmac 830000.ethernet eth0: Link is Up - 1Gbps/Full - flow control rx/tx

$ networkctl status eth0
       State: routable (configured)
Online state: online

eth0 came up with its address and was reachable; nothing had to be rebound by hand.

But the underlying fault is untouched. The PHY still woke ~4.2 s after the driver
configured the link mode, which is the same late-PHY behaviour this issue documents. All that has
changed is that networkd no longer gives up permanently while waiting for it. So this is evidence
the workaround is load-bearing, not evidence the bug is gone - if the mitigation is ever removed,
or the delay grows, the original symptom returns.

Leaving this open, since it tracks the kernel-side cause rather than the mitigation. Worth
deciding whether to re-scope it to the driver fix and let #19 carry the upstream shape, or close
it and open a narrower one - happy to do either.

First genuine cold-boot evidence since the `ActivationPolicy=always-up` mitigation went in. This issue has been sitting on "unverified until the next power-cycle" - the board was fully powered down and powered back on today (2026-08-30), which is that power-cycle. **The mitigation held.** ``` [ 12.061883] sun7i-dwmac 830000.ethernet eth0: configuring for phy/rgmii-id link mode [ 16.263541] sun7i-dwmac 830000.ethernet eth0: Link is Up - 1Gbps/Full - flow control rx/tx $ networkctl status eth0 State: routable (configured) Online state: online ``` eth0 came up with its address and was reachable; nothing had to be rebound by hand. **But the underlying fault is untouched.** The PHY still woke **~4.2 s** after the driver configured the link mode, which is the same late-PHY behaviour this issue documents. All that has changed is that networkd no longer gives up permanently while waiting for it. So this is evidence the workaround is load-bearing, not evidence the bug is gone - if the mitigation is ever removed, or the delay grows, the original symptom returns. Leaving this open, since it tracks the kernel-side cause rather than the mitigation. Worth deciding whether to re-scope it to the driver fix and let #19 carry the upstream shape, or close it and open a narrower one - happy to do either.
Author
Owner

One data point from the mirror, offered as a lead rather than a diagnosis.

The Olimex A20-OLinuXino-Lime2 page records that board revisions fitted with the Realtek
RTL8211E
PHY - the same PHY as ours - lose packets on transmit, in a way tied to that specific
PHY rather than to the SoC. Several other sunxi boards in Table of Allwinner based boards use
the same part.

I am not claiming that is our fault: our symptom is a late PHY wake on cold boot (~4.2 s from
driver configure to link up), not packet loss, and the mitigation now demonstrably holds across a
real power cycle. But it is worth knowing that this PHY has a history of board-level quirks on
sunxi before the kernel-side fix in #19 is shaped for upstream, because reviewers who have seen
those reports may read our delay as the same class of problem.

If a settle-time fix is proposed upstream, it is worth stating explicitly which RTL8211E
behaviour is being compensated for, and that it was measured rather than inferred.


From a deep pass over the offline linux-sunxi.org mirror (a80/linux-sunxi.org/wiki-backup, 4095 pages, snapshot 2026-08-29). linux-sunxi.org content is CC BY-SA; the pages named above are the attribution.

One data point from the mirror, offered as a lead rather than a diagnosis. The **Olimex A20-OLinuXino-Lime2** page records that board revisions fitted with the **Realtek RTL8211E** PHY - the same PHY as ours - lose packets on transmit, in a way tied to that specific PHY rather than to the SoC. Several other sunxi boards in `Table of Allwinner based boards` use the same part. I am **not** claiming that is our fault: our symptom is a late PHY wake on cold boot (~4.2 s from driver configure to link up), not packet loss, and the mitigation now demonstrably holds across a real power cycle. But it is worth knowing that this PHY has a history of board-level quirks on sunxi before the kernel-side fix in #19 is shaped for upstream, because reviewers who have seen those reports may read our delay as the same class of problem. If a settle-time fix is proposed upstream, it is worth stating explicitly which RTL8211E behaviour is being compensated for, and that it was measured rather than inferred. --- *From a deep pass over the offline linux-sunxi.org mirror (`a80/linux-sunxi.org/wiki-backup`, 4095 pages, snapshot 2026-08-29). linux-sunxi.org content is CC BY-SA; the pages named above are the attribution.*
Author
Owner

Dropping this from P2-high to P3-normal, with the evidence rather than as tidying.

This issue has sat at P2 on the strength of "unverified until the next power-cycle". That
power-cycle happened on 2026-08-30 - a full shutdown and cold power-on with the rescue card out -
and the mitigation held:

[  12.061883] eth0: configuring for phy/rgmii-id link mode
[  16.263541] eth0: Link is Up - 1Gbps/Full - flow control rx/tx
State: routable (configured)   Online state: online

eth0 came up with its address, reachable, nothing rebound by hand.

The bug itself is untouched. The PHY still wakes ~4.2 s after the driver configures the link
mode, which is the fault this issue documents. What changed is only that networkd no longer gives
up permanently while waiting, and that is now measured on a real cold boot rather than assumed
from a warm reboot.

So the availability risk is gone and a latent kernel bug remains - which is P3, not P2. It also
stops this competing for attention with #10, where nothing is verified at all.

Two things this does not mean:

  • it is not fixed, and the mitigation is load-bearing. If ActivationPolicy=always-up is ever
    removed, or the PHY's wake grows past networkd's patience, the original symptom returns.
  • the kernel-side fix still wants shaping for upstream - that is #19, which stays where it is.

If the priority is ever revisited, the thing to re-measure is the 12.06 s -> 16.26 s gap, not
whether the interface came up.

Dropping this from **P2-high to P3-normal**, with the evidence rather than as tidying. This issue has sat at P2 on the strength of "unverified until the next power-cycle". That power-cycle happened on 2026-08-30 - a full shutdown and cold power-on with the rescue card out - and the mitigation held: ``` [ 12.061883] eth0: configuring for phy/rgmii-id link mode [ 16.263541] eth0: Link is Up - 1Gbps/Full - flow control rx/tx State: routable (configured) Online state: online ``` eth0 came up with its address, reachable, nothing rebound by hand. **The bug itself is untouched.** The PHY still wakes ~4.2 s after the driver configures the link mode, which is the fault this issue documents. What changed is only that networkd no longer gives up permanently while waiting, and that is now measured on a real cold boot rather than assumed from a warm reboot. So the availability risk is gone and a latent kernel bug remains - which is P3, not P2. It also stops this competing for attention with #10, where nothing is verified at all. Two things this does **not** mean: - it is not fixed, and the mitigation is load-bearing. If `ActivationPolicy=always-up` is ever removed, or the PHY's wake grows past networkd's patience, the original symptom returns. - the kernel-side fix still wants shaping for upstream - that is #19, which stays where it is. If the priority is ever revisited, the thing to re-measure is the 12.06 s -> 16.26 s gap, not whether the interface came up.
Sign in to join this conversation.
No description provided.