Docker: reconcile the daemon's firewall rules with the board's nft ruleset #66

Closed
opened 2026-08-30 11:38:48 +00:00 by tiagoagueda · 1 comment
Owner

Split out of the Docker enablement work.

Measured, before Docker is installed

  • nft list ruleset - a hand-written table inet filter whose input chain is
    type filter hook input priority filter; policy drop, 25 lines total
  • net.ipv4.ip_forward = 0
  • Debian 13, so iptables is iptables-nft by alternatives default

What Docker will do to that

Starting the daemon sets ip_forward=1 and installs DOCKER, DOCKER-USER and
DOCKER-ISOLATION-* chains. The consequence that surprises people:

Published ports are DNAT-ed in nat/PREROUTING and filtered in filter/FORWARD,
not in INPUT.
The policy drop on the input chain does not protect them.
-p 8080:80 is reachable from the LAN even though the host firewall looks closed.

Same class as the well-known Docker/UFW interaction. Not a bug to fix - a posture to
choose deliberately.

The backend choice is genuinely open

The Tier 0 kernel change enabled both paths on purpose:

  • legacy xtables: NETFILTER_XTABLES_LEGACY + IP_NF_IPTABLES_LEGACY and its tables
  • nftables: NFT_COMPAT, NFT_NAT, NFT_MASQ, NFT_FIB_*

So the kernel does not force the decision. Make it explicitly and write it down - the
two backends put rules in different places and a mixed setup is very hard to reason
about six months later.

Options

  1. Publish only to 127.0.0.1 and front everything with a reverse proxy whose own port
    is governed by the existing input chain. Keeps one place to reason about exposure.
  2. Use DOCKER-USER for allow/deny - Docker leaves that chain alone across restarts.
  3. "iptables": false and hand-manage NAT. Maximum control, and you own every bug.

Done when

  • a published test port is verified not reachable from off-board unless intended
  • the hand-written nft ruleset still applies after systemctl restart docker and after
    a reboot - verified, not assumed
  • ip_forward being on is a recorded decision rather than a side effect
  • whatever is chosen lives in provision.sh
Split out of the Docker enablement work. ## Measured, before Docker is installed - `nft list ruleset` - a hand-written `table inet filter` whose input chain is `type filter hook input priority filter; policy drop`, 25 lines total - `net.ipv4.ip_forward = 0` - Debian 13, so `iptables` is **iptables-nft** by alternatives default ## What Docker will do to that Starting the daemon **sets `ip_forward=1`** and installs `DOCKER`, `DOCKER-USER` and `DOCKER-ISOLATION-*` chains. The consequence that surprises people: > **Published ports are DNAT-ed in `nat/PREROUTING` and filtered in `filter/FORWARD`, > not in `INPUT`.** The `policy drop` on the input chain does not protect them. > `-p 8080:80` is reachable from the LAN even though the host firewall looks closed. Same class as the well-known Docker/UFW interaction. Not a bug to fix - a posture to choose deliberately. ## The backend choice is genuinely open The Tier 0 kernel change enabled **both** paths on purpose: - legacy xtables: `NETFILTER_XTABLES_LEGACY` + `IP_NF_IPTABLES_LEGACY` and its tables - nftables: `NFT_COMPAT`, `NFT_NAT`, `NFT_MASQ`, `NFT_FIB_*` So the kernel does not force the decision. Make it explicitly and write it down - the two backends put rules in different places and a mixed setup is very hard to reason about six months later. ## Options 1. Publish only to `127.0.0.1` and front everything with a reverse proxy whose own port is governed by the existing input chain. Keeps one place to reason about exposure. 2. Use `DOCKER-USER` for allow/deny - Docker leaves that chain alone across restarts. 3. `"iptables": false` and hand-manage NAT. Maximum control, and you own every bug. ## Done when - a published test port is verified **not** reachable from off-board unless intended - the hand-written nft ruleset still applies after `systemctl restart docker` and after a reboot - verified, not assumed - `ip_forward` being on is a recorded decision rather than a side effect - whatever is chosen lives in `provision.sh`
Author
Owner

Done, and the board was in a worse state than this issue described.

What was actually happening

The hand-written inet filter ruleset has chain forward { policy drop; }, and Docker installs
its own base chain at the same hook (ip filter FORWARD). Both are evaluated, so a packet
must survive both - and ours had no rules at all. Measured, not reasoned about:

test result
container -> internet (ping 1.1.1.1) 100% packet loss
container DNS bad address
published port, from off-board timed out
published port, from the board itself worked

Container networking was entirely broken. My earlier "it works" check was worthless: the
image pull is host traffic, and host->container over loopback never traverses FORWARD.

And the hazard this issue was opened about was real but inverted. Published ports were not
exposed - they were unreachable, by the same drop that broke the containers. So the obvious
"fix" (relax the forward policy) would have restored container networking and exposed every
published port to the LAN in one stroke
, with the input chain's policy drop looking like it
was still protecting things.

The decision, implemented

Option 1 from the original post. The forward chain now allows egress specifically rather than
relaxing the policy:

ct state established,related accept
ct state invalid drop
iifname "docker0" accept
iifname "br-*" accept        # user-defined networks get their own bridge
counter comment "forward-dropped"

Nothing inbound to a container is accepted. To publish a service: bind it to the loopback
(-p 127.0.0.1:8080:80) and front it with a reverse proxy on the host, whose own port is
governed by the input chain. There is a commented, worked example in the file for deliberately
exposing a container anyway.

Two things had to change for that to hold

  1. /etc/nftables.conf no longer does flush ruleset. Docker keeps its tables in the same
    ruleset; flushing everything deletes them and nothing recreates them - container networking
    simply stays broken while the firewall looks healthy. It deletes only inet filter now.
  2. Debian's nftables.service tears down with nft flush ruleset regardless of the config
    file.
    A drop-in scopes ExecStop the same way. Without it, systemctl restart nftables
    still wiped Docker's tables - observed, then confirmed fixed.

Verified

container egress + DNS works
published on 0.0.0.0:8099, from ouranos blocked
published on 127.0.0.1:8098, from ouranos blocked
ssh port 22 from ouranos (control) reachable
survives systemctl restart docker yes
survives systemctl restart nftables yes - Docker's tables intact
survives a reboot yes - all tables, egress works, port still blocked

The control matters: without it, "blocked" and "the test is broken" look identical. My first
attempt at this test was ambiguous because curl -w printed on both paths.

ip_forward=1 is now a recorded consequence rather than a side effect - Docker sets it, and the
forward chain above is what makes that safe.

Done, and the board was in a worse state than this issue described. ## What was actually happening The hand-written `inet filter` ruleset has `chain forward { policy drop; }`, and Docker installs its **own** base chain at the same hook (`ip filter FORWARD`). Both are evaluated, so a packet must survive both - and ours had no rules at all. Measured, not reasoned about: | test | result | |---|---| | container -> internet (`ping 1.1.1.1`) | **100% packet loss** | | container DNS | **`bad address`** | | published port, from off-board | **timed out** | | published port, from the board itself | worked | **Container networking was entirely broken.** My earlier "it works" check was worthless: the image pull is host traffic, and host->container over loopback never traverses FORWARD. And the hazard this issue was opened about was real but *inverted*. Published ports were not exposed - they were unreachable, by the same drop that broke the containers. So the obvious "fix" (relax the forward policy) would have restored container networking **and exposed every published port to the LAN in one stroke**, with the input chain's `policy drop` looking like it was still protecting things. ## The decision, implemented Option 1 from the original post. The forward chain now allows **egress specifically** rather than relaxing the policy: ``` ct state established,related accept ct state invalid drop iifname "docker0" accept iifname "br-*" accept # user-defined networks get their own bridge counter comment "forward-dropped" ``` Nothing inbound to a container is accepted. To publish a service: bind it to the loopback (`-p 127.0.0.1:8080:80`) and front it with a reverse proxy on the host, whose own port **is** governed by the input chain. There is a commented, worked example in the file for deliberately exposing a container anyway. ## Two things had to change for that to hold 1. **`/etc/nftables.conf` no longer does `flush ruleset`.** Docker keeps its tables in the same ruleset; flushing everything deletes them and nothing recreates them - container networking simply stays broken while the firewall looks healthy. It deletes only `inet filter` now. 2. **Debian's `nftables.service` tears down with `nft flush ruleset` regardless of the config file.** A drop-in scopes `ExecStop` the same way. Without it, `systemctl restart nftables` still wiped Docker's tables - observed, then confirmed fixed. ## Verified | | | |---|---| | container egress + DNS | works | | published on `0.0.0.0:8099`, from ouranos | **blocked** | | published on `127.0.0.1:8098`, from ouranos | blocked | | ssh port 22 from ouranos (control) | reachable | | survives `systemctl restart docker` | yes | | survives `systemctl restart nftables` | yes - Docker's tables intact | | survives a **reboot** | yes - all tables, egress works, port still blocked | The control matters: without it, "blocked" and "the test is broken" look identical. My first attempt at this test was ambiguous because `curl -w` printed on both paths. `ip_forward=1` is now a recorded consequence rather than a side effect - Docker sets it, and the forward chain above is what makes that safe.
Sign in to join this conversation.
No description provided.