How it works¶
Fundamentally, net-dhcp uses the same mechanism as Docker's built-in
bridge driver to wire networking to containers: a bridge on the host
acts as a switch, and veth pairs connect each container's network
namespace to it. Two things differ:
- A bridge on the host, no routing or filtering. Where Docker creates
and manages its own bridges (and routes/filters traffic),
net-dhcpuses an existing bridge on the host, bridged onto the desired local network, or, since docker-net-dhcp v2.3.0, makes one from a spare NIC with-o parent=and removes it with the network (#903). (In macvlan/ipvlan mode the parent is a host NIC instead: see parent-attached modes.) - External addressing. Instead of allocating addresses from a static
pool on the Docker host,
net-dhcprelies on an external DHCP server to provide them.
The parts and what passes between them are drawn in Architecture.
Flow (bridge mode)¶
- A container-creation request is made.
- A
vethpair is created and the host end is connected to the bridge (both interfaces are still in the host namespace at this point). The host end is nameddh-plus twelve hex digits here, whatever the network asked for: the container's name is not known at this point, so a network that sethost_ifnamegets its rename at step 7, where the daemon's answer already is. - A one-shot DHCP acquisition runs on the container end (still in the host namespace). The plugin provides the initial IP address to Docker.
- Docker moves the container end of the
vethpair into the container's network namespace and sets the IP address. At this point that first acquisition is finished with. net-dhcpstarts a persistent DHCP client on the container end of thevethpair whose socket lives in the container's network namespace. The client is not a process: it is the in-tree DHCP library running inside the plugin, so there is nothing for the container to see and nothing to exec. It never configures the link either. The plugin applies the lease via netlink. The link reaches that open as its index and not as a name: the engine renames the container end on its way into the namespace, and a name read before the move can already belong to another link by the time the socket is opened.- The container's name is asked of the daemon once that client is
already leasing, and handed to it when the answer comes; the client
then renews once immediately so the server's table carries the name
within one exchange. The order is this way round because a daemon
that is still starting a container does not answer questions about
it, and the lease does not have to wait for that answer. Two routes
have the name before the client starts and do not take this one: a
register_dnsnetwork, whose option 81 is built when the client is constructed and has no setter for it, and whose DHCPv6 client builds option 39 from the same name at the same time and renews once at start to carry it, since the address was leased atCreateEndpointbefore the name was known; and any attach that had to ask the daemon anyway to find the container's namespace, which is the fallback to the container's process where the sandbox key is refused, and the path that re-adopts a running container after a plugin restart. A container restart is a different thing and is not one of these: Docker drives it as a detach and a re-attach, and the endpoint is rebuilt before the attach begins. The name handed over late goes to the DHCPv4 client only: this route runs withoutregister_dns, and there a DHCPv6 client sends no name, since option 39 would ask the server to register it (#1029). - On a network that set
host_ifnamethe host-side link is renamed now, after the container or after its hostname, at the point the daemon's answer is already in hand and the attach can no longer fail. The generateddh-name stays on the link as an altname, because the plugin re-derives that name and looks the link up by it in several places without ever reading a name back from the kernel, andDeleteEndpointis one of them: a miss there is the normal end of a forced teardown, so a rename with no altname would leave the veth on the bridge for the life of the host, silently. A rename the kernel takes but will not keep the altname on is undone for the same reason. Between the two calls nothing on the host answers to thedh-name, so those lookups go through one reader that waits for a rename in flight; the rename's own lookup is the exception and runs inside it. - The client keeps running, renewing the lease when required, until the container shuts down.
In macvlan and ipvlan mode the shape is the same, with a child interface on a host NIC in place of the veth pair and the bridge; the client lifecycle, the event plumbing, and everything below are identical.
Where the IPAM shape changes it¶
That flow is the --ipam-driver null shape, where step 3 runs inside
CreateEndpoint. Since v2.1.0 the same plugin also answers the
/IpamDriver.* paths on the same socket and the same mux
(pkg/plugin/routes.go),
and a network can name it twice. What moves is where the address is
acquired, and nothing below it:
- The acquisition runs at
RequestAddress, inside the daemon's own IPAM call, before any endpoint exists (pkg/plugin/ipam_reserve.go). It builds a throwaway link of its own, runs the exchange on it, and removes it, so the address is Docker's to hand out by the time the endpoint is created. - The budget is therefore the daemon's plugin-call budget, which comes
from
docker plugin enable --timeoutand which the plugin is never told. A network'slease_timeoutis capped to it for the reservation, and the cap is logged. That is why the flag has to stay at its 30s default; the reference states the rule for operators under Address allocation. - The identity that carries a restarted container's address back is the
client identifier of the previous endpoint, taken from the lease
record. Docker's request carries no hostname and no endpoint id. The
one thing narrower to match on is the MAC, when the user set it: a
kept identity leased under the requesting MAC is that container's own,
and on a
require_mac=truenetwork nothing else is claimed (#1118). With a Docker-generated MAC there is nothing narrower, so the ambiguous case is counted asipam_rebind_ambiguousinstead of guessed. - IPv6 is acquired later, since v2.3.0.
RequestAddressleases the v4 address only; the DHCPv6 exchange runs atCreateEndpointon the endpoint's own link, as the null shape's one-shot does, andCreateEndpointreports the v6 address to the daemon. The v6 lease record is paired with the v4 one: the re-bind that carries a restarted container's v4 address back carries its DUID, IAID and last v6 address too, and a v6 failure that ends the endpoint gives up both records (#960). - The pool binding lives in the network's own state file, and a file
carrying one is stamped schema 2
(
pkg/plugin/state.go). The IPAM handlers read that file and never call Docker: the daemon replaysRequestPooland oneRequestAddressper stored endpoint from insidelibnetwork.New, which runs before the daemon's own API serves, so a handler that asked Docker anything there would deadlock.ipam_replay_hitsandipam_replay_missare what that replay reports. - The persistent client at step 5 is unchanged. So is every path in this document below this section.
The DHCP client is github.com/claymore666/dhcp-golib, the
project's own library, imported as a Go module and pinned to an exact
version in
go.mod.
pkg/dhcp
is the chassis over it: it builds the protocol parameters
(pkg/dhcp/params.go),
owns the namespace and the socket
(pkg/dhcp/chassis.go),
and translates library events into the events
pkg/plugin
has always consumed. Nothing outside
pkg/dhcp
knows which client is underneath.
How the plugin drives the DHCP client¶
- Events arrive on a channel. The library reports each lease change
as a typed event;
pkg/dhcp/chassis.go'stranslategoroutine maps it to the plugin'sbound/renew/nak/leasefailand the per-endpoint goroutine inpkg/plugin/dhcp_manager.goapplies the address, routes, DNS and MTU via netlink; a network that setsmtukeeps that value over the lease's (#1037). There is no argv to build, no environment to scrub, no JSON to parse and no second binary in the image. - The first lease leaves the default route to the engine.
Joinreturns the gateway and starts the client at once, and the engine installs that gateway afterJoinreturns with an add that fails if a default route is already on the link, which fails the container start withfile exists. Until the plugin sees a default route on the link, or the nextrenewevent, a lease or a Router Advertisement that finds none adds nothing. A manager rebuilt at plugin start had noJoinand adds the route as before (#1084). - A renumber inside one subnet puts back what the kernel took with
the old address. The new address is added before the old one is
deleted. With
promote_secondaries=0on the container link, deleting a primary IPv4 address also deletes every secondary in its subnet, and with no IPv4 address left the kernel drops the link's routes. The plugin reads the addresses back after the delete; if the new one is gone it writes it again and restores the routes it listed before the delete, and if that fails it puts the old address and its routes back (#1081). - The emit must not block, and a drop is counted. The only reader is
that per-endpoint goroutine, and it stops reading the moment the
endpoint is torn down. A bare send would park the translate goroutine
forever on the first event that arrives in that window, and a
Leavewhile a renewal is in flight is enough. The counter snapshot the endpoint owes is parked with it. The channel is buffered, the send has a default arm, and a discarded event incrementsDroppedEvents()and logs what was lost. A silent drop and a wedge look identical from outside the package, which is the whole reason the counter exists. - A lapsed lease comes from the client. The plugin derives it from
no deadline of its own. The state machine drops the lease with
ReasonExpiredwhen the expiry timer fires, and reportsReasonNoServerwhen a retransmission budget runs out with nothing usable heard. Both becomeleasefail, which is what movesdhcp_timeouts. The v1.x outage watchdog is gone, and with itOUTAGE_TICK,OUTAGE_GRACEand the lease-lifetime clamp that existed to keep the watchdog's deadline usable.
What has not changed is the physics: a bound endpoint holds a valid
address until its lease runs out, so a LOST address still cannot be
proven before then. What changed is who says it and how precisely.
The 1.x watchdog re-checked on a 30-second tick and then held a
25-second settling time on top of that before it called an outage.
2.0 reports on the client's own schedule. A silent server is visible
before the lease lapses: since v2.1.0 (#940) renewals_unanswered
moves at the first retransmission of a renewal request. RFC 2131
section 4.4.5 sets that wait at "one-half of the remaining time until
T2 ... down to a minimum of 60 seconds", so the 60 seconds is a floor
and the wait is hours on a long lease: about 4.5 hours after T1 on a
24-hour lease, against 12 hours until that lease ends. The address is
still good at that point, which is why the counter is not
healthy-affecting.
- A zero lease lifetime is an infinite lease. The 1.x plugin clamped
an implausibly long option 51 so its watchdog had a usable deadline,
and counted the clamp. The library hands the chassis an explicit "no
expiry" instead, which the caller tests instead of comparing against a
threshold.
- No client keeps state on disk, so nothing is keyed by interface
name. This is where a private mount namespace per client used to be
needed: the 1.x client kept its lease file and its control socket in
directories named after the interface, and two containers both called
eth0 collided on both, with the second client silently handing its
arguments to the first one's socket and never renewing (#332). The
library holds its lease in memory and the plugin holds the durable
copy in its own record, keyed by endpoint. There is no per-client
mount namespace, no tmpfs, and no lease file to go looking for.
- The socket belongs to the endpoint's namespace because of the
thread that created it. newLibClient locks the OS thread, enters
the endpoint's network namespace, opens the client there, and returns
the thread. A raw AF_PACKET socket keeps the namespace it was
created in, which is the property the mount-namespace machinery used
to approximate from outside. If the thread cannot be returned it is
retired and never reused, since a thread left in a container's
namespace would silently give the next caller the wrong one.
How IPv6 is handled¶
Any ipv6_mode but off gives an endpoint a second DHCP client,
in the same shape as its first: one dhcpManager, one library client,
one record. What differs between the modes is what that client does on
the wire. In dhcp it leases the address over DHCPv6, which is what
-o ipv6=true has always meant. In slaac it sends no Solicit and
forms the address from an advertised prefix. In auto it reads the
advertisement and does what it says: the managed-address flag means
DHCPv6, a clear flag means the prefix. It solicits routers and reads
advertisements in all three, which is where the container's IPv6
default route and its link MTU come from. The routes the advertisement
asks for arrive the same way and skip_routes opts out of those; the
default route is not governed by it. Resolvers are a union and not a
choice: the library merges the advertisement's RDNSS and DNSSL into the
same lists a DHCPv6 server's options 23 and 24 fill, with the server's
taking precedence (RFC 8106 section 5.3.1), and propagate_dns decides
whether the result is written into the container at all. Nothing about
the v4 path changes in any of them, which is the whole design. The
maintainer's rule for this milestone was that IPv6 takes the same shape
as IPv4 unless the v4 shape was itself a hack.
The differences that do exist are the ones the protocol forces.
The identity is minted once and stored. DHCPv4 derives its client
identifier from the MAC on every start; DHCPv6 cannot, because RFC 9915
§11 asks for a DUID that "SHOULD NOT change over time if at all
possible" and one mode has no per-endpoint MAC to derive it from. So
resolveIdentity6 mints it at CreateEndpoint and the endpoint's
record carries it (the Identity field, write-once). Bridge and macvlan
get §11.4's DUID-LL over the endpoint MAC, byte for byte the DUID 1.9.0
put on the wire, so an endpoint upgraded from 1.x keeps its address.
ipvlan gets §11.5's DUID-UUID over the endpoint id, because an ipvlan L2
slave inherits the parent link's MAC and every container on one network
would otherwise present the same identity and claim one binding.
The two families do not share a record. A record is keyed on scope
and hardware address, and a dual-stack endpoint has one hardware address
on one network, so the v6 record is filed under Scope6(networkID), the
network id with a #v6 suffix. Without it the two families collide
exactly and whichever bound last answers both resumptions.
Duplicate-address detection happens in the client, and the kernel is
told not to repeat it. RFC 9915 §18.2.10.1 puts the check on the
client; the library runs it and reports the lease only after it passes.
The chassis then installs the address with IFA_F_NODAD, because a
second run costs a tentative window the container cannot use the
address in and can fail where the first passed: RFC 7527 §4.1's
loopback case takes the address out of service entirely. RFC 4429 §3.3
is the same argument from the standards side. The address also carries
RFC 9915 §7.1's two lifetimes, so the kernel can deprecate instead of
deleting (RFC 4862 §5.5.4); expiry itself is still the library's job and
the lifetimes are a belt for a plugin that dies inside the window.
The Router-Advertisement guard is a precondition, and since v2.2.0 it
turns the kernel's own processing OFF. DHCPv6 carries no next hop,
because RFC 9915 §21 defines no router option, and RFC 5942 §4 rule 1
forbids treating the assigned address's prefix as on-link, so somebody
has to process Router Advertisements or the endpoint has an address and
no route. Until v2.2.0 that somebody was the container's kernel. It is
now the plugin's own DHCPv6 client, which reads advertisements off its
socket and reports the gateway, MTU, routes and DNS through
dhcp.Info; the Join answer carries the gateway and the routes, and the
manager rewrites them when a later advertisement changes them (#821).
That answer only reaches an endpoint that HAS a global IPv6 address.
The daemon disables IPv6 on a container link carrying no global IPv6
address, and the kernel refuses every IPv6 route on such a link, so an
answer with an IPv6 half fails the sandbox outright, it does not
degrade, and the plugin cannot clear disable_ipv6 first, because
that runs in the manager goroutine Join spawns, after the daemon has
moved the link and applied the answer. A segment that hands out no
DHCPv6 address therefore gets its MTU, its resolvers on a
propagate_dns network, and no route. On ipv6_mode=slaac and
ipv6_mode=auto the plugin forms the address from the advertisement
and installs it (#818), so where a prefix forms one the link carries a
global address and the route installs beside it. One ending starts an
endpoint without a global address in every mode, those two included:
dhcpv6_not_offered, the verdict for an acquisition that produced no
DHCPv6 address on a segment that never said one was to be had. The
Networks where DHCPv6 offers no address section of
docs/reference.md carries its rows.
ApplyRouterAdvertGuard therefore writes accept_ra=0, autoconf=0
and keep_addr_on_down=1 and reads each back; DHCPClientOptions
refuses a persistent v6 client that does not claim it, and refuses every
other shape that does. accept_ra=0 because a kernel acting on the same
frames would install a second default route beside the plugin's, and
which of the two wins is a metric comparison nobody chose. autoconf=0
because the plugin holds the lease for the address the container uses.
Writing accept_ra=0 purges nothing, which is the part that is easy
to miss. It stops the kernel processing the NEXT advertisement; a route
an earlier one installed stays until its own lifetime runs out, and RFC
4861 §4.2 allows that to be 65535 seconds. The engine brings the link up
in the sandbox at the kernel default before the guard runs, so the
window is real. purgeRouterAdvertRoutes closes it: after the knobs
take, every RTPROT_RA route on the link is deleted. Failures there
fold into router_advert_guard_failures beside the sysctl ones, because
they are one obligation seen twice. The address the kernel may have
formed in the same window is NOT touched; that is #818's.
It all runs in prepareIPv6Link, in one namespace entry. That placement
is a deviation from where the design put it, inside the client's own
setup, and the reason is mechanical: /proc/sys is read-only in the
managed plugin's rootfs, v6_link.go already owns the mount-namespace
unshare that makes it writable, and doing it in the client would mean a
second one. The ORDER the design fixed is preserved exactly:
disable_ipv6 cleared first, then the guard, then the purge, then the
client, which waits for a non-tentative link-local of its own before it
sends anything.
An absent v6 lease is classified. On a stateless or SLAAC segment
there is no DHCPv6 address by definition, and refusing the endpoint
there means no container can start on those networks at all (#868).
classifyV6Absence decides on what the segment said: a Configured
event, the library's own kind for a reply that carried configuration and
no address, is "not offered", whatever the last advertisement's flags
were; otherwise no advertisement at all is "no router", and an
advertisement with the M bit set is fatal. The wire beats the
diagnostic, because a later advertisement on the same link can set M
after the segment has already answered.
How a network chooses its DHCP server¶
dhcp_servers ranks the servers a network may lease from and
dhcp_deny_servers names ones it must never lease from (#111, #669).
The operator-facing rules are in the
driver reference; the shape of the implementation
follows from where the filtering happens.
-
Both lists match the Server Identifier (option 54). The packet's source address is never used. This is the one thing about these options that a 1.x operator has to re-learn. The external client compared the offer's IP source, which meant that behind a DHCP relay every offer looked like it came from the relay and neither list could tell servers apart; option 54 is what the server says it is, and it is also what a renewal is unicast to, so 2.0 filters on it and the relay limitation goes with the change. The two keys agree whenever a server answers directly. They stay DHCPv4-only even now that DHCPv6 is wired in: a v6 entry is refused at
docker network createinstead of applying to nothing, andclientServerListshands a v6 client no lists at all. That is not an oversight deferred: the library'sParams6has no server-list field, so a v6 client that tried to honour one would not compile.The library's predicate is where the edge cases live, and they are decided and not incidental: deny wins over allow for a server named in both; an allow list fails closed on a message that carries no server identifier at all, because "only these servers" that a message can satisfy by omitting the field is not a restriction; and a deny list alone fails open on that same message, because nothing shows it came from a denied server. - The plugin still never sends both lists. The deny list is subtracted from the preference list at parse time, so after
resolveServerPolicythere is one truth about what is allowed and one kind of list to hand down. That subtraction was forced in 1.x, where a configured whitelist switched the blacklist off inside the client and a network setting both would have got a denial nothing enforced; it is kept here because the property it buys is worth more than the redundancy the library would now tolerate: one truth about what is allowed, instead of two composed at the far end. A preference list that denies its way to empty fails the network create, because the alternative is degrading into "accept any server at all", the opposite of what both options were set to achieve. - Ordering is not expressible to the client, so preference is an acquisition-time ladder. The initial acquisition runs one attempt per preferred server, in the operator's order, each restricted to that server alone. The ladder divides the existing acquisition budget instead of extending it, because a preference list must not makedocker runslower and the one-shot atCreateEndpointalready runs against a tight ceiling. The remainder of the division is dropped and never handed to the last tier, so the attempts can only sum to at most the budget. The per-attempt floor (minAttemptBudget, 3s) predates the library and now sits just under its first retransmission at 4s, so a tier that lands on the floor buys one DISCOVER and no retry. The same reading applies to the undivided budget: the library's intervals are 4s, 8s, 16s, 32s with a 64s ceiling and ±1s of jitter, each armed as its packet goes out, so retransmissions land at ~4s, ~12s, ~28s and ~60s, and the defaultlease_timeoutof 34s funds the first three of them. That is not a regression, since the same 3s used to have to pay for a namespace and a process spawn as well, but its original derivation is dead and nothing re-derives it against the library's schedule. At that default the budget pays for eleven attempts (34s divided by 3s, integer division), so a list of eleven or fewer keeps one attempt each, and a longer list has its tail packed into the eleventh. The budget is divided evenly over the attempts it funds, so the slice each attempt gets falls as the list grows, and it stops falling once the packing starts.dhcp_server_tier_fallbackscounts a fall-through to a lower tier, which is the only outside signal that a preferred server has gone quiet while every container still starts;dhcp_server_policy_exhaustedcounts a restricted acquisition where nothing answered, which is otherwise indistinguishable from an ordinary DHCP timeout. - The persistent client gets the whole allowed set. The tier that won does not narrow it. It has to be able to rebind after the preferred server goes away, and an allow list pinned to the winning tier would strand the endpoint with no lease instead of failing over. Preference is an acquisition-time concept; once a lease is held it stays with whoever granted it, because renewal is unicast to that server.
How a lease is checked against the segment¶
The plugin asks whether some other device already holds the address its DHCP server just leased (#524). Since 2.0 the question is asked by the DHCP client itself, as RFC 5227 Address Conflict Detection, from inside the container's own network namespace and on its own link. The chassis no longer asks from the parent. The operator-facing rules and the counters are in the driver reference; this is the mechanism.
- It is part of the acquisition. No separate step runs beside it.
§2.1 sends three ARP Probes with an all-zero sender protocol address,
then waits ANNOUNCE_WAIT before the address may be used; §2.3 sends
two Announcements once it is; §2.4 keeps listening for the whole life
of the lease. A conflict at any point produces a
DHCPDECLINE(RFC 2131 §3.1(5)) and a fresh DISCOVER, which is what makes the DHCP server's own log the outside evidence for the whole thing. - The vantage point moved, and that is what closed the two holes the old check had. The chassis used to send a datagram from the PARENT link to make the kernel resolve the address, and compare the answering MAC with the endpoint's. That could only ever check the address a new endpoint was about to be handed, so an address that changed mid-life was never re-probed, and it needed the parent to carry an address on the leased subnet, because a host answers an ordinary ARP request only if it can route a reply back to the sender. A §2.1.1 Probe carries an all-zero sender protocol address, which Linux answers for any local target without consulting a route, so the bare-parent limitation is gone; and §2.4 covers the rest of the lease's life, so the mid-life hole is gone with it.
- Our own endpoint holds the address too, which is the premise. The old check answered it with macvlan's parent/child isolation plus a MAC comparison. RFC 5227 answers it in the client: a reply whose sender hardware address is the client's own is not a conflict. That is what keeps bridge mode correct, where the host can reach the container and a did-anything-reply check would report every single endpoint as a conflict. The cost is unchanged: a squatter that is another container on the same parent is invisible, excluded by construction and not pending work (#528).
- It costs seconds, and the operator chooses who pays them.
conflict_check=wait(the default) finishes §2.1 before the address is configured, sodocker runwaits 4.0–7.0s;asyncconfigures the address at the DHCPACK and probes behind it, so a conflict found later CHANGES a running container's address;offsends no ARP at all. Thelease_timeoutdefault is derived from the same constants, one DISCOVER retransmission plus the worst probe window, instead of being written down, so the two cannot drift apart. - The phase survives a plugin restart. In
asyncthe address is in use while §2.1 is still running, so the conflict-detection phase is written into the durable lease record and handed back to the next process on resume. Without it a restart inside that window would leave a container holding an address nothing ever finished checking.
How a lease gets handed back¶
By default it does not, and that is deliberate as of v1.9.0 (#800). Two
values are the exceptions: release_lease=on_stop, which releases at
the stop (#962), and release_lease=on_remove, which holds the address
for the restart window first and releases at the end of it (#984).
A lease is a lease. When a container stops, its address stays leased until the lease expires, and if the container comes back before then it asks for the same address and gets it. That is the ordinary DHCP path, and exactly what happens when a physical host on the segment reboots or loses power. A container is a host on this segment and costs the server what one costs.
On a default network neither client releases. The CreateEndpoint
one-shot ends by cancelling its own manager, which drops the lease
locally with ReasonStopped and sends nothing. The record carries the
lease to the persistent client that takes over moments later, which
resumes it as INIT-REBOOT instead of discovering afresh. The persistent
client is stopped at Leave and keeps the address for the container
that may be about to restart. A stop is this process's own shutdown
reported back to it, which is why nothing counts it as a lease loss:
doing so would report one for every container that started
successfully.
What release_lease=on_stop changes. At Leave, and only there,
the endpoint's lease goes back: a DHCPRELEASE (RFC 2131 section 4.4.6)
for IPv4 and a Release (RFC 9915 section 18.2.7) for IPv6, one
datagram per family.
It is built from the lease record, not from a running client, and
that is the difference that makes the option work for the case it
exists for. A container stopped before the plugin's persistent client
attached has no client to ask, and the address it was using came from
the one-shot exchange at CreateEndpoint -- which wrote it into the
same record. So the record holds the address, the identity as sent, the
chaddr and the server, and the release is assembled from those and sent
from the host's own address on the parent interface. Nothing needs the
container's namespace, which may already be gone.
The v6 address comes off the container link first, which section 18.2.7 requires before the exchange may begin; if it cannot be removed, nothing is sent and the address expires on the server's clock instead. The source is the parent's link-local address and never the address being released, which is the same section's second requirement.
Everything else about that teardown follows from the address being
gone. The record is CLOSED, not LEFT, per family, so the next
start cannot resume an address the server has already put back in its
pool. No tombstone is laid, so no other container inherits the MAC and
the addresses beside it. The tombstone is one object carrying both
families' addresses, so either family releasing suppresses it, while the
record of a family whose release did not happen is retained exactly as
under never and stays resumable. releases_sent counts what left the host and
release_failures counts attempts that put nothing on the wire, both
split per family and both moved by the plugin from the outcome of its
own attempt, which is the only place that knows a release was asked for
and did not happen.
Leave is the only path on_stop releases from. Plugin.Close, a
manager displaced by a newer one for the same endpoint, and the cleanup
after docker network rm all stop clients whose containers are still
running, and a release there would tell the server an address is free
while a live container holds it.
What release_lease=on_remove changes. Nothing at Leave: the
endpoint is torn down exactly as under never, tombstone and all, and
one line goes in the log saying the addresses are being kept for the
window. What releases is the record, later, and from a different
place.
Every DeleteEndpoint already retains the endpoint's record with a
deadline one tombstone TTL away, because that is how long a restarting
container may inherit the MAC and address. On an on_remove network
that deadline is also when the address stops being the container's. A
sweeper ticks every 15 seconds, and a retained record whose deadline has
passed by 5 seconds is handed back with the same sender on_stop uses,
built from the same record, and then closed. So the wall clock from
docker stop to the datagram is 65 to 80 seconds, and the window an
operator reasons about is the one they already know from docker
restart. There is no second option, and there is nothing to keep in
step.
Three consequences follow from the deadline living in the record rather than in a timer:
- A plugin that restarts inside the window still releases at the right moment. The deadline was written to the file; the new process rebuilds it and the first sweep after it passes hands the address back. A timer would have died with the process.
docker network rmhands back every address the network still holds, at once, without waiting for deadlines on a network that will not exist. That release runs before the network's stored options are deleted, because it readsrelease_leaseand the parent interface out of them.- The address reserved for an endpoint Docker never created is reached too, which is the paragraph below.
What decides that an address was claimed back is the address, not
the MAC. The sweep looks for another record on the same scope holding
the same address: one in a live phase, or simply a newer one that is not
closed. A container pinned with --mac-address that comes back on a
different address does not hold the old one, and the old one goes
back; a MAC-keyed check would have closed it unsent and leaked it. The
one deliberate exception is an acquisition still in flight under the
same MAC with no address yet: that is treated as a claim, because the
address it is about to be given may be this one. The cost of the
exception is one-sided by design. An in-flight acquisition that lands
somewhere else leaves one address to expire on the server's clock, which
is what never does with every address; the opposite mistake would hand
an address away from under a container that is starting, which is the
duplicate assignment of #524.
And one address never reaches Leave at all. In IPAM mode an
address reserved for an endpoint whose CreateEndpoint then failed is
retained by ReleaseAddress, not released. Retaining it is what
lets a restart policy's next attempt claim the same address back instead
of burning a second lease on the server, and a reservation with no
endpoint reaches no Leave, so nothing on the on_stop path can see
it. On never and on on_stop no DHCPRELEASE goes on the wire for it
and the address is left to expire, exactly as any other host on the
segment leaves one: on an on_stop network that is a real lease the
server granted that nothing hands back, held by the retention deadline
until it expires.
on_remove is what closes that (#984), and it closes it without
touching ReleaseAddress at all. The retention that handler writes
already carries a deadline, and the sweep hands back every retained
record whose deadline has passed. So the retry still gets its window and
its address, and the address is given up afterwards instead of waiting
for the server's clock. Which of the two wins is still a decision: the
retention wins while the window is open, the release wins when it
closes.
Why this changed. Up to v1.8.x the plugin released aggressively. The
external client emitted a RELEASE on a graceful stop, and a background
reclaim handed back the one-shot's address whenever no persistent
client had taken ownership of it (a container that exited before the
attach completed). Both were trying to return an address promptly
instead of letting it sit until expiry. Both raced the tombstone.
A docker restart is a Leave immediately followed by a Join for the
same MAC, and the tombstone exists to promise that Join the same
address. At the moment the release ran, "this endpoint is gone" and
"this endpoint is coming straight back" were indistinguishable, so the
plugin was observed telling the server an address was free in the same
second the container came back to claim it. The reclaim was measured
firing four times on ordinary restarts of live containers.
What was gained was a faster return of an address nobody wanted. What
was risked was an address handed to someone else while a container was
still using it, the duplicate assignment #524 added detection for,
manufactured by the plugin itself. Waiting for expiry has no such
failure mode, so the whole mechanism went: the release itself, the
reclaim, and the orphaned_leases_released and
orphaned_lease_release_failures counters that measured it.
release_lease=on_stop does not bring that mechanism back. What it
sends comes from the endpoint's own live client, inside the container's
sandbox, before anything is torn down, and the tombstone it would have
raced is not written at all for an endpoint that released. The
background reclaim and its synthesised link stay gone.
release_lease=on_remove does send from a background sweep, and it is
the one value that has to answer this paragraph. What the reclaim got
wrong was not that it ran in the background; it was that it could not
tell "this endpoint is gone" from "this endpoint is coming straight
back", because it ran at the moment those two look identical. The
deadline is what tells them apart. Nothing is sent until the window the
tombstone itself promises has run out, so by the time the sweep looks,
a container that was coming back has come back. And the sweep does look:
before sending it re-reads the records and skips any address another
record now holds, which is the restart it would otherwise have raced.
An acquisition in flight under the same MAC with no address yet counts
as a claim for the same reason. What was removed was a release with no
way to see the restart; what is here is a release that waits for it and
then checks.
The surviving teardown counter was renamed to match: what was
lease_release_failures is now client_stop_failures, because a client
that exits badly is all it can still mean.
The cost on a release_lease=never network, the default, is that a
short-lived container's address is unavailable for one lease time. Size
the server's pool and lease time for the churn, the same way you would
for any other population of hosts. release_lease=on_stop is the
setting that buys the address back sooner, at the price of the stop-time
cost above and of the cases where the release cannot be sent, and of
restart stability: an endpoint that released lays no tombstone.
release_lease=on_remove is the setting that buys it back a minute
later and keeps restart stability, at the price of a pool that has to
carry one window's worth of stopped containers, and of a release that is
attempted once and not retried.
How operations on one parent NIC are serialised¶
A parent NIC registers one rx_handler, so it is a macvlan port, an
ipvlan port or a bridge port, and never two of them. Whichever kind asks
second gets EBUSY. That is a kernel rule; one mode per parent stays
the operator-facing constraint. An 802.1Q sub-interface claims no
rx_handler on its parent, so a vlan network's sub-interface sits
beside any of them (#902).
What the plugin can stop is inflicting it on itself. Since v1.6.0
creating an endpoint and the validate_dhcp probe, which holds its link
for a whole DHCP exchange, take a per-parent gate first, so they queue
instead of refusing each other (#486, #549). Every later path that adds
a link to a parent takes the same gate: the IPAM driver's address
reservation (#110), the vlan sub-interface and the two trial children
its removal adds (#902), and the bridge the plugin makes from a spare
NIC (#903).
There used to be a third, the orphaned-lease reclaim, and it was the demanding one: it ran from a goroutine ordered against no Docker request at all. It is gone (#800, see above), which shortens the worst case the gate has to cover but does not remove the need for it. The probe still holds a parent across a DHCP round trip while an endpoint may ask for the other mode.
parent_link_waits counts operations that queued, which is the
mechanism working. parent_link_wait_timeouts counts ones that gave up
and proceeded anyway; they may still succeed, but the budget has stopped
covering the holder's duration.
The gate excludes more than the kernel does, and since v2.1.0 the
reporting says so. Mutual exclusion is per parent and takes no notice of
kind, while the kernel refuses only a pair that both claim the
rx_handler: children of one kind coexist on a parent happily, and a
vlan sub-interface coexists with every kind. So a caller that gives
up waiting for a holder whose kind coexists with its own has spent the
budget and protected nothing, and it goes on to a LinkAdd the kernel
accepts. That case is
counted as a wait, not as a timeout, which leaves the warning counter
meaning what its action text says. It stopped being hypothetical with
the IPAM driver: an address reservation holds a parent across a whole
DHCP exchange, so two containers starting together on one macvlan
network reach the give-up branch every time.
The rule is enforced by two mechanisms, and it is worth being exact about where each one stops, because the guard type exists precisely to replace a prose guarantee about a property nothing checked.
The compiler holds one half: addChildLink takes a guard value, so a
path that never asks for one does not compile. It does not hold the
other half. "Only lockParent makes a guard" is not something Go can
express. The struct's zero value is valid, so
addChildLink(&parentGuard{}, link) compiles and holds nothing, and
lockParent returns exactly that literal on its own no-parent path, so
the shape is already in the file as a pattern to copy. The realistic
route to it is not malice: a new parent-attached call site, a compiler
demanding a guard, and the zero value sitting right there.
That half is enforced by
scripts/check-parent-gate-accounting.sh,
which fails the build on a parentGuard constructed anywhere but
lockParent. A second accounting file,
.github/linkadd-accounting.txt,
covers the way around the type entirely: a direct netlink.LinkAdd,
which the veth pair of bridge mode needs, having no parent to contend
for.
Nor does the guard say which parent it is for, so one taken on one NIC
and handed to a link on another compiles. That is a deliberate non-goal.
Closing it means a runtime comparison, trading a compile error for a log
line, on a mistake no current call site can make. The comment at the top
of
pkg/plugin/parent_gate.go
is the authority on all of this; if this section and that comment ever
disagree, the comment is right.
How state outlives a process¶
Three separate mechanisms keep addresses stable across three different kinds of restart. Their observable behaviour is documented in the driver reference; this is how they are built.
- Per-network options →
STATE_DIR/<network_id>.json. Written atCreateNetworkso the per-endpoint handlers never call back into the Docker API to learn the mode or parent. That callback is precisely what deadlocked the upstream plugin duringdockerdstartup, when the daemon asked it to restore containers using its own networks. On a cache miss the handlers fall back to the API and back-fill the file. - Tombstones → a single file under
STATE_DIR. Written atDeleteEndpoint, consumed at the nextCreateEndpoint, 60-second TTL. Each carries the previous MAC, the last v4 and v6 addresses, and the container hostname. The lookup is keyed by network ID plus hostname, which is why an endpoint keeps its address across a container restart but not across removal of the network itself, since the replacement network has a different ID. Ambiguity is resolved conservatively: when neither side knows the hostname, a tombstone is consumed only if it is the network's single candidate, so concurrent restarts fall back to fresh MACs instead of risking one container's identity being handed to another. - Recovery → a walk of Docker's network list at startup. For every
endpoint on a plugin-served network, a DHCP manager is rebuilt and its
first acquisition requests the address the container already holds
(option 50). It runs synchronously inside plugin construction when the
daemon answers, which is the normal case and finishes before the
socket accepts anything. When the daemon is not serving yet the
walk cannot run there at all: Docker respawns the plugin during its
own startup, so blocking would make us unreachable to the very daemon
we are waiting on. Recovery is handed to
Listeninstead and runs in a goroutine after the socket is up (#383), which puts it in the same window as theJoins a restarting host is issuing.recovery_deferredcounts that postponement; it is not a fault, and only an exhausted retry budget lands onrecovery_failed.
The deferred path is what makes the compare-and-set load-bearing.
Recovery builds a manager and registers it only if no manager is
already registered for the endpoint, in one locked operation. The
check and the registration used to be two, and a Join landing in the
gap had its live manager evicted from the registry while its client
kept running, untracked, unstoppable, and competing with recovery's
fresh client on the same interface. A Join is newer truth than a
recovery walk and may displace it; recovery is older truth and must
yield, which is what a compare-and-set expresses and a
stop-what-I-displaced does not. recovery_already_managed counts an
endpoint left alone. It does not affect health, since that endpoint
has a renewal client, and it is the only outward sign the race
happened at all (#480, #679).
The plugin's identity is a MAC. Both stability mechanisms exist because DHCP servers key on it, and everything above is in service of presenting the same MAC to the server across an event the container did not choose.
How the counters are exposed¶
/Plugin.Health and /metrics are two renderings of one snapshot
(#651). What each of them says is in the
driver reference; the mechanism is that
one function builds that snapshot and both handlers render it and
nothing else.
- One source, because two hand-kept lists rot. A metrics handler
that read the atomics itself would be a second list of every counter,
and this repository has watched that shape decay more than once (#542,
#636), and a stale list is invisible until an alert that should have
fired does not. The exposition is a table keyed by the
HealthResponseJSON tag it renders, and a unit test walks that struct by reflection and fails on a field nobody claimed. Adding a counter without exposing it is a red unit test instead of a hole in somebody's dashboard. - The snapshot is not a single atomic instant, and that is deliberate,
but only for values read one at a time. The counters are read
without a lock, so two of them can be a few nanoseconds apart. For a
counter an operator reads on its own that is harmless: they are
monotonic counters read for rates and alerting, and never an
accounting ledger, and this is what
/Plugin.Healthhas always done. Only the two map lengths take the mutex, because reading a map during a concurrent write is a data race and not a stale number.
It stops being harmless the moment a rendered value is combined
from two of them, and #730 is what that costs. Each family pair is
therefore loaded exactly once into a local, and the aggregate is the
sum of those two locals, and never a second .Load() of a half that
was already read.
- Both family series are stored; neither is derived. Eleven counters
carry a family label. bumpFamily increments exactly one of a
pair, the v4 half or the v6 half, never both and never a third
aggregate, so _v4 and _v6 are peers, and the unsuffixed counter an
operator alerts on is their sum, computed at snapshot time (#212,
#730).
Until v1.8.0 the aggregate was the stored counter and family="ipv4"
was total - v6 at render time, clamped at zero. Two independently
updated counters combined by subtraction can be read in an order
that yields a value below the previous scrape, and Prometheus reads
any counter decrease as a reset, repaying the whole accumulated
value as an increase on the next scrape, so a one-event skew surfaces
as a rate spike of the entire count. The clamp hid the extreme case
and did nothing about the dip. Adding two monotonic counters has no
such failure mode, because neither operand can decrease; subtracting
them does, in either read order.
If a family series ever needs computing instead of reading again, the
arithmetic belongs in healthSnapshot where both halves are loaded
once, and never in the renderer.
- Two exposure paths, and only one of them opens a port. /metrics
is on the plugin socket unconditionally: it costs nothing, and it lets
an operator with a socket-aware scrape path collect metrics without
the plugin listening anywhere. The TCP listener is METRICS_ADDR, and
it is off unless set. The plugin runs with "network": {"type":
"host"} and holds CAP_NET_ADMIN, CAP_NET_RAW, CAP_SYS_ADMIN and
CAP_SYS_PTRACE, so any port it opens is on the host's own network
namespace. Opening one has to be a decision an operator made, and
never something they inherited by upgrading. That listener's mux
carries /metrics and nothing else, so no libnetwork RPC becomes
reachable over TCP; it binds before the call returns, so an unusable
address fails at startup where somebody sees it, instead of in a
goroutine that logs and leaves the plugin running without the endpoint
that was asked for; and a wildcard bind is said out loud instead of
being refused. A wildcard leaks no lease inventory. The exposition is
aggregate counters plus a per-process instance UUID, with no endpoint
IDs, container names, addresses or MACs, which
SECURITY.md
promises and TestMetricsExposition_NoPerEndpointIdentifiers pins. It
is this plugin's operational telemetry, published on every interface
the host has, since the plugin runs in the host's network namespace
(#709).
- The socket's mode is pinned, and never inherited. Serving
/metrics on the plugin socket is unchanged ground only because that
socket is root-only: anything able to read it can already call every
RPC. A UNIX socket is created 0777 &^ umask, so until #687 that
property was whatever umask the plugin runtime happened to hand us:
true under the usual 0022, false under 0002, and nothing said
which. Listen now chmods the socket to 0600 and refuses to serve
if it cannot, since an unknown mode is exactly the state being guarded
against. Only the daemon speaks this protocol and it connects as root,
so nothing legitimate needs group or other.
Running the tests¶
Four loops, cheapest first. Only the last needs root or a plugin.
| what | command | needs |
|---|---|---|
| the Go unit tests | go test ./... |
nothing; seconds |
| the suite's own guards | go test ./test/integration/harness/ |
nothing |
| the whole fast CI lane | make check |
nothing; about a minute |
| both integration suites | sudo make integration-local |
root, Docker |
make check is the one to run before pushing. It runs the same gates as
the Test workflow's two fast jobs: test for build, vet, format, the
race suite and the short fuzz, and policy-gates for every check-*.sh
and the gate self-tests (#829 split them, and both are required
contexts). It needs no privileges and mutates no host state, so the
answer you get locally is the answer CI will give.
The fuzz step runs four native Go targets, one invocation each, on a budget counted in executions and not in wall clock (#324): the resolv.conf renderer, the host-side link name, the DHCPv6 identity blob and the IPAM PoolID. Their seed corpora also run as ordinary tests in the race step, which is not the same thing: a seed cannot find an input nobody has generated yet. The wire codec's own fuzzing lives in the library module.
It was a no-op until #1010. The two targets it named belonged to the 1.x
lease parsers and had been deleted with them, and go test -fuzz over a
package with no matching target prints PASS and exits 0, so a required
check ran for weeks with one possible verdict.
scripts/check-fuzz-budget.sh
now resolves every name in the step against the package beside it, and
refuses a target in the tree that the step never fuzzes. That is why the
step is four spelled-out invocations and not a loop: a name assembled at
run time is a name the gate cannot resolve. It asks both questions of
scripts/local-lane.sh as well, because the lane carries the same four
invocations and this page tells you it gives the answer CI will give: a
rename in one file alone goes red.
The lane's contents live in
scripts/local-lane.sh,
and
scripts/check-local-lane.sh
fails CI if that file lists fewer gates than the workflow runs; a local
target that hand-listed them would quietly cover less the first time a
gate was added (#636, the same shape as #542).
Everything it does not do is declared and never merely absent.
scripts/local-lane.sh --list-exempt prints the list with reasons, and
that is the place to read it instead of a count written here, which has
already gone stale once. The reasons fall into two kinds: gates that
need the pull request that does not exist locally (a commit range, a
title, a body, or the base ref the PR is opened against), and gates that
need the network. In every case a local answer would be a different
answer and never a cheaper one.
A step whose tool is missing (staticcheck, actionlint, shellcheck)
is skipped loudly and named in the summary instead of passing
silently. STRICT=1 make check turns any skip into a failure. Use that
anywhere a green exit is read as coverage instead of by a person who can
see the summary.
CI shards both suites (#381, #468, #877, D41): the main suite across
nine jobs and the failure suite across two, plus one hosted job that
builds the plugin all eleven install. integration-local deliberately
does not shard, because a local run is one machine, so sharding would
serialise the shards and only add overhead. To reproduce a single CI
shard, sudo make integration-test-shard SHARD=1 OF=9 SUITE=main, or
SHARD=1 OF=2 SUITE=failure; SUITE defaults to main.
Both counts live in
.github/workflows/integration.yml's
matrix, beside the measurement that derives them from the five-minute
budget (D41);
scripts/check-durations-table.sh
keeps the weights that partition them honest, over both suites, and
scripts/test-integration-shard.sh
proves the two partitions together cover the roster exactly once.
The partitioner itself refuses, naming the file and line, a test under
test/integration/ that is neither on the roster nor in a package the
shard target runs whole (#866).
Use integration-local.
make integration-test and make integration-test-failure only run go
test. Building and installing the plugin is a separate chain (make
create enable). CI never diverges because the workflow does the build
as its own step before calling either target. A local run has no such
guarantee, so it tests whatever plugin happens to be installed.
That is not a hypothetical. While validating #374 a stale installed
build made two tests fail for reasons unrelated to the branch, and
made the health floor report clean for counters that build could not
publish at all. Wrong in both directions, from one cause. Rebuilding
reproduced CI exactly.
integration-local chains integration-cleanup create enable
integration-test integration-test-failure, so neither a stale plugin
nor a previous run's leftovers can be what you measure. The cleanup
step mirrors the CI job's own first step: local runs are the only
place that state accumulates, because CI's runners are ephemeral.
The two suite targets deliberately do not depend on a rebuild: CI calls them in sequence between its own build and teardown, and a rebuild dependency there would reinstall the plugin between the two suites, recycling it mid-run and resetting the health floor's observation window with it.
Reading the output¶
Both suites tee to test/integration/logs/. At the end of each, the
health floor prints a verdict for the whole run:
HEALTH FLOOR: clean ... over the whole Ns run (plugin up Ms, ...)says the plugin was up throughout and nothing healthy-affecting moved. The run's duration and the plugin's uptime are separate numbers and are always printed as such: locally the plugin often long predates the suite, and where the gap is large the line says by how much, because the counters are cumulative and carry that earlier history too.HEALTH FLOOR: clean over the last Ns of an Ms runsays the plugin restarted mid-suite, so the counters only cover the tail. The whole-run fault census covers the rest.HEALTH FLOOR: clean ... over the plugin's Ns uptimesays the suite's own duration could not be measured, so no coverage claim is made instead of one being invented.PLUGIN FAULTS: N across the whole runis read from the log and not from the counters, so it survives a restart. Any non-zero value fails the run.
A run that cannot read the plugin log fails instead of reporting clean: an unreadable instrument is not a clean result.
If a local run disagrees with CI¶
Check, in order:
- Did you build? Use
sudo make integration-localand not a bare suite target. - Is a previous run's state still around?
sudo make integration-cleanup.integration-localnow does this for you; you only need it by hand after running a suite target directly. A single leftover container fails an unrelated test with a name conflict and reads exactly like a regression.
The interface_name tests (#125) are not a local-vs-CI divergence: they
probe whether the engine applies a remote driver's DstName and skip
when it does not. The probe (engineAppliesIfname, used by
TestInterfaceName_MultiNetworkDeterministic) runs a throwaway
container and checks the interface the engine actually created. There is
no version threshold to hit. The upstream fix (moby/moby#52866,
stopping the remote-driver proxy from dropping DstName) merged to moby
master on 2026-08-26, is milestoned for engine 29.8.0, and that engine
was released on 2026-09-03. The lane's engine is 29.8.1, read from the
run's Fixture engine drift step, so the probe now succeeds there and
the dependent tests run. They still skip on any box whose engine is
older, and a skip there is expected and is not a signal that the run
diverged.
The attach budget under load¶
AWAIT_TIMEOUT caps the attach that follows a Join. When it runs out
with the container still running, the plugin counts
join_start_failures and the container keeps an address nothing
renews. Whether a small, loaded host reaches that state is a
measurement, not a reading of the code, and
scripts/vm-load-test.sh
takes it
(#969,
under
#403).
It builds a throwaway VM on the developer's own machine: 2 vCPU, 2 GB,
plain qemu with KVM as an ordinary user, user-mode networking with ssh
forwarded to loopback, nothing on the LAN and no root on the host. Inside
it, the engine, this tree's plugin, and the bridge fixture shape the
integration suite uses: a Linux bridge with dnsmasq bound to it. It then
starts bursts of 10, 20 and 50 containers at once, three times each,
at idle and under three levels of stress-ng pressure, and prints per
burst the deltas
of the attach buckets from /Plugin.Health beside what the DHCP server's
lease file says, with the VM's steal and iowait over the burst so a row
taken while the host had other tenants is marked as such. When an
endpoint is not bound after the burst settled, the address on the
container's link is read from inside the container and checked against
the lease file, and the entry counts as the container's lease only
under its own id: that is the difference between a stale address, an
address leased to somebody else, and no address. A burst that attached nothing and a counter that went
backwards are refused, not printed as a row; a loaded level whose load
average never left the floor is printed marked refused, the table says
which bursts those were, and the run exits non-zero. The lease column
counts leases whose hostname is the container id the persistent client
sends at Join, beside the size of the lease file, and it is evidence
only once one burst has proven that key.
| what | command | needs |
|---|---|---|
| its own logic | scripts/vm-load-test.sh --self-test |
nothing; seconds |
| the matrix | scripts/vm-load-test.sh |
/dev/kvm readable and writable, qemu, mtools, Docker for the rootfs build; about three hours of bursts at three repeats, before provisioning and the holds between bursts; VMLT_LEVELS, VMLT_BURSTS and VMLT_REPEATS shrink it |
The matrix is not part of CI and not a gate, while its self-test and the gate test run in CI like every gate self-test: a load measurement varies from run to run and would cry wolf as a red check. It is run by hand when the Join path changes, and each run prints its table as markdown.
Request fixtures¶
pkg/plugin/testdata/requests/
holds the raw request bodies the Docker daemon actually sent, recorded
during an integration run. The unit tests in
pkg/plugin/fixtures_test.go
replay them instead of hand-building CreateEndpointRequest /
JoinRequest values.
The difference matters more than it looks. A hand-built request asserts
the code against our model of what libnetwork sends. When the model
and the daemon disagree, every unit test still passes and the
disagreement surfaces on a privileged runner, or in production. That is
not hypothetical: stable_lease was designed against an assumed
CreateEndpoint payload and had to be reverted from v1.3.0 once the
endpoint identity turned out to be unresolvable in the docker run and
Compose flows (#298, #219). The request shape was the defect.
The handler decodes with DisallowUnknownFields, so a field the daemon
adds and we do not model is a 400 at runtime and never a warning. The
fixture tests replay through that same parser, which turns "the engine
started sending something new" into a unit-test failure instead of a
container that will not start.
Layout¶
pkg/plugin/testdata/requests/
macvlan-run/
manifest.json engine, date, commit, flow
0001-NetworkDriver.CreateNetwork.json one raw body per call,
0002-NetworkDriver.CreateEndpoint.json numbered in the order the
... daemon issued them
bridge-run/
macvlan-restart/
Three flows, because the flows are where the payloads differ. That is exactly how #298 got through.
Regenerating them¶
Dispatch the Capture fixtures workflow. It runs on the integration
lane, so the capture happens against the daemon the suite actually talks
to, and it opens a pull request with the re-recorded bodies:
A pull request instead of a push, deliberately. A changed request body
on an engine bump is a finding, the signal #218 and #125 are blocked on,
so the diff wants eyes instead of an automatic commit. Before it opens
anything the job re-runs check-fixture-engine-drift.sh against its own
output, so a capture that recorded nothing fails there instead of on
somebody else's pull request days later. check-capture-lane.sh keeps
the job on the lane, which is the half the drift gate cannot check: on a
hosted runner the recorded engine and the checked engine move together,
agree with each other, and both describe a daemon the suite never speaks
to.
By hand¶
The workflow drives one command, and you can run it yourself on a host with Docker and the integration prerequisites:
Capture against the daemon the suite runs, and never the one your shell
talks to. These are not always the same machine's engine: the
integration job runs inside the CI runner container, against that
container's nested dockerd, which can be several minor versions ahead
of the host's. A capture taken on the host is a recording of a daemon
the suite never talks to, and check-fixture-engine-drift.sh will
reject it, which is exactly what happened the first time these fixtures
met the gate (26.1 recorded, 29.7 running). To record against the lane's
engine, run the capture inside the runner image:
$ docker run -d --name dh-capture --privileged -v "$PWD":/work \
--entrypoint bash ghcr.io/claymore666/dhcp-ci-runner:latest \
-c 'RUNNER_JIT_CONFIG=x /entrypoint.sh || true; sleep infinity'
$ docker exec -w /work -e PATH=/usr/local/go/bin:$PATH \
-e SUDO_UID=$(id -u) -e SUDO_GID=$(id -g) \
dh-capture make capture-fixtures CAPTURE_COMMIT=$(git rev-parse --short HEAD)
The entrypoint brings up the same supervised daemon the suite uses and then
fails its runner exec, leaving that daemon running; go lives in
/usr/local/go/bin, which a bare docker exec does not put on PATH; and
SUDO_UID/SUDO_GID make the recipe hand the regenerated files back to
you instead of leaving them owned by root.
It builds the instrumented (-cover) plugin, sets REQUEST_CAPTURE_DIR
on it, runs one integration test per flow into a cleared directory, and
writes each flow's manifest.json. Pass CAPTURE_COMMIT from the
unprivileged shell as shown: the recipe runs as root against a checkout
you own, and git refuses that as dubious ownership, which would leave the
commit field empty and produce a capture nobody can attribute.
REQUEST_CAPTURE_DIR is declared in
config-cover.json
only, the same place GOCOVERDIR lives, and for the same reason. It is
test instrumentation, so the shipped manifest never grows a setting
whose only use is regenerating this repository's fixtures. With it
unset, captureHandler returns the mux unchanged and the plugin carries
no extra allocation, syscall, or failure mode.
Why they cannot quietly rot¶
A fixture nobody refreshes is a fossilised assumption that agrees with itself forever. It is the same "asserts our model" problem, now with a green test sitting next to it, which is worse, because it looks like evidence. Three things stop that:
- A missing or empty fixture fails.
loadFixtureFlowscallst.Fatalf, nevert.Skip; a suite that replayed nothing would otherwise report green. - A manifest without provenance fails. Empty
engine,captured,commitorflowis an error, because a capture nobody can date cannot be reviewed for staleness. scripts/check-fixture-engine-drift.shcompares each manifest's engine against the daemon the integration suite actually runs, and fails on amajor.minordifference. It runs in the self-hosted suite job, which is the only host that knows what that engine is; patch releases and distro suffixes (26.1.5vs26.1.5+dfsg1) are not drift. Its self-test isscripts/test-check-fixture-engine-drift.sh.
What to do when the unknown-field test fails¶
It is not automatically a defect. A new field may be irrelevant to us. It means the request contract moved and somebody has to decide, which is the point, because today nothing else would say it moved at all. Model the field, or record why it is ignored.
Issue #218 (stable MAC) is waiting on exactly this signal: it needs
netlabel.EndpointName to arrive at CreateEndpoint, and the captures
confirm that field is absent on engine 29.8. The day a capture from a
newer engine carries it, the test names it.
Issue #125 is not covered by this signal, and that is worth
stating because the shape invites the assumption. Its blocker is on the
response side (the engine honouring the plugin's DstName at Join,
moby/moby#52866); the option itself has always been forwarded in the
request. No request capture will ever change when that fix ships, so the
thing that detects it is the behavioural probe in the integration suite
and never these fixtures. The 26.1 -> 29.7 re-record is the worked
example: it introduced com.docker.network.enable_ipv4 on
CreateNetwork, which is carried inside Options (a map) and so costs
nothing, but it arrived unannounced and the fixtures are what showed it.
See also¶
- Driver reference documents every option, counter and behaviour.
- Bridge mode and macvlan / ipvlan cover the host setup each mode needs.