Skip to content

How it works

Fundamentally, net-dhcp uses the same mechanism as Docker's built-in bridge driver to wire networking to containers: a bridge on the host acts as a switch, and veth pairs connect each container's network namespace to it. Two things differ:

  • A bridge on the host, no routing or filtering. Where Docker creates and manages its own bridges (and routes/filters traffic), net-dhcp uses an existing bridge on the host, bridged onto the desired local network, or, since docker-net-dhcp v2.3.0, makes one from a spare NIC with -o parent= and removes it with the network (#903). (In macvlan/ipvlan mode the parent is a host NIC instead: see parent-attached modes.)
  • External addressing. Instead of allocating addresses from a static pool on the Docker host, net-dhcp relies on an external DHCP server to provide them.

The parts and what passes between them are drawn in Architecture.

Flow (bridge mode)

  1. A container-creation request is made.
  2. A veth pair is created and the host end is connected to the bridge (both interfaces are still in the host namespace at this point). The host end is named dh- plus twelve hex digits here, whatever the network asked for: the container's name is not known at this point, so a network that set host_ifname gets its rename at step 7, where the daemon's answer already is.
  3. A one-shot DHCP acquisition runs on the container end (still in the host namespace). The plugin provides the initial IP address to Docker.
  4. Docker moves the container end of the veth pair into the container's network namespace and sets the IP address. At this point that first acquisition is finished with.
  5. net-dhcp starts a persistent DHCP client on the container end of the veth pair whose socket lives in the container's network namespace. The client is not a process: it is the in-tree DHCP library running inside the plugin, so there is nothing for the container to see and nothing to exec. It never configures the link either. The plugin applies the lease via netlink. The link reaches that open as its index and not as a name: the engine renames the container end on its way into the namespace, and a name read before the move can already belong to another link by the time the socket is opened.
  6. The container's name is asked of the daemon once that client is already leasing, and handed to it when the answer comes; the client then renews once immediately so the server's table carries the name within one exchange. The order is this way round because a daemon that is still starting a container does not answer questions about it, and the lease does not have to wait for that answer. Two routes have the name before the client starts and do not take this one: a register_dns network, whose option 81 is built when the client is constructed and has no setter for it, and whose DHCPv6 client builds option 39 from the same name at the same time and renews once at start to carry it, since the address was leased at CreateEndpoint before the name was known; and any attach that had to ask the daemon anyway to find the container's namespace, which is the fallback to the container's process where the sandbox key is refused, and the path that re-adopts a running container after a plugin restart. A container restart is a different thing and is not one of these: Docker drives it as a detach and a re-attach, and the endpoint is rebuilt before the attach begins. The name handed over late goes to the DHCPv4 client only: this route runs without register_dns, and there a DHCPv6 client sends no name, since option 39 would ask the server to register it (#1029).
  7. On a network that set host_ifname the host-side link is renamed now, after the container or after its hostname, at the point the daemon's answer is already in hand and the attach can no longer fail. The generated dh- name stays on the link as an altname, because the plugin re-derives that name and looks the link up by it in several places without ever reading a name back from the kernel, and DeleteEndpoint is one of them: a miss there is the normal end of a forced teardown, so a rename with no altname would leave the veth on the bridge for the life of the host, silently. A rename the kernel takes but will not keep the altname on is undone for the same reason. Between the two calls nothing on the host answers to the dh- name, so those lookups go through one reader that waits for a rename in flight; the rename's own lookup is the exception and runs inside it.
  8. The client keeps running, renewing the lease when required, until the container shuts down.

In macvlan and ipvlan mode the shape is the same, with a child interface on a host NIC in place of the veth pair and the bridge; the client lifecycle, the event plumbing, and everything below are identical.

Where the IPAM shape changes it

That flow is the --ipam-driver null shape, where step 3 runs inside CreateEndpoint. Since v2.1.0 the same plugin also answers the /IpamDriver.* paths on the same socket and the same mux (pkg/plugin/routes.go), and a network can name it twice. What moves is where the address is acquired, and nothing below it:

  • The acquisition runs at RequestAddress, inside the daemon's own IPAM call, before any endpoint exists (pkg/plugin/ipam_reserve.go). It builds a throwaway link of its own, runs the exchange on it, and removes it, so the address is Docker's to hand out by the time the endpoint is created.
  • The budget is therefore the daemon's plugin-call budget, which comes from docker plugin enable --timeout and which the plugin is never told. A network's lease_timeout is capped to it for the reservation, and the cap is logged. That is why the flag has to stay at its 30s default; the reference states the rule for operators under Address allocation.
  • The identity that carries a restarted container's address back is the client identifier of the previous endpoint, taken from the lease record. Docker's request carries no hostname and no endpoint id. The one thing narrower to match on is the MAC, when the user set it: a kept identity leased under the requesting MAC is that container's own, and on a require_mac=true network nothing else is claimed (#1118). With a Docker-generated MAC there is nothing narrower, so the ambiguous case is counted as ipam_rebind_ambiguous instead of guessed.
  • IPv6 is acquired later, since v2.3.0. RequestAddress leases the v4 address only; the DHCPv6 exchange runs at CreateEndpoint on the endpoint's own link, as the null shape's one-shot does, and CreateEndpoint reports the v6 address to the daemon. The v6 lease record is paired with the v4 one: the re-bind that carries a restarted container's v4 address back carries its DUID, IAID and last v6 address too, and a v6 failure that ends the endpoint gives up both records (#960).
  • The pool binding lives in the network's own state file, and a file carrying one is stamped schema 2 (pkg/plugin/state.go). The IPAM handlers read that file and never call Docker: the daemon replays RequestPool and one RequestAddress per stored endpoint from inside libnetwork.New, which runs before the daemon's own API serves, so a handler that asked Docker anything there would deadlock. ipam_replay_hits and ipam_replay_miss are what that replay reports.
  • The persistent client at step 5 is unchanged. So is every path in this document below this section.

The DHCP client is github.com/claymore666/dhcp-golib, the project's own library, imported as a Go module and pinned to an exact version in go.mod. pkg/dhcp is the chassis over it: it builds the protocol parameters (pkg/dhcp/params.go), owns the namespace and the socket (pkg/dhcp/chassis.go), and translates library events into the events pkg/plugin has always consumed. Nothing outside pkg/dhcp knows which client is underneath.

How the plugin drives the DHCP client

  • Events arrive on a channel. The library reports each lease change as a typed event; pkg/dhcp/chassis.go's translate goroutine maps it to the plugin's bound / renew / nak / leasefail and the per-endpoint goroutine in pkg/plugin/dhcp_manager.go applies the address, routes, DNS and MTU via netlink; a network that sets mtu keeps that value over the lease's (#1037). There is no argv to build, no environment to scrub, no JSON to parse and no second binary in the image.
  • The first lease leaves the default route to the engine. Join returns the gateway and starts the client at once, and the engine installs that gateway after Join returns with an add that fails if a default route is already on the link, which fails the container start with file exists. Until the plugin sees a default route on the link, or the next renew event, a lease or a Router Advertisement that finds none adds nothing. A manager rebuilt at plugin start had no Join and adds the route as before (#1084).
  • A renumber inside one subnet puts back what the kernel took with the old address. The new address is added before the old one is deleted. With promote_secondaries=0 on the container link, deleting a primary IPv4 address also deletes every secondary in its subnet, and with no IPv4 address left the kernel drops the link's routes. The plugin reads the addresses back after the delete; if the new one is gone it writes it again and restores the routes it listed before the delete, and if that fails it puts the old address and its routes back (#1081).
  • The emit must not block, and a drop is counted. The only reader is that per-endpoint goroutine, and it stops reading the moment the endpoint is torn down. A bare send would park the translate goroutine forever on the first event that arrives in that window, and a Leave while a renewal is in flight is enough. The counter snapshot the endpoint owes is parked with it. The channel is buffered, the send has a default arm, and a discarded event increments DroppedEvents() and logs what was lost. A silent drop and a wedge look identical from outside the package, which is the whole reason the counter exists.
  • A lapsed lease comes from the client. The plugin derives it from no deadline of its own. The state machine drops the lease with ReasonExpired when the expiry timer fires, and reports ReasonNoServer when a retransmission budget runs out with nothing usable heard. Both become leasefail, which is what moves dhcp_timeouts. The v1.x outage watchdog is gone, and with it OUTAGE_TICK, OUTAGE_GRACE and the lease-lifetime clamp that existed to keep the watchdog's deadline usable.

What has not changed is the physics: a bound endpoint holds a valid address until its lease runs out, so a LOST address still cannot be proven before then. What changed is who says it and how precisely. The 1.x watchdog re-checked on a 30-second tick and then held a 25-second settling time on top of that before it called an outage. 2.0 reports on the client's own schedule. A silent server is visible before the lease lapses: since v2.1.0 (#940) renewals_unanswered moves at the first retransmission of a renewal request. RFC 2131 section 4.4.5 sets that wait at "one-half of the remaining time until T2 ... down to a minimum of 60 seconds", so the 60 seconds is a floor and the wait is hours on a long lease: about 4.5 hours after T1 on a 24-hour lease, against 12 hours until that lease ends. The address is still good at that point, which is why the counter is not healthy-affecting. - A zero lease lifetime is an infinite lease. The 1.x plugin clamped an implausibly long option 51 so its watchdog had a usable deadline, and counted the clamp. The library hands the chassis an explicit "no expiry" instead, which the caller tests instead of comparing against a threshold. - No client keeps state on disk, so nothing is keyed by interface name. This is where a private mount namespace per client used to be needed: the 1.x client kept its lease file and its control socket in directories named after the interface, and two containers both called eth0 collided on both, with the second client silently handing its arguments to the first one's socket and never renewing (#332). The library holds its lease in memory and the plugin holds the durable copy in its own record, keyed by endpoint. There is no per-client mount namespace, no tmpfs, and no lease file to go looking for. - The socket belongs to the endpoint's namespace because of the thread that created it. newLibClient locks the OS thread, enters the endpoint's network namespace, opens the client there, and returns the thread. A raw AF_PACKET socket keeps the namespace it was created in, which is the property the mount-namespace machinery used to approximate from outside. If the thread cannot be returned it is retired and never reused, since a thread left in a container's namespace would silently give the next caller the wrong one.

How IPv6 is handled

Any ipv6_mode but off gives an endpoint a second DHCP client, in the same shape as its first: one dhcpManager, one library client, one record. What differs between the modes is what that client does on the wire. In dhcp it leases the address over DHCPv6, which is what -o ipv6=true has always meant. In slaac it sends no Solicit and forms the address from an advertised prefix. In auto it reads the advertisement and does what it says: the managed-address flag means DHCPv6, a clear flag means the prefix. It solicits routers and reads advertisements in all three, which is where the container's IPv6 default route and its link MTU come from. The routes the advertisement asks for arrive the same way and skip_routes opts out of those; the default route is not governed by it. Resolvers are a union and not a choice: the library merges the advertisement's RDNSS and DNSSL into the same lists a DHCPv6 server's options 23 and 24 fill, with the server's taking precedence (RFC 8106 section 5.3.1), and propagate_dns decides whether the result is written into the container at all. Nothing about the v4 path changes in any of them, which is the whole design. The maintainer's rule for this milestone was that IPv6 takes the same shape as IPv4 unless the v4 shape was itself a hack.

The differences that do exist are the ones the protocol forces.

The identity is minted once and stored. DHCPv4 derives its client identifier from the MAC on every start; DHCPv6 cannot, because RFC 9915 §11 asks for a DUID that "SHOULD NOT change over time if at all possible" and one mode has no per-endpoint MAC to derive it from. So resolveIdentity6 mints it at CreateEndpoint and the endpoint's record carries it (the Identity field, write-once). Bridge and macvlan get §11.4's DUID-LL over the endpoint MAC, byte for byte the DUID 1.9.0 put on the wire, so an endpoint upgraded from 1.x keeps its address. ipvlan gets §11.5's DUID-UUID over the endpoint id, because an ipvlan L2 slave inherits the parent link's MAC and every container on one network would otherwise present the same identity and claim one binding.

The two families do not share a record. A record is keyed on scope and hardware address, and a dual-stack endpoint has one hardware address on one network, so the v6 record is filed under Scope6(networkID), the network id with a #v6 suffix. Without it the two families collide exactly and whichever bound last answers both resumptions.

Duplicate-address detection happens in the client, and the kernel is told not to repeat it. RFC 9915 §18.2.10.1 puts the check on the client; the library runs it and reports the lease only after it passes. The chassis then installs the address with IFA_F_NODAD, because a second run costs a tentative window the container cannot use the address in and can fail where the first passed: RFC 7527 §4.1's loopback case takes the address out of service entirely. RFC 4429 §3.3 is the same argument from the standards side. The address also carries RFC 9915 §7.1's two lifetimes, so the kernel can deprecate instead of deleting (RFC 4862 §5.5.4); expiry itself is still the library's job and the lifetimes are a belt for a plugin that dies inside the window.

The Router-Advertisement guard is a precondition, and since v2.2.0 it turns the kernel's own processing OFF. DHCPv6 carries no next hop, because RFC 9915 §21 defines no router option, and RFC 5942 §4 rule 1 forbids treating the assigned address's prefix as on-link, so somebody has to process Router Advertisements or the endpoint has an address and no route. Until v2.2.0 that somebody was the container's kernel. It is now the plugin's own DHCPv6 client, which reads advertisements off its socket and reports the gateway, MTU, routes and DNS through dhcp.Info; the Join answer carries the gateway and the routes, and the manager rewrites them when a later advertisement changes them (#821).

That answer only reaches an endpoint that HAS a global IPv6 address. The daemon disables IPv6 on a container link carrying no global IPv6 address, and the kernel refuses every IPv6 route on such a link, so an answer with an IPv6 half fails the sandbox outright, it does not degrade, and the plugin cannot clear disable_ipv6 first, because that runs in the manager goroutine Join spawns, after the daemon has moved the link and applied the answer. A segment that hands out no DHCPv6 address therefore gets its MTU, its resolvers on a propagate_dns network, and no route. On ipv6_mode=slaac and ipv6_mode=auto the plugin forms the address from the advertisement and installs it (#818), so where a prefix forms one the link carries a global address and the route installs beside it. One ending starts an endpoint without a global address in every mode, those two included: dhcpv6_not_offered, the verdict for an acquisition that produced no DHCPv6 address on a segment that never said one was to be had. The Networks where DHCPv6 offers no address section of docs/reference.md carries its rows.

ApplyRouterAdvertGuard therefore writes accept_ra=0, autoconf=0 and keep_addr_on_down=1 and reads each back; DHCPClientOptions refuses a persistent v6 client that does not claim it, and refuses every other shape that does. accept_ra=0 because a kernel acting on the same frames would install a second default route beside the plugin's, and which of the two wins is a metric comparison nobody chose. autoconf=0 because the plugin holds the lease for the address the container uses.

Writing accept_ra=0 purges nothing, which is the part that is easy to miss. It stops the kernel processing the NEXT advertisement; a route an earlier one installed stays until its own lifetime runs out, and RFC 4861 §4.2 allows that to be 65535 seconds. The engine brings the link up in the sandbox at the kernel default before the guard runs, so the window is real. purgeRouterAdvertRoutes closes it: after the knobs take, every RTPROT_RA route on the link is deleted. Failures there fold into router_advert_guard_failures beside the sysctl ones, because they are one obligation seen twice. The address the kernel may have formed in the same window is NOT touched; that is #818's.

It all runs in prepareIPv6Link, in one namespace entry. That placement is a deviation from where the design put it, inside the client's own setup, and the reason is mechanical: /proc/sys is read-only in the managed plugin's rootfs, v6_link.go already owns the mount-namespace unshare that makes it writable, and doing it in the client would mean a second one. The ORDER the design fixed is preserved exactly: disable_ipv6 cleared first, then the guard, then the purge, then the client, which waits for a non-tentative link-local of its own before it sends anything.

An absent v6 lease is classified. On a stateless or SLAAC segment there is no DHCPv6 address by definition, and refusing the endpoint there means no container can start on those networks at all (#868). classifyV6Absence decides on what the segment said: a Configured event, the library's own kind for a reply that carried configuration and no address, is "not offered", whatever the last advertisement's flags were; otherwise no advertisement at all is "no router", and an advertisement with the M bit set is fatal. The wire beats the diagnostic, because a later advertisement on the same link can set M after the segment has already answered.

How a network chooses its DHCP server

dhcp_servers ranks the servers a network may lease from and dhcp_deny_servers names ones it must never lease from (#111, #669). The operator-facing rules are in the driver reference; the shape of the implementation follows from where the filtering happens.

  • Both lists match the Server Identifier (option 54). The packet's source address is never used. This is the one thing about these options that a 1.x operator has to re-learn. The external client compared the offer's IP source, which meant that behind a DHCP relay every offer looked like it came from the relay and neither list could tell servers apart; option 54 is what the server says it is, and it is also what a renewal is unicast to, so 2.0 filters on it and the relay limitation goes with the change. The two keys agree whenever a server answers directly. They stay DHCPv4-only even now that DHCPv6 is wired in: a v6 entry is refused at docker network create instead of applying to nothing, and clientServerLists hands a v6 client no lists at all. That is not an oversight deferred: the library's Params6 has no server-list field, so a v6 client that tried to honour one would not compile.

    The library's predicate is where the edge cases live, and they are decided and not incidental: deny wins over allow for a server named in both; an allow list fails closed on a message that carries no server identifier at all, because "only these servers" that a message can satisfy by omitting the field is not a restriction; and a deny list alone fails open on that same message, because nothing shows it came from a denied server. - The plugin still never sends both lists. The deny list is subtracted from the preference list at parse time, so after resolveServerPolicy there is one truth about what is allowed and one kind of list to hand down. That subtraction was forced in 1.x, where a configured whitelist switched the blacklist off inside the client and a network setting both would have got a denial nothing enforced; it is kept here because the property it buys is worth more than the redundancy the library would now tolerate: one truth about what is allowed, instead of two composed at the far end. A preference list that denies its way to empty fails the network create, because the alternative is degrading into "accept any server at all", the opposite of what both options were set to achieve. - Ordering is not expressible to the client, so preference is an acquisition-time ladder. The initial acquisition runs one attempt per preferred server, in the operator's order, each restricted to that server alone. The ladder divides the existing acquisition budget instead of extending it, because a preference list must not make docker run slower and the one-shot at CreateEndpoint already runs against a tight ceiling. The remainder of the division is dropped and never handed to the last tier, so the attempts can only sum to at most the budget. The per-attempt floor (minAttemptBudget, 3s) predates the library and now sits just under its first retransmission at 4s, so a tier that lands on the floor buys one DISCOVER and no retry. The same reading applies to the undivided budget: the library's intervals are 4s, 8s, 16s, 32s with a 64s ceiling and ±1s of jitter, each armed as its packet goes out, so retransmissions land at ~4s, ~12s, ~28s and ~60s, and the default lease_timeout of 34s funds the first three of them. That is not a regression, since the same 3s used to have to pay for a namespace and a process spawn as well, but its original derivation is dead and nothing re-derives it against the library's schedule. At that default the budget pays for eleven attempts (34s divided by 3s, integer division), so a list of eleven or fewer keeps one attempt each, and a longer list has its tail packed into the eleventh. The budget is divided evenly over the attempts it funds, so the slice each attempt gets falls as the list grows, and it stops falling once the packing starts. dhcp_server_tier_fallbacks counts a fall-through to a lower tier, which is the only outside signal that a preferred server has gone quiet while every container still starts; dhcp_server_policy_exhausted counts a restricted acquisition where nothing answered, which is otherwise indistinguishable from an ordinary DHCP timeout. - The persistent client gets the whole allowed set. The tier that won does not narrow it. It has to be able to rebind after the preferred server goes away, and an allow list pinned to the winning tier would strand the endpoint with no lease instead of failing over. Preference is an acquisition-time concept; once a lease is held it stays with whoever granted it, because renewal is unicast to that server.

How a lease is checked against the segment

The plugin asks whether some other device already holds the address its DHCP server just leased (#524). Since 2.0 the question is asked by the DHCP client itself, as RFC 5227 Address Conflict Detection, from inside the container's own network namespace and on its own link. The chassis no longer asks from the parent. The operator-facing rules and the counters are in the driver reference; this is the mechanism.

  • It is part of the acquisition. No separate step runs beside it. §2.1 sends three ARP Probes with an all-zero sender protocol address, then waits ANNOUNCE_WAIT before the address may be used; §2.3 sends two Announcements once it is; §2.4 keeps listening for the whole life of the lease. A conflict at any point produces a DHCPDECLINE (RFC 2131 §3.1(5)) and a fresh DISCOVER, which is what makes the DHCP server's own log the outside evidence for the whole thing.
  • The vantage point moved, and that is what closed the two holes the old check had. The chassis used to send a datagram from the PARENT link to make the kernel resolve the address, and compare the answering MAC with the endpoint's. That could only ever check the address a new endpoint was about to be handed, so an address that changed mid-life was never re-probed, and it needed the parent to carry an address on the leased subnet, because a host answers an ordinary ARP request only if it can route a reply back to the sender. A §2.1.1 Probe carries an all-zero sender protocol address, which Linux answers for any local target without consulting a route, so the bare-parent limitation is gone; and §2.4 covers the rest of the lease's life, so the mid-life hole is gone with it.
  • Our own endpoint holds the address too, which is the premise. The old check answered it with macvlan's parent/child isolation plus a MAC comparison. RFC 5227 answers it in the client: a reply whose sender hardware address is the client's own is not a conflict. That is what keeps bridge mode correct, where the host can reach the container and a did-anything-reply check would report every single endpoint as a conflict. The cost is unchanged: a squatter that is another container on the same parent is invisible, excluded by construction and not pending work (#528).
  • It costs seconds, and the operator chooses who pays them. conflict_check=wait (the default) finishes §2.1 before the address is configured, so docker run waits 4.0–7.0s; async configures the address at the DHCPACK and probes behind it, so a conflict found later CHANGES a running container's address; off sends no ARP at all. The lease_timeout default is derived from the same constants, one DISCOVER retransmission plus the worst probe window, instead of being written down, so the two cannot drift apart.
  • The phase survives a plugin restart. In async the address is in use while §2.1 is still running, so the conflict-detection phase is written into the durable lease record and handed back to the next process on resume. Without it a restart inside that window would leave a container holding an address nothing ever finished checking.

How a lease gets handed back

By default it does not, and that is deliberate as of v1.9.0 (#800). Two values are the exceptions: release_lease=on_stop, which releases at the stop (#962), and release_lease=on_remove, which holds the address for the restart window first and releases at the end of it (#984).

A lease is a lease. When a container stops, its address stays leased until the lease expires, and if the container comes back before then it asks for the same address and gets it. That is the ordinary DHCP path, and exactly what happens when a physical host on the segment reboots or loses power. A container is a host on this segment and costs the server what one costs.

On a default network neither client releases. The CreateEndpoint one-shot ends by cancelling its own manager, which drops the lease locally with ReasonStopped and sends nothing. The record carries the lease to the persistent client that takes over moments later, which resumes it as INIT-REBOOT instead of discovering afresh. The persistent client is stopped at Leave and keeps the address for the container that may be about to restart. A stop is this process's own shutdown reported back to it, which is why nothing counts it as a lease loss: doing so would report one for every container that started successfully.

What release_lease=on_stop changes. At Leave, and only there, the endpoint's lease goes back: a DHCPRELEASE (RFC 2131 section 4.4.6) for IPv4 and a Release (RFC 9915 section 18.2.7) for IPv6, one datagram per family.

It is built from the lease record, not from a running client, and that is the difference that makes the option work for the case it exists for. A container stopped before the plugin's persistent client attached has no client to ask, and the address it was using came from the one-shot exchange at CreateEndpoint -- which wrote it into the same record. So the record holds the address, the identity as sent, the chaddr and the server, and the release is assembled from those and sent from the host's own address on the parent interface. Nothing needs the container's namespace, which may already be gone.

The v6 address comes off the container link first, which section 18.2.7 requires before the exchange may begin; if it cannot be removed, nothing is sent and the address expires on the server's clock instead. The source is the parent's link-local address and never the address being released, which is the same section's second requirement.

Everything else about that teardown follows from the address being gone. The record is CLOSED, not LEFT, per family, so the next start cannot resume an address the server has already put back in its pool. No tombstone is laid, so no other container inherits the MAC and the addresses beside it. The tombstone is one object carrying both families' addresses, so either family releasing suppresses it, while the record of a family whose release did not happen is retained exactly as under never and stays resumable. releases_sent counts what left the host and release_failures counts attempts that put nothing on the wire, both split per family and both moved by the plugin from the outcome of its own attempt, which is the only place that knows a release was asked for and did not happen.

Leave is the only path on_stop releases from. Plugin.Close, a manager displaced by a newer one for the same endpoint, and the cleanup after docker network rm all stop clients whose containers are still running, and a release there would tell the server an address is free while a live container holds it.

What release_lease=on_remove changes. Nothing at Leave: the endpoint is torn down exactly as under never, tombstone and all, and one line goes in the log saying the addresses are being kept for the window. What releases is the record, later, and from a different place.

Every DeleteEndpoint already retains the endpoint's record with a deadline one tombstone TTL away, because that is how long a restarting container may inherit the MAC and address. On an on_remove network that deadline is also when the address stops being the container's. A sweeper ticks every 15 seconds, and a retained record whose deadline has passed by 5 seconds is handed back with the same sender on_stop uses, built from the same record, and then closed. So the wall clock from docker stop to the datagram is 65 to 80 seconds, and the window an operator reasons about is the one they already know from docker restart. There is no second option, and there is nothing to keep in step.

Three consequences follow from the deadline living in the record rather than in a timer:

  • A plugin that restarts inside the window still releases at the right moment. The deadline was written to the file; the new process rebuilds it and the first sweep after it passes hands the address back. A timer would have died with the process.
  • docker network rm hands back every address the network still holds, at once, without waiting for deadlines on a network that will not exist. That release runs before the network's stored options are deleted, because it reads release_lease and the parent interface out of them.
  • The address reserved for an endpoint Docker never created is reached too, which is the paragraph below.

What decides that an address was claimed back is the address, not the MAC. The sweep looks for another record on the same scope holding the same address: one in a live phase, or simply a newer one that is not closed. A container pinned with --mac-address that comes back on a different address does not hold the old one, and the old one goes back; a MAC-keyed check would have closed it unsent and leaked it. The one deliberate exception is an acquisition still in flight under the same MAC with no address yet: that is treated as a claim, because the address it is about to be given may be this one. The cost of the exception is one-sided by design. An in-flight acquisition that lands somewhere else leaves one address to expire on the server's clock, which is what never does with every address; the opposite mistake would hand an address away from under a container that is starting, which is the duplicate assignment of #524.

And one address never reaches Leave at all. In IPAM mode an address reserved for an endpoint whose CreateEndpoint then failed is retained by ReleaseAddress, not released. Retaining it is what lets a restart policy's next attempt claim the same address back instead of burning a second lease on the server, and a reservation with no endpoint reaches no Leave, so nothing on the on_stop path can see it. On never and on on_stop no DHCPRELEASE goes on the wire for it and the address is left to expire, exactly as any other host on the segment leaves one: on an on_stop network that is a real lease the server granted that nothing hands back, held by the retention deadline until it expires.

on_remove is what closes that (#984), and it closes it without touching ReleaseAddress at all. The retention that handler writes already carries a deadline, and the sweep hands back every retained record whose deadline has passed. So the retry still gets its window and its address, and the address is given up afterwards instead of waiting for the server's clock. Which of the two wins is still a decision: the retention wins while the window is open, the release wins when it closes.

Why this changed. Up to v1.8.x the plugin released aggressively. The external client emitted a RELEASE on a graceful stop, and a background reclaim handed back the one-shot's address whenever no persistent client had taken ownership of it (a container that exited before the attach completed). Both were trying to return an address promptly instead of letting it sit until expiry. Both raced the tombstone.

A docker restart is a Leave immediately followed by a Join for the same MAC, and the tombstone exists to promise that Join the same address. At the moment the release ran, "this endpoint is gone" and "this endpoint is coming straight back" were indistinguishable, so the plugin was observed telling the server an address was free in the same second the container came back to claim it. The reclaim was measured firing four times on ordinary restarts of live containers.

What was gained was a faster return of an address nobody wanted. What was risked was an address handed to someone else while a container was still using it, the duplicate assignment #524 added detection for, manufactured by the plugin itself. Waiting for expiry has no such failure mode, so the whole mechanism went: the release itself, the reclaim, and the orphaned_leases_released and orphaned_lease_release_failures counters that measured it.

release_lease=on_stop does not bring that mechanism back. What it sends comes from the endpoint's own live client, inside the container's sandbox, before anything is torn down, and the tombstone it would have raced is not written at all for an endpoint that released. The background reclaim and its synthesised link stay gone.

release_lease=on_remove does send from a background sweep, and it is the one value that has to answer this paragraph. What the reclaim got wrong was not that it ran in the background; it was that it could not tell "this endpoint is gone" from "this endpoint is coming straight back", because it ran at the moment those two look identical. The deadline is what tells them apart. Nothing is sent until the window the tombstone itself promises has run out, so by the time the sweep looks, a container that was coming back has come back. And the sweep does look: before sending it re-reads the records and skips any address another record now holds, which is the restart it would otherwise have raced. An acquisition in flight under the same MAC with no address yet counts as a claim for the same reason. What was removed was a release with no way to see the restart; what is here is a release that waits for it and then checks.

The surviving teardown counter was renamed to match: what was lease_release_failures is now client_stop_failures, because a client that exits badly is all it can still mean.

The cost on a release_lease=never network, the default, is that a short-lived container's address is unavailable for one lease time. Size the server's pool and lease time for the churn, the same way you would for any other population of hosts. release_lease=on_stop is the setting that buys the address back sooner, at the price of the stop-time cost above and of the cases where the release cannot be sent, and of restart stability: an endpoint that released lays no tombstone. release_lease=on_remove is the setting that buys it back a minute later and keeps restart stability, at the price of a pool that has to carry one window's worth of stopped containers, and of a release that is attempted once and not retried.

How operations on one parent NIC are serialised

A parent NIC registers one rx_handler, so it is a macvlan port, an ipvlan port or a bridge port, and never two of them. Whichever kind asks second gets EBUSY. That is a kernel rule; one mode per parent stays the operator-facing constraint. An 802.1Q sub-interface claims no rx_handler on its parent, so a vlan network's sub-interface sits beside any of them (#902).

What the plugin can stop is inflicting it on itself. Since v1.6.0 creating an endpoint and the validate_dhcp probe, which holds its link for a whole DHCP exchange, take a per-parent gate first, so they queue instead of refusing each other (#486, #549). Every later path that adds a link to a parent takes the same gate: the IPAM driver's address reservation (#110), the vlan sub-interface and the two trial children its removal adds (#902), and the bridge the plugin makes from a spare NIC (#903).

There used to be a third, the orphaned-lease reclaim, and it was the demanding one: it ran from a goroutine ordered against no Docker request at all. It is gone (#800, see above), which shortens the worst case the gate has to cover but does not remove the need for it. The probe still holds a parent across a DHCP round trip while an endpoint may ask for the other mode.

parent_link_waits counts operations that queued, which is the mechanism working. parent_link_wait_timeouts counts ones that gave up and proceeded anyway; they may still succeed, but the budget has stopped covering the holder's duration.

The gate excludes more than the kernel does, and since v2.1.0 the reporting says so. Mutual exclusion is per parent and takes no notice of kind, while the kernel refuses only a pair that both claim the rx_handler: children of one kind coexist on a parent happily, and a vlan sub-interface coexists with every kind. So a caller that gives up waiting for a holder whose kind coexists with its own has spent the budget and protected nothing, and it goes on to a LinkAdd the kernel accepts. That case is counted as a wait, not as a timeout, which leaves the warning counter meaning what its action text says. It stopped being hypothetical with the IPAM driver: an address reservation holds a parent across a whole DHCP exchange, so two containers starting together on one macvlan network reach the give-up branch every time.

The rule is enforced by two mechanisms, and it is worth being exact about where each one stops, because the guard type exists precisely to replace a prose guarantee about a property nothing checked.

The compiler holds one half: addChildLink takes a guard value, so a path that never asks for one does not compile. It does not hold the other half. "Only lockParent makes a guard" is not something Go can express. The struct's zero value is valid, so addChildLink(&parentGuard{}, link) compiles and holds nothing, and lockParent returns exactly that literal on its own no-parent path, so the shape is already in the file as a pattern to copy. The realistic route to it is not malice: a new parent-attached call site, a compiler demanding a guard, and the zero value sitting right there.

That half is enforced by scripts/check-parent-gate-accounting.sh, which fails the build on a parentGuard constructed anywhere but lockParent. A second accounting file, .github/linkadd-accounting.txt, covers the way around the type entirely: a direct netlink.LinkAdd, which the veth pair of bridge mode needs, having no parent to contend for.

Nor does the guard say which parent it is for, so one taken on one NIC and handed to a link on another compiles. That is a deliberate non-goal. Closing it means a runtime comparison, trading a compile error for a log line, on a mistake no current call site can make. The comment at the top of pkg/plugin/parent_gate.go is the authority on all of this; if this section and that comment ever disagree, the comment is right.

How state outlives a process

Three separate mechanisms keep addresses stable across three different kinds of restart. Their observable behaviour is documented in the driver reference; this is how they are built.

  • Per-network options → STATE_DIR/<network_id>.json. Written at CreateNetwork so the per-endpoint handlers never call back into the Docker API to learn the mode or parent. That callback is precisely what deadlocked the upstream plugin during dockerd startup, when the daemon asked it to restore containers using its own networks. On a cache miss the handlers fall back to the API and back-fill the file.
  • Tombstones → a single file under STATE_DIR. Written at DeleteEndpoint, consumed at the next CreateEndpoint, 60-second TTL. Each carries the previous MAC, the last v4 and v6 addresses, and the container hostname. The lookup is keyed by network ID plus hostname, which is why an endpoint keeps its address across a container restart but not across removal of the network itself, since the replacement network has a different ID. Ambiguity is resolved conservatively: when neither side knows the hostname, a tombstone is consumed only if it is the network's single candidate, so concurrent restarts fall back to fresh MACs instead of risking one container's identity being handed to another.
  • Recovery → a walk of Docker's network list at startup. For every endpoint on a plugin-served network, a DHCP manager is rebuilt and its first acquisition requests the address the container already holds (option 50). It runs synchronously inside plugin construction when the daemon answers, which is the normal case and finishes before the socket accepts anything. When the daemon is not serving yet the walk cannot run there at all: Docker respawns the plugin during its own startup, so blocking would make us unreachable to the very daemon we are waiting on. Recovery is handed to Listen instead and runs in a goroutine after the socket is up (#383), which puts it in the same window as the Joins a restarting host is issuing. recovery_deferred counts that postponement; it is not a fault, and only an exhausted retry budget lands on recovery_failed.

The deferred path is what makes the compare-and-set load-bearing. Recovery builds a manager and registers it only if no manager is already registered for the endpoint, in one locked operation. The check and the registration used to be two, and a Join landing in the gap had its live manager evicted from the registry while its client kept running, untracked, unstoppable, and competing with recovery's fresh client on the same interface. A Join is newer truth than a recovery walk and may displace it; recovery is older truth and must yield, which is what a compare-and-set expresses and a stop-what-I-displaced does not. recovery_already_managed counts an endpoint left alone. It does not affect health, since that endpoint has a renewal client, and it is the only outward sign the race happened at all (#480, #679).

The plugin's identity is a MAC. Both stability mechanisms exist because DHCP servers key on it, and everything above is in service of presenting the same MAC to the server across an event the container did not choose.

How the counters are exposed

/Plugin.Health and /metrics are two renderings of one snapshot (#651). What each of them says is in the driver reference; the mechanism is that one function builds that snapshot and both handlers render it and nothing else.

  • One source, because two hand-kept lists rot. A metrics handler that read the atomics itself would be a second list of every counter, and this repository has watched that shape decay more than once (#542, #636), and a stale list is invisible until an alert that should have fired does not. The exposition is a table keyed by the HealthResponse JSON tag it renders, and a unit test walks that struct by reflection and fails on a field nobody claimed. Adding a counter without exposing it is a red unit test instead of a hole in somebody's dashboard.
  • The snapshot is not a single atomic instant, and that is deliberate, but only for values read one at a time. The counters are read without a lock, so two of them can be a few nanoseconds apart. For a counter an operator reads on its own that is harmless: they are monotonic counters read for rates and alerting, and never an accounting ledger, and this is what /Plugin.Health has always done. Only the two map lengths take the mutex, because reading a map during a concurrent write is a data race and not a stale number.

It stops being harmless the moment a rendered value is combined from two of them, and #730 is what that costs. Each family pair is therefore loaded exactly once into a local, and the aggregate is the sum of those two locals, and never a second .Load() of a half that was already read. - Both family series are stored; neither is derived. Eleven counters carry a family label. bumpFamily increments exactly one of a pair, the v4 half or the v6 half, never both and never a third aggregate, so _v4 and _v6 are peers, and the unsuffixed counter an operator alerts on is their sum, computed at snapshot time (#212, #730).

Until v1.8.0 the aggregate was the stored counter and family="ipv4" was total - v6 at render time, clamped at zero. Two independently updated counters combined by subtraction can be read in an order that yields a value below the previous scrape, and Prometheus reads any counter decrease as a reset, repaying the whole accumulated value as an increase on the next scrape, so a one-event skew surfaces as a rate spike of the entire count. The clamp hid the extreme case and did nothing about the dip. Adding two monotonic counters has no such failure mode, because neither operand can decrease; subtracting them does, in either read order.

If a family series ever needs computing instead of reading again, the arithmetic belongs in healthSnapshot where both halves are loaded once, and never in the renderer. - Two exposure paths, and only one of them opens a port. /metrics is on the plugin socket unconditionally: it costs nothing, and it lets an operator with a socket-aware scrape path collect metrics without the plugin listening anywhere. The TCP listener is METRICS_ADDR, and it is off unless set. The plugin runs with "network": {"type": "host"} and holds CAP_NET_ADMIN, CAP_NET_RAW, CAP_SYS_ADMIN and CAP_SYS_PTRACE, so any port it opens is on the host's own network namespace. Opening one has to be a decision an operator made, and never something they inherited by upgrading. That listener's mux carries /metrics and nothing else, so no libnetwork RPC becomes reachable over TCP; it binds before the call returns, so an unusable address fails at startup where somebody sees it, instead of in a goroutine that logs and leaves the plugin running without the endpoint that was asked for; and a wildcard bind is said out loud instead of being refused. A wildcard leaks no lease inventory. The exposition is aggregate counters plus a per-process instance UUID, with no endpoint IDs, container names, addresses or MACs, which SECURITY.md promises and TestMetricsExposition_NoPerEndpointIdentifiers pins. It is this plugin's operational telemetry, published on every interface the host has, since the plugin runs in the host's network namespace (#709). - The socket's mode is pinned, and never inherited. Serving /metrics on the plugin socket is unchanged ground only because that socket is root-only: anything able to read it can already call every RPC. A UNIX socket is created 0777 &^ umask, so until #687 that property was whatever umask the plugin runtime happened to hand us: true under the usual 0022, false under 0002, and nothing said which. Listen now chmods the socket to 0600 and refuses to serve if it cannot, since an unknown mode is exactly the state being guarded against. Only the daemon speaks this protocol and it connects as root, so nothing legitimate needs group or other.

Running the tests

Four loops, cheapest first. Only the last needs root or a plugin.

what command needs
the Go unit tests go test ./... nothing; seconds
the suite's own guards go test ./test/integration/harness/ nothing
the whole fast CI lane make check nothing; about a minute
both integration suites sudo make integration-local root, Docker

make check is the one to run before pushing. It runs the same gates as the Test workflow's two fast jobs: test for build, vet, format, the race suite and the short fuzz, and policy-gates for every check-*.sh and the gate self-tests (#829 split them, and both are required contexts). It needs no privileges and mutates no host state, so the answer you get locally is the answer CI will give.

The fuzz step runs four native Go targets, one invocation each, on a budget counted in executions and not in wall clock (#324): the resolv.conf renderer, the host-side link name, the DHCPv6 identity blob and the IPAM PoolID. Their seed corpora also run as ordinary tests in the race step, which is not the same thing: a seed cannot find an input nobody has generated yet. The wire codec's own fuzzing lives in the library module.

It was a no-op until #1010. The two targets it named belonged to the 1.x lease parsers and had been deleted with them, and go test -fuzz over a package with no matching target prints PASS and exits 0, so a required check ran for weeks with one possible verdict. scripts/check-fuzz-budget.sh now resolves every name in the step against the package beside it, and refuses a target in the tree that the step never fuzzes. That is why the step is four spelled-out invocations and not a loop: a name assembled at run time is a name the gate cannot resolve. It asks both questions of scripts/local-lane.sh as well, because the lane carries the same four invocations and this page tells you it gives the answer CI will give: a rename in one file alone goes red.

The lane's contents live in scripts/local-lane.sh, and scripts/check-local-lane.sh fails CI if that file lists fewer gates than the workflow runs; a local target that hand-listed them would quietly cover less the first time a gate was added (#636, the same shape as #542).

Everything it does not do is declared and never merely absent. scripts/local-lane.sh --list-exempt prints the list with reasons, and that is the place to read it instead of a count written here, which has already gone stale once. The reasons fall into two kinds: gates that need the pull request that does not exist locally (a commit range, a title, a body, or the base ref the PR is opened against), and gates that need the network. In every case a local answer would be a different answer and never a cheaper one.

A step whose tool is missing (staticcheck, actionlint, shellcheck) is skipped loudly and named in the summary instead of passing silently. STRICT=1 make check turns any skip into a failure. Use that anywhere a green exit is read as coverage instead of by a person who can see the summary.

CI shards both suites (#381, #468, #877, D41): the main suite across nine jobs and the failure suite across two, plus one hosted job that builds the plugin all eleven install. integration-local deliberately does not shard, because a local run is one machine, so sharding would serialise the shards and only add overhead. To reproduce a single CI shard, sudo make integration-test-shard SHARD=1 OF=9 SUITE=main, or SHARD=1 OF=2 SUITE=failure; SUITE defaults to main.

Both counts live in .github/workflows/integration.yml's matrix, beside the measurement that derives them from the five-minute budget (D41); scripts/check-durations-table.sh keeps the weights that partition them honest, over both suites, and scripts/test-integration-shard.sh proves the two partitions together cover the roster exactly once. The partitioner itself refuses, naming the file and line, a test under test/integration/ that is neither on the roster nor in a package the shard target runs whole (#866).

Use integration-local.

make integration-test and make integration-test-failure only run go test. Building and installing the plugin is a separate chain (make create enable). CI never diverges because the workflow does the build as its own step before calling either target. A local run has no such guarantee, so it tests whatever plugin happens to be installed.

That is not a hypothetical. While validating #374 a stale installed build made two tests fail for reasons unrelated to the branch, and made the health floor report clean for counters that build could not publish at all. Wrong in both directions, from one cause. Rebuilding reproduced CI exactly.

integration-local chains integration-cleanup create enable integration-test integration-test-failure, so neither a stale plugin nor a previous run's leftovers can be what you measure. The cleanup step mirrors the CI job's own first step: local runs are the only place that state accumulates, because CI's runners are ephemeral.

The two suite targets deliberately do not depend on a rebuild: CI calls them in sequence between its own build and teardown, and a rebuild dependency there would reinstall the plugin between the two suites, recycling it mid-run and resetting the health floor's observation window with it.

Reading the output

Both suites tee to test/integration/logs/. At the end of each, the health floor prints a verdict for the whole run:

  • HEALTH FLOOR: clean ... over the whole Ns run (plugin up Ms, ...) says the plugin was up throughout and nothing healthy-affecting moved. The run's duration and the plugin's uptime are separate numbers and are always printed as such: locally the plugin often long predates the suite, and where the gap is large the line says by how much, because the counters are cumulative and carry that earlier history too.
  • HEALTH FLOOR: clean over the last Ns of an Ms run says the plugin restarted mid-suite, so the counters only cover the tail. The whole-run fault census covers the rest.
  • HEALTH FLOOR: clean ... over the plugin's Ns uptime says the suite's own duration could not be measured, so no coverage claim is made instead of one being invented.
  • PLUGIN FAULTS: N across the whole run is read from the log and not from the counters, so it survives a restart. Any non-zero value fails the run.

A run that cannot read the plugin log fails instead of reporting clean: an unreadable instrument is not a clean result.

If a local run disagrees with CI

Check, in order:

  1. Did you build? Use sudo make integration-local and not a bare suite target.
  2. Is a previous run's state still around? sudo make integration-cleanup. integration-local now does this for you; you only need it by hand after running a suite target directly. A single leftover container fails an unrelated test with a name conflict and reads exactly like a regression.

The interface_name tests (#125) are not a local-vs-CI divergence: they probe whether the engine applies a remote driver's DstName and skip when it does not. The probe (engineAppliesIfname, used by TestInterfaceName_MultiNetworkDeterministic) runs a throwaway container and checks the interface the engine actually created. There is no version threshold to hit. The upstream fix (moby/moby#52866, stopping the remote-driver proxy from dropping DstName) merged to moby master on 2026-08-26, is milestoned for engine 29.8.0, and that engine was released on 2026-09-03. The lane's engine is 29.8.1, read from the run's Fixture engine drift step, so the probe now succeeds there and the dependent tests run. They still skip on any box whose engine is older, and a skip there is expected and is not a signal that the run diverged.

The attach budget under load

AWAIT_TIMEOUT caps the attach that follows a Join. When it runs out with the container still running, the plugin counts join_start_failures and the container keeps an address nothing renews. Whether a small, loaded host reaches that state is a measurement, not a reading of the code, and scripts/vm-load-test.sh takes it (#969, under #403).

It builds a throwaway VM on the developer's own machine: 2 vCPU, 2 GB, plain qemu with KVM as an ordinary user, user-mode networking with ssh forwarded to loopback, nothing on the LAN and no root on the host. Inside it, the engine, this tree's plugin, and the bridge fixture shape the integration suite uses: a Linux bridge with dnsmasq bound to it. It then starts bursts of 10, 20 and 50 containers at once, three times each, at idle and under three levels of stress-ng pressure, and prints per burst the deltas of the attach buckets from /Plugin.Health beside what the DHCP server's lease file says, with the VM's steal and iowait over the burst so a row taken while the host had other tenants is marked as such. When an endpoint is not bound after the burst settled, the address on the container's link is read from inside the container and checked against the lease file, and the entry counts as the container's lease only under its own id: that is the difference between a stale address, an address leased to somebody else, and no address. A burst that attached nothing and a counter that went backwards are refused, not printed as a row; a loaded level whose load average never left the floor is printed marked refused, the table says which bursts those were, and the run exits non-zero. The lease column counts leases whose hostname is the container id the persistent client sends at Join, beside the size of the lease file, and it is evidence only once one burst has proven that key.

what command needs
its own logic scripts/vm-load-test.sh --self-test nothing; seconds
the matrix scripts/vm-load-test.sh /dev/kvm readable and writable, qemu, mtools, Docker for the rootfs build; about three hours of bursts at three repeats, before provisioning and the holds between bursts; VMLT_LEVELS, VMLT_BURSTS and VMLT_REPEATS shrink it

The matrix is not part of CI and not a gate, while its self-test and the gate test run in CI like every gate self-test: a load measurement varies from run to run and would cry wolf as a red check. It is run by hand when the Join path changes, and each run prints its table as markdown.

Request fixtures

pkg/plugin/testdata/requests/ holds the raw request bodies the Docker daemon actually sent, recorded during an integration run. The unit tests in pkg/plugin/fixtures_test.go replay them instead of hand-building CreateEndpointRequest / JoinRequest values.

The difference matters more than it looks. A hand-built request asserts the code against our model of what libnetwork sends. When the model and the daemon disagree, every unit test still passes and the disagreement surfaces on a privileged runner, or in production. That is not hypothetical: stable_lease was designed against an assumed CreateEndpoint payload and had to be reverted from v1.3.0 once the endpoint identity turned out to be unresolvable in the docker run and Compose flows (#298, #219). The request shape was the defect.

The handler decodes with DisallowUnknownFields, so a field the daemon adds and we do not model is a 400 at runtime and never a warning. The fixture tests replay through that same parser, which turns "the engine started sending something new" into a unit-test failure instead of a container that will not start.

Layout

pkg/plugin/testdata/requests/
  macvlan-run/
    manifest.json                              engine, date, commit, flow
    0001-NetworkDriver.CreateNetwork.json      one raw body per call,
    0002-NetworkDriver.CreateEndpoint.json     numbered in the order the
    ...                                        daemon issued them
  bridge-run/
  macvlan-restart/

Three flows, because the flows are where the payloads differ. That is exactly how #298 got through.

Regenerating them

Dispatch the Capture fixtures workflow. It runs on the integration lane, so the capture happens against the daemon the suite actually talks to, and it opens a pull request with the re-recorded bodies:

$ gh workflow run capture-fixtures.yml --ref <branch>

A pull request instead of a push, deliberately. A changed request body on an engine bump is a finding, the signal #218 and #125 are blocked on, so the diff wants eyes instead of an automatic commit. Before it opens anything the job re-runs check-fixture-engine-drift.sh against its own output, so a capture that recorded nothing fails there instead of on somebody else's pull request days later. check-capture-lane.sh keeps the job on the lane, which is the half the drift gate cannot check: on a hosted runner the recorded engine and the checked engine move together, agree with each other, and both describe a daemon the suite never speaks to.

By hand

The workflow drives one command, and you can run it yourself on a host with Docker and the integration prerequisites:

$ sudo make capture-fixtures CAPTURE_COMMIT=$(git rev-parse --short HEAD)

Capture against the daemon the suite runs, and never the one your shell talks to. These are not always the same machine's engine: the integration job runs inside the CI runner container, against that container's nested dockerd, which can be several minor versions ahead of the host's. A capture taken on the host is a recording of a daemon the suite never talks to, and check-fixture-engine-drift.sh will reject it, which is exactly what happened the first time these fixtures met the gate (26.1 recorded, 29.7 running). To record against the lane's engine, run the capture inside the runner image:

$ docker run -d --name dh-capture --privileged -v "$PWD":/work \
    --entrypoint bash ghcr.io/claymore666/dhcp-ci-runner:latest \
    -c 'RUNNER_JIT_CONFIG=x /entrypoint.sh || true; sleep infinity'
$ docker exec -w /work -e PATH=/usr/local/go/bin:$PATH \
    -e SUDO_UID=$(id -u) -e SUDO_GID=$(id -g) \
    dh-capture make capture-fixtures CAPTURE_COMMIT=$(git rev-parse --short HEAD)

The entrypoint brings up the same supervised daemon the suite uses and then fails its runner exec, leaving that daemon running; go lives in /usr/local/go/bin, which a bare docker exec does not put on PATH; and SUDO_UID/SUDO_GID make the recipe hand the regenerated files back to you instead of leaving them owned by root.

It builds the instrumented (-cover) plugin, sets REQUEST_CAPTURE_DIR on it, runs one integration test per flow into a cleared directory, and writes each flow's manifest.json. Pass CAPTURE_COMMIT from the unprivileged shell as shown: the recipe runs as root against a checkout you own, and git refuses that as dubious ownership, which would leave the commit field empty and produce a capture nobody can attribute.

REQUEST_CAPTURE_DIR is declared in config-cover.json only, the same place GOCOVERDIR lives, and for the same reason. It is test instrumentation, so the shipped manifest never grows a setting whose only use is regenerating this repository's fixtures. With it unset, captureHandler returns the mux unchanged and the plugin carries no extra allocation, syscall, or failure mode.

Why they cannot quietly rot

A fixture nobody refreshes is a fossilised assumption that agrees with itself forever. It is the same "asserts our model" problem, now with a green test sitting next to it, which is worse, because it looks like evidence. Three things stop that:

  • A missing or empty fixture fails. loadFixtureFlows calls t.Fatalf, never t.Skip; a suite that replayed nothing would otherwise report green.
  • A manifest without provenance fails. Empty engine, captured, commit or flow is an error, because a capture nobody can date cannot be reviewed for staleness.
  • scripts/check-fixture-engine-drift.sh compares each manifest's engine against the daemon the integration suite actually runs, and fails on a major.minor difference. It runs in the self-hosted suite job, which is the only host that knows what that engine is; patch releases and distro suffixes (26.1.5 vs 26.1.5+dfsg1) are not drift. Its self-test is scripts/test-check-fixture-engine-drift.sh.

What to do when the unknown-field test fails

It is not automatically a defect. A new field may be irrelevant to us. It means the request contract moved and somebody has to decide, which is the point, because today nothing else would say it moved at all. Model the field, or record why it is ignored.

Issue #218 (stable MAC) is waiting on exactly this signal: it needs netlabel.EndpointName to arrive at CreateEndpoint, and the captures confirm that field is absent on engine 29.8. The day a capture from a newer engine carries it, the test names it.

Issue #125 is not covered by this signal, and that is worth stating because the shape invites the assumption. Its blocker is on the response side (the engine honouring the plugin's DstName at Join, moby/moby#52866); the option itself has always been forwarded in the request. No request capture will ever change when that fix ships, so the thing that detects it is the behavioural probe in the integration suite and never these fixtures. The 26.1 -> 29.7 re-record is the worked example: it introduced com.docker.network.enable_ipv4 on CreateNetwork, which is carried inside Options (a map) and so costs nothing, but it arrived unannounced and the fixtures are what showed it.

See also