Release runbook¶
How to publish a vX.Y.Z of docker-net-dhcp. Written from the
v0.8.0 cycle, where we hit two operator-side gotchas (GHCR
package-link, Docker Hub token scope) that aren't reproducible
from the workflow file alone, capturing them here so the next
release isn't another archaeology session.
The goal: a clean release is one tag push, no manual steps.
One-time prerequisites¶
These are per-account / per-Hub-repo setup plus the tooling on the box you release from, not per-release. Done once when the publishing chain is first wired up.
Local tooling on the release box¶
The workflow installs everything it needs itself; this is only about the
commands you type. Nothing here is checked by CI, so a missing tool
surfaces as a step you skip and no gate fails. That is exactly how the
v1.3.5 release ended up unable to run step 10's verification, and how
v1.5.0 was tagged before anyone noticed cosign was absent, leaving the
signature unverified locally until afterwards.
Twice is a class, so it has a check now. Run this before step 1:
Exit 0 means every step below can actually be executed on this box. It
verifies gh, cosign major 3, and a configured user.signingkey;
crane is reported but optional. Its own table-driven tests run in CI
(scripts/test-check-release-tooling.sh),
so the check cannot rot into something that always passes.
| Tool | Needed for | Install |
|---|---|---|
gh |
every step: PRs, milestones, run status, gh release view |
distro package or https://cli.github.com |
cosign |
step 10's verify-blob re-verification |
go install github.com/sigstore/cosign/v3/cmd/cosign@latest |
crane |
optional: comparing :latest and :vX.Y.Z digests by hand |
go install github.com/google/go-containerregistry/cmd/crane@latest |
Use cosign v3 or newer. The release signs checksums.txt keylessly
and emits a Sigstore bundle, which is the v3 default; v2's
--output-signature / --output-certificate pair was removed in
favour of it. v3 is what the workflow itself installs and what the
v1.3.5 verification was run with (v3.1.2), and what verified v1.5.0
(v3.1.3); older majors are untested against
checksums.txt.sigstore.json here.
v2 does not fail with anything resembling "your cosign is too old". It
fails with Error: bundle does not contain cert for verification, please
provide public key, which blames the artifact when the toolchain is the
problem (#522). That string is now quoted on Verifying
releases so a search for it lands on the answer.
scripts/check-cosign-docs.sh
keeps every page that prints a cosign command naming the same major as
scripts/check-release-tooling.sh
enforces.
Also needed, but already true on any box that has committed here: a
git signing key, since step 9 tags with -s. Confirm with
git config --get user.signingkey before you get to the tag.
GHCR package link¶
By default a workflow's GITHUB_TOKEN can push to GHCR packages it
created but not packages that already exist under the user/org.
This fork's ghcr.io/claymore666/docker-net-dhcp package was first
published manually before the release workflow existed, so on first
tag push the workflow gets 403 Forbidden from GHCR even though
permissions: packages: write is set.
Fix it once at https://github.com/users/claymore666/packages/container/docker-net-dhcp/settings:
- Manage Actions access → Add Repository → pick
claymore666/docker-net-dhcp. - Set role to Write.
- Save.
Symptom if missed: workflow run logs show error pushing plugin:
unexpected status from POST request to
https://ghcr.io/v2/.../blobs/uploads/: 403 Forbidden at the Push
to GHCR step. The fix takes effect for the next workflow run; no
re-tag needed.
GitHub Pages source¶
The versioned documentation site (mkdocs-material + mike, #133) is
built and published by
.github/workflows/pages.yml.
That workflow pushes the rendered site to the gh-pages branch; GitHub
Pages has to be told to serve from it, once, after the branch first
exists.
The first run on dev (or the first tag) creates gh-pages. Then, at
https://github.com/claymore666/docker-net-dhcp/settings/pages:
- Build and deployment → Source = Deploy from a branch.
- Branch =
gh-pages/(root). Save.
The site then resolves at https://claymore666.github.io/docker-net-dhcp/.
No per-release action: each vX.Y.Z tag publishes its own docs version
and moves the latest alias automatically (rc tags publish a preview
without moving latest, same guard as the image :latest). Until the
first release, the workflow points the site root at the moving dev
version so it isn't a 404.
Docker Hub secrets and scopes¶
The workflow's Hub steps are gated on a job-level
HAS_HUB_CREDS check. They skip cleanly when credentials are
absent (GHCR alone still publishes), so initial setup can be
deferred.
When you do want Hub published:
- Create both repos on Hub (free) at
https://hub.docker.com/repository/create, namespace
claymore666and visibility Public: net-dhcp— the name the workflow pushes and signs.docker-net-dhcp— the alias, the name every external reference to this project uses (#972). The release copies the signed manifest into it; it is not a second build.
The Hub UI doesn't auto-create plugin repos on first push the way it
does for image repos; create both manually first. A missing alias
repo fails the copy step after GHCR and net-dhcp already hold
:vX.Y.Z and after the signature is made, which leaves a
half-published release with no SBOM, no attestation and no release
page.
2. Generate an access token at
https://app.docker.com/settings/personal-access-tokens:
- Description: something descriptive (docker-net-dhcp release CI).
- Access permissions: Read & Write at minimum, but the
description-sync step needs admin scope on the repo; read+write
alone gets 401 on description PATCH. Picking "Read, Write &
Delete" (the broadest permission level Hub offers personal tokens)
covers both image push and description sync.
- The scope has to cover both repositories. A token regenerated
against net-dhcp alone fails the alias copy at the same point a
missing repository does.
3. Add two repo secrets at
https://github.com/claymore666/docker-net-dhcp/settings/secrets/actions:
- DOCKERHUB_USERNAME = claymore666
- DOCKERHUB_TOKEN = the token from step 2.
Symptom if scope is wrong: image push works, but the Sync Docker Hub
description from README step ends with 401 Unauthorized calling
PATCH /v2/repositories/.... Regenerate the token with the broader
scope, then re-run the sync from an rc tag and never from the
release tag:
The description is not version-specific: the step pushes
README.md
to the same Hub repository whichever tag ran it, so the rc dry-run fixes
it while touching no bare release tag and no :latest. Use the rc that
preceded this release; re-dispatching an rc of an already-released
version is permitted and moves only latest-rc.
Dispatching the release tag would work too and is what this line used to say. Don't: it rebuilds the plugin, and the rebuild carries a new digest; see Dispatching an existing tag.
One reason to check this deliberately, ahead of the run failing on it:
the sync step is continue-on-error, so a 401 shows as a failed step
inside a green job.
Workflow file parsing¶
GitHub Actions parses every workflow on every push, including
branch pushes that don't match the trigger. A parse error doesn't
fail loudly. It produces a "failed" run with no jobs and silently
doesn't trigger on tag pushes either. v0.8.0 hit this with
if: ${{ secrets.X != '' }} at step level (rejected; secrets
context isn't allowed in step-level if).
First line of defence: the actionlint job in the Test workflow lints
every workflow file on every PR (and
scripts/test-actionlint.sh
asserts the linter still catches this exact bug class). Second line: the
rc-tag dry-run exercises the whole
publishing chain before every real tag.
Dispatching an existing tag¶
Dispatching release.yml with a bare existing release tag
rebuilds and re-points that tag (a different toolchain gives a
different digest), which mutates artifacts users may have pinned.
Where it used to be genuinely dangerous is :latest: crane tag
<repo>:$TAG latest does not know what :latest already points at,
so dispatching an older tag moved :latest backwards on both
registries, and there is no rollback: docker plugin create
re-tars the rootfs non-reproducibly, so undoing it means publishing
yet another digest and orphaning the previous signature.
The engine evidence ages out. Publishing is gated on the engine matrix row recorded for the tag's own commit, and those rows are kept for 30 days. Dispatching a tag older than that refuses, and so does dispatching a tag cut before that gate existed, because nothing recorded a row for it. The two refuse differently and say so: a tag nobody measured is refused for having no matrix run at all, and a tag whose rows have gone names the run that measured it and reports that it keeps no rows any more. Re-running the matrix on the tag records a fresh row and clears the second.
Use an rc tag instead. The rc dry-run
runs the identical chain and touches no bare release tag and no
:latest, which is what every recovery step in this file now points
at. Two of them used to hand you the release-tag dispatch (the one
action this section is about) to someone whose release was already
broken.
A check now refuses the irreversible half, so a written warning is no
longer the only guard (#736). promote-latest runs
scripts/assert-newest-release-tag.sh
before it touches either registry, and it refuses, never skips: a
skip during an incident is not noticed, and an incident is exactly when
this fires. What follows is what you need to know if you are looking at
a refusal, or deciding whether to dispatch anyway:
- Re-running the tag currently being released is allowed by the
guard. It still rebuilds, so it still drifts the digest; the guard
covers
:latestmoving backwards and says nothing about the rebuild. Since #1008 it also cancels the run that is publishing.release.ymlgroups onrelease-${{ github.ref }}withcancel-in-progress: true, and aworkflow_dispatchagainst a tag carries that tag as its ref, so the dispatch joins the group and the in-flight run for that tag is cancelled wherever it had got to, images and signatures included. The dispatched run publishes every one of them again, so the end state is one consistent set and the intermediate state is a tag whose registries and release page are mid-replacement.integration-arm64.ymlgroups the same way, onintegration-arm64-${{ github.ref }}. - Dispatching an older release is refused, with a non-zero exit
and nothing published. This is the case that used to succeed
quietly and move
:latestback with it. - A genuine backport is refused too: publishing v1.7.2 after v1.8.0
exists. Deliberate: moving
:latestto a backport is then an explicit manualcrane tag, so it cannot happen by accident. - rc tags are compared only against the other rcs of their own
version, because
git tag --sort=-v:refnamesortsv1.8.0-rc1abovev1.8.0unlessversionsort.suffixis set, and a single "newest tag overall" rule would have refused the real v1.8.0 release. - Re-dispatching an rc of an already-released version is not
refused. It moves
latest-rc, which no documented install command names. Left uncovered deliberately, and recorded in the script's header so the next reader does not take it for an oversight.
The rebuild itself is still a rebuild, so this remains something to do on purpose and not by reflex. The difference is that the irreversible half now fails loudly instead of relying on this paragraph being read first.
Pre-release dry-run (rc tags)¶
A tag with a pre-release suffix (v1.0.0-rc1) runs the release
workflow in pre-release mode. The full chain executes: build,
push of :v1.0.0-rc1 to both registries, Hub description sync, and
the four install proofs. :latest is not moved and no bare release
tag is touched. Zero impact on anything a user pulls by
default.
Since #736 an rc does not skip promotion, it is redirected:
promote-latest runs on an rc too, with the same steps and the same
digest assertion, aimed at latest-rc and latest-rc-arm64 instead
of latest and latest-arm64. That matters twice.
- The dry-run now covers the promotion step. Under the old
if: prerelease != 'true'skip, promotion was the one part of this workflow an rc could not reach, so a change to the promote ordering would first execute on a real tag, with no rollback behind it. :latestis still untouched, and that is now asserted from the registry. The assertion no longer infers it from the skip: the run reads:latestbefore it starts and re-reads it at the end. If the redirect were ever mistyped, the rc would fail, and it would not ship itself to everyone pulling:latest.
latest-rc and latest-rc-arm64 are public tags on both
registries, so treat them as published surface: they are moved by
every rc and they point at a build that has not been accepted. No
install instruction in any document may name them. Install
instructions pin vX.Y.Z or latest. If you are ever adding a tag
to a docker plugin install line in these docs, that is the check
to run.
Use it before every real release tag (step 9 below):
git checkout main && git pull --ff-only # the release commit
git tag -s v1.0.0-rc1 -m "v1.0.0-rc1" && git push origin v1.0.0-rc1
Watch the run; every step including verify-install, since v1.7.0 release-arm64 / verify-install-arm64, since #776 verify-install-hub / verify-install-hub-arm64, and since #972 verify-install-hub-alias / verify-install-hub-alias-arm64 must be green, and since #736 promote-latest, which an rc now reaches. Its last step, Assert a pre-release did not move :latest, is the one that proves the dry-run stayed a dry-run.
The rc tag also starts the arm64 integration lane (#531): pushing
it triggers integration-arm64 on its own. Nothing to dispatch, and
nothing to remember: the tag is the gate, so what the rc proves follows
from the tag. The lane still runs only against release candidates
(v*-rc* never matches a bare release tag), because the runner pool it
targets (label dhcp-ci-arm64) is provided for that window and not
for day-to-day PRs.
An arm64 runner has to be online for the rc. Since v1.8.0 (#632) that needs no action: the board carries a standing runner that registers itself at boot and reconnects on its own, so pushing the tag is the whole procedure. Nothing to mint, nothing to launch. If the lane reports no runner anyway, the board is down. The fix belongs to the host, and no release step will start it.
Two jobs to read, and they say different things:
arm64-suite: the verdict. Must be green before the real tag.arm64-lane-present: red means the suite never started: no runner carrieddhcp-ci-arm64within the wait, so this candidate has no arm64 verdict at all. It exists because a job with no runner sits queued for hours and queued renders as "in progress", which is indistinguishable from a lane still working. Treat it as "bring the board up and re-dispatch", never as a flake to wave through.
Inside arm64-suite, one step is worth knowing about before it
surprises you: Verify the host's NFS-outage watchdog (v1.8.0+, #677).
It asks the board whether it actually booted a working
watchdog, which the source-side check cannot know: that root is a
netbooted image and can predate a fix the tree already carries. If it
fails the suite fails, and it means the board is running unwatched: an
NFS outage will wedge it until someone visits it physically. Fix the
image; do not wave it through.
It can also come back only partly verified, which surfaces as a workflow annotation and does not fail the step. That means the kernel ring buffer has aged past this boot, so the watchdog's holder could not be confirmed. The arming verdict comes from sysfs and is never the part that is lost, so an annotation here does not block the tag.
To re-run the lane against a tag that already exists:
The rc window is the enforcement gate for the documentation review
(procedure step 3): every PR on the milestone must be reconciled against
README,
docs/,
and the RELEASE_NOTES section, and they must describe the version about
to ship. If stale text or an undocumented behaviour change surfaces now,
fix it before the real tag. Then tag the real release. Naming: rc of the
upcoming version (v1.0.0-rc1 before v1.0.0). Semver orders it
before the release and it labels the content truthfully. Bump the rc
number for another attempt after a fix; never reuse an rc tag.
Cleanup (optional): rc plugin tags can be deleted from GHCR/Hub after the real release ships; the git tag stays as the audit trail.
Per-release procedure¶
Pre-flight: every issue / PR going into the release should be on
the vX.Y.Z milestone (the workflow leans on this for the
"Closes" list in the release PR).
Pre-flight, second item: re-measure the supported engines before the rc.
Dispatch
engine-matrix.yml
on the release branch and read the floor job:
A workflow_dispatch answers 404 until the workflow is on the
default branch. GitHub exposes workflow_dispatch and schedule
from the default branch only. The lane was new on dev for v2.1.0, so
neither route existed for that release; v2.1.0 carried the workflow to
main and dropped its entry from
.github/dispatch-pending.txt,
and the dispatch above works from v2.1.1 onward. Confirm before
relying on it: an entry naming this workflow means the route is not
there yet. An empty file does not prove the opposite during a
release. Since #977 the entry is pruned on the release branch at
step 2 and reaches main only with the release pull request, so while
this tree pins a later version than main the file is already silent
about a workflow that has not landed. Whether
.github/workflows/engine-matrix.yml is on main is the direct
answer.
The lane also runs on its own push trigger, over
.github/workflows/engine-matrix.yml,
.github/engine-rows.txt,
scripts/engine-baseline.sh,
scripts/engine-floor.sh
and
pkg/plugin/engine_floor.go.
That run is the measurement for any tree in which none of those paths
has changed since, which is the ordinary case for a patch release: read
it and dispatch nothing.
One job per engine line in
.github/engine-rows.txt,
each driving the whole baseline against that engine in a nested daemon.
The floor job reconciles the minimum the plugin refuses below against
the lowest line that passed. A red floor job blocks the rc: the
number it disagrees with is published in README.md and
docs/index.md, and the plugin refuses to start below it. The lane
runs weekly now that it is on the default branch, so a moving 29 tag
is usually caught before a release asks the question. Read the run and
never the schedule: a release is the moment the published number has to
be true.
- Branch off
dev:git checkout -b release/vX.Y.Z origin/dev - Bump install pins:
scripts/bump-version.sh vX.Y.Z(#251). It rewrites every published-image pin (ghcr.io/claymore666/docker-net-dhcp:vPREVin theplugin install/network create/driver:/plugin inspectsnippets acrossREADME.mdanddocs/) to the new tag, and leaves barevX.Y.Zfeature markers and historical prose (As of vPREV every PR...,v1.1.0 onward) alone. The image ref is what tells a pin from prose. Verify withgit diffandscripts/check-version-pins.sh(the same gatetest.yamlruns: every pin must agree on one version). The gate also fails CI if a future hand-edit leaves the pins inconsistent.
On the same commit, empty
.github/dispatch-pending.txt
of entries. Every workflow it lists reaches the default branch with
this release, so every entry is stale the moment the release PR
merges, and
scripts/check-dispatch-reachable.sh
fails the release PR while one is still there (#977). Pruning here,
beside the version bump, is what lets that gate accept the missing
entries for the rest of the route: it compares the pin this tree
carries against the one on main, and while this tree pins the next
version it treats a missing entry as the removal in transit. Prune
without bumping and the gate is right to fail; that is an ordinary
mid-cycle removal of a live entry.
What that comparison does not know, because the release depends
on it and so does anyone reading a green run. The gate reads two
version pins. It does not look for a release branch, a pruned entry
or a release pull request, so a bare scripts/bump-version.sh vNEXT
on any branch buys the same acceptance, and
scripts/check-version-pins.sh
permits a tree to lead the latest tag. While the pins differ the
undeclared-workflow finding is off for every dispatchable
workflow, not only the ones being released, so one merged undeclared
during the release window is not caught until the pins agree again.
The acceptance is bounded by distance, not by time: it holds
while this tree pins the immediate successor of main, and two
steps apart is refused. Nothing in the gate reads a clock, a tag or
the state of a release, so a release parked after step 5 sits
exactly one step ahead and stays accepted until somebody finishes or
unwinds it. One step, and only ahead: a release that skips a version
and a release cut on an older line are both outside it, as is any
tree simply behind main, and on those the #977 deadlock is back,
because the release PR needs the entry and the PR that removes it is
red. Every run that reports an undeclared workflow while the
pins differ says the acceptance did not apply, names both pins and
the readings that case allows, and names a ledger entry as the way
through; a run with nothing to report stays quiet.
3. Documentation review, PR-driven against the milestone. Don't
review from memory; review from the change set. List every PR on the
vX.Y.Z milestone and reconcile each one's user-visible change
against the docs:
gh pr list --state merged --limit 200 \
--json number,title,milestone \
--jq '.[] | select(.milestone.title=="vX.Y.Z") | "#\(.number) \(.title)"'
For each merged PR, confirm the docs reflect what it changed:
new/changed driver-opts land in the option tables
(README.md,
docs/reference.md,
docs/parent-attached-modes.md);
behaviour changes (Health counters, DHCP-client behaviour, identity,
recovery) land in reference.md / parent-attached-modes.md /
internals.md; examples and numbers match. A milestone PR that
changed user-visible behaviour but carries no doc delta is the signal
to look harder. That is exactly the drift that the #205↔#152 case
(busybox→dhcpcd prose surviving a docs restructure; fixed in
#234/#237) slipped through a memory-based read.
Then still read everything user-visible top-to-bottom for anything
the per-PR pass misses:
README.md
(feature list, driver-opt table, examples),
GOVERNANCE.md
and
SECURITY.md,
every file under
docs/
(including this runbook; process changes during the cycle land here
too), and the coverage table if republished. Anything describing the
previous version's behaviour, options, or numbers gets updated on the
release branch now. Everything under
docs/
(plus docs/index.md, the site home) is what the
versioned documentation site publishes for this tag, so the review
is the site review; there's no separate wiki to reconcile.
Read the pages whole, and aim at the ungated prose. The
reference material defends itself: check-option-docs.sh,
check-docs-drift.sh and check-version-pins.sh gate every driver
option, health counter, plugin setting and image pin, so those tables
are the least likely place to find drift. What rots is everything
else: a walkthrough's shell snippet, a troubleshooting row, a
sentence in a Behaviour section, a hand-maintained list. The v1.5.0
pass found eight divergences (#489) and every one of them was in
ungated prose; none would have been caught by grepping for keywords.
Some drift belongs to no milestone PR at all, so the per-PR read cannot reach it by construction. Check these directly:
- Commands and paths that never existed or stopped working.
docker plugin logswas in the README's bug-report checklist and is not a Docker subcommand. Run the commands the docs tell a reader to run. - Text invalidated by a feature in this release. #440 mounted
STATE_DIRfrom the host and left two recipes still routing operators through the plugin rootfs. A feature PR updates the section it is about; it rarely finds the other page that quietly depended on the old behaviour. - Syntax deprecated upstream. Compose, Docker CLI and dhcpcd move
on their own schedule.
docker compose -f <snippet> configprints the deprecation warnings for anything in a Compose example. - Restated lists that live somewhere else. Required CI checks, registries, privileges. Prefer replacing the copy with a pointer at the authority (the way step 5 defers to branch protection) over updating a copy that will decay again.
- Mechanisms this release added that no page describes. The
failure is absence. Nothing reads as wrong and no grep finds it.
Take the release's mechanism changes and ask which section of
internals.mdcovers each; v1.6.0 shipped a lease reclaim and a per-parent gate with no section for either, while every counter table was green. - Standing preconditions that read like old-version prose. The
BREAKING CHANGE IN v1.5.0block (README.md,docs/index.md,docs/reference.md) is a precondition for every install and never a changelog entry, but it names an old version, so a pass tidying stale version references deletes it in good faith. Keep it. -
Counted claims. "Four flip
healthytofalse" is right until a fifth is added, and the sentence still parses. Check any stated count against the code.This bullet's own example went stale while it was being used as the example. It said the claim was "now enforced by
scripts/check-health-contract.sh(#638)", and v1.8.0 added a fifth counter: the gate held every counter list it read, and the doc still shipped rows saying "four" over a list of five, because the gate read the count word in one sentence and the file states it in five (#724). The claim was not enforced. A narrower claim was, and the bullet had rounded it up.So the instruction is not "change four to five". It is that "a gate covers this" is itself a counted claim, and decays the same way. A gate covers the statements it was taught, and adding a statement is exactly the edit nobody thinks to teach it about.
Read what the gate says it read. Do not trust that it read everything.
check-health-contract.shends in a receipt:scripts/check-health-contract.sh # PASS healthy contract agrees in N doc counter-list(s), N doc # count-word(s) and N code term(s): <the counters>The numbers are deliberately not reproduced here. A count quoted into this file is the very thing this bullet is about. Run it and read them: they are what it parsed on that run. The gate asserts no total. If the file states it six times and the receipt says five, the gate is green and the doc is wrong.
Most gates in
scripts/that discover what to check print the same kind of receipt: a count, or a line per item inspected. A few print a barePASS, and for those the coverage is not visible at all without reading the script. Either way: when a gate's coverage is load-bearing for a claim in these docs, quote what it read. The fact that it exists is not the evidence.The rule survives the correction: the drift ran for two releases under this instruction before #638, which is still the argument for a gate over a rule. It just does not end there: the gate is the floor, and the count of what it reads is the part that has to be re-checked, once per release, by running it.
Verify each finding against the artifact. Reasoning is not evidence.
Run the command, config the snippet, ls the path on the test
box, query the API (gh api repos/.../branches/dev/protection). A
confidently-argued divergence that turns out to be wrong costs more
than the one it replaced.
A finding that describes a class ends in a gate. Same rule as
anywhere else in this project: if the same shape of staleness can
recur on the next mount, option or workflow change, add the check. A
promise to remember is not enough. The rootfs-path finding above
became a fourth rule in check-docs-drift.sh, deriving the
bind-mount destinations from
config.json,
so that class now fails loudly.
The badge answers are documentation too.
.bestpractices.json
at the repo root holds this project's OpenSSF Best Practices answers
(one <criterion>_status plus a <criterion>_justification each),
and every justification is a claim about the repository that can go
stale exactly like prose. Reconcile it against the milestone the same
way, then check it against the live entry:
If a milestone PR earned or invalidated a criterion (a new gate, a document that now exists, a policy that changed), update the file on the release branch.
Getting the reviewed answers onto the live entry is manual and
deliberate: the badge site takes them through its own form, one
criteria level at a time,
https://www.bestpractices.dev/en/projects/13229/{passing,silver,gold}/edit,
in a browser you are signed into. A field only appears on the level
that owns it, and the level-less /edit URL 404s for everyone
including the owner, which is worth knowing before it looks like an
expired login.
Hash-check every justification you paste before submitting. The
procedure is in the script's header. Typing an answer by hand is how
this project once produced a 65-field divergence from its own
source of truth; the hash is what makes hand-entry safe. Then run
--diff again: it must report that the entry matches. That
confirmation is the point of the script, and the reason it has no
push mode.
The work happens here, on the release branch. The rc dry-run (step 8)
is the enforcement gate: the real vX.Y.Z tag does not ship
until every milestone PR is ticked off against the docs. By the real
tag, text and code (and the published site) must agree.
4. Add a ## vX.Y.Z section to
RELEASE_NOTES.md,
above the previous version's section. Summarise what's changing
in user-visible terms; the workflow doesn't auto-build this from
commit messages. Include any operator-visible compatibility notes
(e.g. v0.8.0 narrowed the IsDHCPPlugin regex; that needed a
callout).
House style, applied by default on every release. A release note is reference material an operator scans. v1.8.0 shipped at 90KB of prose and was rewritten to 7KB; write the short version first.
- Fixed structure, in this order: a two-to-three sentence lead,
then
### Upgrade notes,### New,### Fixed,### Deferred,### With thanks to. Omit a section that has no content; do not add others.### Deferredlists only work that was on this release's milestone and left it during the cycle, each entry with its issue and where it went. The next milestone's plan is the roadmap, not a note of this release. - Put operator-visible behaviour changes in a table under Upgrade notes: one row per change, "what changed" and "what it does to you". That table is the part most readers need.
- Group
### Fixedby origin (a review, a theme), one line per defect, each ending in its issue number. State the defect and its effect; do not narrate how it was found or how the fix was chosen. - No process commentary, no anecdotes, no meta-commentary about the notes themselves, no HTML comments carrying instructions to the next writer. Those belong in the issue or this runbook.
- Avoid universal claims ("every defect is fixed", "nothing was carried"). They quantify over things a reader cannot check and go false on somebody else's commit. List what is deferred instead.
- Prefer a named list over a count. Write "six (#720, #721, ...)"; "six of the ten" hides a wrong number.
The heading is exactly ## vX.Y.Z, with nothing after it, and this
is checkable. The release workflow extracts the section by an exact
heading match, so a heading carrying a marker (## v2.0.0
(unreleased) is the shape this file has carried between releases)
matches no tag at all. Until #912 that published a generated one-line
placeholder in place of the notes and exited 0. It now refuses, so
remove whatever follows the version and confirm before tagging:
# THE STATUS IS THE VERDICT. Piped straight into `head`, this block
# exited 0 whether the notes assembled or the extractor refused,
# because `head` reports for the pipeline, so the two outcomes the
# paragraph below promises were indistinguishable to a reader's
# eye. Assemble to a file first, and `&&` the preview onto it.
section=$(mktemp) &&
scripts/release-body.sh vX.Y.Z RELEASE_NOTES.md > "$section" &&
head -20 "$section"
It prints the first 20 lines of the section the release page would
carry and exits 0, or refuses naming what is wrong (decorated
heading, no section, empty section, two sections for one version) and
exits non-zero with nothing previewed. Run it for the rc tag too
(scripts/release-body.sh vX.Y.Z-rc1 …): an rc has no section of its
own and publishes this same one, so an rc dry-run whose notes are not
ready refuses at the release page and ships no placeholder.
The maintainer signs off on the release notes before the tag. Not
optional and not implied by approving the release PR: show the
rendered ## vX.Y.Z section and wait for an explicit go. The rc
window (step 8) is the checkpoint: by the real vX.Y.Z tag the notes
must already be signed off, because the tag publishes them. If the
notes change after the tag, edit
RELEASE_NOTES.md
on dev and gh release edit vX.Y.Z --notes-file the published
body, or the two silently diverge.
Credit outside contributions by name, the way the v1.0.0 notes do. Look them up; do not rely on recall. Almost every PR here is the maintainer's or Dependabot's, so an outside one is easy to miss precisely because it is rare:
gh pr list --state merged --limit 300 --json number,title,author \
--jq '.[] | select(.author.login|test("claymore666|dependabot")|not)
| "#\(.number) @\(.author.login) \(.title)"'
Also confirm the merged commit still carries their authorship
(git log -1 --format='%an <%ae>' <sha>). A rebase or squash of a
fork branch is where that quietly becomes the maintainer's.
5. PR release/vX.Y.Z → dev. Required checks on dev are
test, policy-gates, staticcheck, integration (every PR builds
and exercises its own plugin on the integration runner), actionlint,
govulncheck, attribution, docs-site (mkdocs build --strict,
#889), CodeQL's Analyze (go) + Analyze (actions), and the CodeQL
result check, which fails when the analysis finds a new alert. main
requires those plus coverage and coverage-present, which is why
the ratchet first bites at the release PR in the next step and not
before. Merge when green.
Do not trust the list above. Read the authority. It carried a hand-written total ("eight in total") that was correct when written and became a number nothing checked:
gh api repos/claymore666/docker-net-dhcp/branches/dev/protection \
--jq '.required_status_checks.contexts'
gh api repos/claymore666/docker-net-dhcp/branches/main/protection \
--jq '.required_status_checks.contexts'
Branch protection is the authority; if the prose and the settings disagree, the settings win and the prose is the thing to fix.
policy-gates is the half of the old test job that runs the policy
gates; the Go suite is the other half (#829). It became required on
both branches on 2026-08-28; at the time #829 landed it was on
neither list, and for that window those gates ran on every pull
request and blocked nothing. If it is ever missing from the output
above, that is the state to restore; check before trusting a green
release PR.
docs-site became required in #889. Until then nothing required built
the site, so a pull request could break the nav or a strict-mode link
and be green on every other context; the Docs workflow ran and passed,
and a red there blocked nothing. It carries no path filter for the
reason the workflow's own header gives: a path-filtered required check
is absent on the pull requests the filter excludes, and absent blocks.
coverage-present became required in #735. It is the detector that
tells an absent coverage run apart from a pending one, and it was
advisory: it could go red without blocking the merge it exists to
protect, which is the failure it was built for happening to the
guard itself.
6. Open the release PR dev → main with title
Release vX.Y.Z and a Closes #N line for every issue in
the milestone. The list is what auto-closes them when the PR
merges; without it the milestone stays open after the tag.
Because the list is milestone membership, membership has to be
true. An issue that is in the milestone but not done gets closed as
delivered, silently, by the tag. The taxonomy already says backlog
never sits on a milestoned issue for exactly this reason, and the
Milestone scope workflow
(.github/workflows/milestone-scope.yml,
scripts/check-milestone-scope.sh)
checks it daily, so it does not depend on whoever builds the list. It
splits the two cases, because their fixes are opposite: backlog
with in-dev means the work shipped and the label is stale (drop
the label, keep the milestone); backlog without it means the
work has not started (move it off the milestone). Read that run
before opening the release PR: it is a schedule, so a red one waits
quietly. Release PRs additionally run the Coverage workflow with
the coverage ratchet
(scripts/coverage-ratchet.sh
vs
.github/coverage-baseline.txt):
no release ships with less per-package coverage than the previous
one. If a package beat its floor during the cycle, raise the baseline
as part of the release branch.
Read that run with scripts/coverage-read.sh <run-id> and never
by eye (#794). It prints every package's measured number beside
both floor sets (dev's and main's), because main's floors can
be lower, so a package red against dev may still clear the release
PR, and a package comfortable on dev may not. It also refuses on an
incomplete comparison, an empty baseline, or a raw block it cannot
scope to the ratchet step; none of those is reported as a clean read.
Eyeballing the log is how the v1.8.0 read started, and building the
instrument instead found three defects an eyeball would have shipped.
The exit code carries the verdict: 0 a complete reading, 1 the
ratchet's account and the raw covdata numbers disagree, 2 it cannot
judge. What it does not do is decide the release: a package under
main's floor is information for the raise-or-explain decision the
baseline file records; the script cannot settle it.
A package deleted during the cycle reads DROPPED, not as a
number. The floors come from the merge base, so main's baseline
still floors a package the branch deleted, and the branch's own
baseline is what says the deletion was deliberate: gone from the tree
and gone from that file is DROPPED, counted as compared, and it does
not fail the run. Gone from the tree while the branch's baseline still
floors it is a FAIL naming
.github/coverage-baseline.txt,
and the fix is to remove the row in the change that deleted the
package. Read a branch's run with COVREAD_DEV_REF=<that branch>, or
the head floors resolve from dev and a package that branch dropped
is reported against a floor the run never used.
Coverage shares a concurrency group with the release PR's own
integration run, so it normally starts once integration finishes, and
since D41 resharded the lane to nine main shards plus two failure
shards behind one shared build, that is minutes, down from the five
to eight it took under the five-shard layout (#877, D41). Do not
trust any range from this page; gh run list --workflow
integration.yml --status success re-derives it in one command. A
coverage check still showing nothing well past it is worth the next
paragraph.
Release PR blocked on a check that has no run. A required check
that was cancelled looks exactly like one that is pending: the PR
sits at BLOCKED with nothing to click into. It is not a missing
trigger. GitHub keeps one running plus one pending run per
concurrency group, so pushing another commit to the release PR while
coverage is still queued displaces it, and a run displaced before any
job was assigned creates no check run at all, which is why the
coverage context goes absent and never red.
You should not have to notice this yourself: the Coverage
presence check
(.github/workflows/coverage-presence.yml,
#504) watches the head and fails with the run id and the exact
recovery command when the run was evicted. It is a required context
on main as of #735, so a red one blocks the release PR and cannot
sit in the list being ignored. If it is red, do what it says. The
manual form, for a head it did not cover:
gh run list --workflow coverage.yml --limit 5 # look for "cancelled"
gh run rerun <id> # once the group is idle
Wait for the integration run on the same ref to finish before rerunning, or it will just queue and be displaced again. This cost a full debugging session on v1.3.5 (#365). The fix is thirty seconds once you know the shape of it. 7. Assemble the verification evidence, don't hand-write it.
Prints every integration run that tested exactly this tree, with its window and what else was on the privileged pool at the time. Paste it into the release PR. Do not reconstruct it from memory.Read the overlap line literally. An overlap of none, printed with
ran alone after it, and an overlap of unknown are different
claims: the second means the concurrent-run list did
not reach back far enough to judge, which happens once the repo has
been busy since. Do not upgrade an unknown to "ran alone". The
v1.4.0 write-up asserted a concurrency caveat that the data did not
support, in both directions, which is what #432 was filed about.
- Merge the release PR. Squash or merge commit, both fine;
match what's in
git log. -
Pull main, dry-run, then tag: first push
vX.Y.Z-rc1and confirm the workflow run is green end-to-end (pre-release mode,:latestuntouched; see "Pre-release dry-run" above). Then:Usegit checkout main && git pull --ff-only && # Step 4's check, on the tree that is about to be tagged. The release # page's body comes from here, and its refusal would otherwise arrive # after the images are pushed. # # THE `&&` IS THE POINT. Newline-separated, a refusal here printed its # error and the next line pushed the tag anyway, which is the failure # this check exists to prevent. Chained, the block stops at the first # non-zero and nothing below it runs. scripts/release-body.sh vX.Y.Z RELEASE_NOTES.md >/dev/null && git tag -s vX.Y.Z -m "vX.Y.Z: <one-liner>" && # signed (#175) git push origin vX.Y.Z-s(signed) so the release tag shows Verified on GitHub; the dev box hastag.gpgsign=trueso-awould also sign, but spell it out so it holds from any checkout. Confirm withgit tag -v vX.Y.Z(or the green "Verified" on the tag page). The workflow fires ontags: v*. Watch it at https://github.com/claymore666/docker-net-dhcp/actions/workflows/release.yml. Expected steps, under the names the run shows. Tag resolution is its own job: resolve runs first and has one step, Resolve release tag; a releaser watching the run sees two job rows. The release job then runs, in this order: checkout → setup-go → Log in to GHCR → Log in to Docker Hub → Both registries, or say why not → Push to GHCR → Push to Docker Hub (or skip) → Sync Docker Hub description from README (or skip) → Sync the Hub alias description from README (or skip) → Install cosign → Record and gate the cosign version → Sign published images (cosign keyless) → Install oras → Publish the same manifest under the Hub alias (or skip) → Install syft → Generate SBOM (SPDX + CycloneDX) → Package and sign release artifact → Attest release-artifact provenance → Publish and verify the release provenance bundle → Attest image provenance (GHCR) → Check attestation parity across registries → Upload signed artifacts for the release job → Workflow summary.Publish and verify the release provenance bundle runs
scripts/publish-provenance-asset.sh, which attaches the attestation the step before it produced to the release page asprovenance.intoto.jsonl(provenance-arm64.intoto.jsonlin the arm64 job), so provenance is readable from the page and not only from GitHub's attestation store (#1011). It refuses an asset name Scorecard's provenance check would not count, re-reads the subject list out of the bundle it is about to publish, refuses a bundle that does not name the tarball, and runsgh attestation verify --bundleover every subject, which is the command Verifying releases gives users. A red here means the published bundle does not verify the published bytes, and the release stops before the page exists. Every one of those refusals is driven offline on each lane run byscripts/test-publish-provenance-asset.sh, withghstubbed.Publish the same manifest under the Hub alias runs
scripts/publish-hub-alias.sh, which copies the signed manifest and its referrers withoras cp -r, re-reads the digest through the alias name, refuses anything that is not the digest just signed, and verifies the signature under the alias. Docker Hub can list the copied signature later than the manifest (once still missing 1.4 s after the copy), so the verify reads up to six times, 10 s apart, before it fails (#1043). It comes after signing on purpose: the alias is the same manifest, not a second build (#267). The two Hub description steps are separate because the action PATCHes one repository at a time.Since v1.7.0 the run carries a parallel arm64 chain (#507): release-arm64 (native
ubuntu-24.04-armbuild, pushesvX.Y.Z-arm64; per-arch tags, because a Docker plugin cannot install from a manifest list) and verify-install-arm64.Then, as separate jobs:
- verify-install / verify-install-arm64: install the just-published plugin from GHCR on a clean hosted runner and assert it enables. A red verify-install means users can't install what we just shipped.
- verify-install-hub / verify-install-hub-arm64, and since
#972 verify-install-hub-alias /
verify-install-hub-alias-arm64: the same proof for Docker Hub
under each of its two names, which is the other place a user
installs from (#776). The alias proofs are gated on the same
hub_pushedoutput as the other two, not on the copy step's own result, so a skipped copy does not also skip its own proof. Each is its own job on its own runner and that is deliberate: the value of these jobs is a daemon that has never created a network sandbox, which is how v1.6.0-rc2 caught a bind source the daemon creates lazily (#588). A second install appended toverify-installwould run after that property was already spent, anddocker plugin rmdoes not give it back. When a run published no Hub image the steps are skipped and the job records⚠️ Docker Hub install not verifiedin the summary, so a GHCR-only run cannot be mistaken for a both-registries one. -
promote-latest: since #736 this is where every floating tag moves, for both arches and all three published names, and it runs only after all eight of the above are green.
scripts/check-latest-promotion.shasserts that dependency. It does not carry a list of proof names: it derives them from the workflow's own install-verifying jobs, so a ninth proof is required the moment it exists. Steps: Refuse to promote a floating tag backwards → Install crane → the two logins → Record what :latest resolves to before promotion → Promote the GHCR floating tags → Promote the Docker Hub floating tags → Promote the Hub alias floating tags → Verify the floating tags resolve to the signed digests → Assert a pre-release did not move :latest.Two of those are guards whose evidence comes from the registry, and they check different things. Verify the floating tags resolve to the signed digests compares the floating tag against the version tag by digest. That is the #267 guard, that retagging preserved the digest the signature covers. Assert a pre-release did not move :latest re-reads
:latestand compares it to what the Record step saw before anything was touched; it runs only on an rc, and it is the one that proves the rc contract from outside. - github-release: does not wait forpromote-latest; it needs the same eight jobs. Promotion and the Releases page are siblings, so a refused promotion does not suppress the release.
Every green checklist below includes the arm64 jobs.
10. Confirm the GitHub Release. The github-release job now cuts it
automatically once the install proofs are green (so a plugin that
doesn't install never gets an advertised Releases page). It attaches
the cosign-signed artifacts and builds the body as: a generated lead
line naming the project and version (it becomes the page's
og:description, so it is what link previews show, #469), the ##
vX.Y.Z section of
RELEASE_NOTES.md,
a generated Downloads table, and a link to Verifying
releases. Step 4's notes must therefore
already be in place at tag time. rc tags produce a draft
release: the publish path is still exercised, but dry-run builds
stay out of the public list. No manual gh release create; instead
verify:
gh release view vX.Y.Z # body = the RELEASE_NOTES section; assets:
# net-dhcp-plugin-vX.Y.Z-linux-amd64.tar.gz
# net-dhcp-plugin-vX.Y.Z-linux-arm64.tar.gz
# checksums.txt + checksums.txt.sigstore.json
# checksums-arm64.txt + checksums-arm64.txt.sigstore.json
# provenance.intoto.jsonl + provenance-arm64.intoto.jsonl
# Re-verify the signature the way a downstream consumer would:
cosign verify-blob \
--bundle checksums.txt.sigstore.json \
--certificate-identity-regexp '^https://github.com/claymore666/docker-net-dhcp/.github/workflows/release.yml@' \
--certificate-oidc-issuer https://token.actions.githubusercontent.com \
checksums.txt
# And the provenance, taken from the page, with no API call:
gh attestation verify net-dhcp-plugin-vX.Y.Z-linux-amd64.tar.gz \
--bundle provenance.intoto.jsonl \
--repo claymore666/docker-net-dhcp
--clobber). This satisfies OpenSSF Scorecard Signed-Releases,
whose provenance half reads the asset NAME and counts only a
.intoto.jsonl suffix, which is what the two provenance assets are
named for (#1011);
an rc dry-run produces an equivalent pre-release with the same
signed assets, which is how this path is exercised before the real
tag (rc releases never move :latest and are marked pre-release).
10b. Nothing to refresh, and that is a change from 1.x. Through
1.x this step pasted a per-release digest block into
Verifying releases, and since #547 the
release run enforced it. From 2.0 that step is gone, from the doc and
from release.yml both.
Why it had to go. The tag and the
commit are compiled into the binary from 2.0 (VERSION/COMMIT
through -ldflags -X, plan row O-4). The step's own recovery was
"take the block from the failed run, land it on the tagged
commit, re-tag", and landing it changes the commit, which changes
the binary, which changes the digest the block states. The recovery
could not converge, so the gate had no passing state to reach: an
operator following the runbook to the letter would re-tag forever,
each attempt leaving a :vX.Y.Z published and :latest behind it,
because the failing step sat between the GHCR push and the Hub push.
Measured in review: two builds differing only in the commit ldflag
hash differently.
Where the digests are now. In the signed checksums manifest,
produced by the build that makes them true. checksums.txt and
checksums-arm64.txt each cover the tarball, both SBOMs and the
binary inside the tarball, recorded as rootfs/usr/sbin/net-dhcp, so
an operator who extracts the tarball beside the manifest can run
sha256sum --ignore-missing -c checksums.txt and check the binary
against a cosign-signed reference. That path is the same in both
manifests, because a plugin tarball has one layout, so each
architecture is verified in its own directory, which is what
docs/verifying-releases.md tells the
reader. Nothing here is a release step: the release job does it, on
every tag and every rc, with no paste and no re-run.
What stops this from rotting. scripts/check-release-digest-
fixed-point.sh runs in the normal test lane and holds the two halves
of the argument above: it MEASURES that the commit is still part of
the binary's identity, refuses any file in the tree that records a
digest of a binary this tree builds while that is true, and requires
the release workflow's signed manifest to cover the binary. Restore
the block and it goes red before a tag is ever cut; delete the
manifest line and it goes red too.
- Fast-forward
devtomainso the release commit (version pins, RELEASE_NOTES section) lands ondevtoo: Skipping this leaves the next feature branch starting from the previous version's README/docs, and the next release PR has to re-bump them. Forgotten once after v0.9.0. That's whyrelease.yml's header comment carries the same checklist.
Do this before merging anything else into dev. Forgetting is
the mild failure. The one that has actually happened is
foreclosure: the --ff-only is possible only while dev has no
commits of its own since the release, so the window opens when the
release PR lands and closes on the next merge into dev,
permanently, and by doing something otherwise correct.
# is a back-merge OWED? prints a count
git rev-list --count origin/dev..origin/main
# does `dev` carry commits of its own? (--is-ancestor prints nothing)
git merge-base --is-ancestor origin/dev origin/main \
&& echo "dev is contained in main" \
|| echo "dev has commits of its own"
Two different questions, and neither answers the other. The count
is commits on main that dev lacks; it says a back-merge is
owed. The predicate asks whether dev has commits of its own, which
is what --ff-only needs. Both outcomes of both, because a documented
command with only one outcome written down leaves the reader to supply
the other by negation:
| count | dev contained in main? |
state | do |
|---|---|---|---|
| non-zero | yes | owed and available, the normal post-tag state | run the --ff-only above, now |
| non-zero | no | owed, and foreclosed | back-merge PR |
| zero | yes | main and dev identical |
nothing |
| zero | no | dev ahead mid-cycle, the ordinary state |
nothing |
A non-zero count is not bad news: at the v1.6.0 tag it was 4, and
the fast-forward worked. It is the warning the CI advisory prints, in
those words: "While that is true, git merge --ff-only main still
works."
Once foreclosed the recovery is a back-merge PR, main → dev,
merged with a merge commit; a squash does not put main's commits
on dev, which is the entire point of this step.
The failure is silent, which is why this needs a command and not a
reminder. The divergence carries no content: every release merge
has a dev commit as its second parent (because step 5 merged the
release branch into dev before the tag), so git diff main dev is
clean and the whole suite passes while the graph is wrong. Nothing
reads as broken, so nothing prompts the check.
Worked instance: v1.6.0 published at 23:11:30Z on 2026-08-16; three
Dependabot PRs went into dev at 23:26, fifteen minutes later,
before anyone looked at main. That foreclosed it; recovered by #597
at 23:38. Fifteen minutes is not a window anyone watches by hand.
CI warns, but it cannot block.
.github/workflows/release-backmerge.yml
runs
scripts/check-release-backmerge.sh
in two modes:
- Advisory, on every PR into
dev: warns on any divergence and ignores the grace window, because the fresh divergence is the only case this mode exists for. It warns and never blocks, on purpose; reasoning in the workflow header. An annotation can be scrolled past, which is why the command stays in this step. - Enforcing, on the nightly schedule: the backstop for the
permanent omission, and the only mode
BACKMERGE_GRACE_HOURS(default 24) applies to, so a release in flight is not a false red.
One caveat on the recovery: the back-merge diff is not always empty.
#597 was, but 6af0749 carried three real files. Hitting a conflict
there does not mean you have done something wrong.
12. Prune merged branches. The repo has Automatically delete head
branches enabled, so merged PR head branches are removed on merge.
Two things that setting doesn't cover, so clean them now:
# the release branch is merged but was never a PR head:
git push origin --delete release/vX.Y.Z
# sweep for any other branch already merged into dev that lingered:
git fetch --prune origin
git branch -r --merged origin/dev | grep -vE 'origin/(dev|main|HEAD)$'
upstream/* refs (those are the original fork's remote).
Verifying¶
After the workflow succeeds:
curl -sI https://hub.docker.com/v2/repositories/claymore666/net-dhcp/tags/vX.Y.Z/returnsHTTP/2 200, and so does the same call forclaymore666/docker-net-dhcp. Both Hub names are published by the release run; a 200 on one and a 404 on the other means the alias copy did not happen and the run should have been red.curl -sI https://ghcr.io/v2/claymore666/docker-net-dhcp/manifests/vX.Y.ZreturnsHTTP/2 401(auth required). The manifest IS there, GHCR just won't expose it anonymously. To confirm presence authenticated:gh auth token | docker login ghcr.io -u <you> --password-stdin && docker plugin install ghcr.io/claymore666/docker-net-dhcp:vX.Y.Z.- Both Docker Hub pages (https://hub.docker.com/r/claymore666/net-dhcp and https://hub.docker.com/r/claymore666/docker-net-dhcp) show the new tag in the Tags tab and the README content matches GitHub. The workflow syncs the description per repository, so both are covered. Hub categories are set in the web UI only, with no API behind them, so a new Hub repository keeps whatever categories a person gave it and no run will fix them.
- The milestone is closed (every issue moved to Done by the
release PR's
Closeslist). Verify withgh issue list --milestone vX.Y.Z --state open; should be empty. - Anything that was listed in
.github/dispatch-pending.txtis now dispatchable. Exercise it once. Aworkflow_dispatchworkflow is only exposed from the default branch, so one that merged todevduring this cycle has never run, and this release is the first moment it can. Dispatch it and confirm it does what its documentation claims. The entry itself is already gone: it is removed on the release branch at step 2, travels intodevat step 5 and intomainwith the release PR, becausescripts/check-dispatch-reachable.shcounts a workflow that the release PR merges into the default branch as reachable and its entry as stale (#977). While the release branch anddevpin the next version andmainpins the current one, that gate accepts the missing entry, so step 5 and every other pull request intodevduring the release window stay green. It is reading the two pins and nothing else; the bounds that come with that are in step 2. Dropping the entry after the release instead is what turned the gate red onmainat v2.1.0.
Troubleshooting¶
| Symptom | Likely cause | Fix |
|---|---|---|
| Workflow shows zero-job "failed" runs on every push, tag push doesn't trigger anything | release.yml parse error (often secrets context in step-level if) |
Fix the YAML, then verify with an rc tag (gh workflow run release.yml -f tag=vX.Y.Z-rcN --ref main), which exercises the whole chain without touching a bare release tag or :latest. Then retry the real tag. Do not verify by dispatching the release tag: Dispatching an existing tag |
Push to GHCR step ends 403 Forbidden |
GHCR package not linked to repo with Write | One-time fix in package settings (see prerequisites) |
Push to Docker Hub step ends unauthorized: incorrect username or password |
Token revoked / expired / wrong scope | Regenerate at hub.docker.com, update DOCKERHUB_TOKEN repo secret |
Sync Docker Hub description from README step ends 401 |
Token scope is image-push only and lacks admin | Regenerate token with broader scope (see prerequisites) |
| Hub page README is stale after a release | Description-sync step skipped (no Hub creds) or 401'd | Check the workflow run; either set creds or fix the token |
| Tag push succeeded but no Hub publish | HAS_HUB_CREDS evaluated false (secrets blank) |
Set the secrets, then exercise the publish from an rc tag: gh workflow run release.yml -f tag=vX.Y.Z-rcN --ref main. Do not dispatch the bare release tag: it rebuilds and re-points that tag plus :latest, and scripts/assert-newest-release-tag.sh refuses an older one outright; Dispatching an existing tag |
Backports between dev and main¶
When a release-blocking hotfix has to land on main without
going through dev (e.g. v0.8.0's release.yml parser bug), the
flow is:
- Branch off
main, fix, PR tomain, merge. Don't push tomaindirectly; branch protection and the audit trail. - Cherry-pick the same commit onto a branch off
dev, PR todev. This keepsdevfrom regressing on the next release PR.
The v0.8.0 cycle uses #97 (main hotfix) and #98 (dev backport) as the canonical example.