Skip to content

Release runbook

How to publish a vX.Y.Z of docker-net-dhcp. Written from the v0.8.0 cycle, where we hit two operator-side gotchas (GHCR package-link, Docker Hub token scope) that aren't reproducible from the workflow file alone, capturing them here so the next release isn't another archaeology session.

The goal: a clean release is one tag push, no manual steps.

One-time prerequisites

These are per-account / per-Hub-repo setup plus the tooling on the box you release from, not per-release. Done once when the publishing chain is first wired up.

Local tooling on the release box

The workflow installs everything it needs itself; this is only about the commands you type. Nothing here is checked by CI, so a missing tool surfaces as a step you skip and no gate fails. That is exactly how the v1.3.5 release ended up unable to run step 10's verification, and how v1.5.0 was tagged before anyone noticed cosign was absent, leaving the signature unverified locally until afterwards.

Twice is a class, so it has a check now. Run this before step 1:

bash scripts/check-release-tooling.sh

Exit 0 means every step below can actually be executed on this box. It verifies gh, cosign major 3, and a configured user.signingkey; crane is reported but optional. Its own table-driven tests run in CI (scripts/test-check-release-tooling.sh), so the check cannot rot into something that always passes.

Tool Needed for Install
gh every step: PRs, milestones, run status, gh release view distro package or https://cli.github.com
cosign step 10's verify-blob re-verification go install github.com/sigstore/cosign/v3/cmd/cosign@latest
crane optional: comparing :latest and :vX.Y.Z digests by hand go install github.com/google/go-containerregistry/cmd/crane@latest

Use cosign v3 or newer. The release signs checksums.txt keylessly and emits a Sigstore bundle, which is the v3 default; v2's --output-signature / --output-certificate pair was removed in favour of it. v3 is what the workflow itself installs and what the v1.3.5 verification was run with (v3.1.2), and what verified v1.5.0 (v3.1.3); older majors are untested against checksums.txt.sigstore.json here.

v2 does not fail with anything resembling "your cosign is too old". It fails with Error: bundle does not contain cert for verification, please provide public key, which blames the artifact when the toolchain is the problem (#522). That string is now quoted on Verifying releases so a search for it lands on the answer. scripts/check-cosign-docs.sh keeps every page that prints a cosign command naming the same major as scripts/check-release-tooling.sh enforces.

Also needed, but already true on any box that has committed here: a git signing key, since step 9 tags with -s. Confirm with git config --get user.signingkey before you get to the tag.

By default a workflow's GITHUB_TOKEN can push to GHCR packages it created but not packages that already exist under the user/org. This fork's ghcr.io/claymore666/docker-net-dhcp package was first published manually before the release workflow existed, so on first tag push the workflow gets 403 Forbidden from GHCR even though permissions: packages: write is set.

Fix it once at https://github.com/users/claymore666/packages/container/docker-net-dhcp/settings:

  1. Manage Actions access → Add Repository → pick claymore666/docker-net-dhcp.
  2. Set role to Write.
  3. Save.

Symptom if missed: workflow run logs show error pushing plugin: unexpected status from POST request to https://ghcr.io/v2/.../blobs/uploads/: 403 Forbidden at the Push to GHCR step. The fix takes effect for the next workflow run; no re-tag needed.

GitHub Pages source

The versioned documentation site (mkdocs-material + mike, #133) is built and published by .github/workflows/pages.yml. That workflow pushes the rendered site to the gh-pages branch; GitHub Pages has to be told to serve from it, once, after the branch first exists.

The first run on dev (or the first tag) creates gh-pages. Then, at https://github.com/claymore666/docker-net-dhcp/settings/pages:

  1. Build and deployment → Source = Deploy from a branch.
  2. Branch = gh-pages / (root). Save.

The site then resolves at https://claymore666.github.io/docker-net-dhcp/. No per-release action: each vX.Y.Z tag publishes its own docs version and moves the latest alias automatically (rc tags publish a preview without moving latest, same guard as the image :latest). Until the first release, the workflow points the site root at the moving dev version so it isn't a 404.

Docker Hub secrets and scopes

The workflow's Hub steps are gated on a job-level HAS_HUB_CREDS check. They skip cleanly when credentials are absent (GHCR alone still publishes), so initial setup can be deferred.

When you do want Hub published:

  1. Create both repos on Hub (free) at https://hub.docker.com/repository/create, namespace claymore666 and visibility Public:
  2. net-dhcp — the name the workflow pushes and signs.
  3. docker-net-dhcp — the alias, the name every external reference to this project uses (#972). The release copies the signed manifest into it; it is not a second build.

The Hub UI doesn't auto-create plugin repos on first push the way it does for image repos; create both manually first. A missing alias repo fails the copy step after GHCR and net-dhcp already hold :vX.Y.Z and after the signature is made, which leaves a half-published release with no SBOM, no attestation and no release page. 2. Generate an access token at https://app.docker.com/settings/personal-access-tokens: - Description: something descriptive (docker-net-dhcp release CI). - Access permissions: Read & Write at minimum, but the description-sync step needs admin scope on the repo; read+write alone gets 401 on description PATCH. Picking "Read, Write & Delete" (the broadest permission level Hub offers personal tokens) covers both image push and description sync. - The scope has to cover both repositories. A token regenerated against net-dhcp alone fails the alias copy at the same point a missing repository does. 3. Add two repo secrets at https://github.com/claymore666/docker-net-dhcp/settings/secrets/actions: - DOCKERHUB_USERNAME = claymore666 - DOCKERHUB_TOKEN = the token from step 2.

Symptom if scope is wrong: image push works, but the Sync Docker Hub description from README step ends with 401 Unauthorized calling PATCH /v2/repositories/.... Regenerate the token with the broader scope, then re-run the sync from an rc tag and never from the release tag:

gh workflow run release.yml -f tag=vX.Y.Z-rcN --ref main

The description is not version-specific: the step pushes README.md to the same Hub repository whichever tag ran it, so the rc dry-run fixes it while touching no bare release tag and no :latest. Use the rc that preceded this release; re-dispatching an rc of an already-released version is permitted and moves only latest-rc.

Dispatching the release tag would work too and is what this line used to say. Don't: it rebuilds the plugin, and the rebuild carries a new digest; see Dispatching an existing tag.

One reason to check this deliberately, ahead of the run failing on it: the sync step is continue-on-error, so a 401 shows as a failed step inside a green job.

Workflow file parsing

GitHub Actions parses every workflow on every push, including branch pushes that don't match the trigger. A parse error doesn't fail loudly. It produces a "failed" run with no jobs and silently doesn't trigger on tag pushes either. v0.8.0 hit this with if: ${{ secrets.X != '' }} at step level (rejected; secrets context isn't allowed in step-level if).

First line of defence: the actionlint job in the Test workflow lints every workflow file on every PR (and scripts/test-actionlint.sh asserts the linter still catches this exact bug class). Second line: the rc-tag dry-run exercises the whole publishing chain before every real tag.

Dispatching an existing tag

Dispatching release.yml with a bare existing release tag rebuilds and re-points that tag (a different toolchain gives a different digest), which mutates artifacts users may have pinned. Where it used to be genuinely dangerous is :latest: crane tag <repo>:$TAG latest does not know what :latest already points at, so dispatching an older tag moved :latest backwards on both registries, and there is no rollback: docker plugin create re-tars the rootfs non-reproducibly, so undoing it means publishing yet another digest and orphaning the previous signature.

The engine evidence ages out. Publishing is gated on the engine matrix row recorded for the tag's own commit, and those rows are kept for 30 days. Dispatching a tag older than that refuses, and so does dispatching a tag cut before that gate existed, because nothing recorded a row for it. The two refuse differently and say so: a tag nobody measured is refused for having no matrix run at all, and a tag whose rows have gone names the run that measured it and reports that it keeps no rows any more. Re-running the matrix on the tag records a fresh row and clears the second.

Use an rc tag instead. The rc dry-run runs the identical chain and touches no bare release tag and no :latest, which is what every recovery step in this file now points at. Two of them used to hand you the release-tag dispatch (the one action this section is about) to someone whose release was already broken.

A check now refuses the irreversible half, so a written warning is no longer the only guard (#736). promote-latest runs scripts/assert-newest-release-tag.sh before it touches either registry, and it refuses, never skips: a skip during an incident is not noticed, and an incident is exactly when this fires. What follows is what you need to know if you are looking at a refusal, or deciding whether to dispatch anyway:

  • Re-running the tag currently being released is allowed by the guard. It still rebuilds, so it still drifts the digest; the guard covers :latest moving backwards and says nothing about the rebuild. Since #1008 it also cancels the run that is publishing. release.yml groups on release-${{ github.ref }} with cancel-in-progress: true, and a workflow_dispatch against a tag carries that tag as its ref, so the dispatch joins the group and the in-flight run for that tag is cancelled wherever it had got to, images and signatures included. The dispatched run publishes every one of them again, so the end state is one consistent set and the intermediate state is a tag whose registries and release page are mid-replacement. integration-arm64.yml groups the same way, on integration-arm64-${{ github.ref }}.
  • Dispatching an older release is refused, with a non-zero exit and nothing published. This is the case that used to succeed quietly and move :latest back with it.
  • A genuine backport is refused too: publishing v1.7.2 after v1.8.0 exists. Deliberate: moving :latest to a backport is then an explicit manual crane tag, so it cannot happen by accident.
  • rc tags are compared only against the other rcs of their own version, because git tag --sort=-v:refname sorts v1.8.0-rc1 above v1.8.0 unless versionsort.suffix is set, and a single "newest tag overall" rule would have refused the real v1.8.0 release.
  • Re-dispatching an rc of an already-released version is not refused. It moves latest-rc, which no documented install command names. Left uncovered deliberately, and recorded in the script's header so the next reader does not take it for an oversight.

The rebuild itself is still a rebuild, so this remains something to do on purpose and not by reflex. The difference is that the irreversible half now fails loudly instead of relying on this paragraph being read first.

Pre-release dry-run (rc tags)

A tag with a pre-release suffix (v1.0.0-rc1) runs the release workflow in pre-release mode. The full chain executes: build, push of :v1.0.0-rc1 to both registries, Hub description sync, and the four install proofs. :latest is not moved and no bare release tag is touched. Zero impact on anything a user pulls by default.

Since #736 an rc does not skip promotion, it is redirected: promote-latest runs on an rc too, with the same steps and the same digest assertion, aimed at latest-rc and latest-rc-arm64 instead of latest and latest-arm64. That matters twice.

  • The dry-run now covers the promotion step. Under the old if: prerelease != 'true' skip, promotion was the one part of this workflow an rc could not reach, so a change to the promote ordering would first execute on a real tag, with no rollback behind it.
  • :latest is still untouched, and that is now asserted from the registry. The assertion no longer infers it from the skip: the run reads :latest before it starts and re-reads it at the end. If the redirect were ever mistyped, the rc would fail, and it would not ship itself to everyone pulling :latest.

latest-rc and latest-rc-arm64 are public tags on both registries, so treat them as published surface: they are moved by every rc and they point at a build that has not been accepted. No install instruction in any document may name them. Install instructions pin vX.Y.Z or latest. If you are ever adding a tag to a docker plugin install line in these docs, that is the check to run.

Use it before every real release tag (step 9 below):

git checkout main && git pull --ff-only      # the release commit
git tag -s v1.0.0-rc1 -m "v1.0.0-rc1" && git push origin v1.0.0-rc1

Watch the run; every step including verify-install, since v1.7.0 release-arm64 / verify-install-arm64, since #776 verify-install-hub / verify-install-hub-arm64, and since #972 verify-install-hub-alias / verify-install-hub-alias-arm64 must be green, and since #736 promote-latest, which an rc now reaches. Its last step, Assert a pre-release did not move :latest, is the one that proves the dry-run stayed a dry-run.

The rc tag also starts the arm64 integration lane (#531): pushing it triggers integration-arm64 on its own. Nothing to dispatch, and nothing to remember: the tag is the gate, so what the rc proves follows from the tag. The lane still runs only against release candidates (v*-rc* never matches a bare release tag), because the runner pool it targets (label dhcp-ci-arm64) is provided for that window and not for day-to-day PRs.

An arm64 runner has to be online for the rc. Since v1.8.0 (#632) that needs no action: the board carries a standing runner that registers itself at boot and reconnects on its own, so pushing the tag is the whole procedure. Nothing to mint, nothing to launch. If the lane reports no runner anyway, the board is down. The fix belongs to the host, and no release step will start it.

Two jobs to read, and they say different things:

  • arm64-suite: the verdict. Must be green before the real tag.
  • arm64-lane-present: red means the suite never started: no runner carried dhcp-ci-arm64 within the wait, so this candidate has no arm64 verdict at all. It exists because a job with no runner sits queued for hours and queued renders as "in progress", which is indistinguishable from a lane still working. Treat it as "bring the board up and re-dispatch", never as a flake to wave through.

Inside arm64-suite, one step is worth knowing about before it surprises you: Verify the host's NFS-outage watchdog (v1.8.0+, #677). It asks the board whether it actually booted a working watchdog, which the source-side check cannot know: that root is a netbooted image and can predate a fix the tree already carries. If it fails the suite fails, and it means the board is running unwatched: an NFS outage will wedge it until someone visits it physically. Fix the image; do not wave it through.

It can also come back only partly verified, which surfaces as a workflow annotation and does not fail the step. That means the kernel ring buffer has aged past this boot, so the watchdog's holder could not be confirmed. The arming verdict comes from sysfs and is never the part that is lost, so an annotation here does not block the tag.

To re-run the lane against a tag that already exists:

gh workflow run integration-arm64.yml --ref vX.Y.Z-rcN

The rc window is the enforcement gate for the documentation review (procedure step 3): every PR on the milestone must be reconciled against README, docs/, and the RELEASE_NOTES section, and they must describe the version about to ship. If stale text or an undocumented behaviour change surfaces now, fix it before the real tag. Then tag the real release. Naming: rc of the upcoming version (v1.0.0-rc1 before v1.0.0). Semver orders it before the release and it labels the content truthfully. Bump the rc number for another attempt after a fix; never reuse an rc tag.

Cleanup (optional): rc plugin tags can be deleted from GHCR/Hub after the real release ships; the git tag stays as the audit trail.

Per-release procedure

Pre-flight: every issue / PR going into the release should be on the vX.Y.Z milestone (the workflow leans on this for the "Closes" list in the release PR).

Pre-flight, second item: re-measure the supported engines before the rc. Dispatch engine-matrix.yml on the release branch and read the floor job:

gh workflow run engine-matrix.yml --ref release/vX.Y.Z

A workflow_dispatch answers 404 until the workflow is on the default branch. GitHub exposes workflow_dispatch and schedule from the default branch only. The lane was new on dev for v2.1.0, so neither route existed for that release; v2.1.0 carried the workflow to main and dropped its entry from .github/dispatch-pending.txt, and the dispatch above works from v2.1.1 onward. Confirm before relying on it: an entry naming this workflow means the route is not there yet. An empty file does not prove the opposite during a release. Since #977 the entry is pruned on the release branch at step 2 and reaches main only with the release pull request, so while this tree pins a later version than main the file is already silent about a workflow that has not landed. Whether .github/workflows/engine-matrix.yml is on main is the direct answer.

The lane also runs on its own push trigger, over .github/workflows/engine-matrix.yml, .github/engine-rows.txt, scripts/engine-baseline.sh, scripts/engine-floor.sh and pkg/plugin/engine_floor.go. That run is the measurement for any tree in which none of those paths has changed since, which is the ordinary case for a patch release: read it and dispatch nothing.

One job per engine line in .github/engine-rows.txt, each driving the whole baseline against that engine in a nested daemon. The floor job reconciles the minimum the plugin refuses below against the lowest line that passed. A red floor job blocks the rc: the number it disagrees with is published in README.md and docs/index.md, and the plugin refuses to start below it. The lane runs weekly now that it is on the default branch, so a moving 29 tag is usually caught before a release asks the question. Read the run and never the schedule: a release is the moment the published number has to be true.

  1. Branch off dev: git checkout -b release/vX.Y.Z origin/dev
  2. Bump install pins: scripts/bump-version.sh vX.Y.Z (#251). It rewrites every published-image pin (ghcr.io/claymore666/docker-net-dhcp:vPREV in the plugin install / network create / driver: / plugin inspect snippets across README.md and docs/) to the new tag, and leaves bare vX.Y.Z feature markers and historical prose (As of vPREV every PR..., v1.1.0 onward) alone. The image ref is what tells a pin from prose. Verify with git diff and scripts/check-version-pins.sh (the same gate test.yaml runs: every pin must agree on one version). The gate also fails CI if a future hand-edit leaves the pins inconsistent.

On the same commit, empty .github/dispatch-pending.txt of entries. Every workflow it lists reaches the default branch with this release, so every entry is stale the moment the release PR merges, and scripts/check-dispatch-reachable.sh fails the release PR while one is still there (#977). Pruning here, beside the version bump, is what lets that gate accept the missing entries for the rest of the route: it compares the pin this tree carries against the one on main, and while this tree pins the next version it treats a missing entry as the removal in transit. Prune without bumping and the gate is right to fail; that is an ordinary mid-cycle removal of a live entry.

What that comparison does not know, because the release depends on it and so does anyone reading a green run. The gate reads two version pins. It does not look for a release branch, a pruned entry or a release pull request, so a bare scripts/bump-version.sh vNEXT on any branch buys the same acceptance, and scripts/check-version-pins.sh permits a tree to lead the latest tag. While the pins differ the undeclared-workflow finding is off for every dispatchable workflow, not only the ones being released, so one merged undeclared during the release window is not caught until the pins agree again. The acceptance is bounded by distance, not by time: it holds while this tree pins the immediate successor of main, and two steps apart is refused. Nothing in the gate reads a clock, a tag or the state of a release, so a release parked after step 5 sits exactly one step ahead and stays accepted until somebody finishes or unwinds it. One step, and only ahead: a release that skips a version and a release cut on an older line are both outside it, as is any tree simply behind main, and on those the #977 deadlock is back, because the release PR needs the entry and the PR that removes it is red. Every run that reports an undeclared workflow while the pins differ says the acceptance did not apply, names both pins and the readings that case allows, and names a ledger entry as the way through; a run with nothing to report stays quiet. 3. Documentation review, PR-driven against the milestone. Don't review from memory; review from the change set. List every PR on the vX.Y.Z milestone and reconcile each one's user-visible change against the docs:

gh pr list --state merged --limit 200 \
  --json number,title,milestone \
  --jq '.[] | select(.milestone.title=="vX.Y.Z") | "#\(.number) \(.title)"'

For each merged PR, confirm the docs reflect what it changed: new/changed driver-opts land in the option tables (README.md, docs/reference.md, docs/parent-attached-modes.md); behaviour changes (Health counters, DHCP-client behaviour, identity, recovery) land in reference.md / parent-attached-modes.md / internals.md; examples and numbers match. A milestone PR that changed user-visible behaviour but carries no doc delta is the signal to look harder. That is exactly the drift that the #205↔#152 case (busybox→dhcpcd prose surviving a docs restructure; fixed in #234/#237) slipped through a memory-based read.

Then still read everything user-visible top-to-bottom for anything the per-PR pass misses: README.md (feature list, driver-opt table, examples), GOVERNANCE.md and SECURITY.md, every file under docs/ (including this runbook; process changes during the cycle land here too), and the coverage table if republished. Anything describing the previous version's behaviour, options, or numbers gets updated on the release branch now. Everything under docs/ (plus docs/index.md, the site home) is what the versioned documentation site publishes for this tag, so the review is the site review; there's no separate wiki to reconcile.

Read the pages whole, and aim at the ungated prose. The reference material defends itself: check-option-docs.sh, check-docs-drift.sh and check-version-pins.sh gate every driver option, health counter, plugin setting and image pin, so those tables are the least likely place to find drift. What rots is everything else: a walkthrough's shell snippet, a troubleshooting row, a sentence in a Behaviour section, a hand-maintained list. The v1.5.0 pass found eight divergences (#489) and every one of them was in ungated prose; none would have been caught by grepping for keywords.

Some drift belongs to no milestone PR at all, so the per-PR read cannot reach it by construction. Check these directly:

  • Commands and paths that never existed or stopped working. docker plugin logs was in the README's bug-report checklist and is not a Docker subcommand. Run the commands the docs tell a reader to run.
  • Text invalidated by a feature in this release. #440 mounted STATE_DIR from the host and left two recipes still routing operators through the plugin rootfs. A feature PR updates the section it is about; it rarely finds the other page that quietly depended on the old behaviour.
  • Syntax deprecated upstream. Compose, Docker CLI and dhcpcd move on their own schedule. docker compose -f <snippet> config prints the deprecation warnings for anything in a Compose example.
  • Restated lists that live somewhere else. Required CI checks, registries, privileges. Prefer replacing the copy with a pointer at the authority (the way step 5 defers to branch protection) over updating a copy that will decay again.
  • Mechanisms this release added that no page describes. The failure is absence. Nothing reads as wrong and no grep finds it. Take the release's mechanism changes and ask which section of internals.md covers each; v1.6.0 shipped a lease reclaim and a per-parent gate with no section for either, while every counter table was green.
  • Standing preconditions that read like old-version prose. The BREAKING CHANGE IN v1.5.0 block (README.md, docs/index.md, docs/reference.md) is a precondition for every install and never a changelog entry, but it names an old version, so a pass tidying stale version references deletes it in good faith. Keep it.
  • Counted claims. "Four flip healthy to false" is right until a fifth is added, and the sentence still parses. Check any stated count against the code.

    This bullet's own example went stale while it was being used as the example. It said the claim was "now enforced by scripts/check-health-contract.sh (#638)", and v1.8.0 added a fifth counter: the gate held every counter list it read, and the doc still shipped rows saying "four" over a list of five, because the gate read the count word in one sentence and the file states it in five (#724). The claim was not enforced. A narrower claim was, and the bullet had rounded it up.

    So the instruction is not "change four to five". It is that "a gate covers this" is itself a counted claim, and decays the same way. A gate covers the statements it was taught, and adding a statement is exactly the edit nobody thinks to teach it about.

    Read what the gate says it read. Do not trust that it read everything. check-health-contract.sh ends in a receipt:

    scripts/check-health-contract.sh
    # PASS  healthy contract agrees in N doc counter-list(s), N doc
    #       count-word(s) and N code term(s): <the counters>
    

    The numbers are deliberately not reproduced here. A count quoted into this file is the very thing this bullet is about. Run it and read them: they are what it parsed on that run. The gate asserts no total. If the file states it six times and the receipt says five, the gate is green and the doc is wrong.

    Most gates in scripts/ that discover what to check print the same kind of receipt: a count, or a line per item inspected. A few print a bare PASS, and for those the coverage is not visible at all without reading the script. Either way: when a gate's coverage is load-bearing for a claim in these docs, quote what it read. The fact that it exists is not the evidence.

    The rule survives the correction: the drift ran for two releases under this instruction before #638, which is still the argument for a gate over a rule. It just does not end there: the gate is the floor, and the count of what it reads is the part that has to be re-checked, once per release, by running it.

Verify each finding against the artifact. Reasoning is not evidence. Run the command, config the snippet, ls the path on the test box, query the API (gh api repos/.../branches/dev/protection). A confidently-argued divergence that turns out to be wrong costs more than the one it replaced.

A finding that describes a class ends in a gate. Same rule as anywhere else in this project: if the same shape of staleness can recur on the next mount, option or workflow change, add the check. A promise to remember is not enough. The rootfs-path finding above became a fourth rule in check-docs-drift.sh, deriving the bind-mount destinations from config.json, so that class now fails loudly.

The badge answers are documentation too. .bestpractices.json at the repo root holds this project's OpenSSF Best Practices answers (one <criterion>_status plus a <criterion>_justification each), and every justification is a claim about the repository that can go stale exactly like prose. Reconcile it against the milestone the same way, then check it against the live entry:

python3 scripts/badge-sync.py --diff

If a milestone PR earned or invalidated a criterion (a new gate, a document that now exists, a policy that changed), update the file on the release branch.

Getting the reviewed answers onto the live entry is manual and deliberate: the badge site takes them through its own form, one criteria level at a time, https://www.bestpractices.dev/en/projects/13229/{passing,silver,gold}/edit, in a browser you are signed into. A field only appears on the level that owns it, and the level-less /edit URL 404s for everyone including the owner, which is worth knowing before it looks like an expired login.

Hash-check every justification you paste before submitting. The procedure is in the script's header. Typing an answer by hand is how this project once produced a 65-field divergence from its own source of truth; the hash is what makes hand-entry safe. Then run --diff again: it must report that the entry matches. That confirmation is the point of the script, and the reason it has no push mode.

The work happens here, on the release branch. The rc dry-run (step 8) is the enforcement gate: the real vX.Y.Z tag does not ship until every milestone PR is ticked off against the docs. By the real tag, text and code (and the published site) must agree. 4. Add a ## vX.Y.Z section to RELEASE_NOTES.md, above the previous version's section. Summarise what's changing in user-visible terms; the workflow doesn't auto-build this from commit messages. Include any operator-visible compatibility notes (e.g. v0.8.0 narrowed the IsDHCPPlugin regex; that needed a callout).

House style, applied by default on every release. A release note is reference material an operator scans. v1.8.0 shipped at 90KB of prose and was rewritten to 7KB; write the short version first.

  • Fixed structure, in this order: a two-to-three sentence lead, then ### Upgrade notes, ### New, ### Fixed, ### Deferred, ### With thanks to. Omit a section that has no content; do not add others. ### Deferred lists only work that was on this release's milestone and left it during the cycle, each entry with its issue and where it went. The next milestone's plan is the roadmap, not a note of this release.
  • Put operator-visible behaviour changes in a table under Upgrade notes: one row per change, "what changed" and "what it does to you". That table is the part most readers need.
  • Group ### Fixed by origin (a review, a theme), one line per defect, each ending in its issue number. State the defect and its effect; do not narrate how it was found or how the fix was chosen.
  • No process commentary, no anecdotes, no meta-commentary about the notes themselves, no HTML comments carrying instructions to the next writer. Those belong in the issue or this runbook.
  • Avoid universal claims ("every defect is fixed", "nothing was carried"). They quantify over things a reader cannot check and go false on somebody else's commit. List what is deferred instead.
  • Prefer a named list over a count. Write "six (#720, #721, ...)"; "six of the ten" hides a wrong number.

The heading is exactly ## vX.Y.Z, with nothing after it, and this is checkable. The release workflow extracts the section by an exact heading match, so a heading carrying a marker (## v2.0.0 (unreleased) is the shape this file has carried between releases) matches no tag at all. Until #912 that published a generated one-line placeholder in place of the notes and exited 0. It now refuses, so remove whatever follows the version and confirm before tagging:

# THE STATUS IS THE VERDICT. Piped straight into `head`, this block
# exited 0 whether the notes assembled or the extractor refused,
# because `head` reports for the pipeline, so the two outcomes the
# paragraph below promises were indistinguishable to a reader's
# eye. Assemble to a file first, and `&&` the preview onto it.
section=$(mktemp) &&
scripts/release-body.sh vX.Y.Z RELEASE_NOTES.md > "$section" &&
head -20 "$section"

It prints the first 20 lines of the section the release page would carry and exits 0, or refuses naming what is wrong (decorated heading, no section, empty section, two sections for one version) and exits non-zero with nothing previewed. Run it for the rc tag too (scripts/release-body.sh vX.Y.Z-rc1 …): an rc has no section of its own and publishes this same one, so an rc dry-run whose notes are not ready refuses at the release page and ships no placeholder.

The maintainer signs off on the release notes before the tag. Not optional and not implied by approving the release PR: show the rendered ## vX.Y.Z section and wait for an explicit go. The rc window (step 8) is the checkpoint: by the real vX.Y.Z tag the notes must already be signed off, because the tag publishes them. If the notes change after the tag, edit RELEASE_NOTES.md on dev and gh release edit vX.Y.Z --notes-file the published body, or the two silently diverge.

Credit outside contributions by name, the way the v1.0.0 notes do. Look them up; do not rely on recall. Almost every PR here is the maintainer's or Dependabot's, so an outside one is easy to miss precisely because it is rare:

gh pr list --state merged --limit 300 --json number,title,author \
  --jq '.[] | select(.author.login|test("claymore666|dependabot")|not)
             | "#\(.number) @\(.author.login) \(.title)"'

Also confirm the merged commit still carries their authorship (git log -1 --format='%an <%ae>' <sha>). A rebase or squash of a fork branch is where that quietly becomes the maintainer's. 5. PR release/vX.Y.Z → dev. Required checks on dev are test, policy-gates, staticcheck, integration (every PR builds and exercises its own plugin on the integration runner), actionlint, govulncheck, attribution, docs-site (mkdocs build --strict, #889), CodeQL's Analyze (go) + Analyze (actions), and the CodeQL result check, which fails when the analysis finds a new alert. main requires those plus coverage and coverage-present, which is why the ratchet first bites at the release PR in the next step and not before. Merge when green.

Do not trust the list above. Read the authority. It carried a hand-written total ("eight in total") that was correct when written and became a number nothing checked:

gh api repos/claymore666/docker-net-dhcp/branches/dev/protection \
  --jq '.required_status_checks.contexts'
gh api repos/claymore666/docker-net-dhcp/branches/main/protection \
  --jq '.required_status_checks.contexts'

Branch protection is the authority; if the prose and the settings disagree, the settings win and the prose is the thing to fix.

policy-gates is the half of the old test job that runs the policy gates; the Go suite is the other half (#829). It became required on both branches on 2026-08-28; at the time #829 landed it was on neither list, and for that window those gates ran on every pull request and blocked nothing. If it is ever missing from the output above, that is the state to restore; check before trusting a green release PR.

docs-site became required in #889. Until then nothing required built the site, so a pull request could break the nav or a strict-mode link and be green on every other context; the Docs workflow ran and passed, and a red there blocked nothing. It carries no path filter for the reason the workflow's own header gives: a path-filtered required check is absent on the pull requests the filter excludes, and absent blocks.

coverage-present became required in #735. It is the detector that tells an absent coverage run apart from a pending one, and it was advisory: it could go red without blocking the merge it exists to protect, which is the failure it was built for happening to the guard itself. 6. Open the release PR dev → main with title Release vX.Y.Z and a Closes #N line for every issue in the milestone. The list is what auto-closes them when the PR merges; without it the milestone stays open after the tag.

Because the list is milestone membership, membership has to be true. An issue that is in the milestone but not done gets closed as delivered, silently, by the tag. The taxonomy already says backlog never sits on a milestoned issue for exactly this reason, and the Milestone scope workflow (.github/workflows/milestone-scope.yml, scripts/check-milestone-scope.sh) checks it daily, so it does not depend on whoever builds the list. It splits the two cases, because their fixes are opposite: backlog with in-dev means the work shipped and the label is stale (drop the label, keep the milestone); backlog without it means the work has not started (move it off the milestone). Read that run before opening the release PR: it is a schedule, so a red one waits quietly. Release PRs additionally run the Coverage workflow with the coverage ratchet (scripts/coverage-ratchet.sh vs .github/coverage-baseline.txt): no release ships with less per-package coverage than the previous one. If a package beat its floor during the cycle, raise the baseline as part of the release branch.

Read that run with scripts/coverage-read.sh <run-id> and never by eye (#794). It prints every package's measured number beside both floor sets (dev's and main's), because main's floors can be lower, so a package red against dev may still clear the release PR, and a package comfortable on dev may not. It also refuses on an incomplete comparison, an empty baseline, or a raw block it cannot scope to the ratchet step; none of those is reported as a clean read. Eyeballing the log is how the v1.8.0 read started, and building the instrument instead found three defects an eyeball would have shipped.

bash scripts/coverage-read.sh 32623575563

The exit code carries the verdict: 0 a complete reading, 1 the ratchet's account and the raw covdata numbers disagree, 2 it cannot judge. What it does not do is decide the release: a package under main's floor is information for the raise-or-explain decision the baseline file records; the script cannot settle it.

A package deleted during the cycle reads DROPPED, not as a number. The floors come from the merge base, so main's baseline still floors a package the branch deleted, and the branch's own baseline is what says the deletion was deliberate: gone from the tree and gone from that file is DROPPED, counted as compared, and it does not fail the run. Gone from the tree while the branch's baseline still floors it is a FAIL naming .github/coverage-baseline.txt, and the fix is to remove the row in the change that deleted the package. Read a branch's run with COVREAD_DEV_REF=<that branch>, or the head floors resolve from dev and a package that branch dropped is reported against a floor the run never used.

Coverage shares a concurrency group with the release PR's own integration run, so it normally starts once integration finishes, and since D41 resharded the lane to nine main shards plus two failure shards behind one shared build, that is minutes, down from the five to eight it took under the five-shard layout (#877, D41). Do not trust any range from this page; gh run list --workflow integration.yml --status success re-derives it in one command. A coverage check still showing nothing well past it is worth the next paragraph.

Release PR blocked on a check that has no run. A required check that was cancelled looks exactly like one that is pending: the PR sits at BLOCKED with nothing to click into. It is not a missing trigger. GitHub keeps one running plus one pending run per concurrency group, so pushing another commit to the release PR while coverage is still queued displaces it, and a run displaced before any job was assigned creates no check run at all, which is why the coverage context goes absent and never red.

You should not have to notice this yourself: the Coverage presence check (.github/workflows/coverage-presence.yml, #504) watches the head and fails with the run id and the exact recovery command when the run was evicted. It is a required context on main as of #735, so a red one blocks the release PR and cannot sit in the list being ignored. If it is red, do what it says. The manual form, for a head it did not cover:

gh run list --workflow coverage.yml --limit 5   # look for "cancelled"
gh run rerun <id>                               # once the group is idle

Wait for the integration run on the same ref to finish before rerunning, or it will just queue and be displaced again. This cost a full debugging session on v1.3.5 (#365). The fix is thirty seconds once you know the shape of it. 7. Assemble the verification evidence, don't hand-write it.

scripts/run-evidence.sh "$(git rev-parse 'HEAD^{tree}')"
Prints every integration run that tested exactly this tree, with its window and what else was on the privileged pool at the time. Paste it into the release PR. Do not reconstruct it from memory.

Read the overlap line literally. An overlap of none, printed with ran alone after it, and an overlap of unknown are different claims: the second means the concurrent-run list did not reach back far enough to judge, which happens once the repo has been busy since. Do not upgrade an unknown to "ran alone". The v1.4.0 write-up asserted a concurrency caveat that the data did not support, in both directions, which is what #432 was filed about.

  1. Merge the release PR. Squash or merge commit, both fine; match what's in git log.
  2. Pull main, dry-run, then tag: first push vX.Y.Z-rc1 and confirm the workflow run is green end-to-end (pre-release mode, :latest untouched; see "Pre-release dry-run" above). Then:

    git checkout main && git pull --ff-only &&
    # Step 4's check, on the tree that is about to be tagged. The release
    # page's body comes from here, and its refusal would otherwise arrive
    # after the images are pushed.
    #
    # THE `&&` IS THE POINT. Newline-separated, a refusal here printed its
    # error and the next line pushed the tag anyway, which is the failure
    # this check exists to prevent. Chained, the block stops at the first
    # non-zero and nothing below it runs.
    scripts/release-body.sh vX.Y.Z RELEASE_NOTES.md >/dev/null &&
    git tag -s vX.Y.Z -m "vX.Y.Z: <one-liner>" &&   # signed (#175)
    git push origin vX.Y.Z
    
    Use -s (signed) so the release tag shows Verified on GitHub; the dev box has tag.gpgsign=true so -a would also sign, but spell it out so it holds from any checkout. Confirm with git tag -v vX.Y.Z (or the green "Verified" on the tag page). The workflow fires on tags: v*. Watch it at https://github.com/claymore666/docker-net-dhcp/actions/workflows/release.yml. Expected steps, under the names the run shows. Tag resolution is its own job: resolve runs first and has one step, Resolve release tag; a releaser watching the run sees two job rows. The release job then runs, in this order: checkout → setup-go → Log in to GHCR → Log in to Docker Hub → Both registries, or say why not → Push to GHCR → Push to Docker Hub (or skip) → Sync Docker Hub description from README (or skip) → Sync the Hub alias description from README (or skip) → Install cosign → Record and gate the cosign version → Sign published images (cosign keyless) → Install oras → Publish the same manifest under the Hub alias (or skip) → Install syft → Generate SBOM (SPDX + CycloneDX) → Package and sign release artifact → Attest release-artifact provenance → Publish and verify the release provenance bundle → Attest image provenance (GHCR) → Check attestation parity across registries → Upload signed artifacts for the release job → Workflow summary.

    Publish and verify the release provenance bundle runs scripts/publish-provenance-asset.sh, which attaches the attestation the step before it produced to the release page as provenance.intoto.jsonl (provenance-arm64.intoto.jsonl in the arm64 job), so provenance is readable from the page and not only from GitHub's attestation store (#1011). It refuses an asset name Scorecard's provenance check would not count, re-reads the subject list out of the bundle it is about to publish, refuses a bundle that does not name the tarball, and runs gh attestation verify --bundle over every subject, which is the command Verifying releases gives users. A red here means the published bundle does not verify the published bytes, and the release stops before the page exists. Every one of those refusals is driven offline on each lane run by scripts/test-publish-provenance-asset.sh, with gh stubbed.

    Publish the same manifest under the Hub alias runs scripts/publish-hub-alias.sh, which copies the signed manifest and its referrers with oras cp -r, re-reads the digest through the alias name, refuses anything that is not the digest just signed, and verifies the signature under the alias. Docker Hub can list the copied signature later than the manifest (once still missing 1.4 s after the copy), so the verify reads up to six times, 10 s apart, before it fails (#1043). It comes after signing on purpose: the alias is the same manifest, not a second build (#267). The two Hub description steps are separate because the action PATCHes one repository at a time.

    Since v1.7.0 the run carries a parallel arm64 chain (#507): release-arm64 (native ubuntu-24.04-arm build, pushes vX.Y.Z-arm64; per-arch tags, because a Docker plugin cannot install from a manifest list) and verify-install-arm64.

    Then, as separate jobs:

    • verify-install / verify-install-arm64: install the just-published plugin from GHCR on a clean hosted runner and assert it enables. A red verify-install means users can't install what we just shipped.
    • verify-install-hub / verify-install-hub-arm64, and since #972 verify-install-hub-alias / verify-install-hub-alias-arm64: the same proof for Docker Hub under each of its two names, which is the other place a user installs from (#776). The alias proofs are gated on the same hub_pushed output as the other two, not on the copy step's own result, so a skipped copy does not also skip its own proof. Each is its own job on its own runner and that is deliberate: the value of these jobs is a daemon that has never created a network sandbox, which is how v1.6.0-rc2 caught a bind source the daemon creates lazily (#588). A second install appended to verify-install would run after that property was already spent, and docker plugin rm does not give it back. When a run published no Hub image the steps are skipped and the job records ⚠️ Docker Hub install not verified in the summary, so a GHCR-only run cannot be mistaken for a both-registries one.
    • promote-latest: since #736 this is where every floating tag moves, for both arches and all three published names, and it runs only after all eight of the above are green. scripts/check-latest-promotion.sh asserts that dependency. It does not carry a list of proof names: it derives them from the workflow's own install-verifying jobs, so a ninth proof is required the moment it exists. Steps: Refuse to promote a floating tag backwards → Install crane → the two logins → Record what :latest resolves to before promotion → Promote the GHCR floating tags → Promote the Docker Hub floating tags → Promote the Hub alias floating tags → Verify the floating tags resolve to the signed digests → Assert a pre-release did not move :latest.

      Two of those are guards whose evidence comes from the registry, and they check different things. Verify the floating tags resolve to the signed digests compares the floating tag against the version tag by digest. That is the #267 guard, that retagging preserved the digest the signature covers. Assert a pre-release did not move :latest re-reads :latest and compares it to what the Record step saw before anything was touched; it runs only on an rc, and it is the one that proves the rc contract from outside. - github-release: does not wait for promote-latest; it needs the same eight jobs. Promotion and the Releases page are siblings, so a refused promotion does not suppress the release.

Every green checklist below includes the arm64 jobs. 10. Confirm the GitHub Release. The github-release job now cuts it automatically once the install proofs are green (so a plugin that doesn't install never gets an advertised Releases page). It attaches the cosign-signed artifacts and builds the body as: a generated lead line naming the project and version (it becomes the page's og:description, so it is what link previews show, #469), the ## vX.Y.Z section of RELEASE_NOTES.md, a generated Downloads table, and a link to Verifying releases. Step 4's notes must therefore already be in place at tag time. rc tags produce a draft release: the publish path is still exercised, but dry-run builds stay out of the public list. No manual gh release create; instead verify:

gh release view vX.Y.Z   # body = the RELEASE_NOTES section; assets:
                         #   net-dhcp-plugin-vX.Y.Z-linux-amd64.tar.gz
                         #   net-dhcp-plugin-vX.Y.Z-linux-arm64.tar.gz
                         #   checksums.txt + checksums.txt.sigstore.json
                         #   checksums-arm64.txt + checksums-arm64.txt.sigstore.json
                         #   provenance.intoto.jsonl + provenance-arm64.intoto.jsonl
# Re-verify the signature the way a downstream consumer would:
cosign verify-blob \
  --bundle checksums.txt.sigstore.json \
  --certificate-identity-regexp '^https://github.com/claymore666/docker-net-dhcp/.github/workflows/release.yml@' \
  --certificate-oidc-issuer https://token.actions.githubusercontent.com \
  checksums.txt
# And the provenance, taken from the page, with no API call:
gh attestation verify net-dhcp-plugin-vX.Y.Z-linux-amd64.tar.gz \
  --bundle provenance.intoto.jsonl \
  --repo claymore666/docker-net-dhcp
Adjust the title/notes in the UI if the one-liner needs polish. The job is idempotent on a tag re-dispatch (re-uploads assets with --clobber). This satisfies OpenSSF Scorecard Signed-Releases, whose provenance half reads the asset NAME and counts only a .intoto.jsonl suffix, which is what the two provenance assets are named for (#1011); an rc dry-run produces an equivalent pre-release with the same signed assets, which is how this path is exercised before the real tag (rc releases never move :latest and are marked pre-release). 10b. Nothing to refresh, and that is a change from 1.x. Through 1.x this step pasted a per-release digest block into Verifying releases, and since #547 the release run enforced it. From 2.0 that step is gone, from the doc and from release.yml both.

Why it had to go. The tag and the commit are compiled into the binary from 2.0 (VERSION/COMMIT through -ldflags -X, plan row O-4). The step's own recovery was "take the block from the failed run, land it on the tagged commit, re-tag", and landing it changes the commit, which changes the binary, which changes the digest the block states. The recovery could not converge, so the gate had no passing state to reach: an operator following the runbook to the letter would re-tag forever, each attempt leaving a :vX.Y.Z published and :latest behind it, because the failing step sat between the GHCR push and the Hub push. Measured in review: two builds differing only in the commit ldflag hash differently.

Where the digests are now. In the signed checksums manifest, produced by the build that makes them true. checksums.txt and checksums-arm64.txt each cover the tarball, both SBOMs and the binary inside the tarball, recorded as rootfs/usr/sbin/net-dhcp, so an operator who extracts the tarball beside the manifest can run sha256sum --ignore-missing -c checksums.txt and check the binary against a cosign-signed reference. That path is the same in both manifests, because a plugin tarball has one layout, so each architecture is verified in its own directory, which is what docs/verifying-releases.md tells the reader. Nothing here is a release step: the release job does it, on every tag and every rc, with no paste and no re-run.

What stops this from rotting. scripts/check-release-digest- fixed-point.sh runs in the normal test lane and holds the two halves of the argument above: it MEASURES that the commit is still part of the binary's identity, refuses any file in the tree that records a digest of a binary this tree builds while that is true, and requires the release workflow's signed manifest to cover the binary. Restore the block and it goes red before a tag is ever cut; delete the manifest line and it goes red too.

  1. Fast-forward dev to main so the release commit (version pins, RELEASE_NOTES section) lands on dev too:
    git checkout dev && git merge --ff-only main && git push origin dev
    
    Skipping this leaves the next feature branch starting from the previous version's README/docs, and the next release PR has to re-bump them. Forgotten once after v0.9.0. That's why release.yml's header comment carries the same checklist.

Do this before merging anything else into dev. Forgetting is the mild failure. The one that has actually happened is foreclosure: the --ff-only is possible only while dev has no commits of its own since the release, so the window opens when the release PR lands and closes on the next merge into dev, permanently, and by doing something otherwise correct.

# is a back-merge OWED?  prints a count
git rev-list --count origin/dev..origin/main
# does `dev` carry commits of its own?  (--is-ancestor prints nothing)
git merge-base --is-ancestor origin/dev origin/main \
    && echo "dev is contained in main" \
    || echo "dev has commits of its own"

Two different questions, and neither answers the other. The count is commits on main that dev lacks; it says a back-merge is owed. The predicate asks whether dev has commits of its own, which is what --ff-only needs. Both outcomes of both, because a documented command with only one outcome written down leaves the reader to supply the other by negation:

count dev contained in main? state do
non-zero yes owed and available, the normal post-tag state run the --ff-only above, now
non-zero no owed, and foreclosed back-merge PR
zero yes main and dev identical nothing
zero no dev ahead mid-cycle, the ordinary state nothing

A non-zero count is not bad news: at the v1.6.0 tag it was 4, and the fast-forward worked. It is the warning the CI advisory prints, in those words: "While that is true, git merge --ff-only main still works."

Once foreclosed the recovery is a back-merge PR, main → dev, merged with a merge commit; a squash does not put main's commits on dev, which is the entire point of this step.

The failure is silent, which is why this needs a command and not a reminder. The divergence carries no content: every release merge has a dev commit as its second parent (because step 5 merged the release branch into dev before the tag), so git diff main dev is clean and the whole suite passes while the graph is wrong. Nothing reads as broken, so nothing prompts the check.

Worked instance: v1.6.0 published at 23:11:30Z on 2026-08-16; three Dependabot PRs went into dev at 23:26, fifteen minutes later, before anyone looked at main. That foreclosed it; recovered by #597 at 23:38. Fifteen minutes is not a window anyone watches by hand.

CI warns, but it cannot block. .github/workflows/release-backmerge.yml runs scripts/check-release-backmerge.sh in two modes:

  • Advisory, on every PR into dev: warns on any divergence and ignores the grace window, because the fresh divergence is the only case this mode exists for. It warns and never blocks, on purpose; reasoning in the workflow header. An annotation can be scrolled past, which is why the command stays in this step.
  • Enforcing, on the nightly schedule: the backstop for the permanent omission, and the only mode BACKMERGE_GRACE_HOURS (default 24) applies to, so a release in flight is not a false red.

One caveat on the recovery: the back-merge diff is not always empty. #597 was, but 6af0749 carried three real files. Hitting a conflict there does not mean you have done something wrong. 12. Prune merged branches. The repo has Automatically delete head branches enabled, so merged PR head branches are removed on merge. Two things that setting doesn't cover, so clean them now:

# the release branch is merged but was never a PR head:
git push origin --delete release/vX.Y.Z
# sweep for any other branch already merged into dev that lingered:
git fetch --prune origin
git branch -r --merged origin/dev | grep -vE 'origin/(dev|main|HEAD)$'
Delete what that sweep lists. Leave alone: open-PR branches, Dependabot branches (it recreates its own; close via the PR), and the upstream/* refs (those are the original fork's remote).

Verifying

After the workflow succeeds:

  • curl -sI https://hub.docker.com/v2/repositories/claymore666/net-dhcp/tags/vX.Y.Z/ returns HTTP/2 200, and so does the same call for claymore666/docker-net-dhcp. Both Hub names are published by the release run; a 200 on one and a 404 on the other means the alias copy did not happen and the run should have been red.
  • curl -sI https://ghcr.io/v2/claymore666/docker-net-dhcp/manifests/vX.Y.Z returns HTTP/2 401 (auth required). The manifest IS there, GHCR just won't expose it anonymously. To confirm presence authenticated: gh auth token | docker login ghcr.io -u <you> --password-stdin && docker plugin install ghcr.io/claymore666/docker-net-dhcp:vX.Y.Z.
  • Both Docker Hub pages (https://hub.docker.com/r/claymore666/net-dhcp and https://hub.docker.com/r/claymore666/docker-net-dhcp) show the new tag in the Tags tab and the README content matches GitHub. The workflow syncs the description per repository, so both are covered. Hub categories are set in the web UI only, with no API behind them, so a new Hub repository keeps whatever categories a person gave it and no run will fix them.
  • The milestone is closed (every issue moved to Done by the release PR's Closes list). Verify with gh issue list --milestone vX.Y.Z --state open; should be empty.
  • Anything that was listed in .github/dispatch-pending.txt is now dispatchable. Exercise it once. A workflow_dispatch workflow is only exposed from the default branch, so one that merged to dev during this cycle has never run, and this release is the first moment it can. Dispatch it and confirm it does what its documentation claims. The entry itself is already gone: it is removed on the release branch at step 2, travels into dev at step 5 and into main with the release PR, because scripts/check-dispatch-reachable.sh counts a workflow that the release PR merges into the default branch as reachable and its entry as stale (#977). While the release branch and dev pin the next version and main pins the current one, that gate accepts the missing entry, so step 5 and every other pull request into dev during the release window stay green. It is reading the two pins and nothing else; the bounds that come with that are in step 2. Dropping the entry after the release instead is what turned the gate red on main at v2.1.0.

Troubleshooting

Symptom Likely cause Fix
Workflow shows zero-job "failed" runs on every push, tag push doesn't trigger anything release.yml parse error (often secrets context in step-level if) Fix the YAML, then verify with an rc tag (gh workflow run release.yml -f tag=vX.Y.Z-rcN --ref main), which exercises the whole chain without touching a bare release tag or :latest. Then retry the real tag. Do not verify by dispatching the release tag: Dispatching an existing tag
Push to GHCR step ends 403 Forbidden GHCR package not linked to repo with Write One-time fix in package settings (see prerequisites)
Push to Docker Hub step ends unauthorized: incorrect username or password Token revoked / expired / wrong scope Regenerate at hub.docker.com, update DOCKERHUB_TOKEN repo secret
Sync Docker Hub description from README step ends 401 Token scope is image-push only and lacks admin Regenerate token with broader scope (see prerequisites)
Hub page README is stale after a release Description-sync step skipped (no Hub creds) or 401'd Check the workflow run; either set creds or fix the token
Tag push succeeded but no Hub publish HAS_HUB_CREDS evaluated false (secrets blank) Set the secrets, then exercise the publish from an rc tag: gh workflow run release.yml -f tag=vX.Y.Z-rcN --ref main. Do not dispatch the bare release tag: it rebuilds and re-points that tag plus :latest, and scripts/assert-newest-release-tag.sh refuses an older one outright; Dispatching an existing tag

Backports between dev and main

When a release-blocking hotfix has to land on main without going through dev (e.g. v0.8.0's release.yml parser bug), the flow is:

  1. Branch off main, fix, PR to main, merge. Don't push to main directly; branch protection and the audit trail.
  2. Cherry-pick the same commit onto a branch off dev, PR to dev. This keeps dev from regressing on the next release PR.

The v0.8.0 cycle uses #97 (main hotfix) and #98 (dev backport) as the canonical example.