Plan: the image supply chain, from end-of-life migration to production that reports its own failures¶
Prompt¶
Written 2026-09-13, consolidating work that had spread across three concurrent sessions and nine open issues in four repositories. Two of those sessions have been closed; this plan is the single statement of what is outstanding and in what order.
Two threads run through every issue here and they are not the same problem. The first is a migration: Debian 12 reached end of standard support on 2026-06-10 and the fleet still names it in 80 places. That is finite work with an end. The second is that the machinery producing our base images failed five separate times in one week without telling anyone, and was found only because somebody went looking for something else. That has no end unless the machinery changes.
Read docs/audits/eol-distro.md for what the criterion measures
and, importantly, what it cannot see.
Situation¶
The open work, as filed¶
| Issue | Repository | What |
|---|---|---|
| #123 | development | Debian 12 runner labels and container bases, fleet wide |
| #38 | private-ci | Collated inventory of obsolete base image usage |
| #39 | private-ci | Dependencies cache disk still built from debian:11 |
| #40 | private-ci | Retire the unused debian-11 runner image and label |
| #44 | private-ci | A conductor restart silently skips that day's nightly rebuild |
| #45 | private-ci | No debian-gnome-13; the last bookworm label with no successor |
| #826 | 33fl | Two GitLab static runners stay on bookworm until deleted by hand |
| Phases 1-6 | images | PLAN-image-build-modernisation.md; see its own Execution table for what has landed |
Already landed and not repeated below: development#119 (the
eol-distro criterion), actions#66 (debian-13-docker published a
daemon with no client), 33fl#820 (static runners now built on
Debian 13), images#2 (the build fixes) and images#3 (that
repository's own plan).
Five silent failures in one week¶
Each of these was found by hand, and none of them raised an alert:
- shakenfist/images published nothing for sixteen days.
build.shis#!/bin/bash -erun from cron; one failing image aborted the run and cron mailed root, which nobody reads. debian-docker:12,debian-gnome:12anddebian-xfce:12were Debian 11 for two years and two months. They passedDIB_RELEASE=bullseyewhile namingdebian-12-extras, from 2024-07-06 until 2026-09-12. Nothing checked that the image matched its own name.ci-images/debian-13-dockerpublished a working daemon with nodockerbinary. Debian 13 splitdocker.io, and the CI images are built with recommends disabled. Fixed in actions#66 by making the build rundocker versionand fail.- private-ci's nightly rebuild did not fire on 2026-09-12 and nothing said so. A day on which the loop neither builds nor errors is indistinguishable from a day with nothing to do.
debian-11reportsFalsein every nightly cycle result and has done since it stopped being buildable. The cycle summary carries a permanent failure that nobody reads.
The shape is identical every time: something that produces images stopped producing correct images, and no signal existed. Four of the five were found in the same week only because one investigation led to another.
There is a measured cost to the staleness, beyond the missing
images. While debian:13 was frozen for those sixteen days, every
CI image built from it apt-upgraded a fortnight of packages during
the build and carried the superseded versions into the published
blob. Measured on the plain debian-13 label, where nothing
changed but the freshness of the base:
| Version | Size |
|---|---|
| v61 (stale base) | 1320.4 MB |
| v62 (refreshed base) | 1146.3 MB |
174.1 MB, 13.2%. The same comparison on debian-13-docker is
quoted as 15% on #44, but that pair conflates the base refresh
with actions#66's own fix; the plain label above is the clean
measurement.
The structural finding¶
The eol-distro criterion greps workflows for runner labels and
container images. That makes it a check on consumers. Every
place the fleet actually produces an end-of-life image is
invisible to it:
- shakenfist/images decides which releases get built at all. It sits on the audit's excluded list, whose stated reasons are internal tooling, historical archives and non-projects; none of the three fits a repository built from nightly that the fleet's CI depends on.
- private-ci decides which images become runner labels, in
IMAGE_BUILDSandCI_IMAGES. It is in scope for four plan criteria andsfui-vendor, and nothing else. - 33fl's static runners advertise only
self-hostedandstatic, so no workflow anywhere names an operating system. A grep of every consuming repository finds nothing while every one of those jobs runs on a retired release.
So the three definitions that create the exposure sit outside the criterion that bans it, and the 80 findings in #123 are the downstream shadow of decisions the audit cannot read. Fixing the 80 without fixing the three means the count returns at the next end-of-life date.
Mission and problem statement¶
Close the nine open issues in a sequence that does not break CI on the way through, and change the image pipeline so that the next failure announces itself instead of waiting to be noticed.
Not in scope: what goes inside the images, the DIB patches carried in shakenfist/images, and whether the fleet should consume these images at all. 33fl is a different organisation with its own conventions, so this plan tracks its one issue and does not prescribe how it is fixed.
Open questions¶
Q1. One criterion that reads producers, or a second criterion?¶
Phase 6 has to measure the producer definitions that eol-distro
cannot see. Either eol-distro grows the ability to read
IMAGE_BUILDS, CI_IMAGES and a build list, or a second
criterion does it.
Default if nobody answers: a separate criterion. The two
report different defects -- one says "this repository names a
banned label", the other says "this repository offers one" --
and they are fixed by different people. More decisively, an issue
title is the fleet-wide idempotency key for filing and closing,
and scripts/tests/test_metadata.py freezes those titles
precisely because changing one orphans every issue already open
under the old title. Widening eol-distro's meaning changes what
its existing title claims; a new criterion carries a new title and
orphans nothing.
Decisions¶
D1. Signal before surgery¶
Every later phase changes something that produces images, and the whole reason this plan exists is that we cannot currently tell when image production breaks. Making those changes first and the detection last would be running the same experiment that produced the sixteen-day outage.
So detection comes first, even though it closes no migration issue and resolves nothing on #123. Phases 1 and 2 are also the only phases that address "so they don't occur again"; the rest is cleanup that a future end-of-life date will otherwise recreate.
D2. Verify the artifact, not the name¶
Three of the five silent failures were a published artifact that
did not match its own label: bullseye as debian:12, a docker
image with no docker, a runner label whose base no longer builds.
A freshness check catches none of these -- all three were current,
and two were being rebuilt nightly.
So the plan carries a separate phase that asserts what an image
is rather than when it was made. actions#66 already established
the pattern by ending the build with docker version; this
generalises it.
Operative reading, added 2026-09-14. Implementing phase 2 showed that "verify the artifact, not the name" taken literally is the wrong way round: comparing the artifact against the build's own inputs is exactly what would have passed every night for two years. The reading that works is verify the artifact against the name it will be published under, and the name is held by whatever publishes rather than whatever builds. The correction in phase 2 sets out why. Anything else applying D2 -- the private-ci bullet in that phase, and any future producer -- takes this reading.
D3. debian-gnome-13 before the consumer sweep¶
private-ci#45 is the only issue on the critical path of #123.
debian-gnome-12 is the last bookworm label with no successor, so
any repository whose workflows name it cannot be migrated until the
successor label exists. Everything else in #123 is a swap between
labels that both already work.
D4. Retire producers last, and only after their consumers¶
private-ci#40's follow-up retires debian-12 and
debian-12-docker from IMAGE_BUILDS and CI_IMAGES. Doing that
before #123 completes takes the runners out from under jobs that
still request them. The debian-11 half of #40 has no such
constraint -- nothing requests it -- so it can go early.
D5. This plan does not re-plan shakenfist/images¶
That repository has its own plan, merged in images#3, with seven phases of its own. Its phase 1 (per-image failure isolation) and phase 4 (the freshness watchdog) are load-bearing here, so they are named in the phase table below, but the detail stays there rather than being copied.
D6. 33fl gets no detection here¶
The Mission says 33fl is a different organisation with its own
conventions, and that applies to detection as much as to fixes.
Its static runners are named as the third producer in the
structural finding because the exposure is real and worth
recording, but phase 1 builds two detections, not three. 33fl#826
tracks its own rollover, and the note in
group_vars/all/static_runners.yml is where the exposure is
recorded for the next reader.
What belongs here instead is the general lesson: a static runner fleet advertises no operating system, so it is structurally invisible to a label-based audit. Any future fleet of that shape needs the same treatment, and phase 6 says so.
Execution¶
| Phase | Status | Merged |
|---|---|---|
| 1. Alarm on absence | Complete | images acccd2b (#5), images 47ed141 (#6), private-ci ae7b1f8 (#49), private-ci e1f8fb1 (#54), 33fl 6de1764 (#827) |
| 2. Verify the artifact, not the name | Complete | images 6028647 (#7), images 4800c72 (#8), images b872641 (#9), actions 2ac4a94 (#74), 33fl bbbd842 (#836) |
| 3. Unblock the migration | Complete | private-ci 9eace9d (#60), private-ci 2e18c13 (#61), private-ci dbb78ca (#63), actions 8684eec (#78), actions 781d267 (#80), kerbside 79c2506 (#435), kerbside cfef26a (#450) |
| 4. The consumer sweep | Complete | Label half: actions 5a677a9 (#90), agent-python 8303e99 (#140), client-python 93a0999 (#405), client-python-k3s 1576e72 (#67), clingwrap d5eb4ea (#136), divergulent 53136f2 (#117), library-utilities 6f63b95 (#60), ryll 060e649 (#397), sfui 30f5501 (#35), instar a1c09aa (#589), occystrap 4b9d5ff (#143), shakenfist 54b18a0 (#4306), with visual-digest-rust #23 closed as superseded. Guest-image half: actions 8c02ab0 (#97, 4f), shakenfist 2a94e58 (#4379, 4g), shakenfist 984fdd1 (#4385, 4h part 1), actions 227593c (#124, 4h part 2). 4i and 4j carry no commit. |
| 5. Retire the end-of-life producers | In progress | |
| 6. Close the audit's blind spot | Not started | |
| 7. Push audit | Not started |
1. Alarm on absence¶
Closes: part of private-ci#44. Depends on: nothing.
Status: complete, 2026-09-14. Both detections are built, merged and deployed. The shakenfist/images watchdog is images#5; the conductor's persisted nightly and its staleness gauge are private-ci#49; the stale-label issue is private-ci#54; the Grafana rule watching the gauge is 33fl#827. images#6 landed the per-image failure isolation this phase carries as prevention.
Correction, 2026-09-17. This paragraph used to say images#6 "is
audited against the images repository's default branch as part of
that pull request, per phase 7". It is not. That is what the push
audit shared block requires, not a record of something that
happened: shakenfist/images has no PUSH-AUDIT.md, and phase 6 of
its own PLAN-image-build-modernisation -- the phase that would run
the audit -- is Not started. Restating a policy in the past tense
is how a plan comes to believe it has evidence it never collected.
What is true is recorded under phase 2 below, for every landing in
this plan so far.
The conductor deployed at 20:37 on 2026-09-13, carrying private-ci
e1f8fb15. It logged the priming path on startup -- it claimed that
day's slot rather than starting a rebuild in the evening -- and
Prometheus is scraping the gauge, which reads the primed slot rather
than zero. So the fallback that seeds from the claimed slot is
doing its job: without it a fresh deployment would read zero and
alert immediately despite having missed nothing.
How the invariant at the head of this phase is satisfied. The gauge is exported by the conductor process itself, on maui, and scraped by Prometheus on the same host. Taken alone that would re-create failure 4: a conductor that is not running exports nothing, and a rule that only compares a timestamp sees no data and stays quiet. Two things prevent it, and both are load-bearing rather than incidental.
up{job="conductor"} is synthesised by Prometheus when a scrape
succeeds or fails, not exported by the conductor, so it reports 0
for a conductor that is not running. The pre-existing Host down
rule in 33fl watches it. That is the detection which does not
depend on the thing it watches, and it was already in place before
this plan.
33fl#827 also sets noDataState: Alerting, so the absence of the
gauge alerts in its own right rather than reading as health.
Stating the division precisely, because the first draft of this
note got it wrong: private-ci#54's stale-label issues cover a
conductor that is running and reports a label going stale. A
conductor that is absent entirely is covered by Host down and by
the no-data state of 33fl#827, not by anything conductor-side. The
images producer is covered independently of all of it by images#5,
which reads the published site from a hosted runner.
Do not use the 26 hour threshold this plan originally specified.
It was taken from the Nightly report not dispatched rule, where
the thing being timed is a workflow dispatch and is instantaneous.
A full image rebuild is eleven images built serially and takes about
an hour and a half -- six runs between 2026-09-07 and 2026-09-13
took between 1h24m and 1h37m -- and the gauge advances on
completion, not on the slot. So the first completion after priming
lands 25.6 hours after the primed slot, which leaves 24 minutes of
margin against a 13 minute observed spread in build duration. The
rule would have had a real chance of firing spuriously on its first
night, which is the worst possible introduction for an alert.
Use 30 hours. A genuinely skipped night reaches 48, so anything between roughly 28 and 44 separates "skipped" from "slow", and 30 still alerts within about four hours of when completion was due. 33fl#827 landed at 30 hours with that reasoning recorded beside it.
The general point, for any future rule of this shape: a threshold over a completion timestamp has to cover the slot interval plus how long the work takes, not just the slot interval.
What follows is the original specification, kept verbatim because the phase 2 correction below shows what is lost by quietly editing a plan to match what was built. It was departed from in one place. The private-ci bullet specifies "a scheduled job elsewhere that reads the published dashboard or the conductor's API"; what landed reads a Prometheus gauge through a Grafana rule instead, because that estate already existed, already scraped the conductor, and already carried two rules of exactly this shape to copy. The requirement the bullet was protecting -- that the detection not depend on the thing it watches -- is met as set out above. The scheduler-persistence item in the last paragraph is done, in private-ci#49.
Two detections, for the two producers this organisation controls. Neither may depend on the thing it watches -- a component that has stopped running also stops reporting that it has stopped running, which is how failure 4 stayed invisible for a day. 33fl's static runners get no detection here, per D6.
- shakenfist/images: phase 4 of that repository's plan. A
scheduled
HEADagainstimages.shakenfist.com/<image>/latest.qcow2, comparingLast-Modifiedagainst a 72-hour threshold, from a runner with no connection to the build host. This measures what a consumer receives, so it also catches a broken publish step or stale nginx. - private-ci: the conductor already tracks blob age per label
in
_update_status(image_ages=...), so the data exists. It must not be the conductor that reports on it, though: the conductor is the component whose restart silently skipped a nightly rebuild, and a conductor that is not running cannot tell anyone it is not running. The detection is therefore a scheduled job elsewhere that reads the published dashboard or the conductor's API and files an issue when any label's age exceeds N days, or when the cycle summary is absent or stale.
Phase 1 also carries one piece of prevention, because it is what makes the detection actionable: phase 1 of the shakenfist/images plan, so one failing image stops one image rather than the whole run. An alarm that fires for the entire list every time teaches people to ignore it.
Also fix the scheduler bug #44 documents: persist the last
completed nightly rather than recomputing nightly_due into a
local at every loop start. Note that #44 is honest that this
mechanism does not explain the 2026-09-12 miss, so the fourth
checkbox there -- working out what actually happened -- stays open
after this phase and may be a separate defect.
2. Verify the artifact, not the name¶
Closes: nothing on its own. Depends on: nothing.
The phase that would have caught three of the five failures, and the one most likely to be dropped for being nobody's issue.
Status: complete, 2026-09-16. verify-release is in the
element list of all fourteen images published on the night of
2026-09-16, the build host is on images b872641 with a
root-owned checkout, and the desktop and dependency assertions are
on actions main. private-ci#45's first checkbox is ticked with
the artifact read directly -- Debian 13.7, trixie, GNOME 48 -- so
phase 3's stated dependency is recorded rather than remembered.
None of this plan's out-of-repository landings has been push
audited, and phase 7 should not expect to cite one. Checked
2026-09-17 across all ten merged pull requests recorded in phases 1
and 2 -- images#5, #6, #7, #8, #9; private-ci#49, #54; actions#74;
33fl#827, #836 -- plus private-ci#60 from phase 3. Not one carries
an audit in its body or its comments, and none of
shakenfist/images, shakenfist/actions,
Mach33Labs/33fl or shakenfist/private-ci has a PUSH-AUDIT.md
at all -- so no such audit could have been run. The shared block's
"audited as part of the pull request that lands it" has been the
plan's assumption rather than its practice.
The shas themselves are verified. All eleven merge commits
recorded in the Execution table were read on 2026-09-18 with git
cat-file -p <sha> in each of the four repositories: every one has
two parents and a message naming the matching pull request. No
repository here squash-merges, so each recorded sha's diff against
its first parent is the whole of what landed, which is what the
push audit shared block requires of the record rather than merely
of the number.
What that costs phase 7 is specific rather than general. It has
three options and should say which it took: run the accumulated
audit itself against each repository's default branch, which is the
work the shared block says it may cite instead of doing; accept the
three shakenfist/images landings as covered by that repository's
own phase 6 when it runs, which leaves actions#74 and 33fl#836
uncovered by anything; or record the gap and decline it in writing.
The one landing where this is not bookkeeping is images#8, which
broke the nightly build on its first night: an audit of the
accumulated diff is precisely the instrument that would have looked
at a self-update under errexit, and it did not run.
What this phase cost, which belongs in phase 7's audit. The
self-update it shipped (images#8) broke the nightly build on its
first night. build.sh runs under errexit, and
before=$(git rev-parse HEAD) on a line of its own is a simple
command, so when git refused the checkout on ownership the run
ended at 05:00:01 having published nothing. The host has no MTA,
so cron discarded the one line that explained it. It was found by
the completion check at the head of phase 3's planning, not by any
alarm: tools/check-image-freshness.sh uses a 72 hour threshold
and would not have reported it until 2026-09-17.
Two fixes, one per cause: images#9 puts every git command inside
the if condition so any git failure warns and builds anyway, with
two regression tests that exit 128 against the previous script;
33fl#836 owns the checkout as the user cron runs as, and asks
cron's question by stripping SUDO_UID rather than sudo's. The
general lesson for the rest of this plan: a check run under sudo
is not the check cron runs, and git is one of several tools that
behaves differently between them.
- shakenfist/images: after building an image, assert that it
is what it claims. Landed 2026-09-13 as the
verify-releaseelement (images#7). This was listed under Future work in that repository's plan; this plan promoted it, because it is the only control that addresses D2.
Correction, 2026-09-13. This bullet used to say that reading
/etc/os-release and comparing it against "the release the build
asked for" was enough to have caught the two-year bullseye defect
on its first night. That is wrong, and the way it is wrong is the
most useful thing in this phase.
Nothing in those builds disagreed with itself. build.sh passed
DIB_RELEASE=bullseye, diskimage-builder built bullseye, and the
image honestly reported bullseye. Every comparison between the
image and the build's own inputs would have passed, every night,
for two years and two months. What was wrong was the name the
artifact was published under.
So the invariant worth asserting is not "the image matches what
was requested" but "the image matches what it is about to be
called" -- and the name is held by the thing doing the publishing,
which is usually not the thing doing the building. The element
therefore compares against the publish label, which build.sh
passes in, and keeps the DIB_RELEASE comparison only as a
secondary check.
D2 says "verify the artifact, not the name". Read literally that
is the wrong way round: verifying the artifact against the build
is what would have failed here. Read as "verify the artifact
against the name it will be published under", which is what it
was reaching for, it is exactly right. Anything else applying D2
-- the private-ci bullet below, and any future producer -- should
take the second reading.
* private-ci: confirm debian-gnome:13 is genuinely trixie
before wiring it up, per #45's own caveat, and adopt the
actions#66 pattern -- end each image build by exercising the
thing that image exists to provide.
3. Unblock the migration¶
Closes: private-ci#45 (partly -- see decision 3.5), private-ci#39. Depends on: phase 2, satisfied. Planning effort: high, because the sequencing spans two repositories and the gnome snapshot machinery is not where the original sketch said it was.
Two labels are stuck. debian-gnome-12 has no trixie successor, so
the eol-distro criterion bans a label the fleet still needs. And
the dependencies cache disk -- which every runner and inner CI
primary mounts at /srv/ci, and whose absence blocks all CI
provisioning -- is built from debian:11, a base that can no longer
be built at all: bullseye-security's Release expired
2026-09-08.
Status: complete, and verified in production. Both stuck labels
are unstuck. ci-images/debian-gnome-13 first built on 2026-09-17
at 06:53 (529s, blob b9b624c8-713e-440b-ad0f-82b0e5f5ef60), and
the dependencies disk now builds on Debian 13 -- verified in
production rather than at merge, with the conductor deployed at
private-ci ad9863eb and the nightly creating its builder as
disks=['100@debian:13', '50'] on 2026-09-18, against
['100@debian:11', '50'] the night before. private-ci#39 is closed
with that evidence. private-ci#45 keeps its fourth checkbox for
phase 5. kerbside#450 merged on 2026-09-19, correcting the defect
kerbside#435 introduced, which was the last outstanding piece.
Three things this phase got wrong, recorded for phase 7.
The survey scoped the gnome move at seven places across two
repositories. It was about thirteen across four files, in four
repositories. GNOME_LABEL's comment, two docstrings, three tests,
and the marker paragraphs in private-ci's own AGENTS.md and
ARCHITECTURE.md all named the release and would have gone
factually wrong. More importantly, shakenfist/kerbside reads the
cached snapshot off the disk by hardcoded path and was never looked
at, because the survey grepped only the repositories it expected to
touch. A consumer you do not grep for is a consumer you do not have.
A merge is not a deploy, again. private-ci#61 merged on 2026-09-17 at 09:52 and that evening's nightly still built on Debian 11, because the conductor had not been redeployed -- it had even restarted in between, which is not the same thing. This is the same distinction that cost phase 2 a night of images, in a different component, and neither phase had a check that would have caught it.
Two ordering gates the plan set were both crossed. actions#78
merged ahead of its "must not merge until 3c has built" gate, and
private-ci#63 merged ahead of actions#80 despite the playbook-first
rule. Neither caused damage -- the first by luck, since ansible's
auto interpreter discovery handles bullseye unaided, and the
second because the conductor was not deployed in the window -- but a
gate stated only in a plan file and a pull request body is not a
gate. Phase 7 should ask what would actually have enforced them.
The gnome rename was reconsidered mid-phase, and the second answer
was better. The plan called for renaming the cached snapshot from
debian-12-gnome-agents to debian-13-gnome-agents and sequencing
three repositories around it. Review on actions#80 pointed out that
the disk is reformatted from scratch on every build, so a rename has
no transition window at all. The published name is now
debian-gnome-agents, carrying no release, with a transitional
hardlink at the old name; gnome_release governs only the label
lookup and the scratch filename. That removes the release from a
cross-repository interface entirely, so the next desktop bump is not
a fleet change. kerbside#435 was written against the abandoned
design and merged anyway, reading a path that is never published and
silently taking the legacy hardlink on every run; kerbside#450
corrected it to the published name, keeping the legacy fallback for
clusters whose disk predates the stable name.
What the survey found¶
Checked 2026-09-16 against private-ci 84f39cc and actions
edc4b73 -- that repository's main that day, pinned to a sha
rather than written as origin/main, because a moving basis is
what makes a line number unreconcilable a fortnight later. The
original sketch for this phase was written before phase 2 executed.
Most of it survived; one claim was materially incomplete and one
was a near miss.
Two bases, stated once so that every number below can be
placed. Every line number in this survey is as it stands on those
two commits. Every line number in the step briefs is as it stands
after step 3a, that is on private-ci 392700d, where 3a's
seven-line IMAGE_BUILDS entry shifted everything below
conductor/imagebuilder.py:123 down by seven. So the survey's
:143 and step 3b's :149 are the same entry read on different
bases and by different anchors: :143 is its base_image line
pre-3a, :149 its name line post-3a (:142 and :150 being the
other of each pair). conductor/imagebuilder.py is byte-identical
between 84f39cc and 3a's first parent, so the seven-line shift is
the only difference between the two bases. Nothing in this plan
pins actions after the survey, so re-derive anything quoted from
actions/ against its main before editing rather than trusting
the number here.
Confirmed as written.
debian-gnome-13really is absent fromIMAGE_BUILDS. Inprivate-ci,debian-gnome-12sits atconductor/imagebuilder.py:120withplaybook: ansible/ci-image-desktop.yml, exactly as private-ci#45 quotes it.- The dependencies entry really is
base_image: 'debian:11', atconductor/imagebuilder.py:143inprivate-ci. - The two conditions to reword really are in
actions, atansible/ci-dependencies.yml:50and:61, still at those exact line numbers onedc4b73after phase 2 edited that file. debian:13androcky:10really are absent from the cached image list inactions(theCache all minimal images we currently build to reduce network traffictask,ansible/ci-dependencies.yml:124-177).
Measured, because the sketch did not ask: the two new images
fit. Step 3d grows the cached set on a fixed disk and decision 3.4
leaves three frozen entries in place, so the growth is worth a
number rather than an assumption. The cache disk is the builder
instance's second disk, declared as - "50" at
ansible/ci-dependencies.yml:23-26 in actions: 50GB. The eleven
upstream images cached today total 8.9GiB by Content-Length on
images.shakenfist.com (measured 2026-09-17); the two additions
are debian:13 at 0.41GiB and rocky:10 at 0.96GiB, so 1.4GiB of
growth.
Comfortable -- but that bounds the growth, not the headroom. The
disk also carries the github-actions-runner tarball, the gnome
snapshot, and a depth-1 clone of imago-testdata which is
resparsified in place and whose size is not knowable from here. So
step 3c gathers a df reading from a live disk and step 3d stops
rather than adding the two images if the headroom is thin. This
paragraph is not the answer, only the part of it that could be
measured without a cluster.
And the playbook does not check what it caches. Worth stating
because it is tempting to assume it does: actions'
ansible/ci-dependencies.yml has exactly one cache-contents task,
List contents of /srv/ci/cached -- a bare ls -lrth at
:285-286 -- which prints and asserts nothing. get_url fails the
play on a 404 or a full disk, so a missing entry does break the
build, but nothing inspects the set afterwards, and no task
confirms each entry has content. So "the build passed" is weaker
evidence than a definition of done should
lean on, and this phase's says what is actually read instead.
Materially incomplete: the gnome snapshot is not one line. The
sketch treats "should the cache disk snapshot debian-gnome-13
instead" as a decision. It is a decision, but acting on it touches
two repositories: one module constant, plus seven further sites
that spell the label out literally and that no constant reaches.
The master plan did not say so. Counted as sites rather than as
lines, because several sites span more than one line and the count
is meant to be checkable:
- In
private-ci,conductor/imagebuilder.py:225--GNOME_LABEL = 'debian-gnome-12', the constant. It is read on five lines (:514,:519,:538,:871,:1213; six occurrences in the file counting the definition), which are the gnome-less marker (:507-538), the nightly scheduling (:871) and the operator log line (:1213). Those readers need no edit -- they follow the constant. - Four literal sites in that same file that the constant does not
reach, so editing
GNOME_LABELleaves them describing the old behaviour: theIMAGE_BUILDScomment block (:132,:134,:138), the gnome-less marker's own explanatory note (:220), and two docstrings (:533and:569). Step 3f rewrites all four; the definition of done greps for them, so reading them and leaving them will not pass. - Three literal sites in
actions, inansible/ci-dependencies.yml: the jq selector at:220(select(.source_url == "sf://label/ci-images/debian-gnome-12")), the skip message at:230, and the download path/tmp/debian-12-gnome-agentsat:252-253.
GNOME_LABEL and the playbook have to move together. The marker
exists so that a cluster which built dependencies before the gnome
label existed rebuilds it once the label appears; if the constant
watches debian-gnome-12 while the playbook snapshots
debian-gnome-13, that rebuild is triggered by the wrong label's
arrival.
A near miss worth recording so nobody else chases it.
In actions, ansible/ci-image-desktop.yml:151 hardcodes
label: "ci-images/debian-gnome-12", which looks like it would send
a debian-gnome-13 build to the 12 label. It does not: the
conductor passes label in extra_vars (private-ci
conductor/imagebuilder.py:791, 'label': 'ci-images/%s' %
image['label']), which overrides the play's default. The same is
true of the base_image: "debian:11" at actions'
ansible/ci-dependencies.yml:8. Both are stale defaults that only
bite somebody running the playbook by hand, and both should be
corrected while in the file rather than left as traps.
Two findings out of scope, recorded here rather than fixed.
- The cached image list carries
ubuntu:20.04,debian:11andfedora:40.shakenfist/imagesbuilds none of those any more -- its list isubuntu:22.04 ubuntu:24.04 centos:9-stream debian:12 debian-docker:12 debian-gnome:12 debian-xfce:12 debian:13 debian-docker:13 debian-gnome:13 debian-xfce:13 rocky:8 rocky:9 rocky:10. All three URLs still return 200, so CI is quietly caching three frozen artifacts, two of them end of life. That is inventory work for phase 4 and retirement for phase 5, and it belongs to private-ci#38's collated inventory. File it there rather than widening this phase. build.sh's reconciliation summary --Built:/Failed:/Not attempted:atbuild.sh:736-747inimages-- is printed to stdout only. Per-image logs ship to Loki; the summary does not. Under cron on a host with no MTA it is discarded, so the one output that distinguishes "never attempted" from "built fine" reaches nobody. images#6 built that reconciliation precisely to make that state visible. File againstshakenfist/images; it is a phase 1 detection gap rather than a phase 3 migration step.
Decisions¶
Numbered 3.N rather than 1-5, because this plan already has a
top-level Decisions section numbered D1 to D6 and the step briefs are
read one row at a time. "Decision 3.5" and "D5" are different
decisions and now look it.
3.1. Order is: label, then base image, then snapshot. Add
debian-gnome-13 first (step 3a), because step 3f cannot point
the snapshot at a label that does not exist. Move the
dependencies base image second (step 3b). Switch the gnome
snapshot last (step 3f). Each code step that depends on something
having been built is immediately preceded by its own observation
step: 3c gates 3d, and 3e gates 3f.
3.2. The cross-repository ordering constraint is real and is
stated. Deleting the debian:11 interpreter branch from
ci-dependencies.yml must land after the IMAGE_BUILDS base
image move has built successfully, not before. While
dependencies still builds on debian:11, removing that branch
sends bullseye down the auto-detect path -- which is the quirk
the branch exists for. Two pull requests in two repositories,
with a build in between, not one flag day.
3.3. Delete the debian:11 branch rather than reword it. The
master plan says to reword the two when: conditions to name the
bullseye interpreter quirk. Once the dependencies entry is on
debian:13, nothing invokes ci-dependencies.yml with
base_image: debian:11 at all -- it is the only entry that uses
that playbook -- so the condition is not obscure, it is dead. Two
add_host tasks collapse to one with no when:. This is a
deliberate departure from the master plan's wording, on the
grounds that a clearly-named condition for a case that cannot
occur is still a thing the next reader has to rule out.
3.4. Yes, switch the snapshot to debian-gnome-13. private-ci#45
leaves it open. The cache disk exists so CI does not pull from
the network, and a cache of the EOL desktop is a cache of the
thing phase 5 is about to delete. Switching now means one nightly
cycle in which the disk still carries the 12 snapshot, which is
harmless.
3.5. private-ci#45 is not closed by this phase. Its fourth
checkbox -- retire debian-gnome-12 once nothing consumes it --
is phase 5's work, and this phase deliberately leaves both
debian-gnome-12 and debian-11 building so nothing breaks
mid-plan. The issue keeps three of four boxes ticked and closes
in phase 5. This is the decision most likely to be argued
with: it leaves two end-of-life labels building for two more
phases, and eol-distro will keep reporting them the whole time.
The alternative -- retire as we go -- couples this phase to
finding every consumer, which is exactly what phase 4 is for.
Step plan¶
| Step | Effort | Model | Isolation | Brief for sub-agent |
|---|---|---|---|---|
| 3a | medium | sonnet | none | Landed 2026-09-17 as private-ci#60 (392700d). In shakenfist/private-ci, add a debian-gnome-13 entry to IMAGE_BUILDS in conductor/imagebuilder.py, immediately after the debian-gnome-12 entry. Copy its shape exactly: name and label both debian-gnome-13, base_image debian-gnome:13, base_image_user debian, playbook ansible/ci-image-desktop.yml. Do not touch GNOME_LABEL in this step. Update conductor/tests/test_imagebuilder.py -- the build-order and missing-label assertions both enumerate labels. Run tox (or the repo's test command) and confirm green. Commit subject: "Add a debian-gnome-13 CI image." |
| 3b | high | opus | worktree | In shakenfist/private-ci, change the dependencies entry in conductor/imagebuilder.py ('name': 'dependencies' at :149 once 3a has landed) from base_image: 'debian:11' to 'debian:13'. base_image_user stays debian. Read the comment block immediately above it (:133-148) before editing -- it explains the gnome-less marker and the first/last build ordering, and it names debian-gnome-12; leave that naming alone, step 3f moves it. Check whether any test in conductor/tests/test_imagebuilder.py asserts the dependencies base image, and run the suite and confirm green either way -- the check is not the verification. Two debian-11 things are out of scope and are not each other: the IMAGE_BUILDS entry at :58-63 is the image this repository builds and decision 3.5 keeps it until phase 5; the CI_IMAGES entry at conductor/provisioner.py:46-51 is the label runners boot from. Neither is the cache disk. High effort because the dependencies label gates all CI provisioning: if this build fails, nothing provisions. Commit subject: "Build the dependencies disk on Debian 13." |
| 3c | low | haiku | none | The gate for 3d, and the mitigation the risks section names. Observation step, no code change. After 3b merges, confirm the conductor has rebuilt the dependencies label (DEPENDENCIES_LABEL, conductor/imagebuilder.py:233) on debian:13 and that the build succeeded: sf-client --json artifact list, or the conductor's own log. Report the label's blob uuid and its creation time, and confirm that time is after 3b merged -- yesterday's blob still answering the lookup is exactly what this gate exists to catch. While there, take two more readings, both for the risks section rather than for 3d: run df -h /srv/ci/cached on a runner with the disk mounted and report free space, which 3d's item (1) checks against the 1.4GiB the two new images need; and report how long the nightly cycle now takes with twelve images, from the conductor's log, so that phase 1's 30-hour staleness threshold can be re-read against a measurement instead of against this plan's arithmetic. No commit. |
| 3d | medium | sonnet | none | In shakenfist/actions, edit ansible/ci-dependencies.yml. (1) Add debian:13 and rocky:10 to the cached image list in the Cache all minimal images we currently build to reduce network traffic task (:124-177), following the existing - { url: ..., name: ... } shape exactly -- but first check 3c's reported free space against the 1.4GiB the two images need (survey), and stop and say so if the headroom is under 5GiB rather than adding them anyway. (2) Delete the Add to ansible (force python3) task at :40-51 and remove the when: base_image != "debian:11" from the task at :52-62, so one unconditional add_host remains -- see decision 3.3. (3) Change the stale default at :8 from base_image: "debian:11" to "debian:13". (4) Change mkfs.ext4 /dev/vdc at :109 to mkfs.ext4 -O ^orphan_file /dev/vdc, with a comment above it saying only that this one feature is disabled, why, and what the floor is: the disk is mounted read-write by every runner, the oldest of which boots ubuntu2004-ci-template.qcow2 on kernel 5.4, and orphan_file needs 5.15. Do not write that the feature list is pinned -- it is not, this disables one bit, and a comment claiming more than the command does is the next reader's trap. See the back brief. This must not merge until step 3c has reported a successful dependencies build on debian:13 -- see decision 3.2. Verify with tools/ansible-syntax-check.sh, which phase 2 added. Two commits, not one. Items (1) to (3) are the migration and go under "Cache Debian 13 and Rocky 10, drop bullseye."; item (4) is its own commit, subject "Build the cache disk to the kernel 5.4 feature floor.", with the back brief's compat-versus-ro_compat measurement in the body. It has its own reasoning, its own definition-of-done bullet and a different blast radius -- it changes what every runner mounts, against a stated kernel floor -- and the risk story in this phase leans on a revert being one commit. One pull request is fine; one commit is not. |
| 3e | low | haiku | none | The gate for 3f. Observation step, no code change. Confirm ci-images/debian-gnome-13 exists and has a blob: sf-client --json artifact list filtered on sf://label/ci-images/debian-gnome-13, or the conductor's own log. A label that exists with no blob is the failure mode -- 3f points the playbook's jq selector at this label, and a selector that matches nothing skips the snapshot silently. Report the blob uuid. No commit. |
| 3f | high | opus | worktree | Move the dependencies disk's gnome snapshot from debian-gnome-12 to debian-gnome-13, across two repositories, as two pull requests. Merge the shakenfist/actions one first, then the shakenfist/private-ci one: the playbook snapshots whatever label it is told to look up and already handles the lookup failing (the snapshot is skipped with a warning), so a skipped snapshot for one night is recoverable, whereas a GNOME_LABEL watching a label the playbook never snapshots is a marker that never clears. In shakenfist/actions/ansible/ci-dependencies.yml: the jq selector at :220, the skip message at :230-232, the comment at :212-216, and the /tmp/debian-12-gnome-agents scratch path at :252-254 and :269. The destination at :270 is different: /srv/ci/cached/debian-12-gnome-agents is the name the snapshot has on the cache disk, so anything outside this repository that reads the disk reads that name. That grep set was too narrow and its result was wrong. A grep of shakenfist/shakenfist, shakenfist/private-ci and shakenfist/actions on 2026-09-17 found only the audit documentation, but two consumers sit outside those three. shakenfist/kerbside copies the disk by name in .github/workflows/functional-tests.yml:583-589 and boots it as a SPICE test target, which is a real functional dependency and not documentation. And the documentation hit is in this repository, at docs/audits/eol-distro.md:80 and :128 -- shakenfist/docs/components/development/audits/eol-distro.md:128 is the published mirror of it, so editing the path the earlier draft cited fixes nothing. Of the two lines here, :128 calls the disk a real dependency on an old release and becomes wrong on rename; :80 uses the name as an example of a string the audit's token matcher must not read as a runner reference, and an example naming a disk that no longer exists is a weaker example rather than a broken one. So grep shakenfist/images, shakenfist/kerbside, Mach33Labs/33fl and this repository as well before renaming, and if the name stays, say so in the commit message rather than leaving it looking forgotten. Also fix the stale play default at ansible/ci-image-desktop.yml:151. Run tools/ansible-syntax-check.sh. In shakenfist/private-ci (line numbers post-3a, on 392700d, per the survey's note on bases): GNOME_LABEL at conductor/imagebuilder.py:232, its five uses (:521, :526, :545, :878, :1220), and rewrite the four places that carry the literal without going through the constant -- the IMAGE_BUILDS comment block at :133-147, the gnome-less marker note at :226-230, and the two docstrings at :540 and :576 -- so that all of them describe the behaviour in terms of debian-gnome-13. The definition of done greps for debian-gnome-12, so re-reading these and leaving them alone will not make it pass. Grep the tests for GNOME_LABEL and for the literal debian-gnome-12 (conductor/tests/test_imagebuilder.py:58, :93) and run the suite green. High effort because the gnome-less marker decides when a cluster rebuilds its cache disk, and a constant that watches one label while the playbook snapshots another produces a rebuild that never fires. Commit subjects: "Snapshot the Debian 13 desktop image." in each repository. |
| 3g | low | haiku | none | Housekeeping. Tick checkboxes two and three on private-ci#45 and leave it open per decision 3.5, saying in a comment which phase closes it. Close private-ci#39 with the merge commits. Record the two out-of-scope findings from the survey: comment the frozen cache entries onto private-ci#38, which already exists and is not closed here, and file a new issue for the discarded reconciliation summary against shakenfist/images. No commit in this repository. |
| 3h | low | haiku | none | Confirms step 3f actually worked, which nothing else does. Observation step, no code change. Last in the table rather than adjacent to 3f because it has to wait for a nightly dependencies rebuild to complete after 3f merged. The survey establishes two facts that combine badly: the playbook's jq selector skips the snapshot with a warning when the lookup matches nothing, and the playbook asserts nothing about what it cached. So a grep for the absence of the old label passes whether or not the snapshot ever succeeded. On a dependencies disk built after 3f, confirm /srv/ci/cached/ carries the gnome snapshot at non-zero size, that it is the Debian 13 artifact and not a survivor of the old name, and that the playbook run did not log the label does not exist yet skip message. Report the blob uuid it was downloaded from so it can be matched against the one step 3e reported. If the snapshot was skipped, say so rather than filing it: the marker is designed to trigger one rebuild, so the next nightly may fix it, and knowing which happened is the point of this step. Two definition-of-done bullets have no other step that collects their evidence, so collect both here, on the same disk: run dumpe2fs -h on it and report whether orphan_file appears in its feature list, and report the /srv/ci/cached entries for debian:13 and rocky:10 with their sizes, read from the List contents of /srv/ci/cached output at ci-dependencies.yml:285-286 in that run. This step is the only one that looks at a disk built after 3d: 3c reads one built before it, and 3e reads a blob rather than a disk. No commit. |
Risks and mitigations¶
- The dependencies build fails on trixie and CI stops
provisioning. This is the real risk in the phase: a missing
dependencieslabel blocks everything. Mitigated by step 3b being its own pull request with nothing else in it, so a revert is one commit; and by_build_order()already buildingdependenciesfirst when its label is missing, so recovery is the next scan rather than a manual intervention. Checked by step 3c, which is where somebody watches the conductor build the label.
The one-commit revert holds only until 3d item (4) lands.
After that, reverting 3b on its own puts base_image back to
debian:11 while mkfs.ext4 -O ^orphan_file stays in the
playbook -- and the measurement recorded below is that bullseye's
mke2fs 1.46.2 exits 1 on that flag. The revert would fail the
play and leave no dependencies label at all, which is precisely
the outcome the mitigation exists to prevent. So the rollback is
one commit while 3b is the newest thing landed, and two
afterwards: revert the shakenfist/actions commit that added the
flag first, then 3b. Stated here because a mitigation that
fails closed is worse than none, and the moment somebody needs
this is the worst moment to work it out.
* 3d merges before 3b builds. Then bullseye takes the
auto-detect interpreter path and the dependencies build breaks for
the reason the deleted branch existed. Mitigated by decision 3.2
being stated in 3d's own brief rather than only here, and by step
3c being a gate in the table between the two. The earlier draft of
this plan claimed 3b -- the label observation, now 3e -- as that
gate; it was not one, it sat before the wrong step, and the
phase's one CI-stopping risk was resting on a sentence.
* The gnome snapshot switch half-lands. GNOME_LABEL in one
repository and the playbook in another cannot merge atomically.
Mitigated by ordering: merge the playbook first (it snapshots
whatever label it is told to look up, and the lookup failing is
already handled -- the snapshot is skipped with a warning), then
the constant. A skipped snapshot for one night is recoverable; a
marker watching a label nobody builds is not self-correcting. That
ordering is stated in 3f's own brief, not only here, for the same
reason decision 3.2's is: 3f is the step most likely to be executed
by a sub-agent reading one row.
* Step 3a lengthens the nightly cycle that phase 1's staleness
threshold was sized against. Phase 1 chose the conductor's
30-hour threshold from measured duration -- eleven images built
serially, 1h24m to 1h37m across six runs, with the gauge advancing
on completion. 3a makes it twelve, and a desktop image is not one
of the cheap ones. The arithmetic still holds comfortably:
completion lands about 25.6 hours after the primed slot, so 30
hours leaves roughly 4.4 hours of headroom and one image cannot
plausibly consume it. Recorded because the plan should say the
interaction was looked at rather than leave it to be rediscovered,
and because phase 5 later removes images while this phase adds
one. Step 3c's brief asks for the new cycle duration alongside
its df reading, so the arithmetic above gets checked against a
measurement; the request lives in the table row because a
sub-agent executing one row does not read this section.
* Every dependencies disk built between 3b and 3d carries
orphan_file. 3b moves the builder to trixie; 3d item (4) is
what turns the feature back off. Decision 3.2 forces 3d to wait for
3c, so the window is real: disks built in it are snapshotted and
mounted read-write by runners on kernels 5.4, 5.10 and 5.14, all
below the 5.15 floor the feature needs. It spans as many nightly
cycles as pass between the two merges, which is at least one and
should be one.
Item (4) cannot be moved before 3b to close the window, and the
attempt would be worse than the window. Measured 2026-09-18 in
debian:11 and debian:13 containers: bullseye's mke2fs
1.46.2 rejects the flag outright -- mkfs.ext4 -O ^orphan_file
exits 1 with Invalid filesystem option set: ^orphan_file -- and
plain mkfs.ext4 there sets no orphan_file anyway, because the
feature did not exist. Trixie's 1.47.2 sets it by default and
accepts ^orphan_file to remove it. The task runs mkfs.ext4
through shell:, so a non-zero exit fails the play and the
dependencies label -- which gates all CI provisioning -- does
not get built. So item (4) must land in the same pull request as
the trixie move or after it, never before. That is a stronger
constraint than decision 3.2 and is the reason the window cannot be
engineered away, rather than an oversight.
What makes the window acceptable is the back brief's
compat-versus-ro_compat measurement: orphan_file sets a compat
bit, so an older kernel mounts such a filesystem and ignores the
feature. The narrow failure mode is a snapshot taken after an
unclean unmount, where the orphan inode list is in a file the old
kernel will not read. One nightly cycle of that exposure, on a
disk that is rebuilt nightly anyway, is cheaper than a day with no
dependencies label.
* debian-gnome:13 turns out not to boot a desktop under CI even
though the guest image is correct. Phase 2 confirmed the
artifact is trixie with GNOME 48 and that it reaches the gdm3
greeter. ci-image-desktop.yml now proves this at build time
(actions#74), so this fails the build loudly rather than producing
a label that looks fine.
Definition of done¶
IMAGE_BUILDScontains adebian-gnome-13entry andci-images/debian-gnome-13exists as a label with a blob.grep -rn 'debian:11' conductor/imagebuilder.pyreturns only theIMAGE_BUILDSentry that builds thedebian-11image (:58-63before this phase edits the file), which decision 3.5 keeps until phase 5 -- and no dependencies entry. Thedebian-11entry inconductor/provisioner.py'sCI_IMAGES, which is the label runners boot from, is a different thing in a different file and this phase does not touch it.- In
shakenfist/actions,grep -n 'debian:11' ansible/ci-dependencies.ymlreturns only the two lines of the cached image list'sdebian:11entry (itsurland itsname), which the survey's out-of-scope finding deliberately leaves for private-ci#38 to collate. In particular it returns nowhen:clause and nobase_image:play default: the latter is step 3d item (3), the smallest of its four changes and the one that would otherwise have no gate at all. Note that this is a grep whose passing output is non-empty, which is deliberate -- a bullet asking for nothing would be unsatisfiable while those three frozen entries stay, for the same reason the olddebian-gnome-12bullet was.ansible-playbook --syntax-checkpasses viatools/ansible-syntax-check.sh. dumpe2fs -hon a dependencies disk built after this phase does not listorphan_fileamong its features. Collected by step 3h, which is the only step that reads a disk built after 3d.- The cached image list contains
debian:13androcky:10, and adependenciesdisk built after this phase has/srv/ci/cachedentries for both with non-zero size. The playbook does not assert this -- see the survey -- so the evidence is theList contents of /srv/ci/cachedoutput atci-dependencies.yml:285-286in a successful run, read rather than assumed, together with thedfstep 3c reported. Step 3h collects it; 3c cannot, because it runs before 3d adds the two images. - The old label survives only in the entry phase 5 retires. This
is two greps, one per repository, not one across both:
shakenfist/actionshas noconductor/directory, so a singlegrep -rn 'debian-gnome-12' conductor/ ansible/errors and exits non-zero there. - In
shakenfist/private-ci,grep -rn 'debian-gnome-12' conductor/imagebuilder.pyreturns only the lines of theIMAGE_BUILDSentry that phase 5 retires -- two of them, itsnameand itslabel, not one hit. Itsbase_imageis'debian-gnome:12'with a colon and so does not match this pattern at all. No constant, no comment block, no marker note and neither docstring. conductor/tests/is deliberately outside that grep. Decision 3.5 keeps the entry building until phase 5, so the assertions namingdebian-gnome-12there must stay; a gate that demanded they go would be satisfied most cheaply by deleting legitimate test coverage. Step 3f reads them and runs the suite green instead.- In
shakenfist/actions,grep -rn 'debian-gnome-12' ansible/returns nothing at all: no jq selector, no skip message, no comment, no play default. grep -rn 'debian-12-gnome' ansible/inshakenfist/actionsis a separate check, because the/tmpand cache-disk paths spell the label the other way round and can never appear in the grep above however it is scoped. It returns nothing if step 3f renamed the cache-disk destination, or only that destination if 3f decided to leave the name alone -- in which case 3f's commit message says so, per its brief, and this bullet is satisfied by that sentence rather than by an empty result.- Step 3h has reported, from a
dependenciesdisk built after 3f, that the gnome snapshot exists on the cache disk at non-zero size and is the Debian 13 artifact, and that the playbook's run did not log the skip message. The greps above pass whether or not the snapshot ever ran, so this is the bullet that distinguishes a switched label from a silently skipped one. - private-ci#39 is closed; private-ci#45 has boxes one, two and three ticked, box four open, and a comment naming phase 5 as its closer.
- The frozen-entry finding is recorded on private-ci#38, which
already existed before this phase and stays open, and a new issue
exists against
shakenfist/imagesfor the discarded reconciliation summary. Two findings, one new issue.
Back brief¶
Confirm before step 3b is written: that building the dependencies
cache disk on debian:13 is acceptable given the disk is snapshotted
and mounted by every runner, and that nothing consuming /srv/ci
depends on the disk's own filesystem being bullseye-era. The step is
cheap to propose and expensive to get wrong -- a broken dependencies
label stops all CI provisioning -- and the answer lives in what
mounts the disk rather than in what builds it.
Answered 2026-09-16: yes, with one addition that step 3d must carry.
The builder's operating system is not consumed by anything. The
dependencies label is a snapshot of the second disk alone:
ci-dependencies.yml snapshots with all: true, records
cisnapshot['meta']['vdc']['blob_uuid'], and then explicitly
deletes the vda snapshot. So base_image picks the throwaway
builder, not the artifact, and moving it to debian:13 changes
nothing a runner sees.
What it does change is the filesystem, and that is the coupling the
sketch missed. The disk is made by a bare mkfs.ext4 /dev/vdc
(ci-dependencies.yml:108-109), so it inherits whatever the
builder's distribution defaults to. Measured on trixie, e2fsprogs
1.47.2:
Filesystem features: has_journal ext_attr resize_inode dir_index
orphan_file filetype extent 64bit flex_bg metadata_csum_seed ...
orphan_file is new. man 5 ext4: "supported by Linux kernels
starting version 5.15, and by e2fsprogs starting with version
1.47.0." Bullseye's e2fsprogs did not set it; trixie's does, by
default, silently.
The consumers are older than that. Every runner attaches this disk
as its second disk (provisioner.py:1512), and the topology
playbooks mount /dev/vdc read-write with no options
(ci-topology-slim-primary.yml:308-313 and five siblings). One of
those still boots ubuntu2004-ci-template.qcow2 -- kernel 5.4. The
debian-11 boot label is 5.10, and rocky-9 is 5.14. All three
are below the floor.
Measured 2026-09-17, so the risk is bounded rather than unknown. An earlier draft of this answer left open whether the bit is compatible or incompatible. It is settleable, and was settled by building both filesystems on trixie (e2fsprogs 1.47.2) and reading the superblock feature words directly:
mkfs.ext4 -O orphan_file compat=0x0000103c ro_compat=0x0000046b
mkfs.ext4 -O ^orphan_file compat=0x0000003c ro_compat=0x0000046b
The only difference is bit 0x1000 of s_feature_compat, which is
EXT4_FEATURE_COMPAT_ORPHAN_FILE. It is a compat feature: a
kernel that does not know it mounts the filesystem normally,
read-write, and ignores it. mke2fs also rejects -O
orphan_present as an invalid option, which confirms the paired
orphan_present bit is set by the kernel at runtime rather than at
format time; that one is ro_compat, and an unknown ro_compat
bit forces a read-only mount.
So the failure mode is narrow and specific: a pre-5.15 kernel is
refused a read-write mount only if the disk was snapshotted while
the orphan file still held entries. The playbook unmounts the disk
cleanly (ci-dependencies.yml:288-291) before snapshotting, and
every runner mounts a fresh copy of that snapshot, so in the normal
path orphan_present is clear and nothing breaks. The exposure is
a snapshot taken after an unclean unmount -- which is a real state
to be in, and one nobody would connect to a read-only /srv/ci on
a 20.04 runner six months later.
That bounds the risk; it does not remove it, and it does not change the action. A shared cache disk should be built to a stated compatibility floor rather than to whatever the builder's distribution defaults to, because otherwise every future bump of the builder OS re-rolls this dice in silence.
So step 3d also changes mkfs.ext4 /dev/vdc to mkfs.ext4 -O
^orphan_file /dev/vdc and says in a comment what the floor is and
who sets it. Note what that comment must not say: one flag is not
a pinned feature set, and the next e2fsprogs default that turns on
a new bit rolls the same dice again. The comment records that this
one feature is disabled for the 5.4 floor, which is true; pinning
the whole set with an explicit -O list would be the stronger
change, and is deliberately not taken here because a list written
today goes stale silently in the other direction, dropping features
the disk would benefit from. If a second feature ever has to be
disabled, that is the point to reconsider.
4. The consumer sweep¶
Closes: development#123. Contributes to private-ci#38, which closes at the end of phase 5. Depends on: phase 3, satisfied. Planning effort: high, because the inventory the master plan carried is a fortnight stale in both directions, and because the criterion that produced it cannot see the half of the migration that actually gates phase 5.
Status: complete, 2026-10-02. 4j ran on 2026-10-02 and all six of its checks pass; what they found is recorded under What 4j confirmed at the end of this section. One step landed out of band, one marker the phase was going to add turned out to be already there and broken, and a bug in another repository blocked the under-cloud half for three days. All of that is recorded at the end of the survey -- the first two and the block under Amended 2026-09-21, the block's removal under Amended 2026-09-24 -- and the affected steps carry the correction inline. All thirteen label pull requests have resolved -- ten in the first sweep, occystrap#143 and shakenfist#4306 on 2026-09-24, and visual-digest-rust#23 closed as superseded by a change that moved the label anyway -- and 4f merged as actions#97 the same day. 4g merged as shakenfist#4379 on 2026-09-29 (22:03 UTC). Running 4g found a defect in 4h's own safety gate: the gate grepped the artifact's URL spelling, and seven live consumers name it by its bare short name, so it would have returned clean while the step broke them. 4h then moved those consumers in shakenfist#4385 before removing the upload in actions#124, and its brief gives both gates as commands, amended 2026-09-30. 4f's post-merge verification has now passed on both runs, which discharges 4g's hold: see the 2026-09-29 amendment below. The 2026-09-26 amendment recording it as failing is left as the snapshot it was. shakenfist#4309 unblocked the under-cloud half on 2026-09-23 -- the under-cloud rather than the guest-image half, which is a distinction this phase keeps deliberately and which the risks section spells out. Amendment dates in this phase are local time, AEST (UTC+10); timestamps quoted from issues, merges and CI runs are UTC and say so, which is why an amendment can be dated a day after the UTC day its commit landed on.
private-ci#38 is the collated inventory of where the fleet uses obsolete base images. It is a reference rather than a task, and it closes when the thing it inventories is gone: this phase clears the consumer half and phase 5 clears the producer half, so #38 closes at the end of phase 5 rather than when its own checklist is ticked.
Two things make this less mechanical than a reference count suggests:
- Each runner-label move is two files. A label change also needs
the replacement declared in that repository's
.github/actionlint.yamlunderself-hosted-runner: labels:, in the same commit. actionlint fails a workflow naming an undeclared label, so missing this turns a one-line fix into a failing lint. - The runner labels are the visible half. The guest images the
CI clusters boot, and the artifact name a smoke cluster uploads
into itself, are also Debian 12, and the
eol-distrocriterion reads neither. That half is the half phase 5 trips over.
What the survey found¶
Checked 2026-09-19 against each repository's main as fetched that
day, by shallow-cloning all twenty-nine non-archived repositories
in the shakenfist organisation and running this plan's own
criterion over every one of them, including the eight the daily
audit does not cover. The scan is reproducible from a checkout of this
repository:
import sys
sys.path.insert(0, 'scripts')
from audit.checks import distros
from audit.repo import Repo
found = distros.scan(Repo(clone_path, name, 'shakenfist'))
It agrees line for line with docs/audits/compliance.md as
regenerated at 2026-09-19T10:32, which is the artifact to re-read
rather than any number written here.
The count has almost halved, and not because of this plan.
development#123 measured 80 references in 15 repositories on
2026-09-12. Today it is 45 in 13. Thirty-eight references were
cleared by ordinary repository work answering the daily audit's own
issues:
development's four cleared when #119 merged, exactly as #123
predicted; instar moved its whole pool on 2026-09-16 (dbc1cc0,
"Move CI onto the Debian 13 runner pool"), taking 20 of its 21; and
kerbside moved on 2026-09-17 (45af922, "Move CI off Debian 12,
which is end of life"), taking all 14. The count moved the other way
too -- occystrap went from 3 to 5 and shakenfist from 4 to 5 --
because new workflows land naming the label the fleet still
advertises. So the inventory is the compliance page, re-read at the
start of each step, and never a number in this plan.
Every remaining label reference is a literal runs-on: line.
All 44 of them; there is no matrix indirection, no
workflow_dispatch input default and no reusable-workflow input to
chase, which is what makes the sweep a token edit per line rather
than a reading exercise. The forty-fifth finding is an image, not a
label, and is decision 4.3.
| Repository | debian-12 |
debian-12-docker |
The whole actionlint.yaml edit |
|---|---|---|---|
| ryll | 5 | 12 | add debian-13-docker, delete both retired |
| occystrap | 4 | 1 | add debian-13-docker, delete both retired |
| shakenfist | 4 | 1 | add both, delete both retired |
| client-python-k3s | 4 | - | delete debian-12 |
| actions | 3 | - | delete debian-12 and a stray debian-12-docker |
| agent-python | 2 | - | add debian-13, delete debian-12 |
| clingwrap | 2 | - | add debian-13, delete debian-12 |
| sfui | 2 | - | delete debian-12 |
| client-python | 1 | - | delete debian-12 |
| divergulent | 1 | - | delete debian-12 |
| library-utilities | 1 | - | delete debian-12 |
| visual-digest-rust | 1 | - | add debian-13, delete debian-12 |
| instar | - | - | none, and that is checked rather than assumed |
The fourth column is the entire self-hosted-runner: labels: edit for
that repository, read on 2026-09-19: what the workflows will need
declared once they have moved, and what stays declared with nothing
left to use it. actions is the asymmetric one -- it declares
debian-12-docker and no workflow in it names that label, so the
declaration already outlives its last user and a step that deletes
only debian-12 leaves it behind. instar declares debian-13,
debian-13-docker and neither retired label, which is why step 4d
edits one line; step 4j greps all thirteen rather than trusting this
row, because a declaration leaves no finding and so nothing else in
the phase would notice.
Both replacements are verified in production rather than
declared. Last night's nightly published ci-images/debian-13
(blob b70b037f-05e2-47d4-8512-7c661adfc917) and
ci-images/debian-13-docker (7b708330-4489-4181-b703-c93bb2fd6ff8),
and runners of both labels served real jobs on 2026-09-19 -- a
debian-13-docker worker came online and busy at 19:08 for a
Mermaid lint, and debian-13 jobs at xl and s were scheduled
alongside it. /srv/ci/debian:13 is on the dependencies cache disk,
which phase 3 put there (actions
ansible/ci-dependencies.yml:194). Eleven of the twelve
IMAGE_BUILDS entries published a label last night; the one that did
not is debian-11, whose base can no longer be built, which is phase
5's retirement and phase 1's permanent False seen from the other
side.
The half the criterion cannot see, and it is the half that gates
phase 5. The eol-distro specification says so itself -- guest
images and cached disks are "the pre-push reviewer's to raise" --
but nobody has raised them, and phase 5 removes debian-12 from
IMAGE_BUILDS and CI_IMAGES. Every site below then names a label
the conductor no longer builds:
- The guest image label,
sf://label/ci-images/debian-12, at eight live sites:actionsbuild-smoke-cluster/action.yml:29(the composite action's default) and.github/workflows/smoke-cluster.yml:56(the reusable workflow's own input default, which feeds it), andshakenfist.github/workflows/functional-tests.yml:443,:463,:477,:515and.github/workflows/scheduled-tests.yml:42,:52. Two callers pass nobase_imageat all and so take the default today:shakenfist.github/workflows/functional-tests.yml:558andkerbside.github/workflows/sf-e2e-functional.yml:104. - The cached image and the artifact name it is uploaded under, at
actionsbuild-smoke-cluster/action.yml:249andshakenfist.github/workflows/functional-tests.yml:583. Both sites have the same shape and neither is greppable by its verb:sf-client artifact uploadis assembled into a${setup}shell variable a few lines above, and the line that carries the artifact name and the source path reads"${setup} debian-12 /srv/ci/debian:12 --shared --no-checksum". Grep/srv/ci/debian:12to find them; a grep forartifact upload debian-12finds neither. - That artifact name, read back by the test suite, as
sf://upload/system/debian-12: 43 references across 12 files inshakenfist, indeploy/shakenfist_ci/andtests/anddeploy/nodelifecycletests.sh:132. There is a constant for it --CLUSTER_CI_IMAGEatdeploy/shakenfist_ci/base.py:37-- and it reaches one of the 43. This isGNOME_LABELagain: a constant that names the thing, and a couple of dozen literals the constant does not reach. (Noted 2026-09-26: these paths omit the top-levelshakenfist/package directory the source tree nestsdeploy/andtests/under, so they never resolved as written; 4g's item (3) carries the resolving form. The line numbers here stay as the dated snapshot they were.) - The job names, read by a tool.
shakenfist's matrix calls its lanes "Debian 12 cluster" and "Debian 12 tier" (functional-tests.yml:441,:475), andtools/ci_headroom_harvest.py:134and:141keyBUNDLE_TOPOLOGIESoff those exact strings, with the file's own comment warning that the derivation "would break silently if that changed". Renaming the lane without the tool is a silent stop, not a failure.
private-ci's own unit tests match ci-images/debian-12 eight
times in conductor/tests/test_imagebuilder.py -- seven naming the
label and one the -docker variant, which the substring every grep
in this phase uses also returns; those follow the
IMAGE_BUILDS entry in phase 5 rather than moving here.
One claim in the master plan has already been overtaken.
123 recorded two findings that are not label swaps. The first still¶
holds: instar boots debian:12 in a functional-test matrix that
deliberately covers several distributions. The second does not --
kerbside's two bookworm-tagged rust images are already on trixie
(rust/kerbside-proxy/Dockerfile:17 is rust:slim-trixie and
loadtests/latency/Dockerfile:8 is rust:1.97-trixie), and that
repository has been compliant since 2026-09-17.
Eight active repositories are outside the audit's matrix, and
one of them is shakenfist/images, which builds the guest images
this whole plan is about. The others are client-python-ova,
divergulent-reviews, homebrew-tap, performance,
reproducables, sonobouy and uefi-latency-guest. All eight were
scanned by hand for this survey and none has a finding today, so
nothing is being missed right now -- but nothing is watching them
either, and "we grepped it once" is the state phase 3 said was not
good enough. Widening the matrix is phase 6's work, and decision 4.6
says why it is not done here.
Amended 2026-09-21, after this plan merged, and anchored 2026-09-22. Three things moved in the day after #148 landed. The survey above is left exactly as it was read on 2026-09-19, because it is the record of what the phase was planned against; the corrections are here, and the steps they affect carry them inline as well. Every date in this block and in the steps it amends is the day the thing was read, which is why some of them are later than this heading: the 2026-09-21 pass found the three corrections below, and a second pass on 2026-09-22 re-read the counts, the line numbers and the greps before any of it was relied on.
shakenfist/actions migrated itself, twenty-five minutes after
this plan merged. e2a56bd ("Move off Debian 12 runners and guest
images.", 2026-09-20 02:08) answers actions#69 -- the daily audit's
own issue, not this phase. It moved all three runs-on: lines to
debian-13 and deleted both retired declarations from
.github/actionlint.yaml, including the stray debian-12-docker the
survey flagged: step 4b's third repository, done. The fleet is now 42
references in 12 repositories, and actions is off the compliance
page. Step 4b covers two repositories now. This is the fourth time
this plan has watched the daily audit clear work it had scheduled --
after development's own four via #119, instar and kerbside --
which is the argument for 4j grepping the fleet rather than reading
the table above.
instar already carries its marker, and it has never worked. The
survey recorded instar's image: 'debian:12' as needing decision
4.3's exception, and step 4d as written says to add one. One is
already there, with a better reason than 4d proposed, at
.github/workflows/functional-tests.yml:588 -- nine lines above the
finding at :597, written against the whole matrix entry.
is_excepted() reads the finding's own line and exactly one line
above it (scripts/audit/checks/distros.py:275-279), so it has never
applied, which is why instar is still on the compliance page having
moved its runner pool on 2026-09-16. That file last changed
2026-09-18, so this was true when the survey ran: the survey read the
scan's output, which reports the finding, and did not read the lines
around it. A scan that filters exceptions silently cannot tell a
missing marker from a misplaced one, and neither could the survey.
Step 4d moves the marker rather than adding a second.
A trixie under-cloud breaks every instance's agent, which blocks
half this phase. (Superseded 2026-09-24 -- the root cause was qemu
SPICE packaging, not the agent, and shakenfist#4309 closed it; the
canary evidence below still stands, the diagnosis in this heading does
not, and the 1820-second figures were inflated by the
is_powered_on() bug the amendment describes.) shakenfist#4280, filed
2026-09-20 03:53 out of the canary actions ran for its own
migration. Instances booted on a Debian 13 hypervisor never reach
agent ready -- agent_state: "not ready (no contact)",
agent_start_time: null -- and the three test_agentop_deadlines
tests went from 137, 192 and 207 seconds to roughly 1820 each,
consuming the step's whole 45-minute budget so that the rest of the
suite never ran. Two canary runs six hours apart on the same topology,
35471083618 green and 35483225701 red, with the under-cloud image
the only difference between them. actions reverted its own default
in 2e0d32a ("Keep the under-cloud on bookworm.") and wrote the
evidence into the input's comment at
.github/workflows/smoke-cluster.yml:55-65. This plan does not fix
it: D5's reasoning applies, and a plan that closes a migration is not
the place to debug a hypervisor's side channel.
The blocker splits the guest-image half in two, and only one half is stuck. (Superseded 2026-09-24 -- shakenfist#4309 closed the block; the under-cloud/guest separation below survives and is load-bearing, the present-tense claims about being stuck do not. See the amendment under Amended 2026-09-24.) The survey treated "the guest images the CI clusters boot" as one thing. #4280's evidence separates them, and that separation is the useful part of it:
- The under-cloud image --
base_image, what the hypervisor VMs themselves boot -- is blocked. That is 4f's two defaults (build-smoke-cluster/action.yml:32and.github/workflows/smoke-cluster.yml:67, both renumbered since the survey) and 4g's item (1), the sixbase_imagesites inshakenfist. None of them moves until #4280 closes. - The uploaded guest artifact --
sf://upload/system/debian-12, what instances inside the nested cluster boot -- is not known to be blocked. #4280 says so explicitly, and the shape of its evidence is what makes the separation usable: both canary runs uploaded the same bookworm artifact, so the guest was the controlled variable rather than the suspect. Read what that does and does not settle. It rules the guest out as the cause of #4280; it says nothing about whether a trixie guest works, because no canary ran one. So what is unblocked is the rename -- 4f's release-neutral upload, 4g's items (2) to (4), and 4h -- and not, on this evidence, the content change 4f currently carries with it by sourcing the new name from/srv/ci/debian:13. Whether those two should land together is back brief question 4, which is where it is decided rather than here.
Decision 4.5 is what makes that separation survivable.
sf://upload/system/debian asserts no release, so the references
can be renamed while #4280 is open and the artifact's contents can
follow later without a second fleet-wide edit. The decision was
argued for the next desktop bump; it earns its keep sooner than that.
Amended 2026-09-24: shakenfist#4280 is fixed, and the deferral above is withdrawn. shakenfist#4309 ("Start instances on Debian 13 hypervisors.") merged 2026-09-23 19:28 UTC and closed it. Read the root cause before re-reading the deferral, because it is not what the 09-21 amendment assumed and that decides how much of the amendment survives:
- The bug was in provisioning, not in the agent. Debian 13, like
Ubuntu 24.04, ships qemu's SPICE support in a separate
qemu-system-modules-spicepackage.roles/node/tasks/bootstrap.ymlinstalled it only on Ubuntu 24.04, so on a trixie hypervisor libvirt refused every domain definition --spice graphics are not supported with this QEMU, 72 times in the libvirtd journal of run35483225701, which is the same red canary the amendment above cites. No instance ever started, so no agent was ever in a position to make contact. The fix installs the package wherever apt has it rather than enumerating releases, which covers the next one too. - A truthiness bug is what hid it.
Instance.is_powered_on()returned the string'off'when libvirt had no domain, and a non-empty string is true, socreate()marked instances whose every power-on attempt had failed ascreated. That is why the suite waited out the full agent timeout instead of erroring, and it is the 1820-secondtest_agentop_deadlinesfigure the amendment recorded as evidence about the agent. #4309 returnsFalsefor a missing domain and adds five tests, two of which fail when the old return is restored. The further power-state defects that audit found are shakenfist#4307 and are not this plan's problem.
What this restores. The under-cloud half moves in this phase as
originally planned: 4f's two defaults and 4g's item (1). The eight
sites are listed here, re-read on 2026-09-24 against each
repository's default branch, and this list rather than the
definition of done is the record of them --
actions build-smoke-cluster/action.yml:32 and
.github/workflows/smoke-cluster.yml:67, and shakenfist
functional-tests.yml:443, :463, :479, :517 and
scheduled-tests.yml:42, :52. None of them acquired the comment
the deferral would have required, because no step ran; what exists
is the two comments 2e0d32a wrote in actions, which 4f now
deletes with the default they explain rather than levelling up.
Where the phase had got to when this landed. 4a to 4e opened
thirteen pull requests across thirteen repositories. Ten had
merged by 2026-09-24 01:00: agent-python#140, clingwrap#136,
library-utilities#60, client-python#405, divergulent#117, sfui#35,
client-python-k3s#67, ryll#397, instar#589 and
kerbside-patches#1734. Two were still in flight with no failing
check: occystrap#143 and shakenfist#4306. The thirteenth,
visual-digest-rust#23, did not merge and will not: that
repository's CI had never run since its default branch was
renamed, because ci.yml still triggered on main
(visual-digest-rust#24), and its own #25 then fixed the trigger
and moved the runner label in one change, so #23 was closed as
superseded. Read that as the label half being done there rather
than as a step being skipped -- the label moved, by a different
pull request than this plan named, which is exactly the kind of
claim 4j's declaration grep exists to check rather than take on
trust. 4f does not wait for the two in flight: it is in actions
and neither of them touches that repository. 4g does wait for
shakenfist#4306, because both edit functional-tests.yml; the
wait is for review cleanliness rather than for correctness, since
4e changes runner-label tokens in place and renumbers nothing, so
item (1)'s six line numbers survive it either way.
Amended 2026-09-26. All three of those have since landed:
occystrap#143 at 2026-09-24 19:36 UTC, shakenfist#4306 at 2026-09-24
23:23 UTC, and 4f itself as actions#97 at 2026-09-24 19:35 UTC (merge
8c02ab0e7). The label half of this phase is therefore complete in
every repository and 4g's wait for shakenfist#4306 is discharged --
the only wait that sentence means; the next paragraph closes a
different one. The paragraph above is left as the snapshot it was,
because the survey's dated claims are read as evidence of when a thing
was true rather than as current state.
4f merged and its verification has not passed. 4g does not start
until it does. 4f's brief makes two post-merge runs its own
responsibility and says a failure in either is a revert of the
default-move commit. One of the two has run and failed:
kerbside's nightly sf-e2e-functional on develop, run
36110697157 on 2026-09-25 08:01 UTC, the first scheduled run after
4f merged. The two before it, on 09-23 and 09-24, both passed. The
other run, shakenfist's node-lifecycle job on develop, has not
been triggered at all, so it is unknown rather than green.
What the run log establishes, read rather than inferred: the
cluster built and deployed normally, and it took 4f's new default
-- the build-smoke-cluster step shows base_image:
sf://label/ci-images/debian-13 and both uploads, the old
debian-12 name and the new debian one. The failure is later,
in import-instance.sh: a guest booted inside the nested cluster
reached state=error with power_state: off, error_message:
null and video: {model: cirrus, vdi: spice}.
What it does not establish is the cause, and the distinction matters because this looks like shakenfist#4280 and is not it.
4309's fix is in the deployed collection -- version¶
0.8.0-rc5.dev1385+g4404669e6, and de5e3833c is an ancestor of
it -- and the task it added ran and did its work: Install SPICE
modules where they are packaged separately reports changed:
[primary]. So the missing package is not the explanation this
time. The other half of #4309 is a plausible reason this is
visible now: before it, is_powered_on() returned the truthy
string 'off' for a missing domain, so an instance that failed
this way was marked created and the caller waited out a timeout
instead of erroring. A failure that used to present as a hang now
presents as state=error, which is an improvement in reporting
and not evidence of a new fault.
The run's own artifacts cannot take this further, which is its
own finding. kerbside's tools/sf-e2e/gather-artifacts.sh
collects sf-api and sf-console journals only. The domain
definition happens in the node daemon and in libvirtd, and neither
journal is collected, so the error #4280's evidence turned on --
libvirt refusing a domain -- cannot be read out of a failed run at
all. sf-api.journal carries no ERROR-level line for the
instance; what it shows is a scheduler admitting it only after
waiving the demand guard, which is capacity pressure and is
kerbside#284's subject rather than this phase's.
So the state of 4f is: landed, fleet-wide for every consumer
pinning @main, with one of its two verification runs failing for
a reason not yet attributed and the other not yet run. That is a
decision for Michael -- revert 4f's default-move commit, or
diagnose first -- and not one an implementer picks up from this
table. kerbside#482 was auto-filed for the failing nightly on
2026-09-25 and is where the diagnosis belongs.
What it does not restore. The 09-21 amendment separated the under-cloud image from the uploaded guest artifact, and that separation stays: they are two images with two consumers, and 4f is still additive for the artifact because a rename and a content change are still different things. What expires is only the claim that the under-cloud cannot move. Back brief questions 4 and 5 were both premised on #4280 being open and are answered there.
Four numbers drifted while the phase waited, all in
shakenfist. Three are in functional-tests.yml and all moved by
the same two lines: the matrix lane in 4g's item (2) was :475 and is
:477; the node-lifecycle caller in 4f's brief and in the definition
of done was :558 and is :560; the upload in 4g's item (4) was
:583 and is :585. The fourth is deploy/nodelifecycletests.sh,
which moved further and twice: 7e9c1cd45 took the upload reference
from :132 to :187 on 2026-09-22, and f0bf8c67a took it from
:187 to :217 on 2026-09-24 at 20:21 UTC. Every content at those
sites is unchanged. The first version of this paragraph named only the
lane while the same commit silently moved the upload, and left the
node-lifecycle number stale while claiming it had been re-read --
which is the defect this paragraph exists to document, committed
inside the document that documents it. The :558 error is worth its
own sentence because of how it happened: the job's step name is at
:558 today and its uses: line at :560, so a re-read that stops
at the first plausible line confirms the old number instead of
checking it. The nodelifecycletests.sh number is worth its own
sentence for the opposite reason: :187 was correct when this
branch's second commit wrote it at 19:42 UTC on 2026-09-24, and
f0bf8c67a invalidated it thirty-nine minutes later. A re-read is not
a fix when the file it reads is under active development on a
timescale shorter than a review round -- which is why item (3) now
carries the grep that finds the line rather than the line, and why no
step brief in this plan carries a line number for that file -- the
survey's dated snapshot still records :132, as a snapshot rather
than as an address. This is the fourth to seventh time a line number
in this phase has moved under ordinary work in that repository. The
survey's own numbers are left as it recorded them on 2026-09-19 --
:477, :515, :558, :583 and nodelifecycletests.sh:132 --
because that section is a dated snapshot and renumbering half of it
would make it disagree with itself.
Amended 2026-09-29: 4f's verification has passed on both runs, and 4g's hold is discharged. Neither run needed a revert, and the failure the 09-26 amendment could not attribute has not recurred.
kerbside's nightly sf-e2e-functional on develop has passed three
consecutive times since the failure that amendment records -- runs
36228146379 (2026-09-26 07:53 UTC), 36306108473 (09-27 08:24) and
36399738723 (09-28 08:50), against the failure at 36110697157 (09-25
08:01). Three passes over a trixie under-cloud with no change to the
code between them is what makes the 09-25 run transient rather than a
regression from 4f's default move. It is not a diagnosis, and
kerbside#482 stays open for one; what it settles is the decision the
09-26 amendment parked with Michael, because there is now nothing for
a revert to fix.
shakenfist's node-lifecycle job on develop -- the run that
amendment records as never triggered -- has since run twice and
passed twice: merge_group runs 36501556771 (2026-09-29 00:27 UTC) and
36513297089 (02:41), each a full six-host cluster build of about
fifty-eight minutes. It did not need the manual dispatch that
amendment anticipated, because ordinary merge traffic supplied it.
That it ran green at all is the second half of this amendment, and
it took a fix. Between 4f merging and those runs, every
merge_group run of that job failed, and the plan did not know
because of where the job is gated: node_lifecycle_collection runs on
merge_group and workflow_dispatch only, and skips on
code_changed != 'false', so pull request CI structurally cannot
reach it and a docs-only merge skips it. Green pull requests and green
docs merges concealed a failure on every code merge for four days.
The cause was in actions, not in this repository, and it was a
design mismatch rather than a regression. The mesh-interface task in
the three multi-node topology playbooks inferred a netplan renderer
from ansible_distribution_version | int > 11, but the Debian 13 CI
image is deliberately systemd-networkd: trixie dropped ifupdown
from the default install, cloud-init detects that and writes to
/etc/systemd/network/, and debian-13-extras installs no netplan
for the playbook to call. Two factors had to coincide to break a job,
which is why exactly one broke -- taking 4f's unpinned default and
configuring a mesh. The other two unpinned callers, kerbside's
sf-e2e-functional and this repository's canary, build localhost
topologies with no mesh; the four pinned matrix lanes configure a mesh
but boot bookworm or Ubuntu.
actions#112 replaced the inference with detection -- probe
/usr/sbin/netplan, set a mesh_style fact, and carry an ifupdown,
a netplan and a networkd arm -- and added an ungated assertion
that the mesh address is actually up before the play proceeds, which
is the post-condition whose absence let a missing mesh address present
as a healthy play and a cluster unable to reach its database. It
merged at 2026-09-29 00:03 UTC as 6b57319f1. Run 36513297089's log
is the evidence that it works rather than merely passes: the detection
returns ok on all six hosts, the ifupdown and netplan arms skip
on all six, the systemd-networkd arm reports changed on all six,
and the assertion returns ok on its first attempt.
The blind spot this exposed outlives the bug and is phase 6's
business rather than this phase's, so it is recorded and not fixed
here: no workflow in shakenfist/actions builds a slim-primary or
slim-tier cluster, so that repository's own CI cannot exercise a
mesh at all, while every consumer pins @main and takes each merge
live fleet-wide immediately. A green pull request there is not
evidence about the multi-node path, and this is the second time in
this phase that the verification a change needed lived in a different
repository from the change.
Decisions¶
Numbered 4.N for the same reason phase 3's are numbered 3.N:
this plan has a top-level D1 to D6, and "decision 4.2" and "D2"
should not be confusable.
4.1. One pull request per repository, grouped into steps by
size. Each carries its own actionlint.yaml edit and is
reviewed against the workflows it touches, which is what the
master plan asked for. The steps group repositories only so that
one sub-agent can carry several trivial ones; the pull requests
stay separate, because the consistency issue they close is
per repository and so is the CI that proves the move worked.
4.2. The retired label leaves actionlint.yaml in the same
commit as the last workflow line that names it. That list is
the set of labels a workflow may name, so leaving debian-12
declared after the last user is gone lets the next workflow name
a retired label and pass lint -- which is exactly how
occystrap and shakenfist grew new findings this fortnight.
This repository already took that decision for itself and wrote
the reasoning into .github/actionlint.yaml:15-20: "The bookworm
labels are deliberately absent ... Leaving them undeclared is what
gives actionlint something to say about it". So this is the fleet
applying a convention it has already adopted at the source of the
templates, rather than a judgment call being made fresh.
The cost is that a revert needs the declaration back, and a
reviewer may reasonably prefer to keep the declarations until
phase 5 retires the labels themselves. Taken anyway: the lint is
the only thing standing between a fleet-wide convention and the
next copy-pasted workflow, and a revert that needs two lines is
not a hard revert.
4.3. instar is marked, not migrated. Its one remaining
finding is image: 'debian:12' at
.github/workflows/functional-tests.yml:597, test input in a
matrix that deliberately covers several releases. The criterion
describes this case and provides the marker for it; use it, with
the reason on the line.
4.4. The guest-image half is in scope for this phase. It is not
what the master plan's section described, and it roughly doubles
the phase. It is in anyway, because D4 retires producers only
after their consumers, and phase 5 removes debian-12 from
IMAGE_BUILDS and CI_IMAGES: every site listed in the survey
above would then name a label the conductor does not build. The
alternative -- a phase 4a for the invisible half -- was rejected
because it separates two halves of one repository's migration
into two plans, and shakenfist has both.
4.5. The uploaded artifact gets a release-neutral name. The
smoke cluster uploads the cached image into itself as
debian-12 and 43 test references read it back by that name. The
obvious move is debian-13, and the obvious move buys another
43-reference edit at the next release. Phase 3 reached the same
fork with the gnome snapshot and took the neutral name
(debian-gnome-agents), which is why the next desktop bump is
not a fleet change; take it again here. The artifact becomes
sf://upload/system/debian, the release survives only in the
path the action copies from (/srv/ci/debian:13), and the tests
stop naming a release they do not care about. This is the
decision a reviewer is most likely to argue with, because a test
that says debian no longer says which Debian it exercised --
the answer is that it never did: the name said 12 while the
bytes were whatever the cache disk last cached, which for
debian-gnome:12 was Debian 11 for two years (private-ci#38).
4.6. Repositories outside the audit matrix are surveyed here and watched in phase 6. The survey is above; widening the matrix changes a workflow every repository's compliance depends on, and phase 6 is the phase that owns the audit's blind spots.
4.7. The frozen cached-image list is not touched. ubuntu:20.04,
debian:11 and fedora:40 stay in ci-dependencies.yml's cache
list, and debian:12 stays there and becomes frozen alongside
them rather than being removed: the cache is what lets a test boot
an old guest deliberately. No step edits that file.
Retirement is phase 5's and private-ci#38's.
Step plan¶
Every step that edits a repository other than this one opens a pull
request there and waits for that repository's own CI. Steps 4a to
4e are independent of each other and of 4f to 4h; within 4f to 4h
the order is a real constraint and is stated in each brief. No step
prunes or regenerates REVIEWS.md (see the phase landing shared
block in PLAN-TEMPLATE.md).
| Step | Effort | Model | Isolation | Brief for sub-agent |
|---|---|---|---|---|
| 4a | low | sonnet | worktree | Six repositories, one pull request each, all the same shape. shakenfist/agent-python (.github/workflows/functional-tests.yml:25, release.yml:71), client-python (release.yml:73), clingwrap (functional-tests.yml:22, release.yml:73), divergulent (release.yml:71), library-utilities (release.yml:81), visual-digest-rust (ci.yml:16). Every one is a literal runs-on: [self-hosted, ..., debian-12, ...]; change that token to debian-13 and leave the size token and everything else alone. Then .github/actionlint.yaml: add debian-13 to self-hosted-runner: labels: in agent-python, clingwrap and visual-digest-rust (the other three already declare it), and delete debian-12 from all six per decision 4.2. Re-read the repository's consistency issue first for the current line numbers -- the audit refiles daily and the numbers here are from 2026-09-19. Commit subject in each: "Move CI onto the Debian 13 runner pool." actionlint passing is not the verification: wait for the repository's own CI to run a job on the new label and pass, because the label provisions a different image and this is the step that finds out whether anything in it was load-bearing. Do not touch REVIEWS.md. |
| 4b | low | sonnet | worktree | Two repositories, same shape as 4a, separated only because each has more than two lines. shakenfist/sfui (functional-tests.yml:18, :51) and client-python-k3s (functional-tests.yml:103, :225, release.yml:73, supply-chain.yml:67). Both already declare debian-13; delete debian-12 from each actionlint.yaml. Same commit subject and same verification as 4a. shakenfist/actions was the third repository here and is already done -- e2a56bd moved its three runs-on: lines and deleted both retired declarations, the stray debian-12-docker included, on 2026-09-20 in answer to actions#69. Confirm that rather than assume it, because this step's original brief is the record of what was needed. The check is grep -rn 'debian-12' .github/ in a fresh clone, and the one hit it should return is the under-cloud default at .github/workflows/smoke-cluster.yml:67, which is a guest image rather than a runner label and was blocked by shakenfist#4280 at the time -- do not "finish" the migration by moving it. Amended 2026-09-24: shakenfist#4309 closed that bug and 4f moves that default. 4b has merged and this brief is the record of what it needed; read the instruction not to touch the hit as scoped to 4b, not as advice to 4f. That was still one hit on 2026-09-21: 2e0d32a wrote an eleven-line reason above that input at :55-65, and it quotes ci-images/debian-13, the move it is refusing, rather than the -12 default below it -- the shape 4g's item (1) was going to copy to the other six sites while the deferral stood, and no longer does. Read the hit rather than the count, though: a hit in a runs-on: line, or a debian-12 list item in .github/actionlint.yaml, is what a partial label migration looks like and is what this step completes. A hit inside a comment is a reason, not a finding. Do not re-edit what is already correct. |
| 4c | medium | sonnet | worktree | The two *-docker repositories, one pull request each. shakenfist/ryll: five debian-12 (ci.yml:209, :290, :351, :385, supply-chain.yml:48) and twelve debian-12-docker (ci.yml:83, :115, :138, :235, fuzz.yml:60, manual-build.yml:99, mermaid-lint.yml:77, release.yml:76, :272, :319, :392, supply-chain.yml:66). shakenfist/occystrap: four debian-12 (functional-tests.yml:50, python-unit-tests.yml:47, release.yml:71, supply-chain.yml:76) and one debian-12-docker (mermaid-lint.yml:77). Both declare debian-13 but neither declares debian-13-docker; add it, and delete both Debian 12 declarations. Medium rather than low because these are the repositories whose jobs actually use the docker daemon: the debian-13-docker image ships docker.io and docker-cli only since actions#66, and the build runs docker version so a broken image fails rather than publishing -- so if a docker job misbehaves on the new label, report it rather than working around it, because it means that fix regressed. Commit subject: "Move CI onto the Debian 13 runner pool." |
| 4d | low | sonnet | worktree | The repositories no other step edits, as two pull requests. First, one line in shakenfist/instar. .github/workflows/functional-tests.yml:597 is image: 'debian:12', deliberate test input in a matrix that covers several releases (decision 4.3). A marker already exists and does not work; move it, do not add a second. :588-595 is a comment beginning # audit-ok: eol-distro -- Debian 12 is a SUPPORTED TARGET here, which explains the matrix entry better than any wording this plan would have proposed. A range rather than a length, because the range is self-checking against the :596/:597 anchors below it: re-read on 2026-09-21 at instar e3239f5, it is eight lines, and an earlier draft of this brief said eleven, which is actions' block at smoke-cluster.yml:55-65 and would have swallowed both anchors. It has never taken effect, because is_excepted() reads the finding's own line and exactly one line above it (scripts/audit/checks/distros.py:275-279) and this marker sits nine lines up. Keep the prose where it is -- a reader needs it at the top of the entry -- and move a marker carrying its own short reason onto :596, the - name: 'Debian 12' line, or directly above :597, indented to match: # audit-ok: eol-distro -- supported target, see the note above or similar. Not the bare token. docs/audits/eol-distro.md:169-176 specifies the shape as "Mark the line, or the line above it, with the reason", and its worked example carries one; a dangling -- satisfies EXCEPTION_RE and would leave the fleet's canonical instance of this marker as a token with nothing after it. Check the result with the criterion rather than by eye: re-run this plan's scan snippet against the edited clone and confirm instar returns zero findings. Change nothing else: instar moved its runner labels on 2026-09-16 and this is its only remaining finding. Its .github/actionlint.yaml declared debian-13, debian-13-docker and neither retired label on 2026-09-19, so there is nothing to delete -- but read it rather than trusting that sentence, and if a retired label is declared, delete it in this commit and say so, since no other step in the phase touches this repository. Commit subject: "Put the eol-distro marker where the audit reads it." -- the image has been marked deliberate since before this phase was planned, and what this commit changes is where the mark sits, which is also the general lesson and the second time the fleet has hit it. Second, shakenfist/kerbside-patches: it has fourteen workflows, every one of them already on debian-13, and its .github/actionlint.yaml still declares debian-12. It is not one of the repositories the criterion lists, because a declaration produces no finding -- it is decision 4.2's failure state in the repository that most recently migrated, and nothing else in this phase looks at it. Delete the declaration; there is no workflow line to change. Commit subject: "Stop declaring a retired runner label." private-ci also declares it and is deliberately left alone: it has no workflows at all, so the declaration governs nothing, and the repository is excluded from this criterion. Say that in the pull request rather than leaving it looking unnoticed. |
| 4e | medium | sonnet | worktree | shakenfist/shakenfist, runner labels only. Five lines: .github/workflows/functional-tests.yml:536, :726, mermaid-lint.yml:94 (debian-12-docker), pin-indirect-dependencies.yml:55, release.yml:105. .github/actionlint.yaml declares neither replacement: add debian-13 and debian-13-docker, delete debian-12 and debian-12-docker. Runner labels only. The same workflow file also names the guest image sf://label/ci-images/debian-12 at :443, :463, :477, :515, and those are step 4g -- moving them here would put a guest-image change into a pull request reviewed as a runner move. Medium because this repository's functional tests are the heaviest in the fleet and a provisioning failure here is expensive to diagnose from a red matrix. Commit subject: "Move CI onto the Debian 13 runner pool." |
| 4f | high | opus | worktree | 4f has merged as actions#97 and this brief is the record of what it needed; its post-merge verification has now passed on both runs, and the 2026-09-29 amendment in the survey is the record of that -- the 2026-09-26 amendment above it records the failing state it passed through, and is a snapshot rather than current. Additive for the artifact, not for the default, and it must merge before 4g. Restored 2026-09-24: shakenfist#4309 closed shakenfist#4280, so the two guest-image defaults this step was always going to move are back in it, alongside the release-neutral upload. In shakenfist/actions, build-smoke-cluster/action.yml. The action uploads the cached image into the cluster it just built as artifact debian-12, and every sf://upload/system/debian-12 reference in shakenfist reads it back -- 4g's item (3) gives the grep that enumerates them and says why the count is not to be trusted. Decision 4.5 moves that name to debian, with no release in it. The upload line does not read the way a grep for the command would expect. build-smoke-cluster/action.yml assembles sf-client artifact upload into a ${setup} shell variable at :260 and invokes it at :263, which literally reads "${setup} debian-12 /srv/ci/debian:12 --shared --no-checksum". The artifact name and the source path are on :263; the verb is not. Those numbers are from 2026-09-21 and moved by fourteen lines when e2a56bd rewrote the comment above them, which is the reminder that they are a grep target rather than an address. Do it in two landings so neither repository is ever reading a name the other does not write: this step adds a second upload under the name debian, sourced from /srv/ci/debian:13, which phase 3 put on the cache disk (ansible/ci-dependencies.yml:194); 4h removes the old one after 4g has landed. /srv/ci/debian:13 is the right path and /srv/ci/cached/debian:13 is not: the builder mounts the dependencies disk at /srv/ci/cached and writes {{item.name}} into it, while the cluster nodes that run this upload mount the same disk at /srv/ci (the Mount /srv/ci tasks in the ci-topology-*.yml playbooks), so the file the builder wrote as /srv/ci/cached/debian:13 is /srv/ci/debian:13 on the node doing the uploading. Do not "correct" the path to the builder's spelling. Leave the existing debian-12 upload exactly as it is, source path included. Additive has to mean the content too: repointing :263 at /srv/ci/debian:13 would leave an artifact called debian-12 containing trixie, and every reference in shakenfist would start exercising Debian 13 one merge before the repository whose tests would explain a failure. Decision 4.7 keeps /srv/ci/debian:12 cached for exactly this. The two guest-image defaults move in this step. Restored 2026-09-24. build-smoke-cluster/action.yml:32 and .github/workflows/smoke-cluster.yml:67 become sf://label/ci-images/debian-13. Unlike the upload, this is not additive and no transitional window covers it: every consumer pins @main, so the new default is live fleet-wide the moment it merges. Delete the two comments 2e0d32a wrote to explain the old default rather than editing them. They are not the same comment -- the full evidence sits above the workflow input at .github/workflows/smoke-cluster.yml:55-65 and names shakenfist/shakenfist#4280, while a three-line pointer above the action input at build-smoke-cluster/action.yml:28-30 sends the reader to the workflow and does not name the issue -- but both explain a refusal this step withdraws, and a comment saying the under-cloud is deliberately bookworm sitting above a line that says trixie is worse than no comment at all. An earlier version of this brief had 4f adding the issue reference to that pointer, because the definition of done then required each deferred site to carry one; that bullet is gone with the deferral. docs/actions.md:202-216 is a third site and no grep for the image string finds it. It is a prose paragraph, "Two things here are still Debian 12 on purpose", explaining both the default this step moves and the artifact name 4f to 4h rename, and it links #4280. Both of its halves stop being true in this phase, so rewrite it here rather than leaving a document that contradicts the file it documents -- and say in it that the artifact rename is in flight rather than done, because 4g and 4h have not landed when this does. Confirm the sweep with the bare grep -rn --exclude-dir=.git 4280 . over a fresh actions clone, prose included, which must return no reference to the issue -- read each hit, because a coincidental four-digit match (a byte count, a port, an abbreviated sha) is not one. It is the shakenfist# anchor that is being dropped here, not the .git scoping that every other gate in this plan carries: a clone's history still holds 2e0d32a's message and the deleted comments, and grep -rn over a pack file reports a binary match rather than nothing. Bare rather than anchored here, for the reason the deleted text gave and this amendment nearly lost: the epoch-timestamp collision that makes anchoring necessary is in shakenfist, not in this repository, so a bare grep over one clone is the wider net and costs nothing. It matters most at docs/actions.md, where the reference is a markdown link carrying both the shakenfist/shakenfist#4280 text and the issue URL -- an anchored pattern happens to match the text form today, and would stop matching if the link text were ever shortened to the URL alone. The anchored form stays where the plan uses it as a fleet-wide set check. The release-neutral upload above is the guest artifact, which #4280's evidence held constant and still does. Every consumer of this repository pins @main, so what does land here lands for the whole fleet the moment it merges; say so in the pull request. After this merges, trigger kerbside's sf-e2e-functional and shakenfist's node-lifecycle job on their default branches and read both runs to completion. They are the two consumers that take the guest-image default without passing one, and neither repository's own diff shows the move: no step in this phase touches kerbside at all, and 4g's edit list does not reach the node-lifecycle job. A failure in either is a revert of this commit. Both runs are a definition-of-done bullet and this sentence is what causes them; the paragraph below explains why there are two rather than naming them a second time. Each carries the default move as well as the upload, so a failure has two suspects rather than one -- the guest the instances boot is held constant, so the run isolates the guest but not the run: read which image the under-cloud booted out of the run log before concluding anything about the upload. The deferral would have needed one run now and one after the block cleared; this needs two now. The default move is not additive and two consumers take it, which is what the post-merge runs are for. shakenfist functional-tests.yml:560, the node-lifecycle job, passes only topology to build-smoke-cluster@main and so takes the default; it is not in 4g's edit list and no other step in this phase reaches it. kerbside sf-e2e-functional.yml:104 does the same, in a repository no step here touches. That is what the bolded instruction above asks for, and it asks for a deliberate trigger because waiting for whenever a pull request happens to run them is not the same thing. Both line numbers were re-read on 2026-09-24 and again on 2026-09-26. The gate before editing is a published blob: read the conductor's sf-client label update "ci-images/debian-13" line rather than the IMAGE_BUILDS entry, because an entry is not a blob and phase 3 lost a night to exactly that distinction. Keep the upload a single commit so its revert is one commit. The transitional double upload this step creates is bounded by 4h, and 4j check (5) is what proves 4h happened. It is not free while it lasts: every smoke cluster build uploads the cached image twice instead of once, against the same primary, so the window costs one extra image copy per cluster rather than one extra name. That is an argument for 4g and 4h following 4f promptly, not for skipping the transition -- the alternative is a flag day across two repositories that pin @main. The upload and ${setup} numbers in this brief were re-read on 2026-09-21 and the two consumer numbers on 2026-09-24 and 2026-09-26; there is no single date for the brief. No consistency issue lists these sites -- they are the half the criterion cannot see -- so locate them with grep -rn 'ci-images/debian-12' . and grep -rn '/srv/ci/debian:12' ., and treat the one upload and the two defaults as the expected result of those greps rather than as the instruction. Grep the source path rather than the command: the literal string artifact upload debian-12 appears nowhere in the fleet, for the ${setup} reason above, so a grep for it returns nothing and would license the conclusion that this step has nothing to do. Two commits in one pull request: "Upload the cluster base image under a release-neutral name." for the upload, and "Boot smoke clusters on Debian 13." for the two defaults, the two comments they carried and the docs/actions.md paragraph. The paragraph documents both halves while sitting entirely in the second commit, so reverting that commit does not restore a consistent document: it brings back prose calling the artifact name deliberately Debian 12 while the neutral debian upload from the first commit is still in place, which is the contradiction the rewrite exists to prevent, reached by following this brief. A revert of the default-move commit must therefore re-correct the artifact half of the paragraph by hand. Splitting the paragraph edit across the two commits would be tidier and was not done, because the two halves are two sentences of one argument and separating them reads worse than the note does. Keep them apart because their reverts are different sizes and different risks -- the upload is additive and reverts to a no-op, while the default move is live for every consumer pinning @main and is the one a failing kerbside run sends you back to. |
| 4g | high | opus | worktree | The hold is discharged: 4f's verification passed on both runs and 4g may start. 4f merged as actions#97 on 2026-09-24; the 2026-09-26 amendment in the survey records its verification as failing and the 2026-09-29 amendment records it passing, which is the one that is current. The default move is not being reverted, so this brief is followed as written and item (1) needs no re-planning. kerbside#482 stays open for a diagnosis of the single transient failure, and nothing in 4g waits on it. One thing did change under 4g while it was held: the mesh-interface task in actions now detects the network renderer rather than inferring it from the release, which is what makes a trixie multi-node cluster work at all -- so the post-merge run this brief calls the real evidence is now testing 4g's edit rather than that bug. Item (1) was blocked by shakenfist#4280 and was restored 2026-09-24 when shakenfist#4309 closed it. shakenfist/shakenfist, the guest-image half, in one pull request but not one commit. (1) Restored 2026-09-24. The four base_image: 'sf://label/ci-images/debian-12' in .github/workflows/functional-tests.yml (:443, :463, :479, :517) and the two in .github/workflows/scheduled-tests.yml (:42, :52) are the under-cloud the hypervisor VMs boot. They become sf://label/ci-images/debian-13, base_image_user staying debian at each site. All six line numbers were re-read on 2026-09-24 and again on 2026-09-26, and are current as of the later date. shakenfist#4280 blocked this for three days and shakenfist#4309 fixed it by installing qemu-system-modules-spice on trixie, so what had failed was provisioning rather than the agent and nothing about these six lines was ever wrong. Write no comment at any of these sites. The deferral required one at each naming the issue; that requirement went with the deferral, and a comment explaining a bookworm under-cloud above a line that says trixie is worse than none. Do not write an audit-ok token here either. That was true while the deferral stood and is true now for a reason that outlives it: the criterion's exception is the literal token audit-ok: eol-distro (EXCEPTION_RE, scripts/audit/checks/distros.py:197), is_excepted() reads only the finding's own line and the one above it, and docs/audits/eol-distro.md's "What this does not cover" puts guest images outside the criterion deliberately -- so there is no finding here to except and nothing a marker would suppress. If phase 6 widens the criterion, the marker shape for a guest-image site is that phase's decision. This is the first time the suite runs a trixie under-cloud since the canary that failed. #4309 is what makes it expected to work, and it landed with unit tests rather than with a green functional run, so the first merge run after this is the real evidence -- a failure there is a finding about #4309 and not about this edit, and the log to read is the libvirtd journal for the domain-definition error #4309 names. (2) The matrix lane names at functional-tests.yml:441 and :477 are "Debian 12 cluster" and "Debian 12 tier", and tools/ci_headroom_harvest.py:134 and :141 key BUNDLE_TOPOLOGIES off those strings and off the derived GitHub job names in the same entries; that file's own comment says the derivation would break silently if the names changed. Rename lanes and tool in the same commit, and update shakenfist/tests/test_ci_headroom_harvest.py:61-62. (3) Replace every sf://upload/system/debian-12 reference with sf://upload/system/debian -- git grep -n 'sf://upload/system/debian-12' -- shakenfist/ is what enumerates them, 41 lines in 13 files on 2026-09-26; most are under shakenfist/deploy/shakenfist_ci/ and shakenfist/tests/, and one is shakenfist/deploy/nodelifecycletests.sh -- and route them through CLUSTER_CI_IMAGE (shakenfist/deploy/shakenfist_ci/base.py:37) wherever the file already imports from base, so the next release is one line. Amended 2026-09-25 and again 2026-09-26: four path corrections, one line number withdrawn, and a count correction that was itself wrong and is superseded below, all verified against origin/develop. The source tree is nested -- deploy/ and tests/ live under a top-level shakenfist/ package directory, while tools/ is at the repository root -- so the paths given without the prefix did not resolve. That is four of them in this item, plus shakenfist/tests/test_ci_headroom_harvest.py:61-62 in item (2), which took the prefix in the same commit and is confirmed present at that path: the package-level tests/ is the right one even though the script it covers, tools/ci_headroom_harvest.py, is at the root. The count is a count of lines, and it already includes nodelifecycletests.sh rather than standing beside it, so the old "12 files plus one more" reading double-counted it -- and the directories are now given as the grep that enumerates the references rather than as their scope, because nodelifecycletests.sh is a sibling of shakenfist_ci/ and not a member, so an implementer deriving the grep from the two directory names misses it. The line number for that file is withdrawn rather than corrected: it was :132 in the survey, :187 when this branch's second commit corrected it, and :217 thirty-nine minutes later. The survey's drift paragraph carries the detail. A literal that the constant does not reach is the phase 3 failure this step is repeating on purpose; the definition of done greps for the old name, so leaving any is not passing. (4) functional-tests.yml:585 uploads the image itself for the node-lifecycle job, the same command as the action's upload line (:263 on 2026-09-21, and a grep target rather than an address -- e2a56bd moved it by fourteen lines): move it to debian and /srv/ci/debian:13 too. Commit subjects, one per numbered item and all four given verbatim rather than inferred, which is this phase's convention everywhere else: (1) "Boot the cluster lanes on Debian 13.", (2) "Rename the Debian 12 matrix lanes.", (3) "Read the cluster image by its neutral name.", (4) "Upload the node lifecycle image as debian.". Every line number in this brief was read on 2026-09-19 in shakenfist and re-read against origin/develop on 2026-09-26, except the one actions number in item (4), which was re-read on 2026-09-21. No commit in this phase renumbers them -- but ordinary work in that repository does: functional-tests.yml:477 and :515 had become :479 and :517 by 2026-09-22, which is why item (1) lists them at the later numbers. They are a grep target rather than an address, and that matters most in item (1), which edits six specific lines and has no grep of its own beyond ci-images/debian-12. The one actions number, in item (4), was re-read on 2026-09-21 after e2a56bd moved it. None of these sites appears in any consistency issue, because they are the half the criterion cannot see. 4g runs after 4e has edited two of the same files and after however much ordinary work has landed in between, so locate the work with grep -rn 'ci-images/debian-12' .github/workflows/, grep -rn 'sf://upload/system/debian-12' . and grep -rn 'Debian 12 cluster\|Debian 12 tier' ., and read the six base_image sites and the two lane names as the expected result of those greps: those two counts are stable, nothing in this repository's ordinary work has moved either since the survey, and either of them moving at all means something else edited these sites and is the finding worth halting for. The sf://upload/system/debian-12 count is not in that category and is not a gate. Run git grep -l 'sf://upload/system/debian-12' -- shakenfist/ and edit what it lists; it was 13 files and 41 lines on 2026-09-26 and both figures move under ordinary refactoring, so report what you find and carry on rather than halting on a disagreement. What proves the item finished is the definition of done's grep for the old name returning nothing, which is a question about the end state rather than about a count agreeing with this document. Amended 2026-09-26, and this supersedes a wrong correction made on 2026-09-25. The 09-25 amendment said the "48 in 13 files by 2026-09-22" figure "was never true". It was true. It was measured repo-wide and the correction re-measured it under -- shakenfist/, which is a different question, and then reported the disagreement as an error in the original rather than as a difference of scope. Measured properly, with git grep -h 'sf://upload/system/debian-12' <sha> | wc -l for lines and git grep -l ... | wc -l for files, at the tip of origin/develop: 2026-09-19, 43 lines in 12 files repo-wide and the same under shakenfist/; 2026-09-22, 48 in 13 repo-wide and 43 in 12 scoped; 2026-09-25, 50 in 13 and 43 in 12; 2026-09-26, 48 in 14 and 41 in 13 scoped. The repo-wide figures run ahead because this plan is itself synced into that repository at docs/components/development/plans/PLAN-image-supply-chain.md, so every paragraph written here about the string adds to a count measured there. Do not use the count as a tripwire. It has now moved under ordinary work too: 5202774ed (2026-09-25 19:18 UTC) shared the interface hot plug tests via a mixin, which took guest_ci_tests/test_agentops.py and smoke_ci_tests/test_agentops.py from five matches each to three and added shakenfist_ci/instance_hotplug.py with two. A number that a routine refactor moves twice in a week cannot detect a third party editing these sites, which is the only job the halting rule gave it. What can do that job is the file set and the finishing grep, so both are stated below and the number is not an expectation. High effort because item (3) is forty-odd sites in a test suite whose failures are slow to read, and because item (2) fails silently rather than loudly. Item (1) adds to that rather than replacing it: items (2) to (4) rename what the tests boot and what the tool keys off, item (1) changes what the hypervisors themselves run, and a failing run after this lands has more than one suspect -- which is the argument for reading the run log rather than re-reading the diff. |
| 4h | medium | sonnet | worktree | After 4g has merged, and gated on a grep rather than on this sentence. Two pull requests in two repositories, and the order between them is the point -- amended 2026-09-30. When this step was written it was one actions commit; running 4g found that it also has consumers to move first, and 4g had already merged as shakenfist#4379 (2026-09-29 22:03 UTC) without them, so they are this step's. (1) shakenfist/shakenfist, one pull request, merged before (2) is opened. Both repositories are consumed at their default branch -- build-smoke-cluster@main is live fleet-wide the moment it merges -- so removing the upload first breaks every consumer still reading the old name, which is the flag day 4f split its own work to avoid. The consumers are references to the artifact by its bare short name rather than its URL: the URL form sf://upload/system/debian-12 is the spelling of the upload, and these are invisible to a grep for it and were invisible to all three of 4g's definition-of-done greps. On 2026-09-30 gate B below found seven, all on origin/develop: shakenfist/deploy/shakenfist_ci/guest_ci_tests/test_boot.py carries four -- two testscenarios entries, each a scenario name 'debian-12' and a 'base': 'debian-12' that is the actual read -- and shakenfist/deploy/shakenfist_ci/cluster_ci_tests/test_imagefetch.py three, each sf-client artifact download debian-12 .... That is what the grep found on that date, not an expectation: edit what gate B lists when you run it, and report a disagreement rather than halting on it, as 4g's brief says of its own count. Move each read to debian, and rename the two scenario names with them so a test id does not name an image it no longer boots. The /var/www/html/debian-12-... paths in test_imagefetch.py are filenames the test writes, not reads of the artifact, and gate B deliberately does not report them. They pass today only because build-smoke-cluster still uploads both names, so this pull request's own CI proves the new name is read and cannot prove the old one is unread -- gate B is what proves that. Commit subject: "Read the cluster image artifact as debian, not debian-12." (2) shakenfist/actions, after (1) has merged. Remove the pre-existing debian-12 upload that 4f deliberately left in place in build-smoke-cluster/action.yml, leaving only the debian one that 4f added. 4f adds a name and removes none; this step removes the old name. If the diff you are about to write deletes the line that says debian, you have the wrong one. Rewrite the comment block above the upload in the same edit: it explains the two-name window and quotes the old URL, and it is gate A's one expected hit. The rewritten comment names neither the old URL nor the two-name window, so gate A returns nothing after this edit -- the definition of done asks for exactly that, and a comment saying "formerly sf://upload/system/debian-12" fails it at 4j. Before editing (2), run both gates over fresh clones of every non-archived repository in the organisation -- not just shakenfist, which is where 4g worked -- and stop and report instead of editing if either shows a read of the artifact beyond its expected residue. Run them from a directory that contains only the clones, and pass * rather than .: the filters below anchor on repo/path, and whether grep -r . prefixes its output with ./ depends on the grep version. Gate A, the URL: grep -rnI --exclude='*.md' --exclude-dir=.git 'sf://upload/system/debian-12' *. Expected residue on 2026-09-30: one line, the comment in actions/build-smoke-cluster/action.yml (:241 on that date) that this edit rewrites. Gate B, the short name: grep -rnIE --exclude='*.md' --exclude-dir=.git '(^|[^[:alnum:]_/.-])debian-12([^[:alnum:]_.-]|$)' * | grep -vE -e '^(development/scripts|private-ci/conductor)/' -e '^[^:]*/actionlint\.yaml:' -e '^[^:]+:[0-9]+:[[:space:]]*(-[[:space:]]*)?runs-on:' -e '^[^:]+:[0-9]+:[[:space:]]*#'. Expected residue on 2026-09-30 once (1) has merged: two lines, the upload itself at actions/build-smoke-cluster/action.yml:267, which this edit deletes, and a docstring in shakenfist/tests/test_mariadb_capacity_admission.py (:1210) that names the runner. Before (1) merges the seven above are in it too. Read gate B's output by hand rather than comparing it to zero -- a boolean is what hid the seven -- and treat anything that reads the artifact as a halt. What each filter subtracts, so a disagreement can be diagnosed rather than filtered further: the token pattern admits only a whole debian-12, not preceded by / (so neither ci-images/debian-12, the under-cloud label, nor the URL, which is gate A's) and not followed by - or . (so neither debian-12-docker, debian-12-sfagent and the other image and bundle names, nor a .qcow2 filename). The upload is named exactly debian-12, so a longer token is a different label, image or file, not this artifact. /srv/ci/debian:12 does not contain the token at all and is 4j check (5)'s. The two directory excludes are the audit's runner-label checks and fixtures in development and private-ci's runner image builder, whose IMAGE_BUILDS name is a runner label -- 83 lines between them on 2026-09-30, every one a runner label and phase 5's. They are the one place this filter could hide a consumer, which is why there are exactly two and why they are directories that build runners rather than smoke clusters; do not add a third to make the output shorter. actionlint.yaml and a runs-on: key are runner labels and 4j check (4)'s; the runs-on filter is anchored as a YAML key, optionally a list item, so a code line that merely mentions runs-on is still reported (anchored and unanchored returned the same lines on 2026-09-30). A line whose first non-blank character is # is a comment, out of scope for the same reason prose is; a trailing comment on a code line is still reported. Checked 2026-09-30 by eleven mutations over the clone set: a .j2 under actions/ansible/, an extensionless script, a Makefile, a workflow list item, a double-quoted string, a code line with a trailing comment and a file in private-ci outside conductor/ are each reported; a -docker variant, the label URL, a comment line and a markdown file are not. Both gates take -I, so they differ only in their pattern: the artifact is read by source, and a binary fixture that happened to contain the string would print a Binary file ... matches line that is not a read. Gate B sees only a literal debian-12. A name assembled at runtime -- f'debian-{release}', debian-${VERSION}, 'debian-' + release -- is invisible to both gates, so also run grep -rnIE --exclude='*.md' --exclude-dir=.git "debian-[\${'\"]" * and read it by hand the same way; the bracket expression is deliberate, because inside double quotes a \$ in an alternation reaches grep as a bare $, an end-of-line anchor, and silently stops matching debian-${VERSION}. Checked by mutation: f'debian-{release}', 'debian-' + rel, "debian-" + x and debian-${VERSION} are reported, debian-12 and debian-gnome are not. On 2026-09-30 it returned three lines, none a read of this artifact: the Jinja-templated GNOME download name at actions/ansible/ci-dependencies.yml:308, and two 'debian-' prefix tuples in divergulent. A name built further from its parts than that is beyond any grep; the backstop is the CI of the consumers (1) moved. That gap widens over time rather than closing: 4g's advice to route reads through CLUSTER_CI_IMAGE is advice to stop spelling the name literally. Both gates take --exclude='*.md' --exclude-dir=.git: this plan is in one of those clones and names the string a dozen times, and a gate that halts on its own plan file is a gate an agent learns to override. Prose is out of scope here for the same reason eol-distro exempts it -- a document describing a migration is not a dependency on it. Exclude prose rather than allow-listing extensions. An allow-list of .py, .yml, .yaml and .sh reads well and silently drops .j2 -- of which actions/ansible/ is full -- along with extensionless scripts and Makefile. This gate's failure is destructive by omission: it deletes the upload while a consumer it could not see still reads the name. Phase 3's retrospective is that a gate stated in a plan file is not a gate; this one is a command whose output decides the step. Drop the "in flight" qualifier from docs/actions.md in the same commit. Find it by content rather than by its opening words: grep docs/actions.md for in flight on actions' default branch, which is the paragraph 4f wrote about the guest-image upload. (It begins "The guest image the action uploads into the nested cluster" as 4f actually wrote it in actions#97, read 2026-09-26 -- but 4f's brief never dictated that wording, so the quote is a convenience and the grep is the address.) It says the rename was in flight, which was true when 4f landed and stops being true here; after this step there is one upload, under the name debian, and no transitional window left to describe. Leaving it is the same defect 4f rewrote the paragraph to avoid -- a document contradicting the file it documents -- and no other step in this phase reaches that file: 4g is in a different repository and 4j's greps exclude *.md by design. Commit subject: "Drop the transitional cluster image name." |
| 4i | low | haiku | none | Housekeeping, no commit in this repository. Close development#123 with the merge commits, noting that the fleet cleared 38 of its 80 references through ordinary repository work answering the daily audit before this phase began, and grew new ones in the same fortnight. That is a statement about the period before the phase, which no later merge can falsify; do not turn it into a claim about the final split, because the compliance page the next sentence sends you to shows the current count and not who cleared what, and by the time 4i runs this phase will have cleared most of the remainder itself. Read the closing number off docs/audits/compliance.md when this step runs rather than from this plan, per the survey's own rule that the inventory is the compliance page and never a number in this document: the survey counted 45 references in 13 repositories on 2026-09-19, the 2026-09-21 amendment made that 42 in 12 when actions cleared itself, and 4i runs after every other step has merged. Amended 2026-09-24: that was "the eight remaining steps -- sixteen pull requests, by those steps' own briefs" when this was written; 4a to 4e opened thirteen pull requests, eleven of which had resolved by 2026-09-24 -- ten merged and visual-digest-rust#23 closed as superseded -- so what 4i waited for was occystrap#143, shakenfist#4306, and 4f, 4g and 4h at one pull request each. Amended 2026-09-26: occystrap#143 and shakenfist#4306 merged on 2026-09-24 and 4f merged as actions#97 the same day, so 4i now waits for 4g and 4h alone, and the merge commits it closes development#123 with include 8c02ab0e7 for 4f. Amended 2026-09-30: 4g merged as shakenfist#4379, and 4h is now two pull requests, shakenfist then actions; both merge commits belong in the closing comment, because the shakenfist one is the consumer move 4g's discovery added and citing only the actions merge loses it. The net is the less interesting half: a fleet that grows references while clearing them is why decision 4.2 deletes the declaration with the last user. Comment on private-ci#38 with the guest-image inventory this survey found, since #38 is the collated reference and did not have it: the eight sf://label/ci-images/debian-12 sites, the upload name, the seven reads of that upload by its bare short name debian-12 found live in shakenfist on 2026-09-30 after every URL grep had returned clean, and the eight references in private-ci's own conductor/tests/test_imagebuilder.py (seven naming the label, one the -docker variant) that follow IMAGE_BUILDS in phase 5. Leave #38 open; it closes at the end of phase 5. Do not close the per-repository consistency issues by hand -- the audit closes them itself when the repository goes compliant, and closing one by hand hides a repository that did not. actions' own issue will already have been closed that way, on 2026-09-20, before this step runs. |
| 4j | medium | sonnet | none | Confirms the phase, which nothing else does. Medium and sonnet rather than the mechanical pair its first draft had: six checks, three of which read a run log or a live grep across twenty-nine clones and decide whether the answer agrees with a diff. Observation step, no commit, run after every pull request above has merged and at least one morning's audit has run. (1) Read docs/audits/compliance.md on this repository's main and confirm the eol-distro table has no non-compliant row, and that its generation timestamp is after the last merge -- a stale page looks healthy, which the page's own header warns about. (2) Confirm the renamed artifact is both written and read, which takes two different runs because 4f and 4g land in different repositories. Writing: in a smoke cluster build after 4f, an upload line names the artifact debian with nothing after it. Booting: in a completed shakenfist functional-tests run after 4g, an instance boots sf://upload/system/debian with nothing after debian. Anchor both ends, because the old names contain the new ones: debian-12 contains debian and sf://upload/system/debian-12 contains sf://upload/system/debian, so an unanchored read passes on the pre-rename line. Between 4f and 4h the expected state of the build log is two upload lines, the old name and the new one, which is worth counting: finding one is itself a finding, and which one it is says whether 4f has not landed or 4h has landed early. Read both out of run logs rather than out of the workflow files, which only prove what was asked for. Keep the halves apart: 4f adds an upload in actions and renames nothing that shakenfist reads, so a run between 4f and 4g shows the write and cannot show the boot, and reporting the upload line as if it were the boot is a pass on the wrong assertion. The definition of done states the booting half only, and after 4g. Read the under-cloud image too, and expect debian-13. This check was written to do that; the 09-21 amendment inverted it while shakenfist#4280 was open, and 2026-09-24 restored it. After 4f and 4g both logs show the under-cloud booting ci-images/debian-13, and one still reading -12 means a default did not move -- check which commit the run is of before reporting it, because a re-run of a pre-4g commit reads the old value legitimately. (3) Confirm tools/ci_headroom_harvest.py still matches its bundles after the lane rename, by running it against a merge run that completed after 4g. (4) grep -E "^[[:space:]]*-[[:space:]]*['\"]?debian-12" <clone>/.github/actionlint.yaml, the quote optional because - "debian-12" and - 'debian-12' are both valid YAML and these very files already quote their shellcheck entries that way, across fresh clones of every non-archived repository in the organisation, not the ones the criterion lists -- a stale declaration produces no finding in any criterion, so this is the only check in the phase that would catch one, and the repositories most likely to carry one are the ones that already migrated and so are not listed at all. The list-item anchor is load-bearing: this repository and hunkydory both name debian-12 in a comment explaining why it is not declared, and a plain substring grep reports both. private-ci is the one expected hit and is step 4d's stated exception, for having no workflows at all. (5) grep -rn --exclude='*.md' --exclude-dir=.git '/srv/ci/debian:12' . over the same clones must return nothing, and build-smoke-cluster/action.yml must carry exactly one line naming /srv/ci/debian:13. The --exclude scoping is 4h's, for 4h's reason and in 4h's shape -- exclude prose, do not allow-list extensions: this plan file is in one of those clones and names the path several times, and a gate that halts on its own plan file is a gate an agent learns to override. It does not collide with decision 4.7, which was checked rather than assumed: ci-dependencies.yml spells its cache list as name: "debian:12" against an images.shakenfist.com URL, and checked against actions at the same commit as the rest of this survey, build-smoke-cluster/action.yml:249 as the survey read it -- :263 since e2a56bd, and found by grep rather than by line -- was the only line in that repository naming a /srv/ci/debian path at all, so the bookworm cache entry 4.7 keeps is not a hit. Grep the source path, not the command: the upload is assembled from a ${setup} variable, so the literal string artifact upload debian-12 appears nowhere in the fleet and a check looking for it passes whether or not 4h ran -- which is the shape of silent skip this step exists to catch, found by running the grep rather than reading it -- only this grep and gate B below can tell whether 4h ran, because the definition of done's label and URL greps match the reader spelling in shakenfist rather than the writer line in actions, while gate B reports the debian-12 upload line itself until 4h deletes it; this grep is the mechanical half of the pair, compared to zero rather than read by hand, and a skipped 4h leaves every smoke cluster in the fleet uploading a bookworm image forever with no finding anywhere. The exactly one half catches the inverse mistake 4h's brief warns about. Then run 4h's gate B, the short-name read, over the same clones -- the command is in 4h's brief and the definition of done -- and read its output by hand rather than comparing it to zero: after 4h it should show only the runner docstring the definition of done names, and a line that reads the artifact by the name debian-12 fails this check even though every URL grep above passes, which is the blind spot 4g found. (6) Read docs/actions.md on actions' default branch after 4h and confirm it describes the finished state: the under-cloud paragraph says trixie, and the guest-image paragraph describes a single upload under the name debian with no transitional window and no "in flight" qualifier. This is the one definition-of-done bullet whose subject no other check reaches, because every grep in this step excludes *.md by design, so without it the claim is a gate stated in a plan file rather than a gate. Report all six; if (2) to (6) disagrees with the diff, say so rather than filing it, because the phase is not over until they agree. |
Risks and mitigations¶
A label that provisions is not an image that works. The
debian-13 image is a different rootfs: a job that relied on a
package bookworm shipped and trixie does not will fail at the step
that uses it, not at provisioning. Mitigated per repository rather
than centrally -- every step above waits for that repository's own
CI on the new label and treats a green actionlint as no evidence at
all. This is cheap here because the fleet has already done it
fourteen times: instar and kerbside moved 34 references between
them in two days without a follow-up fix.
The guest-image change lands for the whole fleet at once.
build-smoke-cluster@main is what every consumer pins, so 4f is live
everywhere the moment it merges, including for kerbside, which takes
the default. Amended 2026-09-24: the sharp half of this risk is back,
because shakenfist#4309 closed the bug that deferred it. 4f now lands
two things of different shapes. The upload is additive -- the old name
and its bookworm source stay untouched until 4h -- so nothing changes
underneath a consumer, and its residual risk is the cost of the
transitional window. The default is not additive: it moves every
consumer's under-cloud to trixie at the moment it merges, which is the
sharp half and is what this heading is about. Mitigated by 4h closing
the transitional window, by 4j check (5) proving it closed, and by 4f
triggering kerbside's sf-e2e-functional and shakenfist's
node-lifecycle job on their default branches after merge -- the two
consumers that take the default, one of them in the only repository no
step here edits. Those two runs hold the guest constant, because
nothing reads the new artifact name until 4g, so they are a cleaner
experiment than the "both changes at once" an earlier draft of this
paragraph claimed -- for the guest. They do not isolate the run: each
run carries two changes, the default move and the extra upload, while
what the instances boot is unchanged, which is the sense in which 4f's
brief says a failure has two suspects.
Amended 2026-09-26: this mitigation has fired, and half of it is
red. kerbside's nightly sf-e2e-functional after 4f merged, run
36110697157, failed for a reason not yet attributed; the
node-lifecycle run has not been triggered. Whether 4f's default
move is reverted is open. The survey's 2026-09-26 amendment is the
record and kerbside#482 is where the diagnosis lives. Amended
2026-09-29: both halves are now green and no revert is happening --
kerbside passed on 09-26, 09-27 and 09-28, and node-lifecycle
passed twice on 09-29 once actions#112 landed. The mitigation
worked, and it worked late. What it caught, it caught through
kerbside's nightly and not through this fleet's own CI, because
the job that was actually broken runs on merge_group and
workflow_dispatch only: four days of green pull requests hid it.
A mitigation that depends on a job the default verification path
cannot reach is worth less than this paragraph assumed.
Nothing has ever booted a trixie guest artifact. The 09-21
amendment established this and shakenfist#4309 does not change it:
both canary runs uploaded the same bookworm guest, so the guest
was the controlled variable rather than a tested one, and #4309
fixed hypervisor provisioning. The first run that exercises a
trixie guest is the shakenfist functional-tests run after 4g,
and for the four matrix lanes -- which pass base_image
explicitly, so their under-cloud moves in 4g rather than in 4f --
that run moves under-cloud and guest together. Mitigated only
partly, and the partial is the honest word: the node-lifecycle run
after 4f, once someone triggers it -- as of 2026-09-26 it has not
run -- is a data point about the same under-cloud image in the
same repository, which narrows a 4g failure towards the guest
without isolating it, and 4g's brief says to read the libvirtd
journal for the domain-definition error #4309 names before
concluding anything. If that is not enough when 4g runs, the
fallback is back brief question 4's proposal: source the neutral
name from /srv/ci/debian:12, land 4g, and move the content to
:13 afterwards. Recorded here so it is a decision available at
that point rather than a rediscovery.
The blocker closed, and nothing in this plan noticed.
shakenfist#4280 gated 4f's defaults, 4g's item (1) and phase 5's
Debian 12 half, and it was an open bug in another repository with
no owner named here. The failure mode this section named was not
that it stayed open -- it was that it would close and nobody would
notice, leaving eight sites on a retired image with a comment
explaining a reason that had expired. It closed on 2026-09-23 and
was withdrawn from this plan the next morning, before any step had
written one of those comments. Read that as a near miss rather
than as a mitigation that worked: what caught it was someone
asking whether the bug had merged, not a check. The two
mitigations this section claimed would not have. Requiring a
comment at each of the eight makes the sites greppable and says
nothing about the issue's state, and it had not run yet anyway;
phase 5's first step reading the issue would have happened weeks
later, after 4f and 4g had landed comments that were already
wrong. The analysis to inherit is unchanged and still unfunded:
plan-status-vocabulary defines Blocked as "cannot proceed
until something outside the plan changes" and
scripts/audit/checks/plans.py accepts it, but a half-blocked
phase is still In progress, the vocabulary's rule is that the
term is the whole cell, and so there is nowhere to record what
the block is. Phase 7 should ask for a blocked-on column, and this
episode is the argument for it: a cell naming shakenfist#4280 is
something a daily audit could have watched close, which is exactly
what no human remembered to do.
A silent stop, not a failure. ci_headroom_harvest.py keys off
job names, and 4g renames them. Nothing fails if the tool is missed:
it matches nothing and reports empty. Mitigated by putting the
rename and the tool in one commit, and by 4j re-running the tool
against a real merge run rather than reading the diff.
The inventory moves while the phase runs. Two repositories grew new findings during the fortnight this plan sat still. Mitigated by every step re-reading the repository's consistency issue for current line numbers before editing, and by 4j asserting against the regenerated compliance page rather than against this plan's table.
Phase 5 starts before this finishes. D4 orders producers after
consumers, and the guest-image half is the part that makes that
ordering real. Mitigated by 4j being the gate: phase 5's first step
should refuse to start until 4j has reported every check it runs
agreeing. Amended 2026-09-24: 4j agreeing is sufficient again. The
eight under-cloud sites are back in this phase, so when 4j reports
every check agreeing there is no ci-images/debian-12 left anywhere
for phase 5 to trip over, and phase 5's Debian 12 follow-up gates on
4j and nothing else. The debian-11 bullet remains unaffected and
startable now, as it has been since phase 1 completed.
Definition of done¶
All twelve bullets hold as of 2026-10-02. Eight are covered by 4j, whose output is under What 4j confirmed below: bullet 1 is 4j (1), bullet 2 is 4j (4), bullets 4, 5 and 6 are 4j (5), bullet 8 is 4j (6), bullet 10 is 4j (2) and bullet 11 is 4j (3). Bullets 3, 7, 9 and 12 are not among 4j's checks, so each carries its own evidence inline. A checklist and an observation step's brief written weeks apart overlap only in part, which is worth knowing before the next phase writes both.
-
docs/audits/compliance.md'seol-distrotable has nonon-compliantrow, on a page generated after the last merge. -
grep -E "^[[:space:]]*-[[:space:]]*['\"]?debian-12"over the.github/actionlint.yamlof every non-archived repository in the organisation returns onlyprivate-ci, which has no workflows for a declaration to govern. The anchor matters: this repository andhunkydoryname the label in a comment saying why it is absent, and a substring grep reports both. The scope is the whole organisation rather than the repositories the criterion lists, because a repository that has already migrated is where a declaration outlives its last user --kerbside-patcheswas exactly that when this phase was planned. -
grep -rn --exclude='*.md' --exclude-dir=.git "ci-images/debian-12"over fresh clones of every non-archived repository returns nothing outsideprivate-ci/conductor/tests/test_imagebuilder.py, whose eight references -- seven naming the label and one the-dockervariant -- phase 5 moves with theIMAGE_BUILDSentries, with nothing else exempted. Amended 2026-09-24: this bullet listed eight under-cloud sites as expected hits while shakenfist#4280 was open, each required to carry a comment naming the issue. shakenfist#4309 closed that bug, 4f and 4g move all eight, and both the exemption and the comment requirement are withdrawn -- a hit anywhere outside that oneprivate-citest file now fails this bullet, which is what it asserted before the bug was found. The substring rather than thesf://label/form, so that both spellings are caught; theprivate-citest file happens to use the prefixed one. The--excludeis what keeps the grep off prose -- this plan names the string a dozen times, anddevelopmentis in the clone set -- and it is the same exemptioneol-distromakes for a document describing a migration. It excludes prose rather than allow-listing extensions, so a reference in a.j2template or an extensionless script is still read.private-ciis in the clone set deliberately: it is excluded from the criterion, so it is exactly the repository a criterion-shaped check would miss. Run 2026-10-02 over fresh--depth 1clones of all 29 non-archived repositories, with no clone failures: eight hits, all inprivate-ci/conductor/tests/test_imagebuilder.py--:121,:156,:167,:180,:187,:197,:216naming the label and:874the-dockervariant, which is the seven and one this bullet predicted. Nothing anywhere else. This is the bullet no step owned: it is not one of 4j's six checks, and until this was run the phase was being declared complete on a criterion with no result. - 4h's gate A, the same grep for
sf://upload/system/debian-12with-Iadded, returns nothing at all. -
4h's gate B, the short-name read, run over the same clones from a directory holding only them, shows nothing that reads the artifact:
grep -rnIE --exclude='*.md' --exclude-dir=.git \ '(^|[^[:alnum:]_/.-])debian-12([^[:alnum:]_.-]|$)' * \ | grep -vE -e '^(development/scripts|private-ci/conductor)/' \ -e '^[^:]*/actionlint\.yaml:' \ -e '^[^:]+:[0-9]+:[[:space:]]*(-[[:space:]]*)?runs-on:' \ -e '^[^:]+:[0-9]+:[[:space:]]*#'Its output is read by hand rather than compared to zero. The expected residue after 4h, as of 2026-09-30, is one docstring in
shakenfist/tests/test_mariadb_capacity_admission.pynaming the runner, which is not a read. The bullet above cannot stand in for this one: consumers name the artifact by its short name as well as its URL, and on 2026-09-30 seven such reads were live inshakenfistafter all three of 4g's greps had returned clean. What each filter subtracts, and why, is in 4h's brief. - [x] The same grep for/srv/ci/debian:12returns nothing, andactions/build-smoke-cluster/action.ymlcarries exactly one line naming/srv/ci/debian:13. This is one of two checks that can tell whether 4h ran: the label and URL greps above match the reader spelling inshakenfist, not the writer line inactions, while gate B above reports thedebian-12upload line itself until 4h deletes it. This bullet is the mechanical half, compared to zero rather than read by hand, and itsexactly onehalf also catches the inverse mistake of deleting thedebianline instead. It greps the source path rather thanartifact upload debian-12, because the command is assembled from a${setup}variable and that literal exists nowhere -- a bullet asking for it would pass vacuously. Decision 4.7's bookworm cache entry is not a hit:ci-dependencies.ymlspells its cache list asname: "debian:12"against animages.shakenfist.comURL, not as a/srv/ci/path. - [x]kerbside'ssf-e2e-functionalworkflow andshakenfist's node-lifecycle job have each completed successfully on their default branches after 4f merged, triggered by 4f, which is the step that owns this bullet. They are the two consumers that takebuild-smoke-cluster's default without passing abase_image--kerbside.github/workflows/sf-e2e-functional.yml:104andshakenfist.github/workflows/functional-tests.yml:560, both re-read 2026-09-24 -- so 4f moves their under-cloud without either repository's diff showing it, and 4g's edit list reaches neither.kerbsideis additionally the one repository no step in this phase edits at all. Amended 2026-09-24: while shakenfist#4280 was open this bullet coveredkerbsidealone and proved only that the upload had not broken it, because the default was not moving. It covers both and proves both again. Verified in the 2026-09-29 amendment in the survey above, which is where the run IDs are:kerbside's nightlysf-e2e-functionalondeveloppassed three consecutive times (36228146379, 36306108473, 36399738723) after the 09-25 failure, andshakenfist's node-lifecycle job passed twice ondevelop(merge_group runs 36501556771 and 36513297089, each a full six-host build). The amendment also records why the second half took a fix inactionsfirst. - [x]actionsdocs/actions.mddescribes the finished state: the under-cloud paragraph says trixie, and the guest-image paragraph describes a single upload under the namedebianwith no transitional window. 4f writes the first and leaves the second saying the rename is in flight, which is true between 4f and 4h and false after; 4h owns removing it. Nothing else in the phase reads that file -- 4j's greps exclude*.md-- so without this bullet the phase closes with the documentation describing a completed rename as still under way. - [x]instar'sfunctional-tests.ymlcarries theaudit-ok: eol-distromarker within one line of the finding, and with a reason after it, and instar#564 is closed by the audit rather than by hand. The two requirements are independent and this phase only had to add the first: a marker with a reason has been in that file since before this phase was planned, nine lines away, so proximity is what was missing and the reason is whatdocs/audits/eol-distro.md:169-176has always asked for. Dropping either leaves the fleet's canonical instance of this marker wrong in one of the two ways. Verified 2026-10-02 oninstaratdevelop: the marker is.github/workflows/functional-tests.yml:651and the finding it covers isimage: 'debian:12'at:652, one line apart, with the reason-- supported target, see the note aboveon the marker line. instar#564 is closed,COMPLETED, at 2026-09-24T11:29:44Z -- by the audit, which is the half of this bullet the phase could not do itself. - [x] A completedshakenfistfunctional-tests run after 4g shows instances bootingsf://upload/system/debian, read from the run's log, and the under-cloud in that same log readingci-images/debian-13. Amended 2026-09-24: shakenfist#4280 inverted the second half of this for three days and shakenfist#4309 restored it. It is the first green trixie under-cloud this suite will have produced, so read it rather than assuming it. - [x]tools/ci_headroom_harvest.pymatches its bundles on a merge run completed after 4g. - [x] development#123 is closed; private-ci#38 is still open and carries the guest-image inventory. Verified 2026-10-02: #123 closed by 4i, and #38's inventory is the comment of 2026-10-01 08:06 UTC, which lists the label and artifact references the survey found and which the issue did not previously have. #38 stays open until the end of phase 5 by design, which is why it is a bullet about two different states rather than two closures.
What 4j confirmed¶
4j ran on 2026-10-02 (AEST), after the 2026-10-01 12:44 UTC audit run regenerated the compliance page -- 22:44 the previous evening in Canberra, which is still on AEST until daylight saving starts on 2026-10-04. Bare dates in this subsection are AEST and timestamps are UTC, which is the convention the status paragraph states; the two differ by a calendar day for anything before 10:00 local. All six checks pass. Recorded here because the phase is not confirmed by its merges and nothing else in this document says so.
(1) The eol-distro table on main lists no non-compliant row -- 19
compliant, cloudgood and private-ci N/A -- and the page was
generated at 2026-10-01 12:44:43 UTC, after the phase's last merge at
08:04:28 UTC. Both halves, because a stale page looks healthy: the
check was genuinely blocked on 2026-10-01, when the then-current page
predated the last two merges, and it was not run until that cleared.
(2) The renamed artifact is both written and read, read out of run logs
rather than workflow files, with both ends anchored. In the three
merge_group runs completed after 4g (36651832849, 36661716274,
36669390897, each confirmed a descendant of 2a94e582f) the write side
shows the expected two upload lines for that date, the read side
shows base=sf://upload/system/debian unsuffixed, and the under-cloud
shows sf://label/ci-images/debian-13. The post-4h state -- one
upload line -- was then read out of run 36938755972, whose build log
postdates 4h part 2: exactly one upload, debian /srv/ci/debian:13,
and no occurrence of debian-12 or debian:12 anywhere in its 9,090
lines. The brief asks for the source check and the log check as a pair;
both agree.
(3) tools/ci_headroom_harvest.py matches all five BUNDLE_TOPOLOGIES
entries against a merge run completed after 4g, with no
UnknownBundleError. This is item (2) of 4g's own work, whose failure
mode was silent.
(4) The anchored actionlint.yaml list-item grep across fresh clones of
all 29 non-archived repositories returns one line, private-ci, which
is step 4d's stated exception for having no workflows. The anchor is
load-bearing and was checked: an unanchored grep also reports
development and hunkydory, both comments explaining why the label is
not declared.
(5) /srv/ci/debian:12 returns nothing across those clones,
build-smoke-cluster/action.yml carries exactly one /srv/ci/debian:13,
gate A returns nothing, and gate B returns one line: a docstring in
shakenfist/tests/test_mariadb_capacity_admission.py naming a runner,
not the artifact. The runtime-assembly grep returns the same three known
benign lines it did on 2026-09-30.
(6) docs/actions.md describes the finished state: the under-cloud
paragraph says trixie, the guest-image paragraph describes a single
upload named debian, and in flight and transition no longer appear.
One failure in that post-4h run is not this phase's. The Debian 13
tier lane failed on
test_no_unbudgeted_fixed_rate_database_polling, which found four
undeclared fixed-rate polls above its 0.25/s ceiling --
GetBlobAttributes/queues, UpdateBlobTransfer/transfers,
UpdateBlobLastUsed/queues and GetObjectsByState/queues. The same
lane failed the same way in run 36806873191, before 4h merged, so it
predates this phase and is recurring rather than a flake. It is blob and
transfer traffic wanting a database_load_budget.yaml entry, and it
belongs to the database-load work. It has an owner there:
shakenfist#4401, filed automatically from run 36938755972 at 2026-10-02
01:06 UTC, names the same four pairs and the same Debian 13 slim-tier
lane. So this paragraph is a pointer rather than the only
record of a recurring failure on another repository's default branch.
A note for later steps that quote a commit subject. 4h part 1's
brief prescribes the subject Read the cluster image artifact as debian,
not debian-12., which is 57 characters; the commit convention wants 50
or fewer. The step followed the brief, because a plan file carries
Michael's authority and substituting a different subject quietly is
worse than landing a long one. A brief that dictates a subject should
count it first.
Back brief¶
Five things to agree before the remaining steps run, because each is cheap to propose and expensive to redo. The first three were written when the phase was planned. The last two came from the 2026-09-21 amendment and decided what the phase was while shakenfist#4280 was open; both are answered as of 2026-09-24, and both answers are recorded in place rather than deleted, because the reasoning is what a later reader needs when the next blocker lands:
- Decision 4.5, the release-neutral artifact name. It is a
43-site edit in a test suite, and doing it as
debian-13instead is the same edit for a worse result. If the neutral name is wrong, say so before 4f, not after 4g. - Decision 4.2, deleting the retired declaration. The
alternative -- leave
debian-12declared until phase 5 -- is defensible and makes every revert one line. This plan takes the stricter reading because two repositories grew new findings while it waited, and because this repository already made the same call for itself and wrote it into.github/actionlint.yaml:15-20. Agreeing here makes that a fleet convention rather than one repository's habit. - Decision 4.4, the scope. This phase is now two phases' worth of work in one section. Splitting the guest-image half into its own phase is reasonable; what is not reasonable is running phase 5 without it.
- Whether 4f should rename the artifact and change its contents
in one landing. As written it adds the
debianupload sourced from/srv/ci/debian:13, so the rename carries a bookworm to trixie guest change with it. That was fine when the under-cloud was moving in the same phase and a failure had one obvious suspect. With shakenfist#4280 open it conflates two variables in a suite whose failures take an hour to read, and #4280 was found precisely because someone held the guest constant and moved one thing. Both source paths are on the dependencies cache disk (actionsansible/ci-dependencies.yml:186-195), so sourcing the neutral name from/srv/ci/debian:12and moving it to:13in a separate later landing costs one commit and buys a controlled experiment. The plan as amended still says/srv/ci/debian:13; this is the change I would make and did not, because it revises a decision rather than correcting a fact, and 4h and 4j's greps are written against the:13spelling. Priced, so the answer is a decision rather than a flag: 4h's gate, 4j check (5)'s "exactly one line naming/srv/ci/debian:13" and the definition-of-done bullet that pairs with it are all written against:13and move with this choice, decision 4.7's bookworm cache entry stops being a retired path and becomes the one this phase ships, and a later landing is needed to move the artifact's contents from:12to:13once #4280 has been excluded as a suspect. Three cells, one decision, and one extra commit in a phase that does not exist yet. Answered 2026-09-24: no change, the plan keeps/srv/ci/debian:13-- but not for the reason first given. The first version of this answer said the controlled experiment was "not available at any price", because 4f moves the under-cloud default in the same landing. That conflates the two variables the 09-21 amendment was careful to separate. shakenfist#4309 fixed hypervisor provisioning; it says nothing about whether a trixie guest works, and the 09-21 analysis above -- that no canary ever ran one -- is untouched by it. The guest content is still a variable that could be held constant, and this question's proposal is still purchasable at its stated price.
It is not bought because the step ordering already buys most of
it, one step earlier than the question looked. 4f uploads the
trixie artifact under a name nothing reads yet: the
references still say debian-12, which 4f leaves sourced from
/srv/ci/debian:12. So the two runs after 4f move the
under-cloud with the guest held constant -- the run also does
the extra upload, so it isolates the guest and not the run --
which is the definition of done's kerbside and node-lifecycle
bullet, and
4g then adds the guest content for the jobs that read the
renamed artifact.
What that does not cover is recorded in the risks section rather
than argued away here: the four matrix lanes pass base_image
explicitly, so their under-cloud moves in 4g, and for those
lanes 4g changes under-cloud and guest together. The
node-lifecycle run is a data point about the same image in the
same repository, not a substitute for isolating it. 4h, 4j check
(5) and the definition-of-done bullet stay written against
:13, and the fallback if 4g's run is unreadable is this
question's own proposal, which the risks section now states.
5. Where the eight deferred under-cloud sites go. They were
not in this phase any more and not in phase 5, which is
producers. Three options were open: reopen this phase when
#4280 closed, add a phase 4b, or fold them into phase 5's first
step as its entry gate. The third was tempting and wrong -- it
makes a producer phase start with a consumer edit, which is the
exact blurring D4 exists to prevent. Answered 2026-09-24: the
first, reopen this phase. shakenfist#4309 closed the bug
before any step had acted on the deferral, so there was nothing
to migrate between phases and no landed comment to unwrite; the
eight sites go back into 4f and 4g where they were planned, and
phases 5 and 6 are unchanged by any of it. That this was
answerable so cheaply is an accident of timing rather than a
vindication of deferring it -- see the blocker paragraph in the
risks section.
5. Retire the end-of-life producers¶
Closes: private-ci#38, private-ci#40, private-ci#45, 33fl#826.
Depends on: phase 4, complete 2026-10-02. Planning effort: high,
because the three bullets this section carried named two of the four
IMAGE_BUILDS entries that have to go, missed the one phase 3
assigned here by name, and described 33fl's half as hand work that a
tool in that repository has been doing weekly since July.
Status: planned, 2026-10-02. The survey below was run on
2026-10-02, the day phase 4 closed, and it moved the phase in both
directions: the retirement is larger than the section said in
private-ci and in actions, and smaller in 33fl, where the
blocker the issue describes appears to have been automated away two
months ago. It also found that the test suite is not the gate a
sweep of this shape wants it to be -- one test asserts a successful
build of a "known image" with the boundary mocked, so it will keep
passing once the image is not known any more.
One thing this phase cannot do is prove itself with the audit.
eol-distro does not read either list it empties: private-ci is
scoped by only_checks to four plan criteria and sfui-vendor
(scripts/audit/repo.py:88-93), and 33fl is in another
organisation. So compliance.md reads exactly the same before and
after, and a done criterion asking for a green audit would pass
vacuously. Phase 6 is the phase that makes the producers
measurable; until it runs, the evidence here is greps, a test count
and a nightly cycle summary.
What this section asked for, 2026-09-13¶
Left as written, because the survey below contradicts parts of it and the contradiction is the useful record:
- private-ci#40,
debian-11: remove fromIMAGE_BUILDSandCI_IMAGES. No workflow requests it, so this can go as soon as phase 1 lands -- it also removes the permanentFalsein every nightly cycle summary. - private-ci#40 follow-up,
debian-12anddebian-12-docker: only after phase 4, per D4. Amended 2026-09-21 and again 2026-09-24: for three days this also waited on shakenfist#4280, because phase 4 was going to close with eight under-cloud sites still namingci-images/debian-12and removing the producer while they did is the breakage D4 exists to prevent. shakenfist#4309 closed that bug, phase 4 moves all eight in 4f and 4g, and the gate is "phase 4 complete" again -- specifically, 4j reporting every check agreeing. Thedebian-11bullet above is unaffected and has been unblocked since phase 1 completed. - 33fl#826: delete the two GitLab static runner instances so
static_runner.ymlrebuilds them ondebian:13, and confirm the six GitHub runners have rolled over. Consider a retire tool so this is not manual next time.
What the survey found¶
Checked 2026-10-02 against shakenfist/private-ci at master
(8816709), shakenfist/actions at main (d72644f),
shakenfist/shakenfist at develop, shakenfist/development at
main (72d863e) and the 33fl working copy at a7d7803c.
-
Four
IMAGE_BUILDSentries have to go, not two.conductor/imagebuilder.pystill carriesdebian-11,debian-12,debian-12-dockeranddebian-gnome-12. The fourth is not an oversight in the tree, it is an omission in this section: decision 3.5 assigned it here by name -- "its fourth checkbox -- retiredebian-gnome-12once nothing consumes it -- is phase 5's work" -- and private-ci#45 is open with three of its four boxes ticked for exactly that reason. private-ci#40 does not cover it either; its fourth checkbox names onlydebian-12. So without this bullet the phase closes with an end-of-life desktop image building nightly and an issue that cannot be closed. -
CI_IMAGESholds three of the four, and the two lists are not the same set.conductor/provisioner.py:45-100listsdebian-11,debian-12anddebian-12-docker; it has nodebian-gnome-12, because the desktop images are built for nested CI clusters to consume rather than for runners to boot (imagebuilder.py:55-57). A unit test asserts everyCI_IMAGESlabel has anIMAGE_BUILDSentry and not the converse (private-ci's ownAGENTS.md:259at8816709, and the test isMatrixTestCase.test_every_runner_label_has_a_buildinconductor/tests/test_imagebuilder.py-- named rather than cited by line, because the line moves and the name does not), so the asymmetry is intended and the edit is four entries in one file and three in the other, not seven in both. -
Removing all seven entries fails 22 tests, and every one of them is fixture-shaped. Measured rather than predicted, on a clone of
masterwith the entries deleted and nothing else changed:pytest conductor/testsreports 1267 passed, 2 skipped before and 22 failed, 1245 passed, 2 skipped after. The failures are confined to three files --test_provisioner_claims.py(10),test_imagebuilder.py(9) andtest_provisioner_create_workers.py(3) -- and they fail the same way:create_workers()walksCI_IMAGES, so a queue asking for a label that no longer exists provisions nothing and the assertion reads2 != 0orno claim requested for size 'xs'. None of them is a logic failure and none of them wants a code change.
Two constraints on the sweep, both from reading the tests rather
than the failures. test_provisioner_create_workers.py:256-263
needs two different surviving labels and depends on their
order in CI_IMAGES -- its comment says "two runners are two
different labels; CI_IMAGES is walked in order, and debian-11
comes before debian-12" -- so renaming both fixtures to one
label silently removes what the test covers. And the cheap fix
of inserting a synthetic debian-12 entry into the lists in
test setup would make all 22 pass while re-creating the coupling
this phase exists to remove.
-
The test run is not the gate, and one test proves it.
test_web.py:176-182posts{'name': 'debian-12'}to/api/build-imagewithimagebuilder.request_buildmocked to returnTrue, and asserts a 200 and{'requested': 'debian-12'}. The realrequest_buildreturnsFalsefor a name not inIMAGE_BUILDS(imagebuilder.py:637), so after this phase the test calledtest_build_known_imageasserts a successful build of an image that is not known, and it does so without failing. Four more files name a retiree and do not fail --test_db.py(5 lines),test_db_costs.py(2),test_provisioner_costs.py(7) andtest_staticrunners.py(2) -- but those are opaque strings in arunner_oscolumn or an instancemetadatadict, where any label would do.test_web.pyis the one where the fixture is load-bearing and the mock hides it. A sweep driven by the failing list misses it. -
ansible/ci-image.ymlinactionscarries the same dead branch phase 3 deleted fromci-dependencies.yml. Fourwhen:lines --:49,:60,:455and:466-- gate two pairs ofadd_hosttasks onbase_image == "debian:11"versus!= "debian:11", one pair for the rebuild host and one for the test host, differing only in whetheransible_python_interpreteris forced to/usr/bin/python3.IMAGE_BUILDSis the only caller of that playbook -- no workflow inactionsinvokes it and neither does 33fl -- so once thedebian-11entry is gone all four conditions are dead and each pair collapses to one task with nowhen:. That is decision 3.3's reasoning applied to the file phase 3 did not reach, and decision 3.2's ordering applies too, in the same direction: whiledebian-11is still building, deleting the branch sends bullseye down the auto-detect path, which is the quirk the branch exists for. -
The frozen cache entries are not serving anybody, which is a defect rather than an argument for deleting them. Decision 4.7 left
ubuntu:20.04,debian:11,debian:12andfedora:40inci-dependencies.yml's cached image list and said "retirement is phase 5's and private-ci#38's", on the grounds that the cache is what lets a test boot an old guest deliberately. A test does:shakenfist/deploy/ansible_module_ci/004.yml:120and005.ymlcreate instances from10@debian:11, and 005's idempotency assertion is the regression coverage for shakenfist#3669. But it does not boot it from the cache.build-smoke-clusteruploads exactly one cached image into the nested cluster (action.yml:257, thedebianartifact 4h left behind), no topology playbook configures an image mirror, and 004.yml's own comment says the create "pays a cold ~407 MiB internet fetch inside its await budget" which "has been observed to take 578 seconds (issue 4000)" against a 600 second timeout. So the bytes are on the attached/srv/cidisk and the cluster fetches them over the internet anyway. That is a performance defect in how the cache is used, not a retirement, and decision 5.4 keeps the entries and files it. -
33fl's half may already be done, and three statements say it cannot be. 33fl#826 says the two GitLab runners "have no retire tool, so nothing will ever recreate them", and the comment above
static_runner_debian_releaseingroup_vars/all/static_runners.yml:52-53says the same, and this section repeats it as "consider a retire tool so this is not manual next time". All three are stale.tools/retire-gitlab-runners.pyhas been in that repository since743a0200on 2026-07-30 -- before the issue was filed and six weeks before this plan was written -- it pauses the server-side runner record, drains, deletes the backing Shaken Fist instance and resumes, andrundaily.sh:262runs it inside the--weeklyblock besideretire-github-runners.py.static_runner_debian_releasebecame 13 infa0e79d8on 2026-09-12. So if a weekly cycle has run since, both GitLab runners and all six GitHub runners have been rebuilt on trixie already and 33fl#826's work is observation and three text corrections. Not asserted here:--weeklyis operator-invoked and the survey found no log of it reaching Loki ({job="ansible-deploy"}carries the deploy phases, notrundaily's retire lines), so step 5c reads the fleet first and deletes only if the read says bookworm. -
Two more stale sentences, one of them in this repository.
docs/audits/eol-distro.md:57-63saysdebian-gnome-12"is listed although the CI conductor advertises nodebian-gnome-13label yet" and that "what is missing is an entry in private-ci'sIMAGE_BUILDStable, not an image". Phase 3 added that entry, so the paragraph now tells a reader to file a request for something that exists. Andconductor/imagebuilder.py:195-198explainsSTALE_LABEL_SECONDSby naming "the debian-11 case in #40", which stops being an example the moment this phase lands. -
Nothing else in the fleet names the labels. Corroborating 4j rather than repeating it:
/srv/ci/debian:12appears nowhere inactionsexcept the single/srv/ci/debian:13line 4h left, and the onlydebian-12left inprivate-cioutsideconductor/is the declaration in.github/actionlint.yaml, which phase 4's done criterion deliberately exempted as the one repository with no workflows for a declaration to govern. It goes with the entries, per decision 4.2's reasoning.
Scope¶
In: private-ci's two lists, their test fixtures and that
repository's actionlint.yaml; the dead debian:11 branches in
actions/ansible/ci-image.yml; 33fl's static runner fleet state and
the three stale statements about it; the stale paragraph in this
repository's docs/audits/eol-distro.md; and closing private-ci#38,
40, #45 and 33fl#826.¶
Out, and why:
shakenfist/imagescontinuing to publishdebian:11,debian:12anddebian-gnome:12. D5 -- this plan does not re-plan that repository, and private-ci#40 says the published images stay because they are useful for testing older guests. Retiring a runner label is not retiring an image.- Teaching the audit to read
IMAGE_BUILDSandCI_IMAGES. Phase 6, and Q1 is the open question about which instrument does it. - The cache-fetch defect in finding 6. Filed in step 5e, not
fixed: it is a change to how a nested cluster resolves guest
images, it touches the action every consumer pins at
@main, and it has nothing to do with retiring a producer. rocky-9. Not on the end-of-life table this plan works from.- Pruning or regenerating
REVIEWS.mdin any repository, per the phase landing shared block inPLAN-TEMPLATE.md. Editingdocs/audits/eol-distro.mdin step 5d will invalidate that file's review mark; CI prunes it on the default branch and no step prunes it by hand.
Decisions¶
5.1. debian-gnome-12 is retired here, with the Debian 12 runner
labels. Decision 3.5 said so and this section did not, which
is the omission finding 1 records. The consequence is that
private-ci#45 closes in this phase, as 3.5 promised, and that
the fixture sweep in finding 3 covers three labels rather than
two.
5.2. One pull request per repository, and two commits inside
private-ci's. The repository rule is inherited from decision
4.1 and for the same reason. The split inside it is new: the
first commit removes debian-11 and debian-gnome-12, which
nothing in the fleet has ever requested, and the second removes
debian-12 and debian-12-docker, which phase 4 spent itself
clearing. If the Debian 12 half has to come back -- a consumer
the audit cannot see, a static runner, a repository outside the
matrix -- the revert is one commit and does not take the
uncontroversial half with it. The fixture sweep splits along the
same line, which is why this is two commits rather than two
pull requests: the two halves share
test_provisioner_create_workers.py and reviewing them apart
would mean reviewing that file twice.
5.3. The 22 failures are fixed by renaming fixtures to surviving
labels, and test_web.py is renamed although it does not
fail. Two surviving labels, not one, because of the ordering
dependency in finding 3. No synthetic entries in test setup. And
the done criterion is the test count, not a green run: no
fewer passed and the same skipped, so a sweep that deletes
coverage to make the suite pass fails the criterion. No fewer
rather than exactly equal, because there is one test worth adding
-- that request_build returns False for a retired name, which
is the hole finding 4 found -- and a criterion demanding equality
would argue against writing it. Any added test is named in the
commit message so the rise is accounted for. The
reference point is the parent of 5a's merge commit, measured at
the time, not a number written down here. master moves while
a phase runs -- it was 8816709 when the survey ran and
bc00be0 hours later -- so an absolute target would fail the
criterion for an unrelated commit that added a test, or hide
coverage this phase deleted behind one that did. 1267 passed and
2 skipped is what 8816709 gave on 2026-10-02, recorded so the
expected magnitude is known. This is the decision that matters
most to get right, because the cheap alternatives all leave a
green suite.
5.4. The four frozen cache entries stay, and the reason is
written down rather than inferred. Decision 4.7 deferred them
here. Keeping them costs one download per nightly dependencies
build; removing them would foreclose the fix finding 6 points
at, because an image that is not cached cannot later be served
from the cache. So the entries stay and each gains a comment
saying it is deliberately frozen and what would justify removing
it. This is the decision most likely to be argued with: an
unreferenced end-of-life image in a cache looks exactly like
what this plan is about. The answer is that this plan is about
producers of runners -- an image nothing boots cannot run a
job on an unpatched kernel -- and that debian:11 is not
unreferenced anyway: shakenfist#3669's regression coverage boots
it, over the internet, which is the defect rather than the
justification.
5.5. 33fl is read before it is edited, and the read may close the issue without a deletion. Finding 7 is three stale statements deep, so step 5c's first act is to ask the fleet what release it is on rather than to delete two instances on the strength of an issue written in September. If they are already on trixie, the step corrects the three statements and reports the evidence; if they are not, it deletes them, waits for the reconcile, and then does the same. Either way 5c does not close #826 -- 5e does, with 5c's evidence, because closures are gathered in one step.
5.6. private-ci#38 closes after the phase is confirmed, not after the last pull request merges. It is the collated inventory and it closes when the thing it inventories is gone, which phase 4's status section said when it deferred it here. "Gone" is what 5f establishes, so #38 closes as 5f's final action rather than in 5e, and the "after" half of its comment quotes 5f's output. It was 5e when the phase was first written, which would have closed the inventory before anything confirmed the phase -- the defect 4i committed on development#123, reporting a compliance page as current status when it predated the merges it was describing.
Step plan¶
Every step that edits a repository other than this one opens a pull
request there and waits for that repository's own CI. Among the four
editing steps, 5a gates 5b and nothing else: 5c and 5d are independent
of both and of each other. The two observation steps are not free of
order either -- 5e follows all four, because it closes issues with
their merge commits, and 5f follows 5e and at least one nightly cycle
after 5a. No step prunes or regenerates REVIEWS.md.
| Step | Effort | Model | Isolation | Brief for sub-agent |
|---|---|---|---|---|
| 5a | high | opus | worktree | The phase's substance, in shakenfist/private-ci, one pull request and two commits (decision 5.2). Commit one removes the debian-11 and debian-gnome-12 entries from IMAGE_BUILDS in conductor/imagebuilder.py and the debian-11 entry from CI_IMAGES in conductor/provisioner.py (:45-100; there is no gnome entry there, see finding 2), and rewords the STALE_LABEL_SECONDS comment at :195-198 which explains itself by naming "the debian-11 case in #40". The reworded comment must not put any retired label in single quotes -- name them bare or not at all. 5b's ordering gate and 5f check (1) both grep this file for 'debian-11' and friends, so a comment keeping 'debian-11' as a historical example makes both report a hit forever after. A false halt rather than a missed regression, but one that would be diagnosed at 5b rather than here. Commit two removes debian-12 and debian-12-docker from both lists and deletes debian-12 from .github/actionlint.yaml (finding 9; decision 4.2's reasoning). The test sweep is the work, not the deletion. Baseline first, on your own branch point, and record the numbers: pytest conductor/tests gave 1267 passed, 2 skipped on 8816709 on 2026-10-02, and if your branch point gives something else then that is your target rather than this -- then expect 22 failures in test_provisioner_claims.py, test_imagebuilder.py and test_provisioner_create_workers.py, every one of them a fixture naming a label that no longer provisions. Rename fixtures to surviving labels; do not add synthetic entries to the lists in test setup, and do not delete a test to make the count go green. test_provisioner_create_workers.py:256-263 needs two different surviving labels and depends on their order in CI_IMAGES, so read its comment before choosing: debian-13 and debian-13-docker satisfy it, one label used twice does not. Then sweep the files that do not fail: test_web.py:176-182 asserts a 200 for a "known image" with request_build mocked, so it keeps passing while asserting something false (finding 4) -- rename it; test_db.py, test_db_costs.py, test_provisioner_costs.py and test_staticrunners.py use the label as an opaque runner_os or metadata string, so rename them for honesty and say in the commit message that nothing there was load-bearing. Finish with no fewer passed than your baseline and the same skipped; if you add a test -- request_build returning False for a retired name is the one worth adding -- name it in the commit message so the count rising is accounted for. Verify with pre-commit run --all-files and tox -e flake8. Commit subjects: Retire the debian-11 and gnome-12 images. and Retire the Debian 12 runner images. Do not touch ansible/ in any repository -- that is 5b. |
| 5b | medium | sonnet | worktree | After 5a has merged and one nightly cycle has built with the shortened list, and gated on a grep rather than on this sentence: git grep -n "'debian-11'" conductor/imagebuilder.py in a fresh private-ci clone must return nothing, and the conductor's cycle summary for the night after must show no debian-11 build. That grep is known to fire rather than assumed to: on 8816709, the survey commit, it returns two lines -- the entry's name at :59 and its label at :62, the base_image spelling debian:11 deliberately not matching -- so zero means 5a landed rather than meaning the pattern was wrong. The quoted-label form is chosen over one naming the 'name': key for the same reason: a pattern tied to the dict layout returns nothing if the layout is ever reshaped, which is a gate that cannot tell the two worlds apart. Then, in shakenfist/actions, delete the four dead debian:11 conditions in ansible/ci-image.yml (:49, :60, :455, :466) and collapse each pair of add_host tasks into one. Two pairs: the rebuild host at :39-60 and the test host at :445-466. They differ only in ansible_python_interpreter: /usr/bin/python3, which the surviving task must not carry -- the detect-system-python arm is the one that stays, which is what the != "debian:11" condition meant. IMAGE_BUILDS is the only caller of this playbook, so after 5a nothing invokes it with base_image: debian:11; this is decision 3.3's reasoning applied to the file phase 3 did not reach, and decision 3.2's ordering is why it cannot land first. Afterwards grep -n 'debian.11' ansible/ci-image.yml returns nothing and grep -c 'name: Add to ansible' ansible/ci-image.yml returns 2, which is the check that catches deleting the wrong arm of a pair. That target is anchored rather than asserted: the same grep returns 4 on d72644f, at :39, :51, :445 and :457, so two afterwards is one surviving task per pair. Run tools/ansible-syntax-check.sh and pre-commit run --all-files. The conductor reads this repository at main (imagebuilder.py:41, ACTIONS_BRANCH = 'main'), so the change is live on merge with no pinning to protect you -- the same flag-day property 4h had. Second commit in the same pull request: in ansible/ci-dependencies.yml, add a comment to each of the four frozen cached image entries -- debian:11, debian:12, ubuntu:20.04 and fedora:40 -- saying it is deliberately frozen per decision 5.4 and naming what would justify removing it (nothing boots it any more, or the cache starts serving reads so that keeping it stops being free). Remove no entry, and do not touch debian:13, rocky:10, ubuntu:22.04, ubuntu:24.04 or centos:9-stream. This is the only step that writes those comments, which 5f check (4) and the sixth done bullet both require; it is here rather than in 5a because 5a opens no pull request against actions. Nothing about the comments depends on 5a, but they wait for it anyway because they share this step's pull request and decision 5.2 keeps one pull request per repository. That costs nothing: 5f needs 5a and a nightly cycle regardless. An earlier draft called this half "gate-free", which was wrong -- the whole of 5b is gated. Commit subjects: Drop the dead bullseye python branches. and Say which cached images are frozen. |
| 5c | medium | sonnet | worktree | 33fl, and read before you edit (decision 5.5). The premise of 33fl#826 is that the two GitLab static runners have no automated path off bookworm. tools/retire-gitlab-runners.py is that path, it has existed since 2026-07-30, and rundaily.sh:262 runs it in the --weekly block; static_runner_debian_release became 13 on 2026-09-12 (fa0e79d8). So first establish what the eight runners in static_runner_fleet are actually running -- the static-runners namespace on sfcbr, via sf-client instance list and the instance creation times, or ansible_distribution_release from a one-off fact gather -- and report it before changing anything. If all eight are on trixie: delete nothing, and fix the three stale statements -- the comment at group_vars/all/static_runners.yml:52-53 ("The gitlab runners have no retire tool") must name the tool and say the roll happens on the weekly cycle. If the two GitLab instances are still bookworm, delete them so the next reconcile rebuilds them, wait for it, then make the same correction. Either way leave static_runner_debian_release: 13 alone and leave the paragraph about the audit's blindness to this fleet alone -- it is still true and phase 6 is where it goes. Do not close #826; 5e does that. Commit subject: Say the gitlab runners do roll over. |
| 5d | low | sonnet | worktree | One paragraph in this repository. docs/audits/eol-distro.md:57-63 says debian-gnome-12 "is listed although the CI conductor advertises no debian-gnome-13 label yet" and that "what is missing is an entry in private-ci's IMAGE_BUILDS table, not an image". Phase 3 added that entry, so rewrite the paragraph to say the successor label exists and that a finding naming debian-gnome-12 is now a request to retire the old entry rather than to add a new one. Leave the end-of-life table at :31-32 exactly as it is -- it is the registry of banned labels and it must keep listing all five of debian-11, debian-11-docker, debian-12, debian-12-docker and debian-gnome-12 after this phase, because its job is to name what a workflow may not ask for. Five banned labels against four retirements is not an inconsistency to resolve: debian-11-docker was never produced, and debian-gnome-12 is a desktop image rather than a runner boot image. Do not touch FROZEN_METADATA or FROZEN_ISSUE_TITLES; the criterion's id, spec path and issue title do not change. Do not prune REVIEWS.md. Verify with pre-commit run --all-files. Commit subject: Say debian-gnome-13 exists. |
| 5e | medium | sonnet | none | Housekeeping, no commit, but everything it writes is outward-facing and hard to retract, which is why it is not the mechanical pair its first draft had: 4i ran low/haiku and produced two defects in one closing comment -- it double-counted a set of pull requests, and reported a compliance page as current status when that page predated the merges it was describing. File one issue in shakenfist/shakenfist: the nested cluster fetches debian:11 from images.shakenfist.com although the bytes are on the attached /srv/ci cache disk, costing a cold ~407 MiB download inside ansible_module_ci/004.yml's await budget, observed at 578 seconds against a 600 second timeout (shakenfist#4000). Give it the evidence from finding 6 -- build-smoke-cluster/action.yml:257 uploads one cached image and no topology playbook sets a mirror -- and say explicitly that decision 5.4 keeps the cache entries and that this is about how they are served, not whether they exist. Then close private-ci#40, private-ci#45 and 33fl#826 with their merge commits. Leave private-ci#38 open: it closes in 5f, after the phase is confirmed, because its closing comment asserts an "after" state and nothing has checked that yet (decision 5.6). Do not claim the audit confirms any of this -- eol-distro does not read either list, and saying otherwise is the defect 4i produced in development#123's closing comment. Never write a bot trigger phrase in any comment. |
| 5f | medium | sonnet | none | Confirms the phase, which nothing else does, and cannot lean on the audit (see the status paragraph). Run after every pull request above has merged and at least one nightly cycle has completed. Every check below, not a counted subset -- this brief said "six" while listing ten for one round of review, which is enough for a step to stop at (6) and skip the only check the plan says has no automatic backstop: (1) in a fresh private-ci clone, grep -n "'debian-11'\|'debian-12'\|'debian-12-docker'\|'debian-gnome-12'" conductor/imagebuilder.py conductor/provisioner.py returns nothing, and grep -rn 'debian-12' .github/ returns nothing; (2) pytest conductor/tests on 5a's merge commit gives no fewer passed and the same skipped as on its parent -- run it on both, because master moves and an absolute figure drifts; fewer passed means coverage was deleted, and more is fine if 5a's commit message names what it added. For scale, 8816709 gave 1267 passed and 2 skipped; (3) in actions, grep -n 'debian.11' ansible/ci-image.yml returns nothing, grep -c 'name: Add to ansible' ansible/ci-image.yml returns 2 against 4 on d72644f, and tools/ansible-syntax-check.sh passes; (4) the cached image list in ansible/ci-dependencies.yml still contains debian:11, debian:12, ubuntu:20.04 and fedora:40, each carrying the freeze comment -- a check whose passing output is non-empty, deliberately (decision 5.4); (5) the conductor's nightly cycle summary for a night after 5a shows builds for the surviving labels only, with no debian-11, debian-12, debian-12-docker or debian-gnome-12 line and no permanent False, read from the summary rather than inferred from the absence of a failure; (6) all eight entries of static_runner_fleet report Debian 13, read from the fleet; (7) in private-ci, grep -n "'debian-1[12]'\|'debian-12-docker'\|'debian-gnome-12'" conductor/tests/test_web.py returns nothing, and test_build_known_image names a label that IMAGE_BUILDS still contains -- this is the check with no automatic backstop, because request_build is mocked there and the suite passes either way (finding 4). The grep is known to fire: on 8816709 it returns four lines, :179, :181, :182 and :195, so zero means 5a swept the file rather than meaning the pattern was wrong; (8) in 33fl, grep -n 'no retire tool' group_vars/all/static_runners.yml returns nothing and grep -n 'retire-gitlab-runners' group_vars/all/static_runners.yml returns a line; (9) in this repository, grep -n 'advertises no' docs/audits/eol-distro.md returns nothing and its end-of-life table still lists all five banned labels; (10) compliance.md's eol-distro section is byte-identical to the same section on the commit before 5a merged, ignoring the *Generated ...* line -- the phase changes no verdict, and this check is what distinguishes "unchanged" from "nobody looked". Report each check with its output. A stale read looks healthy here too: check (5) wants a cycle summary generated after 5a merged, not the most recent one you can find. Then, and only if every check above agrees, close private-ci#38 (decision 5.6) with the before-and-after comment: what the fleet named when that inventory was written against what it names now, the "after" half quoting those outputs. If any check disagrees, leave it open and report. |
Risks and mitigations¶
A label nothing appears to request may still be requested.
Phase 4 cleared every runs-on: reference the fleet has, verified
across twenty-nine clones, and private-ci#40 says no workflow has
ever asked for debian-11. Both are greps of repositories, and a
static runner advertises only self-hosted and static, so a job
running on one names no operating system at all -- which is the
blindness 33fl#826 records. Mitigated by what the failure looks
like: a workflow asking for a label the conductor no longer offers
does not fail, it queues, so the signal is jobs pending with no
runner rather than a red check. 5f check (5) reads the cycle
summary, and the operator-visible symptom is a queue that does not
drain. Decision 5.2 is the other half of the mitigation: the
Debian 12 half is its own commit, so a revert is one commit.
The test sweep can be satisfied by removing coverage. Twenty-two
failing tests and a deadline is how a suite loses assertions. The
mitigation is a count rather than a judgement: no fewer passed and
the same skipped on 5a's merge commit as on its parent, named in 5a's
brief and re-checked in 5f check (2). Against the parent rather than
against 1267, because master moves while the phase runs. The
subtler version is the mocked boundary in finding 4,
which no count catches -- that one is mitigated by naming the file
and the line in the brief, because nothing else would find it.
actions is live on merge. The conductor reads
ACTIONS_BRANCH = 'main', so 5b's playbook edit reaches the next
image build immediately, and every consumer of
build-smoke-cluster pins @main as well. This is the property
that turned 4h into two ordered pull requests. Mitigated by 5b's
ordering gate being a grep of the merged private-ci tree plus a
nightly cycle, not a sentence in this plan, and by the surviving-arm
check (grep -c 'name: Add to ansible' returns 2) being mechanical.
The 33fl read may be unavailable when 5c runs. It needs the
static-runners namespace on sfcbr to answer, and the survey could
not confirm from this host that rundaily.sh --weekly has run since
2026-09-12. If the read cannot be made, 5c stops and reports rather
than deleting instances on the strength of a stale issue -- deleting
a runner that is already on trixie costs a rebuild and an idle
queue for no gain.
This phase cannot be confirmed by the audit. Stated in the
status paragraph because it is the structural risk, not a caveat:
compliance.md reads the same before and after, so every check in
5f is a grep, a count or a log read. The phase that closes this gap
is phase 6, and the ordering is deliberate -- D4 retires producers
before they are measured, which means the measurement cannot be the
evidence that the retirement worked.
Definition of done¶
Thirteen bullets, and each one names its owner, because phase 4 closed with four bullets no step had checked and that is recorded a thousand lines above. 5f's ten checks cover eleven of them: bullets 1 and 2 are 5f (1), bullet 3 is 5f (2), bullet 5 is 5f (3), bullet 6 is 5f (4), bullet 7 is 5f (5), bullet 8 is 5f (6), bullet 4 is 5f (7), bullet 9 is 5f (8), bullet 10 is 5f (9) and bullet 13 is 5f (10). Bullet 12 and the first half of bullet 11 -- the three issues closed with their merge commits -- are 5e's own actions, verified by their being done. The private-ci#38 half of bullet 11 is 5f's final action, after its checks agree, per decision 5.6. If a bullet is added to this list, 5f gets a check for it in the same commit -- the mapping is the mechanism rather than the decoration. Four of these bullets had no check when the phase was first written, which is the same gap phase 4 closed with, and it was found here only by counting them against 5f (development#207).
- In a fresh
private-ciclone,conductor/imagebuilder.py'sIMAGE_BUILDScontains nodebian-11,debian-12,debian-12-dockerordebian-gnome-12entry, andconductor/provisioner.py'sCI_IMAGEScontains none of the first three. Four entries and three entries, not seven and seven: the desktop images are not runner boot images (finding 2). -
grep -rn 'debian-12' .github/inprivate-cireturns nothing. This is the one repository phase 4's equivalent bullet exempted, because it has no workflows for the declaration to govern; the declaration goes with the entries. -
pytest conductor/testson 5a's merge commit gives no fewer passed and the same skipped as on its parent. A green run with fewer passed fails this bullet; more passed is fine if 5a's commit message names what was added. Measured against the parent rather than against a figure written here, becausemastermoves while the phase runs; for scale,8816709gave 1267 passed and 2 skipped on 2026-10-02. The count is the criterion precisely because 22 tests fail on the deletion alone and the cheapest ways to make them pass are all wrong (decision 5.3). -
conductor/tests/test_web.py'stest_build_known_imagenames a label thatIMAGE_BUILDSstill contains. It does not fail when it stops doing so --request_buildis mocked -- which is why it is a bullet of its own rather than part of the one above. - In
shakenfist/actions,grep -n 'debian.11' ansible/ci-image.ymlreturns nothing andgrep -c 'name: Add to ansible' ansible/ci-image.ymlreturns 2. The second half catches the inverse mistake of keeping the force-python3 arm instead of the detect arm, which a grep for the condition alone cannot see.tools/ansible-syntax-check.shpasses. -
ansible/ci-dependencies.yml's cached image list still containsdebian:11,debian:12,ubuntu:20.04andfedora:40, and each carries a comment saying it is deliberately frozen and what would justify removing it. A grep whose passing output is non-empty, for the same reason phase 3'sdebian:11bullet was: a bullet asking for nothing would be satisfied by deleting the thing decision 5.4 keeps. - A conductor nightly cycle summary generated after 5a merged
lists builds for the surviving labels only, with no
debian-11,debian-12,debian-12-dockerordebian-gnome-12line, and no permanentFalse. Read from the summary, not inferred from the absence of a failure issue. - All eight entries of
33fl'sstatic_runner_fleetreport Debian 13, read from the fleet rather than fromstatic_runner_debian_release. The variable has said 13 since 2026-09-12 and said nothing about what is running. -
group_vars/all/static_runners.ymlno longer says the GitLab runners have no retire tool, and namestools/retire-gitlab-runners.pyand the weekly cycle instead. The paragraph abouteol-distrobeing unable to see this fleet stays: it is still true and it is phase 6's. -
docs/audits/eol-distro.mdno longer says the conductor advertises nodebian-gnome-13label, and its end-of-life table still lists all five labels it bans --debian-11,debian-11-docker,debian-12,debian-12-dockeranddebian-gnome-12, which is not the set of four entries this phase retires:debian-11-dockerhas never been produced anddebian-gnome-12is a desktop image rather than a runner boot image. Both halves: the table is the registry of what a workflow may not name, and emptying it as the labels go would make the criterion unable to report a regression. - private-ci#40, private-ci#45 and 33fl#826 are closed with their merge commits, in 5e. private-ci#38 is closed by 5f, after every one of its checks agrees, with a comment whose "after" half quotes their output -- not in 5e, because that comment asserts a state only 5f has established (decision 5.6).
- An issue exists in
shakenfist/shakenfistfor the cache that is not serving reads (finding 6), namingbuild-smoke-cluster/action.yml:257,ansible_module_ci/004.yml:120and shakenfist#4000. -
compliance.md'seol-distrotable is unchanged by this phase. Stated as a done criterion rather than omitted, because "the audit went green" is the evidence a reader will reach for and it is not available here: neither list this phase empties is read by any criterion until phase 6.
Back brief¶
Four things to agree before 5a starts, because each is cheap to settle now and expensive to redo.
debian-gnome-12is in this phase (decision 5.1). If it is not, private-ci#45 cannot close here, decision 3.5's promise moves again, and 5a's fixture sweep splits differently -- the gnome fixtures intest_imagebuilder.pywould have to stay. This is the finding that changed the phase's size, so it is the first thing to confirm.- The frozen cache entries stay (decision 5.4). The opposite
call is defensible and it is a four-line deletion, but it
forecloses the fix in finding 6 and it contradicts decision 4.7,
which deferred the question here rather than deciding it. If they
are to go, say so before 5b, because the comment it writes is the
thing that would have to be unwritten. (5b's second commit, not
5a's: 5a opens no pull request against
actions. This said 5a when the phase was first written, which left the comments with no owning step at all; development#207 is where that was caught.) - Two commits in
private-ci, not two pull requests (decision 5.2). The revert granularity is the point; the cost is one review covering both halves. - 5c reads the fleet before deleting anything (decision 5.5). The alternative is to follow 33fl#826 as written and delete the two GitLab instances, which costs a rebuild if they are already on trixie. This is only a gate because the issue, the config comment and this plan all say something the repository's own tooling contradicts.
6. Close the audit's blind spot¶
Closes: nothing filed. Depends on: phases 3-5, so the producers are compliant before they are measured.
The phase that stops the 80 findings coming back. Per the structural finding above, all three producers sit outside the criterion that bans what they produce.
- shakenfist/images: decide whether it joins the audit matrix,
and on what terms. Scope is stated in three places and a change
has to make all three agree in one commit, or
scope-coveragereports the repository as decided nowhere andAuditScopeIsStatedOnceTestfails: therepo:matrix in.github/workflows/consistency-audit.yml, which is what actually runs, and the in-scope and excluded lists indocs/audits/README.md, which are what a reader is told.REPO_OVERRIDESinscripts/audit/repo.pyis not a scope statement -- it narrows a repository already in the matrix -- so it is only edited if images is to be scoped to a subset.
Joining is otherwise all-or-nothing: in the matrix means all 52
criteria apply. Measured against the current clone on
2026-09-13, images would report 8 fail, 6 pass, 38 not
applicable. The eight are llm-context-lint-ci, renovate,
ci-review-automation, pre-commit-config, export-repo-config,
default-branch-naming, github-security and
delete-branch-on-merge. That is a morning of consistency
issues rather than a wall, and every one of them is already an
item in phase 5 of that repository's own plan, so the decision
is whether to file them or do them first -- not whether the
repository can survive being measured. Re-measure before acting;
this number is from before phase 5 runs.
* private-ci: extend only_checks to cover a criterion that
reads IMAGE_BUILDS and CI_IMAGES for end-of-life bases. It
is currently scoped to four plan criteria and sfui-vendor --
five only_checks entries in total.
A partial scope is stated a fourth time, as the sentence in the
partial-scope paragraph of docs/audits/README.md that
scripts/audit/scope.py parses and holds against only_checks.
Its docstring records that this is the statement with the worst
track record: only_checks was widened once with the sentence,
and two other documents, left behind and no test noticing. So
the same commit updates the sentence in docs/audits/README.md,
the REPO_OVERRIDES comment block in scripts/audit/repo.py,
and the comment above - private-ci in the audit matrix -- which
is already stale, still reading "Scoped to the sfui-vendor
check".
* 33fl: outside this organisation and this tooling. #826
records the exposure in group_vars/all/static_runners.yml
where the next reader will find it, which is the available
answer; note here that a static runner fleet is structurally
invisible to a label-based audit, so any future fleet of the
same shape needs the same treatment.
And images is not the only repository outside the matrix.
Eight non-archived repositories are outside it: images itself,
plus client-python-ova, divergulent-reviews, homebrew-tap,
performance, reproducables, sonobouy and
uefi-latency-guest. None of them is invisible -- every one is on
the excluded list in docs/audits/README.md, and scope-coverage
reconciles that list and the matrix against the organisation every
morning, which is the criterion written precisely so that a
repository in neither list stops being silent. What is true is
narrower: an excluded repository gets no verdict from any
criterion, and these exclusions were decided before several of the
criteria that would now apply to them existed, eol-distro
included. Phase 4's survey ran the criterion over all eight by hand,
and the result narrows the question further than expected: seven of
them satisfy all three of the criterion's skip conditions -- no
.github/workflows/, no container build file anywhere in the tree
and no top-level templates/ directory -- so eol-distro would
report not applicable for every one of them even if they were in
the matrix. All three were checked, not just the first: EolDistro
skips only when the three are empty together
(scripts/audit/checks/distros.py:450). images is the only one of
the eight where the exclusion costs a verdict, and it is the one this
phase was already about. So phase 6 is revisiting a recorded
decision rather than discovering an omission, and the other seven
need a sentence in docs/audits/README.md saying they are
excluded for having nothing to audit rather than for the reasons
the list currently gives -- which is a smaller and more honest
change than adding them. That sentence goes in as prose. The
excluded list sits inside the span scripts/audit/scope.py parses
between the literal phrases are **excluded** and The `actions`
repository, and bulleted_block() collects only *-prefixed
lines: a plain prose line is ignored and is safe, but the same
sentence written as a bullet raises ScopeParseError for not being
a repository name, and a sub-heading introducing it raises it for
running the list past a heading. Either failure takes scope-coverage
down fleet-wide rather than locally, and AuditScopeIsStatedOnceTest
in scripts/tests/test_registry.py is what will say so at commit
time.
Which instrument does the measuring is Q1 above, and the default there is a separate criterion.
7. Push audit¶
Per the push audit shared block in PLAN-TEMPLATE.md. Runs
PUSH-AUDIT.md over the accumulated diff of every phase against
main, not the diff of the last phase alone.
There are no out-of-repository audits for this phase to cite.
Phases landing elsewhere record <repo> <sha> (#pr), and the
shared block allows citing an audit run as part of the pull request
that landed them. The finding under phase 2 checked all eleven
landings and found that not one carries such an audit, and that
none of the four repositories has a PUSH-AUDIT.md to have run.
An earlier draft of this paragraph asserted the citing as fact;
restating a policy in the past tense is how a plan comes to believe
it has evidence it never collected, which is the defect phase 1's
correction records.
Of the three options that finding names, this phase takes the
first: it runs the accumulated audit itself, once per repository,
against that repository's default branch. Deferring the three
shakenfist/images landings to that repository's own phase 6 would
leave actions#74 and 33fl#836 covered by nothing, and declining the
gap in writing buys nothing when running the audit is the work this
phase exists to do. There is no single accumulated diff spanning
four repositories, so the record must name which repository each
audit covered and the sha it was taken against.
Which runbook, given that none of the four has one. The shared
block's carve-out applies and is invoked here deliberately: a
repository with no PUSH-AUDIT.md still carries the phase, and the
phase says the runbook does not exist yet and what was done instead.
So this phase does not pretend to run a runbook that is absent, and
it does not borrow this repository's. That one is written for this
repository's blast radius -- AGENTS.md here is explicit that a
defect breaks sixteen other repositories quietly rather than
breaking a service -- which is the wrong question to ask of an image
build host or a Grafana repository.
Instead phase 7 carries its own checklist, written into this plan, applied to each of the four repositories' accumulated diff:
- Every landing's recorded sha is the merge commit of its pull request, with two parents, and the diff read is against the first parent. Already done once for all eleven, under phase 2 -- so this step is a re-read only if the Execution table has gained rows since.
- Anything the diff added that runs unattended -- cron, a systemd
timer, a nightly workflow -- either reports its own failure
somewhere a human sees, or is named as not doing so. This is the
check images#8 would have failed: a self-update under
errexitwhose one explanatory line went to a cron mailer that does not exist. - No secret, token or private hostname entered a public repository.
- Anything the diff removed has no consumer left, checked by grep across the four repositories rather than by assumption -- the same check step 3f's cache-disk destination needed.
- For
shakenfist/imagesandshakenfist/private-ci, whether the change can fail closed: a broken build that publishes nothing is recoverable, a broken build that publishes something wrong is not.
Landing a PUSH-AUDIT.md in each of the four repositories would be
the better answer and is explicitly not this phase's work: four
runbooks written for four blast radii is its own plan, and phase 7
is not the place to discover that. Phase 7 files an issue proposing
it, against this repository, and says in the plan that it did.
Audit images#8 first. It is the one landing where this is not bookkeeping: it broke the nightly build on its first night, and an audit of the accumulated diff is precisely the instrument that would have looked at a self-update running under errexit. The rest is catch-up; that one is the demonstration that the catch-up is worth doing.
Only phase 6 and this phase land in this repository, so the
accumulated diff against main here is the small part of the work.
Agent guidance¶
Deliberately short. Six of the seven phases land in other
repositories, under their own conventions and, for
shakenfist/images, its own PLAN-TEMPLATE.md and the
project-specific checks in it. Per-step effort levels, model
recommendations and review checklists belong in the plan of the
repository the work lands in, not here, so this plan does not
carry the shared blocks PLAN-TEMPLATE.md offers for them.
The exception is phase 6, which lands here. It changes what the
fleet is measured against and reaches scripts/audit/, so it is
high effort per this repository's own guidance, and it is
exercised with --dry-run before anything files an issue.
Risks and mitigations¶
The migration is done and the detection is not. The most likely failure of this plan is that phases 3-5 close visible issues, phases 1, 2 and 6 close none, and the plan is declared finished. That is the ordering D1 exists to prevent, and it is why phases 1 and 2 come first despite closing nothing.
Phase 4's label half is long and boring, and that is the half that
moves. Twelve repositories with an actionlint edit each, and the
daily audit files and closes the per-repository issues itself, so
progress is externally visible rather than tracked by hand -- the
fleet cleared 38 of the original 80 references that way before the
phase started, grew three new ones while doing it, and cleared a
thirteenth repository, shakenfist/actions, twenty-five minutes after
this plan merged. The risk moved next door twice. Phase 4's survey
found a second class of consumer the criterion cannot see, and being
boring is what made it easy to believe the criterion's count was the
whole job; then shakenfist#4280 stopped that second class moving at
all for three days, and shakenfist#4309 closed it before any step had
acted on the deferral -- so the phase was back to closing with the
count it was planned to close with, and phase 5 waiting on a phase
rather than on a bug. Amended 2026-09-26: that holds only if 4f's
default move survives. Its post-merge kerbside run failed, and a
revert would put the under-cloud sites back on ci-images/debian-12
with phase 5's Debian 12 half waiting on the diagnosis again. The
label half is still going well. It is no longer the whole phase, and
it was never the part worth watching. Amended 2026-09-29: the default
move survives. The kerbside failure did not recur and the second
verification run passed once actions#112 fixed the renderer inference
behind it, so phase 5 is again waiting on a phase rather than on a
bug. The risk that fired was not the one this paragraph names: what
broke was a playbook that had been reading the distribution version to
decide how to configure a network, in a repository this phase was not
editing, exposed by a default this phase moved.
Retiring a label early breaks CI fleet-wide. D4 and the phase 5
ordering exist for this. A debian-12 retirement before phase 4
completes takes runners out from under jobs still requesting them.
debian:11 is already unrecoverable. It cannot be rebuilt, so
the published copies are all that will ever exist. Phase 3 moves
the dependencies disk off it; until then, do not delete those
copies -- they are load-bearing and were deliberately kept when the
mislabelled images were deleted on 2026-09-12.
Administration and logistics¶
Success criteria¶
- Every one of the nine issues is closed, explicitly declined in writing here, or reduced to a named residual recorded in this plan. The residual case is not a loophole: phase 1 closes only part of private-ci#44 and says so, because that issue's fourth checkbox -- what actually caused the 2026-09-12 miss -- is a separate defect that the scheduler fix does not explain.
- An image that fails to build, or stops being built, produces a GitHub issue without a human looking for it -- demonstrated for each of the three producers.
- A published image that does not match its own name fails its build rather than publishing, demonstrated by a deliberate test.
eol-distroreports zero findings across the fleet, and the producer definitions are inside something that measures them.- No runner label is retired while any workflow still requests it.
pre-commit run --all-filespasses, andscripts/audit-check.pyreports no new failures for this repository.
Documentation index maintenance¶
One row in docs/plans/index.md, dated 2026-09-13, status kept
current as phases land. The row reaches Complete only when every
phase has completed, been abandoned or been superseded.
Future work¶
- Boot-testing published images. Phase 2 asserts that an image matches the name it is published under; nothing asserts that it boots. That is the largest remaining quality gap in the pipeline and it is a plan of its own.
- A Fedora image that builds.
fedora:43andfedora:44have never built -- both fail ongrpcio-toolsneeding a C++ compiler -- so Fedora is absent from the nightly list entirely and thefedoraconvenience symlink points at an end-of-lifefedora:42. - The convenience symlinks.
debianpoints atdebian:12and should probably point atdebian:13;debian-dockerpoints atdebian-docker:12, which served Debian 11 for two years. - Checksums or signatures for published images. Consumers
fetch
latest.qcow2over HTTPS with no way to verify what they received. - The ~70 MB of unexplained variance between two
debian-13-dockerbuilds made hours apart from the same refreshed base (1243.3 MB then 1314.1 MB). It does not threaten the 174 MB staleness measurement, but it means single size measurements are noisier than they look.
Bugs fixed during this work¶
Fixed as phases landed:
- The conductor skipped a nightly rebuild on restart
(private-ci#49, phase 1).
builder_looprecomputednightly_duefrom the clock into a local at every start, so a conductor coming back after the nightly hour set the next rebuild to tomorrow and nothing recorded that tonight had not happened. The slot is now claimed in the database. This does not explain the 2026-09-12 miss that prompted private-ci#44 -- that issue's fourth checkbox stays open as a possibly separate defect. - A nightly that failed halfway could be retried in full by every subsequent restart (private-ci#49, phase 1). Found while fixing the above; the slot is claimed when the cycle starts rather than when it completes, while the staleness metric reads only completions.
Already fixed before this plan was written, and recorded here
because they are the evidence for it:
images#2 (the sixteen-day outage, the grub filesystem
incompatibility and the two-year bullseye mislabelling) and
actions#66 (debian-13-docker publishing without a docker client).
Back brief¶
Before executing any step of this plan, please back brief the operator as to your understanding of the plan and how the work you intend to do aligns with that plan.