Phase 3: explicit saturation coverage¶
Parent plan: PLAN-ci-cloud-sizing.md.
Planning effort: high, as the master plan specifies. The arithmetic in this phase is trivial; the judgement is entirely in where the boundary is asserted. The obvious reading of the master plan -- "fill a cluster to its ledger" -- turns out to be hostile to the suite it would run inside, and the survey below is what changes the shape of the phase.
Decision numbering continues the plan-wide sequence: phase 0 used D1-D8, phase 1 D9-D15 and phase 2 D16-D22, so this phase begins at D23.
Prompt¶
Before responding to questions or discussion points in this document, explore the shakenfist codebase thoroughly. Read relevant source files, understand existing patterns (object lifecycle, state machines, MariaDB storage via the three-layer direct/gRPC/public pattern, Pydantic schemas, daemon architecture, operation queue system, event logging), and ground your answers in what the code actually does today. Do not speculate about the codebase when you could read it instead. Flag any uncertainty explicitly rather than guessing.
Key references for this phase are
shakenfist/scheduler.py (the four capacity stages, and
summarize_resources() which publishes what they refuse on),
shakenfist/external_api/instance.py (the 507 branches of
POST /instances), shakenfist/operations/node_inst_netdesc_op.py
(the asynchronous re-place path and its
Requested node lacks resources abort),
shakenfist/deploy/shakenfist_ci/cluster_ci_tests/test_nodes.py
(the zero-cost placement primitive this phase reuses),
shakenfist/deploy/shakenfist_ci/cluster_ci_tests/test_namespace_claims.py
(the impossible-request refusal pattern, and the retry policy that
sits beside it) and
shakenfist/deploy/shakenfist_ci/retries.py.
Plan file conventions (shared block; do not edit -- the canonical
copy lives in shakenfist/development at
templates/shared-blocks/plan-file-conventions.md):
- All planning documents live in
docs/plans/. - Detailed planning gets one plan file per phase. Phase files are
named for their master plan, sit in the same directory as it,
and append
-phase-NN-descriptivebefore the.mdextension. - The master plan tracks its phases in a table under its Execution section:
| Phase | Plan | Status |
|---|---|---|
| 1. Schema migration | PLAN-thing-phase-01-schema.md | Not started |
| 2. Public API | PLAN-thing-phase-02-api.md | Not started |
- One commit per logical change, and at minimum one commit per phase. Unrelated changes are not batched into a single commit. Each commit is self-contained: it builds, passes tests, and has a message explaining what changed and why.
Context¶
Phase 0 built the scarcity inventory: every distinct failure
signature the current clouds produce because they are small, each
classified as a defect to fix, behaviour to assert, or a test bug.
Phase 2 measured what "small" means -- slim-primary carries a 27
vCPU ledger over six nodes and slim-tier a 12 vCPU ledger over
three, and on slim-tier at least one node sat at or above its own
ledger ceiling in 100% of the 50 job-runs sampled, including
runs that passed.
Phase 3's job, in the master plan's words, is to convert the coverage we get from scarcity by accident into coverage we get on purpose, so that phase 4's bigger clouds cannot quietly close a defect by hiding it. It gates phase 4.
The phase is worth doing for a second reason phase 2 made visible. The suite's exposure to the capacity boundary today is ambient: whether any given merge run exercises a refusal depends on how the five stestr workers happen to interleave. Coverage that depends on timing is coverage that reports differently on every run, which is indistinguishable from a flake, and the response to a flake is to retry it. Deterministic assertions at the boundary are what let phase 4 change the cloud without anyone having to argue about whether the suite got quieter because the system got better.
Scope¶
In scope:
- Functional tests that reach each capacity refusal stage deliberately and assert the refusal contract the server offers today.
- Unit coverage for the stage predicates that a functional test cannot force deterministically (D25).
- A single helper expressing the refusal contract, so that the sibling plan's phase 4 changes it in one place (D27).
- Widening the capacity guard census filter in
shakenfist/actionsby the two events that say what became of a denial -- phase 2 handed this over explicitly (D29). - Re-checking phase 0's inventory against phase 2's census and against the issue tracker as it stands today, and updating the dispositions in place (D30).
- Issues filed for anything these tests expose.
Out of scope, deliberately:
- Any topology change. That is phase 4, which this phase gates.
- Any change to what the scheduler admits, or to the refusal
contract itself. Making a capacity refusal transient -- a
Retry-After, a machine-readable body, a client retry, server side queueing -- belongs to PLAN-transient-capacity-refusals.md, and the admission ledger and claims belong to PLAN-scheduler-reservations.md. This phase asserts what exists and changes none of it. - The headroom band and any gate built on it. Phase 5 owns those.
- Documenting the sizing model. Phase 6 owns that; new tests get their own docstrings and nothing more.
- Chasing #3565's scheduler question. It closed as a test bug on a test change (scheduler-reservations phase 6), and phase 0's row was corrected to say so on 2026-09-01.
- The deterministic affinity reproduction phase 0's #3565 row still asks phase 3 for, and this exclusion was missing until step 3f found it. The bullet above excludes chasing the scheduler question, which is not the same thing: phase 0's row, even after its 2026-09-01 correction, says "the deterministic reproduction phase 3 owes it is still worth having -- it is now coverage of a documented guarantee rather than a hunt for a bug". This phase does not write it, and the reason is D23. A reproduction of the kind phase 0 describes -- fill the affinity target, then place -- has to fill a named node chosen by the affinity test rather than the roomiest one, so it cannot use D26's "skip unless this node is comfortably free" escape: if the target is busy the test has nothing to do but wait or fail. That is a different risk profile from 3c, and pricing it needs the merge-run evidence 3c is about to produce about how often a single-node fill skips. Phase 4 inherits it as an explicit debt, not as an oversight.
- The master plan's other Future work entries -- the under-cloud
probe, per-suite concurrency,
test_coalescing's burst, the two uninstrumented cluster jobs.
What the survey found¶
The master plan's phase 3 section was written on 2026-08-27, before phases 1 and 2 executed and before four of the five issues it argues from were closed. Nine findings, and the four that are false claims rather than new information were corrected at their source in the planning commit -- do not redo them.
F1 -- Four of the inventory's five issue anchors are now closed,
including the umbrella the phase is framed around. #3772 closed
2026-09-09T13:09:41Z as COMPLETED, #3496 on 2026-08-29,
3696 on 2026-09-05; #3813, #3565 and #3907 were already closed¶
when phase 0 wrote the table. Of the six, only #3718 (the second shape of the runner-communication family, out of scope per D8) was still open when this plan was written. The master plan's phase 3 section says "the issue records what it should do, so growing the clouds cannot quietly close it" -- and at that point no open issue held that record. #3772 has since been reopened; see the resolution below.
Two further facts about #3772's closure, because they change what this phase can assume:
- No commit claims it. No commit on
developcarriesFixes #3772. The nearest change is09a8975db("Force a capacity reconcile when a hypervisor is unguarded", the #4087 warm-up fix), whose message says only "Related to #3772 and #4087". The issue was closed by hand with no closing comment, roughly six hours after that fix merged. - Its last two comments are recurrences on the day it closed --
2026-09-09 at 10:08 and 11:22 UTC, both
507 sufficient_idle_cpuin merge groups, one of them with the post-run census showing a node's committed vCPU at peak fraction 1.000. So the signature was still occurring hours before the issue was closed.
Resolved 2026-09-10: the closure was an error, and #3772 is
reopened. The back brief asked the operator which reading was
intended, because if #3772 had been closed as fixed then this
phase's tests would be asserting a contract someone expects to
change. It was not. Three recurrences postdate the closure --
34398770295
on fd617fe72,
34426395717
and
34433668437
on cd379c627 -- all three in merge groups for changes that cannot
reach admission (a keyed cluster_config read, and two Renovate
type-stub bumps). The second is decisive, because its census
separates the two candidate causes: capacity-row coverage was
complete 0 seconds after the first sample, so #4087's warm-up
fix was live, and yet two of the three hypervisors sat at
committed-vCPU peak and p90 fraction 1.000 for the whole
1665-second window. That is genuine exhaustion of the admission
ledger, not a node recorded above its limit during an unguarded
window. The issue was reopened carrying that evidence, and the
automated-fix-attempted label was left in place so the issue-fix
workflow does not race this phase.
So the record is held by the issue and by PLAN-transient-capacity-refusals.md, which owns making a refusal transient and which this plan does not touch; that plan's Related issues section describing #3772 as open is correct again, and needs no edit. This phase's tests assert today's behaviour unchanged, and 3a's helper docstring is right about what will change it.
One thing in the evidence did move, and it weakens an assumption
made elsewhere in this plan. The 2026-09-08 research into #3772
found that every post-#4106 507 was a force_placement
single-candidate create onto one of the 3-vCPU infra nodes
(NODE_CPU_RESERVATION_THREADS=4 on a 4-thread node yields
cpu_schedulable=1, hence limit 3) while the cluster still held
3-9 of its 12 vCPU. The runs above have far less slack than that:
two nodes pinned for a whole window. D23 sizes its fill from the
node's own cpu_limit - cpu_committed and so still works, but the
risks table's reading -- that one node can be filled without
disturbing the other four workers -- is less safe than when it was
written, and D26's skip is doing more of the work than the table
credits it with.
F2 -- A stale line reference. Phase 0's inventory cites
_has_idle_disk_bandwidth() at shakenfist/scheduler.py:329; it is
at shakenfist/scheduler.py:495. Corrected at source.
F3 -- The fill primitive already exists, and it is cheap.
test_cluster_resources_charges_unbooted_placements
(cluster_ci_tests/test_nodes.py:94) reads /admin/resources
through the admin client, picks the hypervisor with the most
cpu_available, and creates a 1 vCPU / 128 MB instance with a
single empty 1 GB disk, no base image and force_placement= that
node -- then deliberately does not wait for it to boot, because
the point is to read the ledger before any domain exists. Nothing
is downloaded and nothing boots, so each unit of fill costs a
database write and a vCPU of ledger. Phase 3's fill loop is that
call repeated, which means its cost is already known rather than
guessed at.
F4 -- The suite can size a fill exactly.
summarize_resources() publishes, per node,
cpu_limit (the capacity row's own limit_cpus, or None when the
row is absent -- scheduler.py:1078), cpu_committed,
cpu_available, cpu_hard_max and cpu_committed_row_present, and
publishes total.capacity_degraded (phase 2's D19). The client
exposes it as get_cluster_resources()
(client-python/shakenfist_client/apiclient.py:1646), and
base.BaseTestCase.system_client is an admin client. So a test can
compute exactly how many one-vCPU instances a node will take, and
can tell "the table says this node is full" from "the table could
not be read" -- which is the distinction phase 2 added the flag for.
F5 -- sufficient_idle_disk has no deterministic functional
trigger. _has_idle_disk_bandwidth() refuses when the node's
DISK_BUSY_PER_SECOND_METRIC exceeds 1200 ms/s
(scheduler.py:495-512). That metric is measured, republished on
the resources daemon's 60 second cadence, and is not carried in
/admin/resources at all, so a test can neither force it nor
observe it. Phase 0's instruction that "phase 3's saturation test
must cover this stage directly" is not achievable as written.
Corrected at source, pointing at D25.
F6 -- None of the four stage predicates has a unit test naming
it. shakenfist/tests/test_scheduler.py mentions
_has_sufficient_disk only inside a comment (line 1199) and does
not name _has_sufficient_cpu, _has_sufficient_ram or
_has_idle_disk_bandwidth at all. The cheapest coverage in this
phase is the coverage nobody has written.
F7 -- There is no serialisation seam in the suite, and this is
the finding that reshapes the phase. cluster-ci.conf sets only
test_path and top_dir, so there is no group_regex: stestr
distributes the cluster suite freely across five workers on one
cluster. A test that fills the cluster to its ledger starves the
other four workers for as long as it holds the fill, and the
failure they would then report is 507 sufficient_idle_cpu -- the
exact signature this plan exists to stop being accidental. On
slim-tier it is worse than hostile, it is close to a no-op: phase
2 measured a node at or above its ledger ceiling in 100% of
slim-tier job-runs already, so "fill the cluster" there means
"fill what four other workers have not filled yet", which is not a
quantity a test can name.
F8 -- One inventory row is already discharged.
test_claim_lifecycle_and_refusals
(cluster_ci_tests/test_namespace_claims.py:440) already asserts
507 on a claim the cluster cannot promise, using an
IMPOSSIBLE_CPUS request that consumes nothing, and the file's
header documents why success assertions retry 507 while refusal
assertions must not. That is precisely the "keep a test that
exercises the refusal path directly" the #3907 row asked phase 3
for. Phase 3 should record it as met rather than write a second
one -- and, more usefully, should copy its shape (D24).
F9 -- Sibling-plan bookkeeping, reported not corrected.
PLAN-transient-capacity-refusals.md lists its phase 1 (close the
warm-up window) as Not started, but that change is on develop
as 09a8975db, having landed through the issue-fix-4087 branch,
and a transient-capacity-refusals-phase-01-warm-up worktree still
exists. Not this plan's table to fix.
Everything else the survey checked held: the four stage names in
scheduler.py are as phase 1 corrected them, the topology matrix
and its concurrency of 5 are as phase 2 described
(.github/workflows/functional-tests.yml:440-482), and the ledger
figures in the master plan's Situation section still match the
committed dataset.
Decision items¶
D23 -- Assert at the node boundary, not by filling the cluster¶
The master plan says "fills a cluster to its ledger". Phase 3 fills
one hypervisor to its ledger instead, by repeated
force_placement creates of the F3 shape, sized from that node's
published cpu_limit - cpu_committed.
Three reasons, in order of weight:
- Blast radius. F7: filling the cluster starves the other four workers and manufactures the failure signature the plan is trying to make deliberate. Filling one node of three or five leaves the rest of the suite a working cluster.
- It is where admission actually binds. Phase 2's headline finding is that the cluster-wide committed fraction and the per-node maximum can differ by a factor of two, and that it is the per-node figure the scheduler refuses on. A test that fills "the cluster" is asserting against the statistic that does not decide anything.
- It is nameable. A node's remaining ledger is a number the API publishes. The cluster's remaining ledger, in a suite where four other workers are creating and deleting, is not.
The cost of the decision: the phase does not directly assert "every node in the cluster refused, therefore the create failed with 507". D24 covers that case by another route.
D24 -- The cluster-wide refusal is asserted with an impossible request¶
Copy F8's shape. A create whose requested resource exceeds any node's published ceiling is refused at the corresponding stage, and consumes nothing while being refused -- so it is deterministic, it is independent of ambient load, it runs in every job on every topology, and it can never starve another worker.
This gives deterministic coverage of three stages:
| Stage | Impossible request | Sized from |
|---|---|---|
sufficient_idle_cpu |
vCPUs greater than max(cpu_limit) over all nodes |
/admin/resources |
sufficient_idle_memory |
memory greater than max(ram_max) |
/admin/resources |
sufficient_free_disk |
a disk larger than max(disk_available) + sum(disk_available) |
/admin/resources |
The sizes are read from the API rather than hardcoded, because a hardcoded "impossible" number is a number phase 4 can make possible.
Correction (2026-09-10, found implementing 3b): the memory and
disk rows of that table originally named max(ram_available) and
max(disk_available). Those are headroom figures, and headroom
moves upward whenever a sibling stestr worker deletes an instance
or a blob. A request sized one unit beyond the largest headroom at
read time therefore becomes satisfiable the moment any worker frees
more than one unit on that node -- the create is admitted, no
exception is raised, and the test fails. That is the "the new tests
become the flake source" risk in the table below, introduced by this
decision's own sizing rule. The CPU row was already correct because
cpu_limit is a ceiling: it is what the node is guarded to,
whether it is idle or full.
The rows above are the corrected sizings. Memory has a published
ceiling, ram_max (scheduler.py:1109-1111,
memory_max * RAM_OVERCOMMIT_RATIO), and a guarded node can be
bounded below it by its row's limit_memory_mb, recoverable as
ram_available + ram_committed; the test takes the larger of the
two so the request exceeds whichever ceiling binds. Disk has no
published ceiling at all -- nothing in summarize_resources()
publishes a node's total disk -- so the disk test cannot be made
load-proof the way the other two are. It uses the cluster's whole
free-disk total as a margin, which exceeds what the largest node
could reach even if every other node's free space were released onto
it, and its docstring says plainly that this is a generous margin
rather than a proof of impossibility.
D23 and D24 are complementary and both are needed: D24 proves the refusal contract (which status, which stage name, which event), D23 proves the ledger is what produces it -- that a node with real capacity, filled with real placements, starts refusing at exactly the published limit. Only D23 would catch a regression where admission stopped charging placements.
D25 -- sufficient_idle_disk gets unit coverage, not functional¶
This overrules phase 0's inventory, which said phase 3's saturation test "must cover this stage directly". F5 is the reason: the predicate reads a measured rate on a 60 second cadence that the API does not publish, so no functional test can force it and none can observe why it did or did not fire. A test that tries would be a test that passes for the wrong reason most of the time.
Instead: a unit test of _has_idle_disk_bandwidth() pinning both
sides of the 1200 ms/s threshold and the shape of the reason dict,
which is the part a later refactor can silently change. Phase 0's
row is corrected at source to say this.
The honest cost, written down rather than glossed: the stage that phase 0 singled out as the one sizing cannot fix is the stage this phase covers least. If it recurs after phase 4, the trace will come from the phase 1 census, not from a test.
D26 -- Every new test skips rather than fails when the cluster is already busy¶
A saturation test that fails because another worker took the capacity first is exactly the flake this plan exists to remove, and it would be a flake we introduced. Every test added by this phase:
- Reads
/admin/resourcesfirst and callsskipTest()with a message naming the figure it wanted, when the headroom it needs is not there -- following the existing idiom attest_nodes.py:111('No hypervisor with two vCPUs of headroom'). - Skips when
total.capacity_degradedis true, or when the target node reportscpu_committed_row_presentfalse. A refusal from an unreadable ledger is not the refusal being asserted, and phase 2 built the flag precisely so this is answerable rather than guessed. - Releases its fill before returning, and relies on
BaseNamespacedTestCase's namespace teardown as the backstop for the case where it does not return. - Carries a deadline.
retries.py'sretry_while_transient()is the existing primitive for "wait while the answer is transient", and refusal assertions must not use it -- the claims file's header (test_namespace_claims.py:35-52) already explains why a caller asserting a refusal must not retry the refusal away.
D27 -- One helper expresses the refusal contract¶
Today's contract, from scheduler.py:540 and the create path:
- HTTP 507, from
external_api/instance.py:901-906, raised asexceptions.LowResourceException. - Message
No nodes remaining at scheduling stage <stage>.
Correction (2026-09-10, found implementing 3c): there is a
second 507 with a different message. When a stage pre-filter
prunes every candidate the message above is raised; when the
pre-filter passes a candidate and the atomic capacity guard inside
Instance.place_instance() then refuses it, the create path returns
507 carrying no node had capacity for this instance, N candidates
refused it (external_api/instance.py:975-982). The helper matches
only the first, deliberately -- the two mean different things, and
conflating them would let a guard refusal satisfy an assertion about
a stage refusal. See D28 for what a test does when it meets the
second.
* An audit event schedule has no candidates at stage <stage>,
aborting, carrying candidates and dropped in its extra.
Every assertion in this phase goes through one helper --
assertRefusedAtStage(response, stage) or equivalent -- rather
than repeating the string. This is not tidiness. The sibling
plan's phase 4 adds Retry-After and a machine-readable transient
refusal to exactly this contract, and its phase 2 changes what the
suite does when it meets one. When that lands it should edit one
helper, not eight tests, and the helper is the place where "what
the contract was in September 2026" is written down for whoever
changes it.
D28 -- The targeted-create-at-a-full-node test records the answer, it does not predict it¶
Phase 0's #3496 row asks phase 3 to "assert the documented behaviour of a targeted create against a full node". There are two paths and the plan deliberately does not guess which one a filled node produces:
- The synchronous create path can refuse with 507 before the instance exists.
- The asynchronous re-place path aborts with
Requested node lacks resources(operations/node_inst_netdesc_op.py:207), which surfaces as an errored instance rather than as an HTTP status.
Correction (2026-09-10, found implementing 3c): there are
three paths, not two. The first bullet above is really two, per
D27's correction: the stage pre-filter's 507 and the capacity
guard's 507 carry different messages. The distinction is not
cosmetic. _has_sufficient_cpu()'s docstring
(scheduler.py:321-350) says it is "a cheap CPU pre-filter (P2) ...
not the admission decision", and that it reads the capacity row's
limit_cpus and charges max(measured_cpus, committed_cpus)
precisely so it sees what the guard sees. A node deliberately filled
to its cpu_limit therefore fails the pre-filter, and the stage
message is the expected answer. Reaching the guard instead means the
pre-filter believed there was room -- that the ledger the test sized
itself from was already stale. That is not a refusal to assert and
not a failure: it is an invalid premise, and 3c skips on it, saying
so. The third path stays as this decision describes it.
The implementing step observes which occurs, asserts that, and writes a sentence in the test explaining which path produced it. If it turns out to be the second, that is worth a paragraph in the phase's close-out: an operator asking for a specific node and getting an errored instance several seconds later is a materially different user experience from a refusal at request time, and the sibling plan would want to know.
D29 -- Widen the census filter in this phase¶
Phase 2's What phase 3 inherits hands this over by name, and the
master plan's Future work carries the entry. Two more
alternations in the LogQL filter in the shakenfist/actions
repository, beside the guard denial the filter already matches:
no candidate admitted and some refused on demand alone, waiving demand guardschedule failed, every candidate refused by capacity guard
The second is the event that actually says the cluster refused a create at the guard, which is what makes a saturation test's observation checkable against the census rather than only against its own assertions. Phase 2 measured 3,480 denials across 32 job-runs and could not say how many ended in a failed create; after this it can.
This step lands in another repository, cannot be tested before it merges, and only the operator can push it -- the same seam phase 2 met at its step 2e. It is sequenced early so a merge run has used it before the phase closes.
D30 -- The inventory is re-checked and updated in place¶
Phase 2's inheritance asks for this explicitly: phase 0 built the inventory from triage history, and phase 2's census is the first chance to see whether the frequencies match. Every row gets its issue state refreshed (F1), its disposition confirmed or changed against the census, and -- for the rows that name a test phase 3 owes -- a link to the test that now discharges it. A row whose disposition survives contact with the data is a result worth one sentence; a row that does not is the more valuable finding.
Step plan¶
| Step | Effort | Model | Isolation | Brief for sub-agent |
|---|---|---|---|---|
| 3a | medium | sonnet | none | Add the refusal-contract helper (D27) to shakenfist/deploy/shakenfist_ci/base.py, beside the other assert* helpers at base.py:1055-1099. It takes the InsufficientResourcesException (or status/body pair) the client raised and the expected stage name, and asserts status 507 and the message No nodes remaining at scheduling stage <stage> produced at scheduler.py:540. Give it a docstring recording that this is the contract as of 2026-09, that PLAN-transient-capacity-refusals.md phase 4 will add Retry-After and a machine-readable body to it, and that this helper is the single place to change when it does. Do not use retries.py here: a refusal assertion must never retry the refusal away, for the reason cluster_ci_tests/test_namespace_claims.py:35-52 gives. No test changes in this step. |
| 3b | medium | sonnet | none | New file shakenfist/deploy/shakenfist_ci/cluster_ci_tests/test_saturation.py implementing D24: three tests, one per stage in D24's table, each reading self.system_client.get_cluster_resources(), computing a request one unit beyond the largest value any node publishes, issuing the create, and asserting through 3a's helper. Model the impossible-request shape on IMPOSSIBLE_CPUS in test_namespace_claims.py:440-467. These consume nothing, so they need no fill and no cleanup beyond the namespace teardown BaseNamespacedTestCase already does. Apply D26's skip rules: skip when total.capacity_degraded is true, and when per_node is empty. Use base.BaseNamespacedTestCase. |
| 3c | high | opus | worktree | The node-fill test (D23), added to test_saturation.py. Pick the hypervisor with the largest cpu_available from /admin/resources, skip per D26 if it does not have at least three vCPU of headroom or if cpu_committed_row_present is false for it. Fill it with cpu_limit - cpu_committed one-vCPU instances using exactly the zero-cost create shape at test_nodes.py:116-122 (1 vCPU, 128 MB, no base image, a single {'size': 1, 'type': 'disk'}, force_placement= the node name, and no wait for boot). Re-read /admin/resources and assert the node now publishes cpu_available at or below zero and cpu_committed at its cpu_limit; then assert that one more force_placement create at that node is refused, through D28's rule -- observe which of the two paths answers (507 at request time, or an errored instance carrying Requested node lacks resources from operations/node_inst_netdesc_op.py:207), assert that, and say in a comment which one it was and why. Delete the fill in the test body, not only in teardown, and assert the ledger returns. This is the step where getting it wrong costs a merge-queue flake for everyone, so mutation-test each assertion: break the thing it claims to prove and confirm it fails. |
| 3d | medium | sonnet | none | Unit tests in shakenfist/tests/test_scheduler.py for the four stage predicates, which have none today (F6): _has_sufficient_cpu (scheduler.py:321), _has_sufficient_ram (scheduler.py:391), _has_sufficient_disk (scheduler.py:467) and _has_idle_disk_bandwidth (scheduler.py:495). For each, pin both sides of the boundary and the shape of the reason dict it returns, since the reason dict is what the audit event publishes and what a later refactor can silently change. _has_idle_disk_bandwidth is the one D25 makes load-bearing: pin the 1200 ms/s threshold from both directions. Follow the existing fixture style in that file; mock metrics the way the surrounding tests do. |
| 3e | low | sonnet | none | In the shakenfist/actions repository, widen the capacity guard census LogQL filter by the two alternations named in D29. Change nothing else. The operator pushes this; it cannot be tested before it merges. Report the exact diff for review rather than assuming it can be verified locally. |
| 3f | high | opus | none | Re-check phase 0's scarcity inventory against phase 2's dataset and today's issue tracker (D30), and rewrite the dispositions in docs/plans/PLAN-ci-cloud-sizing-phase-00-decisions.md:253-298 in place. For each row: refresh the issue state, say whether the census frequency in docs/plans/data/ci-cloud-sizing-baseline/ supports the disposition phase 0 gave it, and link the test from 3b/3c/3d that discharges it. Record that #3907's row was already discharged by test_claim_lifecycle_and_refusals (F8) rather than by anything this phase wrote. File issues for anything 3c exposed. Do not re-litigate #3565, which is out of scope. |
| 3g | medium | sonnet | none | Close-out: set this phase Complete in both the master plan's Execution table and the docs/plans/index.md row (3 of 8 becomes 4 of 8), write the What phase 4 inherits section against what was actually found, run python3 tools/check-plan-status.py and pre-commit run --all-files, and confirm every Definition of done item below by running it rather than by reading it. |
Steps 3a, 3d and 3e are independent of each other. 3b needs 3a. 3c needs 3a and should follow 3b, because 3b proves the helper against a refusal that costs nothing before 3c depends on it during a fill. 3f needs 3b and 3c to have merged and a real merge run to have used them. 3g is last.
Risks and mitigations¶
| Risk | Mitigation | Who checks |
|---|---|---|
| The fill in 3c starves the other four stestr workers and manufactures the very 507s this plan exists to remove. | D23 bounds the fill to a single hypervisor, and D26 skips when that node is not comfortably free. On slim-tier (three nodes, 12 vCPU) filling one 6 vCPU node is half the cluster, which is the worst case; if the merge-run evidence after 3c shows a rise in other tests' 507s, the test is restricted to slim-primary by a topology skip rather than kept and tuned. Amended 2026-09-10: F1's newest runs show two of three slim-tier hypervisors pinned at their ceiling for a whole run, so the headroom this row assumed is not reliably there and D26's skip, not the single-node bound, is the load-bearing mitigation. Expect 3c to skip often on slim-tier; a 3c that never skips there is evidence the skip predicate is wrong, not that the cluster is roomy. |
The operator, over the merge runs following 3c, against the phase 1 census. First run, 2026-09-12: no rise. Five to seven schedule aborts per job, approximately the tests' own, and every guard refusal on the demand dimension rather than on vCPU ledger -- see What the first merge run answered. One run, so the row stays open. |
| The new tests become the flake source. | Every test skips on ambient shortage (D26), and the Definition of done requires clean merge runs after landing, not just a green branch. First run, 2026-09-12: all four ok in all three cluster jobs, none skipped. That is one run against a prediction of frequent skipping, so it says the tests are not fragile rather than that the predicate is calibrated -- D26 has not yet been seen to fire in anger. |
3g, and the operator. |
| A test asserts a refusal produced by an unreadable capacity table rather than by a full one, and passes for the wrong reason. | D26 skips on capacity_degraded and on a missing capacity row -- the exact distinction phase 2's D19 added the flag to make answerable. |
3b and 3c briefs; mutation testing in 3c. |
| The refusal contract changes under us when the sibling plan's phase 4 lands, and eight tests need editing. | D27's single helper, and its docstring naming the plan that will change it. | 3a. |
| ~~#3772 was closed as fixed rather than superseded, so this phase asserts behaviour someone believes has changed.~~ Retired 2026-09-10: the closure was an error and the issue is reopened, so the contract this phase asserts is not expected to change. See F1. | n/a. | Resolved before 3a. |
| ~~3e cannot be verified before it merges, the same seam that made phase 2's 2e awkward.~~ Resolved 2026-09-12: run 34681505274 carried it and its census reports guard refusal counts in all three cluster jobs. | Sequenced early, reviewed as a diff, and confirmed from a real merge run's census section. | The operator, done. |
Definition of done¶
Falsifiable, in order:
shakenfist/deploy/shakenfist_ci/base.pycarries one helper asserting the refusal contract, andgrep -c 'No nodes remaining at scheduling stage'overshakenfist/deploy/shakenfist_ci/returns 1.cluster_ci_tests/test_saturation.pyexists and contains a test per stage in D24's table, each of which passes on a cluster with no free capacity at all -- because an impossible request is refused either way.- Every test added by this phase has an explicit
skipTest()path for ambient shortage and forcapacity_degraded, verified by reading the diff and not by the test passing once. - The node-fill test releases its fill in the test body, and a run
with the release deleted leaves the node's
cpu_committedabove its starting value -- that is, the release is asserted, not assumed. - Each assertion added by 3c has been mutation-tested: the assertion fails when the property it claims to prove is broken. The step reports which mutation it applied for each.
_has_sufficient_cpu,_has_sufficient_ram,_has_sufficient_diskand_has_idle_disk_bandwidthare each named by at least one unit test, both sides of their boundary.grep -covershakenfist/tests/test_scheduler.pyfor each name is at least 1, where today three of the four are 0.- A merge run after 3e shows the Capacity guard census section
reporting a count for
schedule failed, every candidate refused by capacity guard, not a "not collected" notice. The operator confirms this; it cannot be checked from a worktree. - Phase 0's inventory has no row whose issue state is stated wrongly, and no row citing a source line that has moved. Checked by re-reading each citation.
- Every inventory row classified "behaviour to assert" names the
test that now asserts it, or says explicitly why no test does
(which for
sufficient_idle_diskis D25). - No test added by this phase can fail because of another worker's load: for each, the failure modes are enumerated in its docstring and each is either asserted or skipped.
- The master plan's phase 3 section describes what this phase did
rather than what it was expected to do in August, and every
statement it makes about #3772's state matches
gh issue view 3772 --json state,stateReasonon the day the phase closes out. python3 tools/check-plan-status.pypasses, andpre-commit run --all-filespasses in the main repository.
Outcome¶
Complete. Every step has landed: 3a, 3b, 3c, 3d, 3f and 3g in
this repository, 3e in shakenfist/actions (0c87e8d through
5fe6292, pushed by the operator 2026-09-05 to 09-08). The two
Definition of done items which by the plan's own ordering could not
be satisfied from a worktree were answered by the first merge run to
carry the phase, and that run's evidence is recorded below rather than
summarised as "green".
The phase deliberately stayed In progress for a day after #4170
merged, because marking it done on a green branch would have been
exactly the "a status column is a claim, not evidence" failure the
phase-completion check exists to catch. What closed it is a merge run,
not a branch.
Each Definition of done item was run rather than read:
| # | Result | How it was checked |
|---|---|---|
| 1 | pass | grep -rc over shakenfist_ci/ returns the stage string once, in base.py alone |
| 2 | pass | The file exists with four tests covering all three of D24's stages, and its sizing arithmetic is unit tested here (test_ci_saturation.py, 17 tests, mutation tested). All four tests reported ... ok -- not SKIPPED -- in all three cluster jobs of run 34681505274, on clusters whose censuses record 30, 13 and 7 nodes dropped for would exceed hard max CPUs in the same window |
| 3 | pass | Read from the diff: 19 skipTest() paths, 8 capacity_degraded references |
| 4 | pass | Mutation tested: with the in-body release deleted, the ledger assertion fails rather than the test passing |
| 5 | pass | 3c reported 33 mutations across both CI topologies and a fractional-limit one, with no unexpected outcome |
| 6 | pass | All four predicates are named by unit tests; three of the four were zero before |
| 7 | pass | Run 34681505274 carried 3e. Its Capacity guard census reports Placements refused by the guard: 176, 139 and 163 in its three cluster jobs, not a "not collected" notice |
| 8 | pass | Every file:line citation in the inventory re-resolved, and every issue state refreshed against gh |
| 9 | pass | 3f links a discharging test per row, or states why none exists |
| 10 | pass | Read from the diff: each test's docstring enumerates its failure modes, each marked asserted or skipped |
| 11 | pass | gh issue view 3772 reports OPEN/REOPENED, which is what the master plan's phase 3 section now says |
| 12 | pass | Re-run at close-out on develop: check-plan-status.py reports agreement, and pre-commit run --all-files passes all ten configured hooks |
What the first merge run answered, 2026-09-12¶
Run 34681505274
is the merge-queue run which produced f3b245304, the commit that put
this phase on develop. It carried 3e. Every functional test in all
three of its cluster jobs passed; the only failing test in the whole
run was the one described in The raw-create collision below, which
is not a saturation test. Each owed item, against that run:
- Item 7 -- the census reports a count. Capacity guard census, per job, counted over the whole test window:
| job | hypervisors seen | guard refusals | measured-load alone | D13 carried it | worst shortfall |
|---|---|---|---|---|---|
| Debian 12 tier | 3 | 176 | 138 | 38 | 11.649 |
| Ubuntu 24.04 cluster | 5 | 139 | 45 | 94 | 9.906 |
| Debian 12 cluster | 5 | 163 | 54 | 109 | 11.649 |
Every refusal in all three was at stage node and on the demand
dimension alone, with 0 forced ground-truth writes past the guard
(P5). Two things in that table matter to phase 4. The three-node
job refuses most often, which is the expected direction. But the
attribution inverts with size: on the small cluster measured load
alone was already over the bound in 78% of refusals, while on the
five-hypervisor jobs it is the D13 feedforward estimate that carries
two thirds of them over. Adding hypervisors does not make this bound
recede, it changes which half of it binds -- and the feedforward half
is a rate prediction, not a cloud out of room. The "hypervisors seen"
column is what the samples could see, which as the census itself
warns is not the same as how many the cluster had.
* Item 2's other half, partly. The fill test passed on every
topology, so a real full node does refuse. Which of the three
refusal paths it gave is still unrecorded, and the reason is a seam
this phase should have seen: 3c records the path with
addDetail(), and stestr prints details only on a non-ok
verdict. A test which passes is therefore silent about the one thing
the run was supposed to tell us. The census's stage tallies are
consistent with the pre-filter having fired -- 2, 3 and 4
sufficient_idle_cpu aborts across the three jobs against one
deliberate abort per impossible-request test -- but that is inference
over aggregate counts, not the test's own evidence. The artifact
bundle does not rescue it either: the only per-test files in
traces/ are _emit_tracing_event()'s timing records, and the
full-node test does not write one. Phase 5 should not
repeat this shape: a fact worth a merge run is worth an assertion
or a census field, not an addDetail() on the happy path.
* The skip rate on slim-tier: zero of four. The plan predicted
3c would skip often there and said a 3c which never skips is
evidence the predicate is wrong rather than the cluster roomy. One
run is not a rate, and this one does not settle it -- but it is the
opposite of what F1's amendment led the risk table to expect, and it
happened on the hardest job: the three-node one dropped 30 nodes for
would exceed hard max CPUs where the five-hypervisor jobs dropped
13 and 7, so the fill completed there (35.5 s) under more ledger
pressure rather than less. Watch it over several runs before
concluding either way.
* The fill did not disturb the other workers. Five to seven
schedule aborts per job across a whole window -- at most 4 at
sufficient_idle_cpu and one each at sufficient_idle_memory,
sufficient_free_disk and affinity_constraints -- which is
approximately the four tests' own deliberate refusals and no blast
radius. Every guard refusal in all three jobs was on the demand
dimension, which is pre-existing D13 noise: this phase claims vCPU
ledger, and the ledger dimensions recorded nothing.
* The ownership listing's cost was not separately measured, but it
did not breach the budget: test_no_unbudgeted_fixed_rate_database_polling
passed in both jobs that ran it (246.8 s, 246.9 s). That is weaker
evidence than a measurement, and it is weak for a specific reason --
the listing is reached only from the skip and failure branches, and
nothing skipped or failed, so the expensive path was barely
exercised. A run where 3c does skip is the one that tests this.
Code review, 2026-09-11 (PR #4170)¶
CLAUDE.md asks for a code review at the end of a plan. The automated
reviewer raised one [FIX] and eight [CONSIDER] items on #4170;
eight were taken and one was declined. What changed, because three of
these moved a decision rather than just the code:
- A forced fill create is now asserted to have landed on the node it
named (
_create_fill_instance()). This was the[FIX], and it is the most serious thing the review found: a wrong-node placement is issue #3496, and without the assertion it degraded to a skip -- the ours/foreign split would count zero vCPUs of this test's own on the target node and read the shortfall as ambient load, reporting nothing about the one defect the test sits closest to, while leaving a create budget of stray instances on the siblings D23 exists to protect. - D23's ledger-return assertion now skips on a delete backlog. It fails only once the cluster has finished deleting what the test released, asked of the instance listing rather than assumed. The assertion's text claims a ledger regression, and a queue backlog on the shared cluster produces the identical shape -- so as written it was the file's most likely flake, making a claim the evidence did not support. A ledger still charging an instance the cluster has not finished deleting is correct; one charging an instance that is gone is the regression, and that is now what fails.
- D24's sizing arithmetic moved to
shakenfist_ci/sizing.pyand is unit tested byshakenfist/tests/test_ci_saturation.py, which loads it by path -- the arrangementretries.pyandtest_ci_claims_headroom.pyalready use. This is the one part of the phase whose correctness never needed a cluster, and it was the part with no test at any level. 17 unit tests now cover it, each mutation tested: dropping the increment, dropping the minimum clamp, reading a nullcpu_limitas zero and sizing memory from headroom are all caught. The same file carries the AST guard the repository already applies to the claims suite, pinning theapiclientnames and theasync_requestkeyword the saturation suite depends on and nothing here can import.
The five smaller items: the CPU test now skips when any node's
published cpu_max_per_instance is below the computed request, rather
than letting an earlier stage's refusal surface as a stage-name
mismatch; the whole-cluster instance listing is rate limited to once
per 30 s plus one read at each deadline, rather than once per 5 s poll,
which was the unmeasured database-load cost the Outcome section above
already conceded; polled /admin/resources readings are attached with
addDetailUniqueName() so a failed run keeps the ledger's whole
trajectory instead of only its last read; a tautological assertIn()
inside the branch whose condition it restated is gone, with the
evidence left to the addDetail() above it; and
_resolve_unexpected_admission()'s docstring no longer claims it never
returns normally, which was false for exactly the path D28 asked it to
record.
Declined, with reasoning: the review asked for unit coverage of the
placement_filter() re-place discount in both ledger predicates. That
branch is already covered, at the find_candidates() level rather than
the predicate level, by test_scheduler.py's
test_prefilter_does_not_charge_an_instance_for_itself (CPU) and
test_ram_prefilter_does_not_charge_an_instance_for_itself (RAM), both
of which sit exactly on the boundary the discount moves. Duplicating
them at the predicate level would add no failure the existing pair does
not already catch. The other half of that item was a real gap and was
taken: nothing asserted the no-capacity-row refusal's reason dict
(capacity_row_present: False, limit_cpus falling back to the live
hard maximum), and the memory predicate had no P7 coverage at any level
-- three tests were added for those, taking the file to 108.
One defect the review did not find, fixed here: one_unit_beyond()
justified its math.floor by claiming a truncating int() cast could
round a negative reading back into servable range. It cannot --
truncation rounds a negative value towards zero, so it returns a
figure which still strictly exceeds the reading, and with the minimum
clamped at 1 the two are indistinguishable for every input. The code
was right and its stated reason was wrong, which is worse than a
comment being absent. The docstring now says what is actually true, and
the unit tests assert the two properties a caller depends on -- the
result strictly exceeds the reading, and stays a size a create could
ask for -- rather than enshrining particular return values.
The raw-create collision, 2026-09-12 (PR #4186)¶
This phase broke develop for most of a day -- from #4170's merge at
04:47 on 2026-09-12 to the fix at 15:31. The full account, including
why markers rather than routing was the right fix, is in
PLAN-transient-capacity-refusals-phase-02-suite-wait.md
under A defect this phase's own guard found on develop -- the
sibling plan owns the guard, so it owns the story, and this section
deliberately does not restate it.
In short: phase 2's #4166 added an AST guard requiring every
create_instance() in the functional suite either to go through
BaseTestCase's waiting wrapper or to carry a # raw-create:
<reason> marker; this phase's #4170 added test_saturation.py with
four unmarked calls five hours later. Neither was wrong and neither
could see the other. #4186 added the markers at 15:31 and run
34698074952 confirms Sanity checks green again.
Three things belong here rather than there, because they are about this plan's remaining phases:
- The blast radius was larger than the pair.
Sanity checksis not a required status check, so the queue went on merging over the breakage:f654d4a05,f8c801ebeand6435e3fdeall failed the same test and two of them merged anyway. It was briefly invisible as well -- the documentation-sync run in between skippedSanity checkson the changed-paths filter rather than passing it, so a green tick sat on a red tree. Whether that check should be required is a repository-configuration question, recorded rather than decided; it would have stopped the follow-on batches but not the original pair. - Phase 5's guardrails are the same shape of change -- an assertion every topology must satisfy, landing across a tree with other work in flight. A guard which constrains how tests are written needs a story at design time for branches cut before it existed, because no author of such a branch can satisfy it however careful they are. Phase 5 should say what that story is before it writes the assertion.
- This phase's own sequencing hid the risk. 3a through 3d were
planned as independent of the sibling plan, and in terms of the code
they touch they were. The coupling was a convention, which the
step plan had no column for. Phase 2's own What later phases
inherit had even named the dependency -- "a
# raw-create:convention the sizing plan's phase 3 saturation tests need in order to assert a refusal" -- and this plan still did not read it as something 3c had to do.
What phase 4 inherits¶
- Deterministic assertions at three of the four capacity stages that run in every cluster job on every topology, so a topology change that accidentally makes a refusal unreachable fails a test rather than making the suite quieter.
- One place -- the D27 helper -- where the refusal contract is written down, so phase 4 can see at a glance what it must not change and the sibling plan can see what it is changing.
- A test that proves the per-node ledger is what admission refuses
on, which is the property phase 4's resizing is arithmetic over.
If phase 4 changes
cpu_limitand the fill test still passes at the new number, the resize did what it claimed. - An inventory whose dispositions have been checked against measured frequencies rather than against triage memory, and a census that can say what became of a guard denial.
- A written answer on
sufficient_idle_disk: it is the stage sizing cannot fix and the stage no functional test can force, so if it recurs after phase 4 the evidence will be census data. - One debt, named rather than dropped: the deterministic affinity reproduction phase 0's #3565 row asks for, which this phase declines in its Scope section with the reasoning. It is the only inventory row that names a test phase 3 owes and still has none; everything else in the table is discharged by a test that now exists, or explicitly cannot be.
- The first merge run's answers to 3f's two open questions, and
one of them is still open. How often the demand guard's waive
fails to save a create (the #3813 row) is answered and is larger
than expected: every guard refusal in all three cluster jobs was on
the
demanddimension alone -- 176, 139 and 163 of them, nothing on an allocation dimension -- so phase 4's resizing arithmetic has to account for a bound its resize does not move. Which of the three refusal paths a real full node gives 3c (the #3496 row) is not answered, because 3c records it withaddDetail()and a passing test prints no details -- see What the first merge run answered. Phase 4 either reads it from a census field or asserts it; it should not wait for another run to volunteer it.
Back brief¶
Before executing any step, back brief the operator on the understanding of this plan, and in particular on:
- F1 -- why #3772 was closed. Settled 2026-09-10, before 3a: the closure was an error and the issue is reopened. The question was whether the #4087 warm-up fix was thought to have fixed it, in which case some of these tests would assert a contract expected to change and 3a's helper docstring would be wrong about what changes it. Three recurrences postdating the closure answer it -- one with the warm-up fix demonstrably live and two hypervisors at their ledger ceiling for the whole run. The tests assert today's behaviour unchanged, as planned. No gate remains here; F1 carries the evidence and the one assumption it weakened.
- D23 and D24 together, which replace the master plan's "fill a
cluster to its ledger" with "fill one node, and prove the
contract with a request no cluster could satisfy". This is the
decision that changes what the phase is, and it is the one a
reviewer is most likely to disagree with -- the objection being
that nothing then asserts the whole-cluster exhaustion path end
to end. The counter-argument is F7: on
slim-tierthat path is ambient in 100% of runs already, and a test which tries to own it is competing with four other workers for a quantity it cannot name. - D25, which overrules phase 0 on the one stage phase 0 singled out. Worth an explicit yes or no, because it is a reduction in scope against a decision the operator previously approved.
- D29's
shakenfist/actionschange, which only the operator can push and which cannot be verified before it merges.