Scheduler reservations phase 4a: a satisfiable demand guard¶
Prompt¶
This is a phase plan under PLAN-scheduler-reservations.md. The master plan's Prompt section applies unchanged; the decisions it refers to as D-numbers live in PLAN-scheduler-reservations-phase-00-decisions.md and the P-numbers in PLAN-scheduler-reservations-phase-03-primitive.md. This phase's own decisions are numbered E1..E7 so they collide with neither.
Planning effort: high. The phase changes the shape of a
guard inside the admission transaction that phase 3 spent a
whole step stabilising against innodb_snapshot_isolation,
and it settles a tuning constant that two earlier documents
disagree about. Review effort: high.
Situation¶
Phase 4a exists because of issue #3813, and because phase 4 cannot be closed out honestly without it.
Phase 3 shipped the D13 demand feedforward clause inside the
node guard of _direct_admit_instance_placement(). The clause
asks
The budget on the right is denominated per schedulable thread.
The charge demand_add on the left is denominated per
requested vCPU, at cpus x SCHEDULER_DEMAND_PER_VCPU with the
constant seeded at 2.5. The two were never reconciled, so at
the seed constants a node needs cpu_schedulable >= 3.34
before it can admit a 1-vCPU instance at zero measured load
and zero expected demand. A 2-vCPU instance needs seven
threads; a 4-vCPU instance fourteen.
The CI hypervisors publish cpu_schedulable: 2. Every
candidate is refused, every time, on demand alone.
This phase makes the clause satisfiable, and then discharges the two soaks that phases 3 and 4 left outstanding -- in that order, because soaking phase 4's claim accounting on sfcbr while every placement takes the waiver path would soak the wrong system.
Mission and problem statement¶
Make the D13 demand clause a spreader that can never refuse a placement a node has real room for, at every node size this project supports; correct the constant it was seeded with and the documents that disagree about its provenance; then run the outstanding phase 3 and phase 4 soaks and close phase 4.
The phase is done when a single-thread node admits an 8-vCPU
instance at idle, when a second placement into a burst is
spread rather than piled, and when the master plan's phase 3
and phase 4 rows both read Complete without a footnote.
Scope¶
In scope
- The shape of
_demand_guard_clause()(shakenfist/mariadb.py:24976). - The default and description of
SCHEDULER_DEMAND_PER_VCPU(shakenfist/config.py:353). - Correcting D13's provenance sentence, the master plan's phase 6 correction, the phase 3 row flag and the master plan's Future work entry.
- Operator-facing documentation of what the clause now means.
- Phase 3's outstanding step 9 sfcbr soak and phase 4's outstanding step 10 operator review and sfcbr soak.
- Closing #3813, and flipping the phase 3, phase 4 and phase 4a statuses.
Out of scope
- Flipping
CLAIM_ENFORCEMENT_HARD, migrating theScheduler()callers, or extracting the duplicatedplace_walkhelper. That is all phase 5, and this phase must not anticipate it. (Scope guard: no diff tomariadb.py:24815, and the twoplace_walkcopies stay two copies.) - Removing the P9 waiver. See E3.
- The
SCHEDULER_DEMAND_DECAY_SECONDSdecay model, the per-namespace learned demand value D13 defers, and the affinity questions phase 6 owns. - Phase 00a's outstanding post-deploy validation. It is a separate observation on a different question (does the network+database node still take a disproportionate share) and it belongs to phase 00a's own close-out, though the sfcbr burst this phase runs is the natural occasion to collect it. Noted in Future work rather than claimed here.
What the survey found (2026-08-22)¶
Four findings, three of which change what this phase does.
1. The seed constant was transcribed from the wrong row of
its own measurements. D13
(PLAN-scheduler-reservations-phase-00-decisions.md:439-443)
says SCHEDULER_DEMAND_PER_VCPU is "seed 2.5 from the 00a-1
measurements". The 00a-1 Measurements appendix
(PLAN-scheduler-reservations-phase-00a-load-aware-ordering.md:316-320)
records the observed demand-per-vCPU -- cpu_load_1 /
allocated_vcpus -- as 0.12-0.35 in steady CI with a burst
peak estimated at ~0.6, and says in terms: "Seed constant
for open question 13's expected-demand model: ~0.33 steady /
0.6 conservative."
The only 2.5-shaped number in that appendix is on a different
row entirely: "busy plain nodes run 2.3-3.0 allocated vCPUs
per thread", the packing figure that produced
CPU_OVERCOMMIT_RATIO = 3.0. That is a vCPUs-per-thread
quantity being used as a load-per-vCPU quantity. The seed is
not merely untuned; it is four to eight times the measured
conservative figure, and the units error is the whole of
3813's arithmetic.¶
This upgrades the fix from "a tuning decision pending the phase 0 step 3 data analysis" -- which is how the master plan currently defers it -- to "the data analysis already answered this and the answer was mis-copied". The plan corrects the master plan at source.
2. The P9 waiver does cover scheduled creates, so #3813
does not 507 on its own. Both walkers -- the create path at
shakenfist/external_api/instance.py:881-921 and the
queue-worker reschedule at
shakenfist/operations/node_inst_netdesc_op.py:194-232 --
run place_walk(True), and if every denial was
demand_only, re-run place_walk(False). So the live
symptom is not a refused create. It is:
- every create paying two full candidate sweeps, one RPC per
candidate per sweep, plus a denial-detail read per refusal
(
_admission_denial_dimensions(),mariadb.py:25354) and an audit event per refusal; - the spreader never operating, because the clause cannot pass and the waiver ignores it;
expected_demandstill being incremented on every admission (theSETatmariadb.py:25694is unconditional, outside thenode_guardedbranch), so the column is maintained at write cost and read by nothing that can act on it.
3. The master plan's phase 6 correction is wrong about the
mechanism. It says that after every candidate is refused
"the create places through a single forced candidate and the
affinity stage has nothing left to rank"
(PLAN-scheduler-reservations.md:497-509). The re-walk
iterates the same candidates list in the same order, so the
scheduler's ranking -- affinity included -- is preserved
exactly; the waived walk takes the top-ranked candidate.
What is actually lost is the spreading. Because the clause
never passes, nothing makes the top-ranked candidate less
attractive to the next create in a burst, and the ranking
it competes against is cpu_load_1 / cpu_schedulable from a
metrics row up to 60 seconds stale (scheduler.py:131). A
burst therefore piles onto one node until a real allocation
dimension bites. That is a plausible contributor to #3565 and
it is a different claim from the one the master plan makes,
so phase 6 inherits a corrected premise rather than a wrong
one.
4. Nothing outside the admission transaction reads
expected_demand. git grep finds no occurrence in
shakenfist/scheduler.py. The term is written by the
reconciler and by admission, and read only by
_demand_guard_clause() and the denial-detail builder. So
changing the clause's shape cannot perturb candidate ranking,
summarize_resources(), or the admin resources API -- which
is what makes E2 a contained change rather than a scheduler
rework.
The survey found no other stale claim in the master plan's phase 5 or phase 6 stubs.
Corrected already, in the planning commit -- do not redo
these in a later step: D13's provenance sentence and its
2026-08-22 amendment
(PLAN-scheduler-reservations-phase-00-decisions.md), the
phase 6 correction's mechanism and the Future work entry's
deferral (PLAN-scheduler-reservations.md), the phase 3 and
phase 4 status notes, the new Execution row, and
docs/plans/index.md's phase arithmetic (4 of 10 becomes 4
of 11). One further drift was found and fixed while
registering: the master plan's phase 4 note said the
management review was outstanding, but the phase 4 plan
records step 9 as complete.
Finding 4 is recorded here only. What step 3 still owes is the post-fix half: statements that are only true once the code has changed.
Decisions¶
E1. Retune SCHEDULER_DEMAND_PER_VCPU from 2.5 to 0.6, and
say where 0.6 comes from. The 00a-1 appendix's conservative
burst figure, not its steady-state 0.33: the term exists to
cover the actuation-to-observation gap during correlated
bursts, which is exactly the regime the 0.6 estimate was taken
from, and a spreader that under-charges stops spreading. The
description in config.py stops saying "a provisional seed
pending the scheduler reservations phase 0 step 3 data
analysis" -- that analysis is the 00a-1 appendix and it has
landed -- and cites the appendix instead.
SCHEDULER_DEMAND_DECAY_SECONDS keeps its provisional wording.
Nothing measured it, and this phase does not.
E2. Test the node's existing state, not the incoming placement, against the budget. The clause becomes
with demand_add removed from the comparison but still added
to expected_demand by the same UPDATE, exactly as today.
The reasoning, and this is the decision a reviewer is most likely to want to argue with:
- It is dimensionally honest. Both sides are now node state in units of runnable threads. The old form added a per-request term to a per-node budget, which is the defect, and E1 alone does not remove it -- at 0.6 a 4-vCPU instance still charges 2.4 against a 2-thread node's budget of 1.5 and is refused on an idle node. Retuning moves the unsatisfiability threshold; it does not abolish it -- 22 of the 80 cells in E5's sweep still fail with the constant corrected and the clause unchanged. Only a form with no per-request term on the left is unsatisfiability-proof for every combination of node size and instance size, which is what the master plan's success criterion demands.
- The question it now asks -- "is this node already at or above its target load?" -- is the question a spreader should ask. The old form asked whether the node would be over target after this placement, which conflates spreading with bounding, and D13 is explicit that the term is a spreader and never a capacity bound.
- Check-then-charge is safe here because the guard is a real
serialisation point. The comparison and the
expected_demandincrement are the same guarded UPDATE in the same transaction, so two concurrent admissions against one node serialise: the second sees the first's increment. The window this form permits is one over-target placement per node per decay period, not per burst. - Real capacity is still bounded. The three allocation
dimensions in the same WHERE clause are untouched, and
CPU_OVERCOMMIT_RATIOstill caps vCPUs per schedulable thread. Nothing here lets a node accept work it has no room for; it lets a node accept the first piece of work when it is idle, which it must.
The runner-up was flooring the budget --
... <= max(target_load x schedulable, demand_add). It is a
smaller diff and it does fix the idle-node case, but it keeps
the per-request term on the left, so a 4-vCPU instance is
still refused by a 2-thread node carrying any measured load
at all, and the floor has to be re-derived every time either
denomination changes. Rejected as a patch over a units error
rather than a correction of it.
E3. Keep the P9 waiver. It is still reachable and still
correct: when every candidate is genuinely above target
load, refusing the create would turn a spreader into a rate
limit, which is what P9 exists to prevent. What changes is
that it stops being the only path a placement ever takes. The
waiver's audit event ('no candidate admitted and some
refused on demand alone, waiving demand guard') becomes the
signal that the cluster is actually saturated rather than the
signal that the guard is broken, and step 4's soak reads it
that way.
E4. No new configuration knob. The temptation is a
SCHEDULER_DEMAND_ENFORCE boolean so an operator can switch
the clause off. SCHEDULER_TARGET_LOAD <= 0 already disables
it (mariadb.py:24999), that path is tested
(test_demand_clause_is_skipped_for_a_non_positive_target_load),
and a second switch for one clause is a knob that exists
because we were unsure, not because an operator needs it.
E5. The regression test is a property test over sizes, not
an example. #3813 is a statement about a family: for every
node size and every instance size, an idle node admits. One
example at cpu_schedulable: 2 would have passed against the
pre-phase-3 code and would pass against a floor that is
wrong at some other size. The test sweeps
cpu_schedulable in 1..16 against instance sizes in
{1, 2, 4, 8, 16} vCPU on an idle node and asserts admission
in every cell, and it must be mutation-tested against the
current clause rather than asserted to fail by inspection.
Computed at planning time: the current clause and seed admit
26 of the 80 cells, so the sweep must fail 54 of them before
the fix and none after. Retuning the constant alone (E1
without E2) admits 58 and still fails 22 -- which is the
arithmetic behind E2's claim that a retune moves the
threshold rather than abolishing it.
E6. Phase 4's close-out is this phase's step 4, and phase 3's soak rides with it. Both outstanding soaks want the same sfcbr deployment and the same CI burst, and neither is meaningful before E2 lands: soaking claim accounting while every placement takes the waiver path measures the waiver, not the accounting. Running them as one management-session step is cheaper and more truthful than running them twice.
E7. Close #3813 in this phase, not in phase 5. The master plan's success criteria make the whole plan uncloseable while it is open, and the phase 6 correction makes phase 6 unplannable while the mechanism is unsettled. It is closed by step 5, after the soak has been observed and not before -- the arithmetic is provable in a unit test but "the spreader actually spreads on real hardware" is not.
Design¶
The clause¶
_demand_guard_clause() (shakenfist/mariadb.py:24976-25013)
keeps its signature, its None returns and both of its
fail-open behaviours:
target_load <= 0still returnsNoneand skips the clause entirely (the proto3 unset-double case, and E4's disable path).schedulable IS NULLstill passes, for a node whose resources daemon has not yet published typed columns.- A NULL
cpu_load_1with a known thread count still coalesces to zero.
Only the comparison changes:
return sa.or_(
schedulable.is_(None),
sa.func.coalesce(load, 0.0) + capacity.c.expected_demand
<= target_load * schedulable)
demand_add is no longer read by the clause. It stays a
parameter of _direct_admit_instance_placement() and of the
RPC, because the UPDATE's SET still adds it to
expected_demand; the parameter simply stops being consulted
by the WHERE. Do not remove it from the signature, and do not
remove it from _admission_denial_dimensions(), where it
remains the requested figure the demand dimension reports.
What a denial now reports¶
_admission_denial_dimensions() (mariadb.py:25413-25429)
builds the demand dimension as limit = target_load x
cpu_schedulable, used = cpu_load_1 + expected_demand,
requested = demand_add, and _capacity_dimension()
recomputes exceeded as used + requested > limit. Under E2
that would report exceeded for a denial the new clause did
not make, and -- worse -- CapacityAdmissionDenied.demand_only
(shakenfist/exceptions.py:134-149) is derived from the
exceeded set, so a mis-set demand flag changes whether the
P9 waiver fires.
So the demand dimension's exceeded must be computed the way
the clause is: used > limit, with requested reported for
diagnosis but not added. That is a deliberate divergence from
the three allocation dimensions, and it needs a comment
saying why, because every other dimension in that function is
a before-and-after triple.
What does not change¶
The canonical statement order, the guarded-UPDATE-first
ER_CHECKREAD invariant, the node_guarded / cluster_guarded
/ claim_guarded split, the floored decrements, and the
reconciler's decay recompute. This phase touches one boolean
expression, one constant, and the exceeded derivation for
one dimension.
Execution¶
| Step | Effort | Model | Isolation | Brief for sub-agent | Status |
|---|---|---|---|---|---|
| 1 | high | opus | worktree | The clause and the constant. In shakenfist/mariadb.py, change _demand_guard_clause() (:24976) per the Design section: drop demand_add from the comparison, keep it as a parameter, keep both fail-open branches and the target_load <= 0 skip. Rewrite the docstring's opening formula sentence to the new form and say why the incoming charge is not tested (E2). In _admission_denial_dimensions() (:25354, demand dimension at :25413-25429), compute the demand dimension's exceeded as used > limit rather than letting _capacity_dimension() add requested -- read _capacity_dimension() first and either pass a zero requested with the real figure carried separately or set exceeded explicitly, whichever keeps the reply shape unchanged for its consumers (exceptions.py:134-149 derives demand_only from it, and test_node_denial_reports_the_demand_dimension_too at tests/test_mariadb_capacity_admission.py:586 pins the triple). Comment the divergence. In shakenfist/config.py, change SCHEDULER_DEMAND_PER_VCPU (:353) from 2.5 to 0.6 and rewrite its description per E1, citing the 00a-1 Measurements appendix rather than "pending the phase 0 step 3 data analysis"; leave SCHEDULER_DEMAND_DECAY_SECONDS alone. Tests: the E5 property sweep (cpu_schedulable 1..16 x instance sizes {1,2,4,8,16} vCPU, idle node, admission expected in all 80 cells) as a new test in tests/test_mariadb_capacity_admission.py, mutation-tested against the pre-change clause; a spreading test that a second placement into a node already carrying expected_demand above its budget is refused with demand_only true; and the existing demand tests at :440-465, :586-598 and :1467 updated where the formula changed and left alone where it did not. Do not touch CLAIM_ENFORCEMENT_HARD or either place_walk. Commit subject: scheduler: make the demand guard satisfiable. |
Complete -- the E5 sweep moved to step 2, see below |
| 2 | medium | opus | worktree | Live coverage, in tests/test_mariadb_capacity_admission_live.py beside test_the_demand_clause_refuses_on_measured_load (:416) and test_the_demand_clause_passes_on_null_metrics (:436). Two tests against a real server: a node with cpu_schedulable = 1 and zero load admits an 8-vCPU instance (the #3813 case at the smallest supported size), and a node whose expected_demand already exceeds target_load x cpu_schedulable refuses, with the reply's demand dimension reporting exceeded true and the allocation dimensions all false so demand_only is true. Follow the file's existing fixture pattern for seeding node_metrics and scheduler_node_capacity rows; note the suite reports the server regime it ran under, and issue #3759 means CI does not exercise MariaDB 11 here. Commit subject: tests: live coverage for the demand guard. |
Complete -- two existing live tests needed their premises restated, and the sweep landed here |
| 3 | low | sonnet | worktree | The post-fix half of the plan corrections; the provenance corrections already landed in the planning commit and must not be redone (see What the survey found). In PLAN-scheduler-reservations.md, move the #3813 Future work entry -- including its 2026-08-22 correction -- into "Bugs fixed during this work", condensed to what a reader needs after the fact: the units error, the fix, and the phase that made it. Clear the "carries an outstanding defect" clause from the phase 3 status note now that it is false. Check the master plan's success criterion for D13 (The D13 demand clause admits placements on a node that has real room for them, at every node size this project supports) reads true against what shipped, and say so rather than deleting it. Commit subject: docs: record the demand guard defect as fixed. |
Complete |
| 4 | low | sonnet | worktree | Operator and developer documentation. In docs/operator_guide/scheduler.md, state what the demand term now does in one paragraph: it spreads correlated bursts by refusing nodes already at or above SCHEDULER_TARGET_LOAD per schedulable thread, it never refuses a node with real allocation room, and when every node is over target the waiver admits anyway rather than failing the create. Give the two constants and their measured provenance. In docs/developer_guide/subsystem_internals.md, update the admission-transaction description beside the placement one to carry the new clause and the exceeded divergence from step 1. Check CLAUDE.md's scheduler capacity paragraph for anything the change falsifies and correct it if so; ARCHITECTURE.md and AGENTS.md only if the component inventory or a convention actually changed, which it should not have. Commit subject: docs: the demand guard is a spreader, not a bound. |
Complete |
| 5 | n/a | management session | none | Deploy to sfcbr and soak, discharging three outstanding obligations at once (E6): phase 3's step 9 soak, phase 4's step 10 operator review and soak with a real claim on a real namespace, and this phase's own validation. Run a CI burst and record, in this plan's Soak observations section: whether the demand clause now passes for some candidates and refuses others (read the 'schedule candidate refused by capacity guard' audit events and check enforce_demand is true on refusals that were then admitted elsewhere); how often the P9 waiver event fires, which under E3 should be rare and only under genuine saturation; whether a burst spreads across hypervisors rather than piling on the top-ranked node; and that the reconciler reports zero drift across the burst. Then the phase 4 claim soak proper: create a claim for a namespace, run instances in it, confirm the drawdown and that /admin/resources and the tables agree. Phase 00a's own post-deploy question -- whether the network+database node still takes a disproportionate share -- can be observed from the same burst; record it in phase 00a's plan, not this one. |
Complete -- deployed and soaked 2026-08-22 to 2026-08-24; the claim half needed a deliberate exercise, see the step notes |
| 6 | low | sonnet | worktree | Close-out, after step 5 has been recorded. Set the phase 3, phase 4 and phase 4a rows to Complete in the master plan Execution table, remove the phase status notes that describe the soaks as outstanding, and confirm docs/plans/index.md's row arithmetic is right for the new phase count (the phases column is arithmetic over the Execution table; adding 4a changes the denominator). Close #3813 with a comment naming the fix and the soak observation. Commit subject: scheduler: close out phases 3, 4 and 4a. |
Complete |
Step notes¶
- Step 1 was planned to carry the E5 satisfiability sweep. It could
not:
test_mariadb_capacity_admission.pyruns against a mocked connection and asserts on compiled statement text, so it can show the charge has left the WHERE clause but cannot evaluate whether the resulting arithmetic admits anything. The sweep moved to step 2 and is evaluated as SQL against a real server, which is where the arithmetic that broke actually lives. What step 1 proves instead is thatdemand_addbinds nowhere in the compiled comparison and still binds in the SET. - Step 2 found two existing live tests whose premises the fix
invalidated, neither of which was in the plan.
test_the_demand_clause_refuses_on_measured_loadpublished a load that only exceeded the budget once the placement's charge was added, andtest_admit_release_cycling_returns_to_the_seeded_counterscrossed the budget on its fifth round under the old arithmetic. Both were restated rather than relaxed: they assert the same facts, at loads and round counts that are true of the new clause. This is the reason the phase wanted live coverage at all -- the unit suite passed throughout, because the live modules skip without a database. - Step 2 also ran the whole 364-test capacity suite against MariaDB
11.8, which is past the 11.6.2 boundary where
innodb_snapshot_isolationturns a transaction's leadingSELECTintoER_CHECKREAD. CI runs 10.11 and cannot see that (#3759), so this is the first time this phase's transactions have been exercised under the regime the ER_CHECKREAD invariant exists for. - The planning commit's mutation prediction held: the sweep refuses exactly 54 of the 80 cells against the pre-fix clause and seed, which is the figure computed from the arithmetic before any code was written.
- Step 5 could only half happen by waiting. The demand-guard
observations came from a 48-hour sfcbr window, as planned, and are
recorded in Soak observations. The phase 4 claim half could not:
nothing on sfcbr creates a namespace capacity claim on its own, so
the window produced zero claim requests, zero over-limit audit
events and zero expiries -- an absence, not a result. The pathway
was exercised deliberately instead, with
tools/exercise-namespace-claims.py. Waiting longer would not have helped, and the plan should have said so when it wrote the step. - Review round 1 (2026-08-22) raised ten items, of which five
changed the code. The mock database's demand comparison still
implemented the pre-fix arithmetic, so the caller-side P9 waiver
tests were exercising a walk production no longer takes; it is now
check-then-charge like the real clause, with a caller-side test that
fails against the old mock.
_demand_guard_clause()no longer takes the charge it does not read, which turns "do not wire this back in" from a docstring into a signature._capacity_dimension()'schargedflag is keyword-only. The live sweep's instance-size axis was degenerate -- the clause cannot vary by instance size any more -- so it split into a clause-level node-size sweep and a behavioural instance-size sweep through a real admission, and the prose in both plans that claimed an 80-cell grid was corrected to say what the tests actually run. Two wording items (the<=boundary, and demand residue surviving release) were documentation fixes. One item was declined with reasons recorded in Future work: a smoke-CI assertion that no waiver event fires would flake whenever the CI cluster is legitimately at target.
Risks and mitigations¶
- The new clause lets a burst over-commit a node before the reconciler catches up. Check-then-charge admits one placement onto an at-target node, and only the next admission sees the increment. Bounded by the guarded UPDATE serialising within a node, and by the three allocation dimensions which are unchanged. Checked by: step 1's spreading test, and step 5 reading the burst distribution on real hardware rather than trusting the unit test.
- The
exceededdivergence silently changes waiver eligibility.demand_onlyis derived from theexceededset, so getting the demand dimension's derivation wrong makes the waiver fire when it should not (masking real denials) or not fire when it should (507ing creates the cluster has room for). Checked by: step 1 updatingtest_mariadb_capacity_admission.py:1627's waiver- eligibility tests explicitly, step 2's live assertion that a demand-only refusal reports exactly that, and the management review reading_capacity_dimension()andexceptions.py:134-149together. - The denial-detail re-read can now suppress a waiver that
should fire (raised in review).
_admission_denial_dimensions()reads its numbers after the transaction rolled back, andcharged=Falsemakes the demand dimension'sexceededa strictly tighter test than before -- the placement's charge used to act as accidental slack against drift between the guard and that read. Ifcpu_load_1orexpected_demandfalls in between, a genuine demand-only refusal reports nothing exceeded,demand_onlyis False (it requiresexceeded == {'demand'}, and the empty set is not that), the walker does not re-walk, and an exhausted candidate list 507s a create the waiver would have admitted. Narrow: both inputs move on 60-second and five-minute cycles against a read milliseconds later, andexpected_demandis only ever increased except by the reconciler. Checked by: step 5's soak looking for unexplained 507s. If any appear, the cheap fix is to treat an emptyexceededset at the node stage as waivable -- such a denial already means the detail read disagrees with the guard, and the allocation dimensions still bound the re-walk. Recorded in Future work. - 0.6 is still wrong. It is an estimate from one incident on one cluster, and this phase promotes it from provisional to cited. Mitigated by it now being dimensionally consistent, so being wrong makes the spreader too eager or too lax rather than inert; and by step 5 measuring the achieved figure again during the burst. If the soak contradicts it, the constant moves and the clause does not.
- The soak conflates three questions. Step 5 discharges obligations from three phases at once, and a single "it looked fine" observation would let all three through without evidence. Checked by: the step brief enumerating what must be separately recorded, and step 6 refusing to flip a status whose observation is not written down.
- Scope creep into phase 5. A satisfiable guard makes
flipping the hard ceiling look easy. Checked by: the Scope
section's explicit guard, and a
git diffreview thatmariadb.py:24815and bothplace_walkcopies are untouched.
Definition of done¶
Falsifiable, and mostly runnable:
- An idle node with zero
expected_demandadmits, at everycpu_schedulablein 1..16 (clause-level, evaluated as SQL) and at every instance size in {1, 2, 4, 8, 16} vCPU on a two-thread node (behavioural, through a real admission). These are two tests rather than an 80-cell grid, and deliberately so: after the fix the clause does not take the placement's charge at all, so a grid over both axes would evaluate 16 distinct expressions five times each and claim more evidence than it produces. The instance-size axis is therefore exercised where it can still vary the outcome, which is a real admission.
Mutation-tested rather than asserted by inspection. Against
the pre-change clause and seed the combined property fails
54 of those 80 combinations; against the corrected seed with
the old clause shape it still fails 22.
* A node whose cpu_load_1 + expected_demand already exceeds
SCHEDULER_TARGET_LOAD x cpu_schedulable refuses, the
denial's demand dimension reports exceeded true, every
allocation dimension reports false, and
CapacityAdmissionDenied.demand_only is true.
* git grep -n "demand_add" shakenfist/mariadb.py shows no
occurrence inside _demand_guard_clause()'s returned
expression, and the parameter is still in its signature.
* grep -n "2\.5" shakenfist/config.py returns nothing.
Checked at planning time: the only occurrence in the file
today is SCHEDULER_DEMAND_PER_VCPU's default at line 354,
so this is an absolute check rather than a field-scoped
one. Its description must also contain no occurrence of
"provisional seed pending".
* git diff develop -- shakenfist/mariadb.py | grep -E
'^[+-].*CLAIM_ENFORCEMENT_HARD' returns nothing -- changed
lines only, since context lines around an unrelated edit
would otherwise trip it (phase 5 scope guard).
* A live test admits an 8-vCPU instance onto a node with
cpu_schedulable = 1 at zero load, against a real server.
* No fact about the demand clause is stated differently in
docs/operator_guide/scheduler.md,
docs/developer_guide/subsystem_internals.md, CLAUDE.md,
the master plan, and the phase 0 decisions document.
* The phase 0 decisions document no longer attributes 2.5 to
the 00a-1 measurements, and the master plan's phase 6
correction no longer claims the affinity stage has nothing
to rank.
* Soak observations for the demand clause, the P9 waiver
frequency, the burst distribution, reconciler drift, and
the phase 4 claim drawdown are each written into the Soak
observations section below -- five separate observations,
not one summary.
* The master plan's phase 3, phase 4 and phase 4a rows read
Complete, the phase status notes no longer describe an
outstanding soak, and docs/plans/index.md's phase
arithmetic matches the Execution table.
* #3813 is closed.
* pre-commit run --all-files passes.
Soak observations¶
Pre-fix baseline (48h to 2026-08-23 06:19 UTC)¶
Recorded before the phase 4a deploy, and not something the plan anticipated having. PR #3843 merged at 2026-08-22 11:26 UTC but sfcbr was not redeployed until 2026-08-23, so the two days of production traffic either side of the merge all ran the old clause. That makes this window a genuine control arm rather than a recollection, and the post-deploy numbers below are a measured delta rather than a judgement call.
The window was confirmed to be pre-fix from the audit events themselves, not from deploy records. A refusal at 2026-08-23 06:19:17Z read:
Both fields date the code. requested: 10.0 for a 4-vCPU instance is
2.5 per vCPU, the pre-retune SCHEDULER_DEMAND_PER_VCPU; the shipped
default of 0.6 would read 2.4. And exceeded: True while
used 9.36 < limit 16.5 can only be used + requested > limit, the
charged comparison this phase replaced with charged=False. Under the
shipped code that node admits.
Every refusal was a demand refusal. Across 3,904
schedule candidate refused by capacity guard events, the count whose
exceeded set contained any of cpus, memory_mb or disk_gb was
zero:
| node | cpu limit | refusals | demand-only |
|---|---|---|---|
6046afdf |
66 | 686 | 686 (100%) |
f6b7e913 |
66 | 682 | 682 (100%) |
bed5996d |
30 | 669 | 669 (100%) |
963d4df9 |
30 | 639 | 639 (100%) |
f4ba9b6c |
24 | 651 | 651 (100%) |
7ce66641 |
18 | 577 | 577 (100%) |
| total | 3,904 | 3,904 (100%) |
This is #3813 stated as a measurement. The demand term was not one input to placement among several -- it was the only input that ever refused anything, on every node, at every size -- including the two largest, refusing a 4-vCPU request while holding 20 of an admissible 66 vCPUs. The allocation dimensions the capacity tables exist to enforce never once bound.
Consequently the P9 waiver, designed for genuine saturation, carried the majority of all traffic:
| measure | pre-fix |
|---|---|
| placements | 1,052 |
P9 waivers (enforce_demand false) |
648 (62%) |
| hourly waiver rate | 25% -- 86% |
The hourly spread is worth keeping, because it sets the bar for what counts as evidence after the fix: an unchanged clause varies by a factor of three from hour to hour, so a one- or two-hour post-deploy sample cannot distinguish the fix from ordinary variance. The discriminating hours are the busy ones -- 2026-08-22 22:00Z placed 72 instances at an 82% waiver rate.
Placement concentrated on the two largest nodes, which is the spreading failure the term exists to prevent:
| node | cpu limit | placements | share |
|---|---|---|---|
6046afdf |
66 | 330 | 31% |
f6b7e913 |
66 | 303 | 29% |
bed5996d |
30 | 130 | 12% |
963d4df9 |
30 | 125 | 12% |
f4ba9b6c |
24 | 102 | 10% |
7ce66641 |
18 | 62 | 6% |
Two caveats on reading the concentration figure. It is not a clean
measure of the demand term's spreading. With the guard refusing every
candidate, 62% of placements were made by the P9 waiver's second
place_walk(False), which iterates the same ranked candidate list as
the first (external_api/instance.py:881-921) but with the clause
waived -- so it simply takes the highest-ranked candidate, and the
demand term contributed nothing to where those instances went. The 60%
on the top two nodes therefore reflects the scheduler's affinity-then-
load-bucket ranking (scheduler.py:653-691) preferring the largest
nodes, which is partly correct behaviour. What the fix should change is
the mechanism before it changes the distribution: placements should
be made by the first walk with enforce_demand true. Second, the node column here is capacity limit
(vCPUs admitted at the 3.0 overcommit ratio), not physical size, so the
66/30/24/18 spread is roughly a 22/10/8/6 thread spread.
Post-deploy observations¶
Deployed to sfcbr on 2026-08-23 between 06:19Z (last event showing the old arithmetic) and 16:56Z (first showing the new). Measurements below cover 16:00Z 2026-08-23 to 18:00Z 2026-08-24, about 26 hours and 438 placements, against the 48-hour 1,052-placement baseline above.
The deploy is confirmed from the events, not from deploy records.
Refusals now charge 0.6 per requested vCPU (requested reads 0.6 for a
1-vCPU request, 1.2 for 2, 2.4 for 4), and across all 88 refusals there
is no case of exceeded being true while used <= limit -- the
comparison is used > limit, the charged=False form. Under the old
constant those same requests would have been charged 2.5, 5.0 and 10.0.
Observation 1 -- the demand clause discriminates. It does, though
the honest headline is a rate change rather than a composition change.
All 88 refusals are still demand-only; no refusal in either window was
caused by cpus, memory_mb or disk_gb, because sfcbr never
approaches those bounds. What changed is how often the clause fires at
all:
| pre-fix | post-fix | |
|---|---|---|
| refusals | 3,904 | 88 |
| placements | 1,052 | 438 |
| refusals per placement | 3.71 | 0.20 |
An 18-fold reduction. Every one of the 88 carries enforce_demand:
true, so they are first-walk refusals, and only 18 of them escalated
to a waiver -- the rest were refused on one candidate and admitted on
another, which is exactly the discrimination the clause was supposed to
provide and never did. Every node still refuses sometimes and admits
most of the time, including the 18-vCPU node that the old clause could
never satisfy.
Observation 2 -- the P9 waiver is now rare. This is the clearest result:
| pre-fix | post-fix | |
|---|---|---|
| waiver rate | 648/1,052 = 62% | 18/438 = 4.1% |
| hourly range | 25% -- 86% | 0% in 17 of 20 hours |
Fourteen of the eighteen waivers fell in a single hour (2026-08-24 06:00Z, 49 placements, 29%), with the remainder in two hours at 6% (3 of 49) and 3% (1 of 31). The baseline's threefold hourly variance was the reason for insisting on a long sample, and it is what makes this readable: the post-fix distribution is not a quiet-hour artefact, because the busiest post-fix hours (50 placements at 19:00Z, 49 at 09:00Z) ran at 0% and 6%, against a pre-fix busy hour of 72 placements at 82%. The waiver has gone back to being an exception under genuine burst pressure, which is what E3 kept it for.
Observation 3 -- burst distribution. Placement moved off the top-ranked node and onto the smallest, which is the spreading the term exists to produce:
| node | cpu limit | pre-fix | post-fix |
|---|---|---|---|
f6b7e913 |
66 | 28.8% | 30.6% |
6046afdf |
66 | 31.4% | 23.1% |
bed5996d |
30 | 12.4% | 13.9% |
963d4df9 |
30 | 11.9% | 14.6% |
f4ba9b6c |
24 | 9.7% | 8.7% |
7ce66641 |
18 | 5.9% | 9.1% |
The top-ranked node shed 8.3 points and the smallest node gained 3.2, taking it from well under its capacity share to slightly over. Two honest qualifications. First, total absolute deviation from a capacity-proportional split actually rose, 7.5 to 13.5 points -- but capacity-proportional is not the target and never was. The clause spreads by measured load plus expected demand, so a node running hot is skipped whatever its size, and a deviation metric anchored on capacity will read that as a regression. Second, 438 placements across six nodes is a small sample and some of this spread is noise. The mechanism change is the durable finding; the distribution is corroborating, not load-bearing.
Observation 4 -- reconciler drift. 299 reconcile passes over the
window, every one seeing all six nodes, with nodes_added,
nodes_removed and claims_expired all zero throughout and durations
of 27ms/91ms/3787ms (min/median/max). No pass logged a correction.
This is weaker evidence than the Definition of done implies, and the
gap was in the instrumentation rather than the result.
reconcile_scheduler_capacity recomputes the counters wholesale and
used to log membership and timing but no before/after delta, so a
silent correction of a used_* counter left no trace in the log.
Proving zero drift needed the Prometheus gauges, which were not
reachable from the analysis host. What can be said of the window is
that the reconciler ran healthily every five minutes throughout, saw a
stable cluster, and expired nothing -- which is consistent with zero
drift but does not demonstrate it.
The gap turned out to be smaller than it looked and is closed by this
close-out rather than deferred. The per-node delta_used_cpus /
delta_used_memory_mb / delta_used_disk_gb values were already in
the ReconcileSchedulerCapacity reply; only the log line dropped them.
The pass now warns per drifting node and carries drifted_nodes plus
per-dimension drift_* magnitudes on the summary line
(daemons/cluster/scheduled_tasks.py). Magnitudes are summed absolute
rather than signed, so drift in opposite directions on two nodes adds
instead of cancelling into an apparently healthy total.
This means the phase 3 zero-drift criterion was accepted on healthy-pass evidence, not on measured deltas. A reader should not take phase 3's Complete as "checked". The next soak can check it properly, which is the point of fixing the instrumentation now rather than filing an issue against a plan already closed.
Observation 5 -- phase 4 claim drawdown: done deliberately, not by
waiting. The 48h window produced no claim activity at all: zero
requests to /auth/namespaces/<namespace>/claims, zero placement
admitted over namespace capacity claim audit events, and
claims_expired zero in all 299 reconcile passes. That is not a soak
result, it is an absence of one, and no amount of further waiting would
have changed it -- nothing on sfcbr creates claims on its own. The
pathway was therefore exercised on purpose, with
tools/exercise-namespace-claims.py against sfcbr on 2026-08-24: 33
checks, all passing, covering request validation, create, duplicate
refusal, read and cross-namespace non-disclosure, field-masked update,
re-dating, shrink-to-zero-and-back, drawdown against a real instance
(used_cpus=1 used_memory_mb=1024 used_disk_gb=8 after placement),
the below_usage shrink refusal, expiry, and delete.
One correction to that record, found in review. The run did not
demonstrate the 507 capacity refusal, although an earlier draft of this
section said it did. The impossible-claim check ran after the real
claim already existed, and create_namespace_claim probes for an
existing claim before it reaches the guarded UPDATE against
cluster_capacity, so the request could only ever return the 409 for
exists. The check had been widened to accept either status, which
made it a no-500 smoke test rather than the capacity assertion its name
promised. The check now runs before the claim is created, where
capacity is the only thing that can refuse it, and asserts 507
specifically. The 507 path itself was never untested -- test_claims.py
and test_mariadb_capacity_claims_live.py cover it -- so what was wrong
was this record, not the code.
The expiry result is the one worth recording, because it is the only
one that could not be reasoned out from the code. coverage_state is a
stored column swept by the reconciler rather than computed on read, so
a claim stays active past its own expires_at until a pass runs. The
observed sweep latency was 210s, 280s and 321s across three runs
against a five-minute reconcile interval -- the interval behaving as
designed rather than a stall, and the spread is simply where in the
cycle the claim happened to expire. Through that window the claim reads
state: created while refusing a grow with 409 not_active -- D2's
two-facts distinction visible from the outside, which is what the phase
4 soak was for.
Four bugs surfaced, all in the exercise script, none in the claims
code. Three were assumptions about asynchronous object lifecycle: a
non-existent default image, an instance created before its network left
initial, and a namespace deleted before its network had finished
going away. The fourth is worth recording as a testing lesson rather
than a defect: the harness returned each check's detail string as its
result, and callers used that as a "did this pass?" guard, so two
checks whose implementation had no detail to report were silently
skipped instead of run. They were never reported as skipped and never
failed -- the run was green because the assertions did not execute.
A passing check now returns True when it has no detail. The claims
assertions themselves passed on their first run.
Future work¶
- Phase 00a's post-deploy validation -- whether the
network+database node still takes a disproportionate share
of a CI burst -- remains outstanding against phase 00a, and
is the last thing keeping that phase
In progress. Step 5's burst is the natural occasion to collect it, but it is recorded there, not here. - The per-namespace learned demand value D13 defers. Nothing in this phase makes it harder; the constant it replaces is now at least dimensionally correct.
SCHEDULER_DEMAND_DECAY_SECONDS = 600is still an unmeasured provisional seed. If step 5's burst gives a usable time-to-visible-load figure, it belongs in the 00a Measurements appendix.- The duplicated
place_walkinexternal_api/instance.pyandoperations/node_inst_netdesc_op.pyis still two copies with a comment asking that changes be made in both. Phase 5 owns extracting it; this phase deliberately leaves it, having touched neither. - Issue #3759 (a MariaDB 11 CI job for the ER_CHECKREAD invariant) is unchanged by this phase but gates how much the step 2 live tests actually prove in CI.
- An empty
exceededset at the node stage is not waivable (see the risk above). Whetherdemand_onlyshould treat it as waivable turns on whether step 5's soak sees any unexplained 507s; it is a behaviour change to the walkers and does not belong in a phase whose job is to make the clause satisfiable. expected_demandis not credited back on release, so a cluster under rapid create/delete churn can read as over target on demand accumulated by instances that are already gone, until the next five-minute reconcile pass. This is pre-existing and deliberate -- the contribution has partly decayed, so crediting the original figure would over-credit -- but the clause binding for the first time is what makes it observable, and the operator guide now documents how to tell residue from load. Step 5 should measure how often it happens; if it is frequent, the options are crediting a decayed figure on release or shortening the reconcile interval.- No functional-CI coverage of the end-to-end property
(raised in review): that a create on a two-thread hypervisor
succeeds on the first walk with no waiver. Deliberately not
added. The obvious form -- assert no
waiving demand guardevent on a create in the cluster suite -- is a flake generator, because a CI cluster genuinely at target during a burst fires that event correctly and the test would fail for the right reason. The property is covered at the database level by the live suite (which does run in CI viatools/ci-enum-widening-test.sh) and at the caller level bytest_an_idle_node_admits_a_large_instance_on_the_first_walk, which asserts the absence of the waiver event against the mock. Real-hardware proof stays with step 5.
Back brief¶
Before executing, back-brief the operator on:
- E2, the clause's new shape. This is the decision most likely to be argued with: dropping the incoming placement's charge from the comparison changes the question the guard asks, and the runner-up (flooring the budget) is a smaller diff. Confirm the reasoning about check-then-charge being safe inside a guarded UPDATE before any code moves.
- E1, 0.6 rather than 0.33. Conservative burst figure over steady-state, on the grounds that the term exists for bursts.
- E6, folding three soaks into one step. Confirm the operator is willing to run one sfcbr burst that discharges phase 3, phase 4 and phase 4a, and to record five separate observations from it.
- E7, closing #3813 only after the soak. The alternative is closing it when step 1 lands, which is defensible and faster.