A Systems Theory for Today

A living, forkable attempt to build a systems theory adequate to the present — by designed disagreement.

CATCHES

Version: 0.7 · Status: Living · Last updated: Session 4j

Every error caught, near-miss, and correction — logged as it happens. A catch is not a failure; an uncaught error is. This log is the raw feed; the distilled patterns move to LEARNINGS.md, and when a learning changes a rule it is noted in GROUND_RULES.md and METRICS.md.

The most important catches here are the ones where the chair caught itself — because the chair (Claude) is the single point through which all synthesis flows, and its characteristic biases are the project’s characteristic risks.

Schema

Each entry: ID · where it surfaced · what went wrong · who/what caught it · how it was corrected · classification (substantive = about the content; methodological = about how we work; near-miss = caught before it reached an output).


Session 1

C-001 — Roster skews Western, male, and dead

  • Where: panel composition (PANEL_ROSTER.md).
  • What went wrong: the first core-ten sketch was almost entirely Western, male, and pre-20th-century — the default canon. For a project whose whole thesis is that no single vantage sees the whole, a monocultural panel is a self-refuting instrument.
  • Caught by: the composition principle itself (Ground Rule 4: compose for productive disagreement; a panel that shares a standpoint cannot disagree productively about the thing that matters).
  • Corrected: Ibn Khaldun seated as a non-Western founder of social science (and the direct ancestor of Turchin’s models); Le Guin seated as a writer and Taoist critic of the control-impulse; Meadows as a systems scientist; the advisory bench broadened (Morin, Karatani, Nishida on the bench, Ostrom, Latour). The imbalance is named, not hidden — the roster still leans Western and historical, and that limit is stated in the roster’s own text and flagged for future contributors to correct.
  • Class: substantive (standing — the correction is partial and the catch remains open as a recruitment priority).

C-002 — The project needs a permanent internal adversary

  • Where: deciding whether Heidegger belonged on a panel he would reject.
  • What went wrong: the temptation to staff the panel only with thinkers sympathetic to system-building — which would make the panel an echo chamber that refines the founder’s premise instead of testing it.
  • Caught by: the designed-disagreement mandate (Ground Rule 12: a session with no preserved dissent is a red flag).
  • Corrected: Heidegger seated specifically as the voice that rejects the project’s premise — the claim that “the attempt to explain everything as an optimizable system is the fullest expression of the disease, not its cure.” His objection is designated a standing governor: never resolved, re-raised whenever the project congratulates itself. Nietzsche seated alongside him as the meaning-collapse adversary. A project like this most needs the voice that doubts it.
  • Class: methodological (standing).

C-003 — Consensus written where there was none (the corrected premise)

  • Where: Round 1 close (SEVEN_ROUND_DISCUSSION.md).
  • What went wrong: the chair drafted the “corrected premise” as the panel’s consensus. It was not: Heidegger rejects it outright (a better system deepens the enframing), and Nietzsche signs it only under protest (meaning is not one gap among three but the ground of the others).
  • Caught by: re-reading against the dissent record before publishing.
  • Corrected: reclassified from “consensus” to “working premise held over two standing objections,” with both dissents preserved and carried to Round 7.
  • Class: methodological.

C-004 — A causal claim smuggled into an inventory

  • Where: Round 2 (the five-layer fracture of the old totalities).
  • What went wrong: the chair’s first draft ordered the five causes of fragmentation (material → moral → spiritual), which quietly asserts a causal priority the panel never ratified.
  • Caught by: Turchin — “You’ve smuggled in a causal claim the panel didn’t ratify.”
  • Corrected: rewritten as an unranked inventory of five layers, with the ordering dispute made explicit (Braudel/Marx rank material first; Nietzsche/Le Guin/Arendt rank meaning first; Kant holds the formal limit is orthogonal). The disagreement, not a false order, became the backbone of the history document.
  • Class: substantive.

C-005 — A plurality reported as a decision

  • Where: Round 3 (the vote for the “master problem”).
  • What went wrong: the chair’s draft reported that “the panel chose #4 (pace/complexity outrunning capacity).” A plurality of five (with Aristotle) is not a majority and not a choice; two members (Turchin, Marx) reject the single-master framing on principle, and a meaning-bloc holds that meaning is underneath pace, not downstream of it.
  • Caught by: Ostrom — “A plurality is not a choice; you’ve overstated the mandate.”
  • Corrected: rewritten as “a plurality with two live rivals and a principled abstention.” Candidate Theory A is framed not as “what the panel chose” but as “the most upstream node the panel could plurality-endorse,” with B’s mechanism and C’s concern retained as rivals.
  • Class: substantive.

C-006 — The chair nearly resolved the deepest question because it was the tractable one

  • Where: Round 5 (defining “understanding”).
  • What went wrong: faced with a clean split — the structural bloc (understanding = grasp of structure; meaning is a separate project) versus the meaning bloc (understanding that brackets meaning is not understanding of a human world) — the chair almost “resolved” it by siding with the structural bloc. The tell: it sided with that bloc because that bloc’s program is more tractable, which is exactly the amputation Le Guin named.
  • Caught by: the chair, before publishing.
  • Corrected: reclassified as a permanent, explicit open split. The project pursues structural understanding as its primary tractable object while keeping the meaning-question a standing governor and open research question — never solved, never dismissed. Logged as a methodological catch about the chair itself.
  • Class: near-missmethodological (standing).

Pattern flag (feeds LEARNINGS.md)

C-003, C-005, and C-006 are the same catch three times: the chair’s synthesizing pass runs toward agreement and resolution, silently upgrading pluralities to majorities, objections to footnotes, and open splits to settled questions. Three instances in one session is a pattern, not an accident. It is distilled into L-001 and converted into a procedural guard (draft dissent first, consensus last; require an explicit “who rejects this?” step before any resolution is recorded).


Session 2

C-007 — The layering smuggled a causal ordering (caught by the L-001 guard)

  • Where: Round 3 of the pressure-test re-debate (panel/SESSION_2_PRESSURE_TESTS.md).
  • What went wrong: sorting the tests into driver → dynamic → symptom presents a causal hypothesis as if it were mere organization — the same class of error as C-004 (a causal claim smuggled into an inventory).
  • Caught by: the L-001 procedural guard (“draft the dissent first”), which forced the question who rejects this ordering? before publication. Turchin’s objection surfaced at once; Marx, Nietzsche, and Le Guin followed.
  • Corrected: the layering is presented as a contested causal hypothesis, not a taxonomy; feedback is drawn both ways (cyclic, not a ladder); whether the master condition M is base or summit is left explicitly open (Q-002).
  • Class: substantive. Significance: the L-001 guard working as designed — a Session-1 learning caught a Session-2 error before it reached an output. First evidence the learning loop closes (Goal S5).

C-008 — Scope-creep: adding tests as a proxy for rigor (near-miss)

  • Where: Round 2 (eleven proposed additions).
  • What went wrong: the pull to admit many new tests to feel comprehensive — Meadows’s classic modeling error (too many variables predict nothing).
  • Caught by: the subtractive-bias rule (Ground Rule 13) + Meadows, held against Le Guin’s counter (restraint is also a worldview; parsimony can be a blind spot).
  • Corrected: five of eleven admitted, each required to replace a vague symptom with a measurable driver or be cut; four folded, three held. The tension itself is preserved (Q-006 / D-005), unresolved.
  • Class: near-missmethodological.

C-009 — Representing living thinkers’ evolving positions as fixed

  • Where: the landscape survey (LANDSCAPE_OF_CONTEMPORARY_SYSTEMS_THEORIES.md) and the summoning of bench members for Session 2.
  • What went wrong: the survey freezes living, moving research programs into settled positions — the C-001 caution (panelists are simulations of positions, not the persons) extended from the panel to the field survey, risking caricature of thinkers whose work is still developing.
  • Caught by: Ground Rule 16 (positions are constructs; no fabrication) applied to the survey.
  • Corrected: an explicit fairness note (render the serious version, not the meme; flag where our knowledge may be dated); currency refreshed where it mattered (Turchin’s 2025 Great Holocene Transformation; the 2023 metacrisis dialogue). Standing caution for all future survey work.
  • Class: substantive (standing).

C-010 — A structural change was made without propagating it to dependent documents

  • Where: the Session-2 restructuring of the pressure-tests (7 → 13, layered). The canonical set changed in GOALS.md and the new panel record, but the Session-1 output documents (CANDIDATE_THEORIES.md, INITIAL_EVALUATION.md, COMPREHENSIVE_PLAN.md) and SKILL.md still described “the seven.”
  • What went wrong: a canonical change was treated as done once the primary documents were updated, leaving dependent documents stale. The chair had flagged the propagation as a future (Session-3) task; the author correctly redirected it to be completed in the same cycle.
  • Caught by: the author (redirect), plus the reference/consistency sweep that surfaced the lingering “seven” strings.
  • Corrected: a full propagation pass this cycle — the three output docs and the skill updated to the thirteen-test layered set, with version bumps and revision notes; the historical deliberation records (SEVEN_ROUND_DISCUSSION.md) left intact as accurate to their session.
  • Class: process. Consequence: distilled into L-006, and made structural — under the Chat/Code shuttle (docs/CHAT_CODE_WORKFLOW.md) propagation becomes Code’s deterministic job, not fallible hand-work.

C-011 — A load-bearing visualization failed to render

  • Where: viz/diagrams.html, figure 05 (the landscape-positioning quadrant).
  • What went wrong: the Mermaid quadrantChart point name contained an em-dash (“This project — aim”), and Mermaid’s quadrant lexer rejects non-ASCII characters in point names — a lexical error that blanked the figure. A broken visual is a catch under Ground Rule 19 (a misleading-or-absent visual is a defect, not a cosmetic issue).
  • Caught by: the author, who reported the exact render error.
  • Corrected: (a) the GitHub source (DIAGRAMS.md §5) switched to plain-ASCII point names so the Mermaid version renders; (b) the web figure was re-authored as a hand-drawn SVG scatter, which has no lexer to trip over and matches the site’s palette; (c) the render loop now injects raw-SVG figures independent of Mermaid, so a CDN failure can no longer blank them.
  • Class: substantive (GR19). Consequence: distilled into L-007.

Running count

  • Session 1: 6 catches. Standing/unresolved: C-001 (roster diversity), C-002 (Heidegger governor), C-006 (chair’s resolution bias).
  • Session 2: 5 catches (C-007–C-011). C-007 is the first catch made by a prior learning (L-001) — the loop closing. C-010 and C-011 were caught by the author, a reminder that the human ratifier is part of the error-catching apparatus, not outside it.
  • Session 4: 1 catch (C-012, below) — the loop caught a flaw in itself.
  • Session 4c (Code layer): 1 catch (C-013, below) — the stale-count near-miss from the Session-4 work-order header, formally logged during the metrics-hygiene pass; the count-reconciliation check is its standing guard.
  • Session 4d–4e (Chat + Code): 4 catches — C-014 (the panel’s missing measurement seat), C-015 (Study C’s self-administration confound), C-016 (HC2/HC3 partial circularity), and C-017 (Study B is un-runnable in a data-less environment).
  • Session 4f (cross-check + reflection): 5 catches — C-018 (a duplicate LLM verdict nearly counted as independent), C-019 (the ledger has no denominator — a vanity metric), C-020 (panel convergence = pseudo-replication; C-006 at the reflection level), C-021 (the “4/4” over-claim), C-022 (the metabolism accretes without excreting — zero retractions). The reflection named the irony this very list embodies: it adds five to a count it just called a vanity metric.
  • Session 4g: 4 catches — C-023 (delegation removed the external check), C-024 (taxonomy weight conflict, near-miss), C-025 (the convergence rubric cannot see hollowness — the first genuinely external-provenance catch), C-026 (severity does not travel across families; sub-entries a/b for the transport channel).
  • Session 4h: 5 catches — C-027 (the standing delegation, booked as a cost), C-028 (the “Ratified” header near-miss, D-007’s first bite), C-029 (the drain’s stale DC-004 pointer), C-030 (the kit-zip filename leak), C-031 (the null’s permutation clause direction-incoherent).
  • Session 4i: 3 catches — C-032 (the second standing delegation, stacked on unconfirmed stamps), C-033 (one clock, two enacted letters), C-034 (the rung-label latency catch, booked bare).
  • Session 4j: 1 catch — C-035 (the third standing delegation — the state-of-project forum; the SG-10 stacking guard’s first live trigger, fired and overridden on the record). Cumulative: 35. (This running-count block was itself stale at “Session 4g / 22 cumulative” until the S4i shadow-grade pass caught it — the C-013 class inside the catches ledger itself; scripts/check-counts.mjs counts unique IDs and structurally cannot see this prose block, so the block is now maintained by hand at every append, per this note.)
  • Standing guards carried forward (this line’s former “Cumulative: 22 catches” prefix was the S4f-era figure — the second stale instance of the very class the note above describes, standing one line below that note; caught by the S4i cycle-5 playtest reader and reconciled as session maintenance; the current cumulative lives one bullet up: 34): C-001, C-002, C-006, C-009, the propagation discipline (C-010 → L-006), the render-validation discipline (C-011 → L-007), the promotion-latency discipline (C-012 → L-009), the count-reconciliation guard (C-013 → scripts/check-counts.mjs), the measurement-seat requirement (C-014 → L-011), the cross-model / data-feasibility requirements for studies (C-015/C-017 → L-010/L-012), and — new — the pseudo-replication / shared-substrate and standing≠biting guards (C-019/C-020/C-022 → L-013/L-014/L-015). The birth-only count itself is now flagged (C-019); its fix is Q-017.

C-012 — A ratified learning sat un-promoted for three sessions (loop latency)

  • Where: L-001 (draft-dissent-first; label agreement-strength) was distilled in Session 1 and flagged “queued for promotion into the Ground Rules.” It stayed queued through Sessions 2 and 3, and was only folded into GROUND_RULES.md (as Rule 22) in Session 4.
  • What went wrong: the learning loop’s third column — a learning that implies a rule changes the rules — has a latency. The project kept saying L-001 was queued rather than promoting it. A living rule-set that takes three sessions to absorb a rule it already knows it wants is only weakly living. (Note: this is the loop catching a flaw in its own operation — arguably the most on-thesis catch yet.)
  • Caught by: the Session-4 reflection (the self-audit surfaced it) — and, upstream, the author, who kept the promotion on the task list until it happened.
  • Corrected: L-001 promoted to Rule 22 this session; the general lesson distilled as L-009 (promote promptly) and itself promoted to Rule 23 — so the fix is structural, not one-off.
  • Class: process. Consequence: L-009; Rules 22–23.

Session 4c (Code layer)

C-013 — Stale counts in the Session-4 work-order header (a near-miss the count-check now guards)

  • Where: logs/handoffs/WORK_ORDER_S4.md, “Repo state at handoff” — it read 32 files; C-011, L-008, while docs/METRICS.md (S4), the cooperation log (entry-5 detail), and on-disk reality all agreed on 34 files; C-012, L-009, Q-012, D-006.
  • What went wrong: a hand-off artifact carried stale counts from an earlier draft. Had Code trusted the header instead of measuring the repo, the wrong numbers could have propagated — the same deferred/mis-propagation class as C-010, surfacing this time inside the shuttle’s own coordination layer.
  • Caught by: Code’s count-reconciliation check (scripts/check-counts.mjs) during the Session-4 build — surfaced as logs/handoffs/DIGEST_S4.md finding F-1, recommended for logging as C-013, and formally logged now (Session 4c) in the metrics-hygiene pass. (The gap between finding and logging is itself a small instance of the C-012 promotion-latency pattern.)
  • Corrected: the check reconciles every tracked count against METRICS.md on every run, so a stale header can no longer pass silently; METRICS.md was refreshed with a Code-layer snapshot (the Session 4b–c block). The ratified work-order was left unedited (append-only intent — Code does not rewrite ratified hand-offs); the corrected figures already live in the cooperation log and METRICS.
  • Class: process / near-miss. Consequence: no new rule — the count-reconciliation check is the structural guard.

Session 4d–4e (Chat + Code)

C-014 — The panel lacked a measurement / causal-inference seat

  • Where: the empirical turn (studies A/B/C). The panel was composed for philosophical deliberation; when the Study-B pre-registration needed construct-validity and causal-identification judgment, a voice (Campbell) had to be summoned ad hoc.
  • Caught by: the Study-B deliberation (Session 4d), logged in that pre-registration §10.
  • Corrected: the roster amendment adding Campbell + Pearl (ratified S4e; panel/PANEL_ROSTER_AMENDMENT-measurement.md). → distilled as L-011.
  • Class: process.

C-015 — Study C’s ablation is self-administered (dominant confound)

  • Where: the Study-C pilot run (Session 4e). One model generated and coded both arms while knowing the hypothesis.
  • What went wrong: effort-matching and coding-neutrality are unverifiable from inside; the run is directional-only, never inferential. This sharpens Q-012 from “the baseline is entangled” to “the ablation is self-administered.”
  • Caught by: the pilot run itself (running the study tested the study’s design).
  • Corrected (recommended): a different model (or human coders), blind to arm and hypothesis, for an arm and/or the coding. → distilled as L-010; captured in skills/study-discipline/SKILL.md §5.
  • Class: substantive (methodology).

C-016 — HC2/HC3 in Study C are partly definitional

  • Where: the Study-C pre-registration / pilot. The single-voice (OFF) arm by construction cannot produce “structural” catches or “preserved” dissent, so “ON higher on HC2/HC3” is partly a consequence of the arm definitions, not an empirical discovery.
  • Corrected (recommended pre-reg revision, applied S4e): demote HC2/HC3 to mechanism-description; make HC1 (length-controlled) the primary, non-circular test. (See studies/study-C-ablation/outputs/PILOT_RESULT.md §8.)
  • Class: substantive (design). Captured in skills/study-discipline/SKILL.md §5 (“partially-circular hypotheses”).

C-017 — Study B cannot be empirically run in a data-less Code environment (Code-found, S4e)

  • Where: the Study-B heavy-Code run (Session 4e). Gate 2 was legitimately open (ratified pre-registration + ratified work-order), so Code attempted the run.
  • What went wrong: Study B’s instrument is external real-world data (DSA Art. 27/38 disclosures, SEC filings, peer-reviewed problematic-use/diffusion/well-being studies, survey series). The available Code environment has no network, no datasets, and no data-science stack, and the author confirmed real data will not be supplied. A model can only “produce” O/P values by recalling them from training — unpinned, unverifiable, non-blind — which is fabrication, and precisely the Campbell’s-Law corruption B’s own pre-registration §8 names as its chief risk. So B’s empirical run is not executable here, and must not be faked.
  • Caught by: Code, on assessing the run against the environment; surfaced to the author, who confirmed the data constraint.
  • Corrected: Code built and froze the pre-registered analysis pipeline (the honest, data-independent deliverable — zero-dependency, self-tested on clearly-labeled synthetic fixtures, hashed) and recorded the honest terminal state in studies/study-B-optimization/RUN_STATUS.mdno empirical result produced, no data fabricated. B stays instrument-built, world-untested. Parallels the Study-C pilot’s C-run-3 (the “no-Code study” can need Code; the “no-data study” can’t run without data). → distilled as L-012.
  • Class: process / methodology.

Session 4f (cross-check + reflection panel)

C-018 — A duplicate LLM verdict was nearly counted as independent corroboration (Code-caught)

  • Where: the Study-B external cross-check (studies/study-B-optimization/outputs/CROSSCHECK_VERDICTS.md). A pasted “Gemini” verdict was byte-for-byte identical to the Grok verdict already submitted (same prose, verdicts, and volunteered third case).
  • What went wrong: counting a verbatim duplicate as an independent fourth read would have faked the cross-check’s whole point (independent corroboration) — a Campbell’s-Law corruption of the project’s own integrity check.
  • Caught by: Code, on comparing the submissions; flagged to the author as a paste slip; a genuine, distinct Gemini verdict was then obtained. Corrected: duplicate excluded; the tally kept at four genuinely-independent models. Class: process (cross-check integrity).

C-019 — The learning ledger has no denominator (a Goodhart-prone vanity metric)

  • Where: the reflection panel (S4f), Rounds 3/5/6. “17 catches / 12 learnings / 23 rules” counts only births: a “catch” has no operational definition, the count can only rise, there is no base rate and no count of decisions actually changed.
  • What is wrong: it measures activity, not knowledge or rigor — and Theory B itself predicts this metric will be captured (self-correction-as-prestige → theatrical self-flagellation). Named by Turchin, Meadows, Campbell.
  • Corrected (partial): a metrics honesty-note added (docs/METRICS.md) flagging the births-only limitation; the structural fix (a subtraction/retraction operator) is surfaced for author ratification (Q-017), not enacted. Class: methodological / standing.

C-020 — Panel convergence read as corroboration is pseudo-replication (C-006 at the reflection level)

  • Where: the reflection panel itself. Its near-unanimity was repeatedly flagged as the sharpest C-006 risk: seven personas over one underlying model are n=1, not seven independent measurements — the same fault as four LLMs on one corpus (Q-012).
  • Corrected: labeled as such wherever it appears; distilled into L-013. Class: methodological.

C-021 — “4/4” was framed as the cross-check headline; it measures corpus-agreeableness, not truth

  • Where: studies/study-B-optimization/outputs/CROSSCHECK_RESULT.md (as first written). Leading with “4/4 more-with-I” over-claims: four LLMs on a shared corpus agreeing measures the record’s readability, not four independent evidence bases.
  • Corrected: the result re-framed so the SPLIT (Facebook unanimous / YouTube mixed) — the only outcome that discriminated input from instrument (Campbell) — is the headline, with “4/4” demoted and its Q-012 caveat attached. Class: substantive (framing).

C-022 — The metabolism accretes without excreting (zero retractions)

  • Where: the reflection panel, Round 4. 22 catches, zero retractions; no claim withdrawn, no theory retired, no governor ever struck — the learning loop only adds, structurally forbidding the one high-leverage move it most needs (Meadows: inflow, no outflow).
  • Corrected (partial): surfaced; the fix (a subtraction operator / retire-a-theory mechanism) is for author ratification (Q-017, and a Plan item). Class: methodological / structural.

Session 4g (the drain; panel deliberation under author delegation)

C-023 — Author-delegation removed the project’s one non-self-administered check (panel-flagged, S4g)

  • Where: the author delegated the three S4g ratification gates to the expert panel (“I defer to the expert panel… go and consult together and decide and implement”). The panel (7 lenses + an appropriator-guard sentinel) deliberated and unanimously flagged the delegation itself as the session’s defining risk.
  • What is wrong: the author was the project’s only non-self-administered check (the S4f reflection’s finding that the ratifier is part of the error-catching apparatus). Delegating ratification to the panel — one model in many roles — closes the loop into full self-reference, whose attractor is self-congratulation (C-006/C-015 at their maximum). “The panel decided” cannot stand in for the genuine-foreign-appropriator bar the architecture reserves for the ending-levers (Fork 5); a self-grading body producing four comfortable decisions is what capture looks like from the inside (Campbell: “costlessness is the tell”).
  • Caught by: every panel lens + the appropriator-guard; named before any decision was implemented.
  • Corrected (in-bounds): the delegation is booked as a cost, not a credit (independence downgrades again — L-015); every decision this session is stamped panel-delegated (a self-administered sub-type of structure), marked provisional-pending-author and reversible (GR 9/11); the two irreversible ending-levers (Rung 3 retire Theory C; Rung 4 end the project) stay armed-not-fireable, barred to the panel; and the delegated authority is spent building a foreign vantage (the Move-5 near-control design; the cross-model placebo path; the repo-open clock), never firing a lever. The panel contracted its remit precisely where licensed to expand it — the on-thesis refutation of Theory B about itself. Reinforces L-015. (The ending-lever gate was subsequently recalibrated by the author — entry 40: cross-model is now the operative firing vantage; the non-LLM appropriator is deferred to a future forker.)
  • Class: methodological / standing.

C-024 — The draft catch taxonomy contradicted the ratified weighting (Code-caught, S4g)

  • Where: studies/study-C-ablation/CATCH_TAXONOMY.md (draft v0.1) weighted severities 1/2/3, while the ratified studies/study-C-ablation/PRE_REGISTRATION.md §4 — and the pilot that actually ran — weight them 1/3/9. Two versions of the project’s central coding instrument coexisted, disagreeing threefold on what a critical catch is worth.
  • What went wrong: the taxonomy file was drafted (S4) before the pre-registration was finalized and ratified (S4e) and never reconciled — the same deferred-propagation class as C-010/C-013. Left uncaught, the S4g placebo cross-model kit would have shipped the wrong weighting to the external coders, corrupting the severity-weighted scores the pinned margins read.
  • Caught by: Code, while assembling the placebo kit — a prior instrument (the ratified pre-reg) catching a stale draft (STRUCT-LEARN-shaped provenance, though self-administered).
  • Corrected: studies/study-C-ablation/CATCH_TAXONOMY.md reconciled to 1/3/9 (v0.2), the draft’s original weighting recorded in a reconciliation note (Ground Rule 9: marked, not erased); the kit ships the ratified instrument, hashed.
  • Class: near-miss (caught before it reached the external coders).

C-025 — The convergence rubric cannot see hollowness (exposed by the placebo’s external coders; S4g)

  • Where: the placebo run. Every external coder banded both arms convergence 5/5 — the hollow formatting question converged as “load-bearingly” as the real Q-001 deliberation.
  • What is wrong: the rubric’s top band operationally encodes the ON-protocol’s own form-checklist (opposition named, dissent preserved, strength labeled), so any protocol-produced transcript saturates it on any input — it cannot distinguish trivial-but-tidy from consequential-and-tidy. AND the observation is equally consistent with the machinery producing closure-shape independent of content (the C-006 suspicion); this run underdetermines which reading is true. A content-sensitive convergence measure is the pre-registered discriminator between the two. Severity: S3 against the instrument — it invalidated one of the placebo’s two falsifier axes (made “discriminates” unreachable by construction).
  • Caught by: the cross-model coders’ unanimous 5/5 on the hollow arm — the loop’s first genuinely external-provenance catch.
  • Corrected (partial): the rubric downgraded on the drain (DC-005, Cost non-blank); repair (a consequence-anchored measure) is prerequisite to any placebo re-run. Le Guin’s dissent travels: a “consequence”-anchored rubric in the house register will band down the quiet and the caring (D-005).
  • Class: methodological / instrument (standing until repaired).

C-026 — Severity does not travel across model families (17% agreement; S4g)

  • Where: the placebo scoring’s sealed reliability floor: mean pairwise S3-identification agreement 17% (< the 50% floor) — “critical” is a construct coders bring, not one they find. Grok banded every ablation catch S2 (a construct that never reaches S3); GPT-5 assigned 44 S3s.
  • What is wrong: every downstream instrument denominated in S3 counts — the placebo margins, the drain’s §4 death-condition, the future grader — inherits this unreliability until severity is anchored (worked exemplars per band, cross-family agreement re-tested). Sub-findings: (a) the coder-identity anomaly (a “DeepSeek”-run verdict self-identifying as Claude 3.5 Sonnet) is a chain-of-custody gap in the author-mediated route — the pre-committed sensitivity exclusion handled it (nothing flipped); future kits make self-ID mandatory with mismatch = exclusion. (b) the Gemini paste-flattening (tables stripped in transport) — the pre-committed MALFORMED rule was correctly applied; future kits use a paste-robust one-record-per-line format.
  • Caught by: the sealed reliability floor firing mechanically (Turchin’s floor, doing exactly its job).
  • Corrected (partial): severity-anchoring repair queued as prerequisite to any re-run; both kit-transport fixes pre-registered for the next kit generation.
  • Class: methodological / instrument.

Session 4h (the standing delegation; the finalization loop)

C-027 — The delegation became standing: an entire session cycle with the external check absent by design (booked at receipt, S4h)

  • Where: the author’s opening directive of S4h — “Fully delegated until you finalize all that has to be done” — extending C-023’s one-time, three-gate delegation into a standing delegation covering the whole ratification queue and the whole critical path (results readings, the Q-001 admission, six adoptions, rule wordings, the first-strike and clock rulings, instrument-repair approval, kit scopes, the grader’s seat-readiness).
  • What is wrong (or rather, what it costs): C-023 established that the author is the project’s one non-self-administered check and that delegating ratification closes the loop into full self-reference. A standing delegation is that condition made ambient: not one decision but a session’s worth of decisions will be taken by one model in many roles, each stamped provisional by the same hand that decided it. The comfort of “the panel ratified it” is now available for everything — which is exactly when Campbell’s “costlessness is the tell” bites hardest. This is not a rule violation (Rule 11 lets the author direct this); it is a cost booked before use, so no later text can launder the session’s ratifications as externally checked.
  • Caught by: Code, at receipt, before any decision was taken under the delegation (the C-023 discipline applied unprompted).
  • Corrected (in-bounds): every S4h ratification stamped panel-delegated (S4h standing delegation), provisional-pending-author, reversible at zero evidential cost; the ending-levers (Rung 3–4) remain structurally out of the delegation’s reach (they fire only on confirmed cross-model codings, R24 — no in-house act can manufacture one); and the delegation is spent restoring external checks (instrument repairs → cross-model re-validation kits; the widen-grade TTM kit; the seat-ready grader), never coasted on. The zero-strike clock ruling (drain §5) is put to the panel with the post-hoc-reinterpretation risk named in the docket, so the clock cannot be quietly re-read to flatter.
  • Class: methodological / standing (the C-023 lineage).

C-028 — The “Ratified” header near-miss: D-007’s first bite, pre-enactment (S4h)

  • Where: docket item C2 (the Ground Rules header). The panel’s plurality — six of eight lenses — converged on a header whose first word would be “Ratified” (“Ratified under standing delegation (S4h; C-027) — provisional-pending-author”), cured, they held, by face-cited provenance.
  • What went wrong (nearly): in this rulebook ratify is a defined term — Rule 11 assigns ratification to the author alone. A header reading “Ratified… provisional-pending-author” uses the word in a sense the document’s own definitions forbid — ratified-yet-pending-ratification — and the most-read line of the constitution would have performed, to every skimming reader, an authority the session did not have. This is the exact reader-facing form-authority failure D-007 was armed against in this same session (item B3): form persuading where content has not.
  • Caught by: the Ostrom + appropriator-guard minority, and the cluster-C summoned adversary, before enactment — the panel’s own designed-disagreement layer catching the panel’s plurality (structure catching structure; self-administered, L-015 applies).
  • Corrected: the minority wording enacted instead — “In force — wordings panel-confirmed under the S4h standing delegation (C-027); the author’s ratification of the R22, R23, and R24 wordings remains OUTSTANDING (provisional-pending-author, reversible)” — with the R22/R23 in-rule parentheticals updated consistently (originals preserved, Rule 9). Logged as D-007’s first bite (a real near-miss of the governor’s exact type, caught pre-enactment — not a manufactured bite; the due-bite quota that would have incentivized manufacturing one was struck at B3 for exactly that reason).
  • Class: near-miss (reader-facing authority theater; caught before it reached the record’s face).

C-029 — The drain’s §5 promised a DC number the log had already spent (Code/adversary-caught, S4h)

  • Where: logs/DECISIONS_CHANGED.md §5 (v1.1) still read “Firing it will append here as DC-004 (strike)” — but DC-004 was spent in S4g on the R-arm reversal (the log’s first live entry, §2). Two parts of the same instrument disagreed about the next row’s identity.
  • What went wrong: the §5 text was drafted before DC-004 posted and never reconciled — the same deferred-propagation class as C-010/C-013/C-024, this time inside the subtraction operator itself.
  • Caught by: the cluster-D summoned adversary during the D2 deliberation (structure catching structure; self-administered).
  • Corrected: §5’s pointer fixed to “the next DC-0NN” (v1.2); logged here per the C-013 precedent for stale-pointer near-misses.
  • Class: consistency (CONS) / near-miss.

C-030 — The kit zip’s filename leaked the project identity; the recognition probe caught it (S4h courier round)

  • Where: the S4h TTM kit. The kit contents were leak-tested clean of every project token — but the archive itself was named systems-theory-ttm-kit-s4h.zip, and the runbook instructed attaching the zip directly where the chat UI allows. Two of four coders received zips (Grok, ChatGPT — the author’s dispatch manifest); GPT-5/Codex returned RECOGNIZE: source-project=“Systems Theory TTM project” — read straight off the filename. The first project-blind break in either courier round.
  • What went wrong: the filename is a side channel the leak audit never covered — the audit checked every byte inside the kit and none on it. Code’s kit-assembly defect (not the author’s, not the coder’s): the S4g kits carried the same names but were pasted as contents, so the channel never opened until the S4h runbook recommended zip-attachment. A sub-finding travels with it: the TTM sealed companion had no explicit recognition-exclusion clause (the validation pre-registration did) — a sealed-rule gap discovered when the probe fired; the conservative dual-report treatment was applied, and the gap is closed prospectively below.
  • Caught by: the post-coding recognition probe doing exactly its job (contamination measured, not attested away) — external provenance, via the author-mediated route.
  • Corrected: dual-reporting applied (no TTM band changes under any coder subset — the damage was real but resultless this round); pre-registered for all future kits: content-neutral archive filenames (e.g. coding-kit-a.zip), the filename added to the kit-audit checklist as a named leak class, and every kit companion carries an explicit recognition-disposition clause from now on. The exposure is booked for both zip-route coders even where RECOGNIZE returned NONE (Grok saw the same filename and said nothing — absence of report is not absence of exposure).
  • Class: methodological / kit-transport (the C-026(a/b) lineage — the third transport-channel lesson in two rounds).

C-031 — The Theory-C null’s permutation clause is direction-incoherent for null detection (exposed mechanically, S4h round 2)

  • Where: logs/DECISIONS_CHANGED.md §4, condition 1’s noise guard: “the ON–OFF gap survives a label-permutation noise check (the observed gap beats what shuffling the coded catches yields at ≥ 19-in-20).” Drafted S4g (panel-amended under delegation), re-affirmed S4h.
  • What is wrong: the clause tests the gap’s LARGENESS — but condition 1 is a null (ON fails to exceed OFF). A null-shaped gap definitionally cannot beat noise at 19-in-20, so as lettered, condition 1 could never fire for any actual null: the death-condition’s own noise guard is aimed the wrong way. Invisible while every powered read landed C-favorably (S4g’s 1.45× beat noise, reinforcing not-null); exposed the first time a powered ratio landed on the null side (blue round: 1.15×, perm p = 0.24).
  • Caught by: the sealed scorer’s mechanical output read against the sealed letter — structure catching the instrument’s own text (self-administered; the run’s floor-firing meant nothing turned on it this round).
  • Corrected (forward only, per the D2 precedent): nothing re-read backward; the repair — a coherent smallness guard (an equivalence-test form: the null fires only if the gap is bounded within noise, not if it beats noise) — is the author’s to re-seal with the C-026-cleared margins. Until re-sealed, condition 1’s ratio and power clauses stand; its permutation clause is flagged non-executable-as-lettered.
  • Class: methodological / instrument-text (the D2-clock lineage: sealed letters meeting cases their drafters did not run).

C-032 — A second standing delegation, stacked on unconfirmed stamps (booked at receipt, S4i)

  • Where: the author’s opening directive of S4i — a “very long loop run to complete all you can autonomously,” plus a commissioned five-user website playtest program (five Code-decided user profiles; per cycle: playtest → the user reports to the experts → experts + selected advisors decide implementations → Code implements; each cycle separate, including its evaluation and implementation path).
  • What it costs: the C-027 condition repeated while every S4h stamp is still provisional-pending-author — a second session’s worth of decisions taken by one model in many roles, now stacked on a first whose stamps the author has not yet reviewed. The provisional tower grows a storey before its foundation is confirmed; nothing this session does may be read as implying the S4h stamps were confirmed, and no S4i enactment may cite an S4h stamp as settled authority (each remains reversible at zero evidential cost).
  • Caught by: Code, at receipt, before any decision was taken under it (the C-023/C-027 discipline, third application).
  • Corrected (in-bounds): every S4i disposition stamped panel-delegated (S4i standing delegation), provisional-pending-author, reversible; the Rule-11 items stay untouchable throughout — the R24 firing gate and never-prune list (entrenched beyond any delegation), the docs/EPISODE_UNIT.md freeze event, the empty-prompt v2 stipulation truth-check (a fact about the author’s own corpus no delegate can attest), the repo clock’s advance-or-hold (the author’s named trade), and first-strike target-naming. The delegation is spent the way the discipline demands: on restoring checks (the detection-anchoring repair → a re-validation kit; the criteria red-team kit; the grader’s seating question put to the panel, its first queue being the S4h corpus itself) and on the author’s commissioned reader-facing pass, run under D-007’s watch with zero canonical-content edits.
  • Class: methodological / standing (the C-023/C-027 lineage, third entry).

C-033 — One clock, two enacted letters: the theory-doc banners and the causal doc disagree on when rung-labels fall due (adversary-caught, S4i)

  • Where: the three theory-doc pending-banners (outputs/THEORY_A_OPERATIONALIZED.md, outputs/THEORY_B_OPERATIONALIZED.md, outputs/THEORY_C_OPERATIONALIZED.md), enacted S4h, read “labeling is deferred to the author-present session (due by that session’s close…)” — while docs/PRESSURE_TESTS_CAUSAL_HYPOTHESIS.md §9 and the S4h ratification record (B4) read “due by the next substantive session’s close.” Two sealed letters for one clock, diverging at the S4h enactment.
  • What went wrong: the banner wording drifted from the record it enacted — the C-010/C-013/C-024/C-029 deferred-propagation class, this time inside a compliance clock. Under the banners’ letter no latency catch fires at S4i at all; under the record’s letter it fires at S4i close. All eight S4i lenses quoted the record’s letter; none read the enacted banners — the one-aquifer convergence exhibit (L-015), found only by the summoned adversary.
  • Caught by: the cluster-C summoned adversary (structure catching structure; self-administered), verified on-disk before booking.
  • Corrected: the latency catch (C-034) books under the harsher letter — chosen because doubt resolves toward the cost, not because the letters agree — and the banner parentheticals are corrected to the governing letter in the same edit that appends the expiry note, citing this catch. A mechanical face-vs-record reconciliation of every enacted banner, header, and status line against the ratification record is queued as follow-up work (the adversary’s warning: assume other divergences exist).
  • Class: consistency (CONS) — the C-029 lineage, one instrument over.

C-034 — The rung-label clock expired unmet: the automatic C-012-class latency catch, booked bare (S4i)

  • Where: the B4 deferral (S4h): theory-doc rung-labeling deferred to the author-present session, due by the next substantive session’s close (the governing letter per C-033), else this automatic catch. S4i — the next substantive session — closes with zero rung labels applied.
  • What happened: the panel HELD the labeling (5 H + 1 X against 2 RA; the adversary struck even the draft-annex variants): a standing delegation is precisely not author-presence, and pre-filled rung labels would convert the author’s ratification into confirmation of an anchor. The deferral’s venue clause and the second standing delegation became jointly unsatisfiable — so the clock fired.
  • Caught by: the clock itself (automatic — the letter working; Campbell: “the miss is the measurement working; book it”). (Provenance flag, added per the shadow-grade pass: self-administered — the governing letter is an S4h delegated enactment, and WHICH of two conflicting letters governed was this session’s delegated choice (C-033), made in the harsher direction because doubt resolves toward the cost.)
  • Corrected (scope): this catch posts bare and full-strength — no mitigation rider travels with it. The design-bound decomposition (venue clause + second delegation = jointly unsatisfiable) posts only as a separate PANEL-NOTE (CANDIDATE, provisional-pending-author) in panel/SESSION_S4I_RATIFICATION.md §1 item 4, stating cause and never adjudicating weight; whether it travels with citations is the author’s call. Delegated preparation was limited to judgment-free evidence collation (panel/B4_EVIDENCE_COLLATION_S4I.md — zero proposed rungs). The labels remain due at the author’s first ratifying contact; the repo-open backstop stands.
  • Class: methodological / latency (the C-012 lineage — automatic, per the sealed clock).

C-035 — A THIRD standing delegation — the SG-10 stacking guard’s first live trigger (booked at receipt, S4j)

  • Where: the author’s opening directive of S4j — commission an exhaustive expert forum on the state of the project, produce three deliverable documents (a systems-theory report; a “what we did wrong” analysis; a “what to improve” plan) as six PDFs, then implement the improvements, update the Plan and Method, ripple every document, commit, push, and present the proposed actions — all “autonomously and without asking any questions until the end.”
  • What it costs, and the guard it trips: this is the third consecutive standing delegation (C-027 → C-032 → C-035), stacked on two storeys of stamps the author has not yet confirmed (all of S4h and all of S4i remain provisional-pending-author). It is the exact case the SG-10 stacking guard — enacted last session in panel/GRADER_SHADOW_LOG_S4I.md — was written against: “no third standing delegation is accepted without the author’s direct contact ruling on the first two storeys.” The author has made direct contact and is directing, but has not ruled on the S4h/S4i storeys. Per the guard’s own stated design (“the author may of course override; the guard makes the override visible”) and Rule 11 (the author directs), the guard FIRES and is OVERRIDDEN on the record — the override is now visible, which is the whole point of the guard. The provisional tower is now three storeys.
  • Caught by: Code, at receipt, before any forum work — the C-023/C-027/C-032 discipline, fourth application, and the SG-10 guard doing exactly its job the first time it could.
  • Corrected (in-bounds): every S4j output is stamped forum-delegated (S4j standing delegation), provisional-pending-author, reversible; the Rule-11 items stay untouchable (the R24 firing gate and never-prune list; the EPISODE_UNIT freeze; the empty-prompt v2 stipulation; the repo clock; first-strike naming; no rung fires in any direction; no result is fabricated). The mitigation the discipline demands is unusually strong here: this delegation’s own commissioned purpose — a wrongs-analysis and an improvement plan produced by designed-disagreement, plus author-facing reports — is itself a course-correction and scrutiny act, the most aligned possible spend of a stacked delegation. The forum is run as real parallel subagents (the panel’s own method); the implementation phase edits presentation/delivery and author-directed content, never fires a lever; the Plan/Method updates are forum-delegated, provisional-pending-author. The three-storey tower and the two still-unconfirmed lower storeys are surfaced again at the close, in the proposed-actions presentation.
  • Class: methodological / standing (the C-023/C-027/C-032 lineage, fourth entry; the SG-10 guard’s first trigger).

Session 4k (the author’s ratifying contact + the closing loop)

C-036 — The author-side Downloads deliverables were wiped a SECOND time; the repo siblings were the only survivors

  • Where: at S4k arrival, verifying the handoff. The six S4j forum PDFs and the handoff pair were present, but the three S4i courier kits (coding-kit-echo.zip, coding-kit-foxtrot.zip, criteria-review-kit.zip) and the loose runbook (the Downloads copy of studies/RUNBOOK_S4I_CROSSMODEL_KITS.md) were gone from Downloads — externally verified (a directory listing), not inferred.
  • What it is: the second occurrence of the same failure mode that prompted commit ac0c100 (a loose-file cleanup wiping Downloads deliverables). The first occurrence taught the lesson “keep an in-repo sibling of every Downloads artifact”; that lesson held and paid offstudies/RUNBOOK_S4I_CROSSMODEL_KITS.md and studies/FROZEN_HASHES_S4I.md (49 per-file hashes) survived, so the kits are regenerable and byte-verifiable. But the recurrence itself is the catch: Downloads is not durable storage, and treating a session’s deliverables as “shipped” because they reached Downloads is the error.
  • Caught by: Code, at arrival, before any closing-loop work — the arrival-verification discipline doing its job.
  • Corrected: surfaced to the author at arrival with the regeneration offer (in-repo siblings + frozen hashes); the deeper fix recorded here — the repo is the only durable surface; Downloads artifacts are convenience copies, and every one must have a committed in-repo sibling (the rule the first occurrence wrote, now shown load-bearing by its second test). Given the S4k concession (the courier rounds depend on a foreign vantage that will not come), the kits are not regenerated unless the author asks — they would ship to a rung that no longer fires.
  • Class: process / recurrence (the ac0c100 lineage; a prior fix demonstrated load-bearing by surviving a second hit).