B4 EVIDENCE COLLATION (S4i) — judgment-free preparation for the author’s rung-labeling
Status: EVIDENCE COLLATION ONLY — panel-delegated (S4i standing delegation), provisional-pending-author. Per the S4i item-4 hold and its adversary correction: NO proposed rung, NO proposed test-status tag, and no grading adjective appears anywhere in this file; a file produced under delegation containing a proposed rung is a breach of the B4 deferral and books its own catch (see logs/CATCHES.md C-034). The author labels; this file only gathers what the labeler will need. · Last updated: Session 4i (Code)
Scope: the three theory docs (outputs/THEORY_A_OPERATIONALIZED.md, outputs/THEORY_B_OPERATIONALIZED.md, outputs/THEORY_C_OPERATIONALIZED.md), walked section by section. Each carries the identical banner — “Rung-labels pending (B4, S4h) — no rung status may be cited from this document” — which each banner states is “the document’s only rung statement”; this collation does not alter that. Columns: the claim (quoted or tightly paraphrased); the receipts the doc itself offers; the stated N where a number is claimed; the provenance category (in-house deliberation / cross-model coding / external public record / none); and what has run against it, citing the result files below with their own status lines verbatim. Where a cell says “no listed result file reports a run against this claim,” that is a statement of absence in the checked files, not a characterization. All quotation marks in the tables enclose the source documents’ own words.
Result-file register (statuses quoted verbatim from each file’s own status line; “…” marks trimming only)
| Key | File | Own status line (verbatim) |
|---|---|---|
| R1 | studies/study-B-optimization/outputs/CROSSCHECK_RESULT.md |
“RESULT — a cross-model-robust reading of the public documentary record; NOT the quantitative O→P test (that stays frozen-unrun). Preliminary until author/Chat ratification; the ratified pre-registration stands.” |
| R2 | studies/study-B-optimization/outputs/TTM_RESULT_S4H.md |
“RUN COMPLETE (S4h; 4 external coders). THE CODABILITY BAR PASSED AT ITS MAXIMUM: exact-band agreement 30/30 pairs×cases = 100% (bar: ≥0.6) — kill condition 3 does NOT fire; TTM is codable blind, cross-family, unanimously. Yield exactly as pre-registered: ex-ante bandings + codability evidence ONLY — no confirmatory language about the H3 conditional exists or may exist until the separate outcome kit returns. Cross-model-confirmed, n=1 of a kind (L-013), revisitable by a future fork. panel-delegated (S4h standing delegation), provisional-pending-author.” |
| R3 | studies/study-B-optimization/outputs/GREEN_RESULT_S4H.md |
“RUN COMPLETE (2026-07-13; 3 clean coders — GPT-5, Grok, Gemini; DeepSeek headline-excluded per the sealed identity-mismatch rule + cross-kit contamination, dual-reported). As pre-noted in the sealed companion BEFORE any verdict existed: with zero LOW-band ex-ante cases, NO confirmatory H3 arithmetic is possible from this round — the yields are per-case, and they are real. Cross-model-confirmed, n=1 of a kind (L-013), revisitable by a future fork. panel-delegated (S4h standing delegation), provisional-pending-author.” |
| R4 | studies/study-C-ablation/outputs/PILOT_RESULT.md |
“PILOT RESULT — ratified pre-registration executed in Chat; a reduced-N demonstration with a fundamental self-administration confound. Directional only; NOT an inferential result.” |
| R5 | studies/study-C-ablation/outputs/CROSSMODEL_RESULT.md |
“RUN COMPLETE (S4g). No null armed; Theory C is NOT retired; the death-condition stays armed-not-firing. Cross-model-confirmed, n=1 of a kind (L-013), revisitable by a future fork. Canonical under the S4h standing delegation (C-027): panel-delegated, provisional-pending-author, reversible. Guard note (the sentinel, S4h): this reading’s direction is C-favorable and was canonized by the apparatus Theory C describes — the Rung-0 tripwire on any drift toward ‘the ablation supports C’ applies to this ratification’s own downstream citations.” |
| R6 | studies/study-C-ablation/outputs/BLUE_RESULT_S4H.md |
“RUN COMPLETE (2026-07-13; 3 clean coders — GPT-5, Grok, Gemini; DeepSeek excluded per the sealed identity-mismatch rule, its partial verdict archived + dual-report note standing). FINAL: INDETERMINATE — INSTRUMENT UNRELIABLE (mean pairwise S3-identification agreement 1% on live material, versus the sealed 50% floor). No null posts; no rung moves; nothing at run level is read. Scored by the pre-dispatch sealed scorer (src/score_blue_verdicts.mjs, fixture-tested before any verdict existed). Cross-model-confirmed, n=1 of a kind (L-013), revisitable by a future fork. panel-delegated (S4h standing delegation), provisional-pending-author.” |
| R7 | studies/placebo-control/outputs/RESULT.md |
“RUN COMPLETE (S4g). FINAL: INDETERMINATE — instrument unreliable (the sealed reliability floor fired). No rung fires. Cross-model-confirmed, n=1 of a kind (L-013), revisitable by a future fork. Canonical under the S4h standing delegation (C-027): panel-delegated, provisional-pending-author, reversible — the S4g panel-ratified stamp is retained beneath so the two delegation layers stay distinguishable. S4h amendments: point 3 demoted to a non-claimable annex; the may-NOT list extended (panel/SESSION_S4H_RATIFICATION.md A1).” |
| R8 | studies/placebo-control/outputs/VALIDATION_RESULT_S4H.md |
“RUN COMPLETE (S4h; 4 external coders — Gemini, Grok, DeepSeek, GPT-5/Codex; dispatch manifest author-supplied: Gemini/DeepSeek received files, Grok/ChatGPT received zips). SEVERITY ANCHORING: PASSED EVERY GATE. CONVERGENCE INSTRUMENT: mechanically reliable, but the sealed Q1 map landed 4/4 unanimous on the FABRICATION-CHECK CELL — instrument-acceptance FAILURE for the placebo-reading purpose, no exculpatory gloss permitted. Q2 (closure-shape manufacture): NOT SUPPORTED, 3/4 modal — the pattern lost through the corridor built to let it lose. NO RUNG FIRES IN EITHER DIRECTION (the pre-registered non-firing clause, verbatim). Cross-model-confirmed, n=1 of a kind (L-013), revisitable by a future fork. panel-delegated (S4h standing delegation), provisional-pending-author.” |
Theory A (outputs/THEORY_A_OPERATIONALIZED.md)
| Doc section | Claim | Receipts cited | Stated N | Provenance | What has run against it (verbatim status) |
|---|---|---|---|---|---|
| Preamble (italic head-note) | Theory A (“the Adaptation Gap”) is “the first Phase-1 work product”; the doc “does not claim the theory is true — it claims to make the theory checkable”; “Theory A has the best integrative reach and the weakest testability” (the doc’s own comparative self-description); “a gap you cannot measure is a story, not a science (Turchin’s standing objection)” | CANDIDATE_THEORIES.md; OPEN_QUESTIONS.md Q-001; Turchin named as source of the standing objection |
no N stated | in-house deliberation | No listed result file reports a run against this claim. |
| §1 | G = R_c − R_a, “measured independently of its symptoms, positively predicts the severity of the symptom-set (fertility collapse, anomie, populism, epistemic breakdown, coordination failure), cross-sectionally and over time, and retains explanatory power after controlling for rival drivers — chiefly the wealth pump (D3)” | internal cross-refs to §5 (falsification conditions) and §7 (pre-registration); CANDIDATE_THEORIES.md |
no N stated | in-house deliberation | No listed result file reports a run of this proposition. |
| §2 | Candidate proxy sets for R_c (technology diffusion, frontier compute, firm turnover in the top index, information volume/velocity, occupational churn) and R_a (regulatory lag, legislative cycle time, curriculum/skill half-life, institutional-trust recovery, norm stabilization); design notes: R_a is “adaptive throughput, not institutional quality”; inverse lags need sign correction | generic named data sources only (e.g. “S&P 500 / equivalents”); no named studies | no N stated | in-house deliberation (proxy candidates); none for any specific proxy value | No listed result file reports a run against these proxies. |
| §3 | The gap is “only meaningful after normalization”; composite by pre-registered rule (equal weights default); “if reasonable alternative proxy sets and weightings yield materially different G trajectories … that is evidence that G is an artifact of measurement choices” | cross-ref C-004 (“the near-tautology risk, C-004’s cousin”) | no N stated | in-house deliberation | No listed result file reports a run of the normalization/robustness procedure. Touched as coded material: R4 §4 Q4 [ON] used “strongest objection to measuring G = R_c − R_a from proxies” as an ablation question (catches recorded there include “incommensurable,” “‘tunable to fit anything,’” “R_a-aggregation”); R4 status: “PILOT RESULT — … Directional only; NOT an inferential result.” |
| §4 | Cross-sectional prediction: units with wider G show more severe symptoms; longitudinal prediction: “as G widens, symptoms intensify with a lag”; “Theory A predicts G leads symptoms; if symptoms lead G, the causal story is wrong”; symptom-severity index as pre-registered composite | none named beyond the composite’s component types (loneliness/anomie indices, populist vote share, affective polarization, institutional trust, fertility “with care”) | “≈30 OECD countries, or sectors” (cross-sectional units) | in-house deliberation | No listed result file reports a run of either prediction. |
| §5 | Four falsification conditions: (1) narrow-gap severity; (2) a rival dominates (G adds nothing once the rival is controlled); (3) measurement artifact (alternative constructions give materially different G); (4) wrong lag direction | internal cross-ref §6 | no N stated | in-house deliberation | No listed result file reports a check of any of the four conditions. |
| §6 | Discriminator vs the wealth pump: “Theory A earns its place only if G carries significant, independent explanatory power with the wealth pump controlled. If G collapses to noise once inequality is included, D3 is the real driver and Theory A is redundant”; if both carry independent weight, “A and Turchin are complementary drivers” | Turchin’s structural-demographic theory (named body of work, no specific study cited in this doc); D-002, D-006, L-005; Gini / top-income shares / elite-overproduction proxies as control candidates | no N stated | external public record (the named Turchin corpus) + in-house deliberation (the discriminator design) | No listed result file reports a run of the horse-race model. Touched as coded material: R4 §4 Q2 [ON] used “is ‘D3 drives S3 more than D1 does’ falsifiable; distinguish them?” as an ablation question (catches recorded there include collinearity/under-identification and the common-cause/DAG point); R4 status: “PILOT RESULT — … Directional only; NOT an inferential result.” |
| §7 | First study spec: ≈30 OECD countries; stated R_c and R_a proxy triplets; outcome composite; Gini control; tests (a)–(d); pre-registration before outcomes; “this study is also the heavy-Code trigger” | DIAGRAMS.md §7 (Gate 2); companion pre-registration exists at studies/study-A-adaptation-gap/PRE_REGISTRATION.md (cross-reference; not a result) |
“≈30 OECD countries” | in-house deliberation | No listed result file reports a run of Study A. (R1’s status line records, for the parallel B pipeline, “frozen-unrun”; no listed result file states a run status for Study A.) |
| §8 | “D1 (the gap) moves from ‘not yet’ toward a stated, testable claim”; “It does not yet count as tested; Q-001 stays open”; the A-vs-Turchin question is “reframed as an empirical horse-race”; the causal loop “is now drawable” | METRICS.md §4; DIAGRAMS.md §9; D-002/D-006; Q-001 |
no N stated | in-house deliberation | No listed result file reports a run against this claim. |
| §9 | Four self-stated weaknesses of the operationalization: composite tunability; adaptive-throughput ≠ adaptive-success; “OECD-only is a narrow world” (D-005 vantage); “correlation is not the mechanism” | D-005 | no N stated | in-house deliberation | No listed result file reports a run against these statements. |
Theory B (outputs/THEORY_B_OPERATIONALIZED.md)
| Doc section | Claim | Receipts cited | Stated N | Provenance | What has run against it (verbatim status) |
|---|---|---|---|---|---|
| Preamble (italic head-note) | B can be “tested largely with data that already exists”; “engagement-optimization is already measured, by the companies doing it and by researchers studying them”; B is “the natural place to run the project’s first empirical study”; “B’s greatest danger is sliding into a conspiracy theory” | CANDIDATE_THEORIES.md; “researchers studying them” (unnamed) |
no N stated | in-house deliberation | No listed result file reports a run against these framing claims as such; the study-level runs are collated at §§3, 5, 7 below. |
| §1 | “O positively predicts P across systems and over time; systems deliberately re-optimized for a human end show reduced P (a natural experiment already partly running); and — the integrity condition — outcomes track a selection pressure rather than any designer’s intent” | CANDIDATE_THEORIES.md; internal cross-refs §5, §6 |
no N stated | in-house deliberation | R1 (the integrity-condition clause, coded on two documented cases plus four model-volunteered cases): “RESULT — a cross-model-robust reading of the public documentary record; NOT the quantitative O→P test (that stays frozen-unrun). Preliminary until author/Chat ratification…”. R2 (codability of the threat-to-metric construct feeding the H3 conditional): “THE CODABILITY BAR PASSED AT ITS MAXIMUM … no confirmatory language about the H3 conditional exists or may exist until the separate outcome kit returns.” R3 (outcome codings, per-case): “with zero LOW-band ex-ante cases, NO confirmatory H3 arithmetic is possible from this round — the yields are per-case…”. The O→P clause itself: per R1’s status line, “NOT the quantitative O→P test (that stays frozen-unrun).” |
| §2 | O is “in part, directly observable”; five candidate O proxies (objective function, personalization depth, iteration speed, capture tightness, autonomy from human ends); P as a pre-registered composite (compulsive use, affective polarization, false-vs-true spread differential, well-being decrements) | none named (proxy candidates only) | no N stated | in-house deliberation | No listed result file reports a measurement of O or P. Touched as coded material: R4 §4 Q3 [OFF] used “strongest objection to operationalizing O from disclosures?” as an ablation question (catches recorded there include the construct-validity and non-comparability points); R4 status: “PILOT RESULT — … Directional only; NOT an inferential result.” |
| §3 | Cross-system: “higher O predicts higher P, controlling for audience and content type.” Longitudinal/natural experiment: after re-optimization toward a human end “P falls”; after harder engagement re-optimization “P rises”; “several such events have already happened,” permitting retrospective testing | named event types only (“time-well-spent” objective; chronological-feed / design-code mandates); no study citations | no N stated | in-house deliberation (design); external public record invoked generically for the events | The cross-system O→P regression: per R1’s status line, “NOT the quantitative O→P test (that stays frozen-unrun).” De-optimization events as coded cases: R3 blind-coded outcomes on five cases (its own table: A “goal — 3/3”; B “engagement-selection — 2/3 (Grok: mixed-or-unclear)”; C “revenue — 3/3”; D “regulatory — 3/3”; E “engagement-selection — 3/3”); R3 status: “…NO confirmatory H3 arithmetic is possible from this round — the yields are per-case, and they are real.” |
| §4 | A-vs-B discriminator: “A predicts pathology tracks the rate-gap … B predicts pathology tracks optimization intensity”; discriminating cases “wide gap, low O” and “narrow gap, high O”; running both in one model yields the decomposition | L-005 | no N stated | in-house deliberation | No listed result file reports a run of the A/B decomposition. |
| §5 | Selection-not-design integrity test: outcomes following drift-toward-capture against benign intent = “strong evidence for B (selection)”; outcomes tracking a designer’s stated plan = “evidence against B, toward conspiracy”; deliberate de-optimization reversing P bounds the fatalistic reading | “(This pattern is common: teams that set out to ‘connect people’ and produced outrage machines.)” — no citation given in the doc | no N stated | in-house deliberation | R1 coded this test on the documentary record — its own table: Case A (Facebook MSI 2018) “4/4 more-with-I, high confidence”; Case B (YouTube 2019) “4/4 mixed, medium confidence”; its own §0: “Read the split, not the tally”; its own §3 caveat carried: four LLMs “plausibly share training corpora” (L-013). R1 status: “RESULT — a cross-model-robust reading of the public documentary record; NOT the quantitative O→P test (that stays frozen-unrun). Preliminary until author/Chat ratification…”. R2 banded five cases ex-ante on threat-to-metric; status: “TTM is codable blind, cross-family, unanimously. Yield exactly as pre-registered: ex-ante bandings + codability evidence ONLY…”. R3 coded the same five cases’ outcomes; status: “…NO confirmatory H3 arithmetic is possible from this round — the yields are per-case…”; its own §3 bars “any pooling of these outcome verdicts with the S4f cross-check’s G/I verdicts (different rubrics; different pool).” |
| §6 | Four falsification conditions: (1) P rises independent of O; (2) de-optimization does nothing; (3) intent beats selection; (4) O is not measurable non-circularly | none beyond internal cross-refs | no N stated | in-house deliberation | Condition (1): no listed run (the O→P test — per R1, “frozen-unrun”). Condition (2)'s subject matter appears in R3’s per-case outcome codings (status as above; “NO confirmatory H3 arithmetic is possible from this round”). Condition (3)'s subject matter is what R1 coded (status as above). Condition (4): no listed run; touched as coded material in R4 §4 Q3 (status as above). |
| §7 | First study spec: units = comparable digital platforms/products; O and P composites; design (a) cross-system correlation, (b) natural experiment around de-optimization events, (c) selection-vs-design case analysis; pre-registration before outcomes; “B is the cheapest and fastest of the three studies to run” (the doc’s own comparative claim) | named event types as in §3; pre-registration cross-referenced (ratified studies/study-B-optimization/PRE_REGISTRATION.md per R1’s header) |
no N stated (“a set of comparable digital platforms/products”) | in-house deliberation | Part (c) has a run: R1 (via studies/study-B-optimization/APPROACH_B2_DOCUMENTED_CASES.md), status: “RESULT — a cross-model-robust reading of the public documentary record; NOT the quantitative O→P test (that stays frozen-unrun)…”. Parts (a) and (b) as quantitative designs: per the same status line, “frozen-unrun.” Construct-side rounds on the widened pool: R2 (status as in register) and R3 (status as in register). |
| §8 | “D2 (optimization) and S2 (attention economy) move from ‘partial’ to a stated, testable design”; “the A/B relationship becomes empirical”; “the project gets a runnable first study” | METRICS.md §4 |
no N stated | in-house deliberation | No listed result file reports a run against these claims as such; R1 §5 (its own text) proposes coverage-tracker handling “leave the coverage-tracker status as-is until ratification.” |
| §9 | Four self-stated weaknesses: O “is hard to measure from outside”; the conspiracy line “is psychologically slippery”; functionalism risk; “the documented cases and datasets skew toward large Western platforms — the D-005 vantage problem” | D-005 | no N stated | in-house deliberation | No listed result file reports a run against these statements. (R1’s own §4 records, in its own words, “Vantage (D-005): all cases are large Western consumer-attention platforms”; R2’s pool includes one case its own table labels “non-Western” (Douyin).) |
Theory C (outputs/THEORY_C_OPERATIONALIZED.md)
| Doc section | Claim | Receipts cited | Stated N | Provenance | What has run against it (verbatim status) |
|---|---|---|---|---|---|
| Preamble (italic head-note) | C “is not a claim about the world but a claim about the form a theory must take”; operationalizing C = “turning the project’s falsifiability discipline on the project itself” (Goal S5); C “must therefore be held to a higher evidentiary bar”; the meaning half “is left explicitly outside measurement” | CANDIDATE_THEORIES.md; Goal S5; Q-012 |
no N stated | in-house deliberation | The study-level runs are collated at §§3, 5, 6 below. |
| §1 | “A process with designed disagreement + preserved dissent + a catches→learnings→rules loop will surface a higher rate of caught errors and retained live objections, per unit of work, than a single-author process on the same questions; and a forkable, commons-governed theory will improve along tracked metrics … faster than a closed one. If the plural process does not out-catch the individual, C’s method claim is false.” | CANDIDATE_THEORIES.md; internal cross-ref §7 (meaning excluded) |
no N stated | in-house deliberation | R4: “PILOT RESULT — ratified pre-registration executed in Chat; a reduced-N demonstration with a fundamental self-administration confound. Directional only; NOT an inferential result.” R5: “RUN COMPLETE (S4g). No null armed; Theory C is NOT retired; the death-condition stays armed-not-firing…” (with its own guard note: “the Rung-0 tripwire on any drift toward ‘the ablation supports C’ applies to this ratification’s own downstream citations”). R6: “RUN COMPLETE (2026-07-13…). FINAL: INDETERMINATE — INSTRUMENT UNRELIABLE (mean pairwise S3-identification agreement 1% on live material, versus the sealed 50% floor). No null posts; no rung moves; nothing at run level is read.” The forkable-commons clause (improvement vs a closed process): no listed result file reports a run. |
| §2 | Five internal measurables: catch rate, catch provenance (“the load-bearing measure”), objection-retention, rule-change rate, improvement trajectory; “if catches collapse onto the author, the plurality is cosmetic — the C-006 nightmare” | logs/CATCHES.md (“11 so far” at time of writing); OPEN_QUESTIONS.md; LEARNINGS.md → GROUND_RULES.md; METRICS.md; the L-001 → C-007 example |
N = 11 (catches logged at the doc’s writing) | in-house logs / in-house deliberation | Instrument-side runs on whether these measurables can be read at all: R7: “FINAL: INDETERMINATE — instrument unreliable (the sealed reliability floor fired). No rung fires.” R8: “SEVERITY ANCHORING: PASSED EVERY GATE. CONVERGENCE INSTRUMENT: mechanically reliable, but the sealed Q1 map landed 4/4 unanimous on the FABRICATION-CHECK CELL — instrument-acceptance FAILURE for the placebo-reading purpose… NO RUNG FIRES IN EITHER DIRECTION.” R6 (on live material): “INSTRUMENT UNRELIABLE (mean pairwise S3-identification agreement 1% on live material, versus the sealed 50% floor),” with its own §2 diagnosis quoted: “Severity-banding was repaired; catch-DETECTION was never anchored.” |
| §3 | Core test: internal ablation — disagreement-ON vs single-synthesizer arms on comparable questions, pre-registered severity-weighted catch coding; “if the two arms catch the same, designed disagreement adds nothing and the project’s central method is theatre.” Two supporting tests: external red-team; fork-integration | internal cross-ref to the pre-registration (studies/study-C-ablation/PRE_REGISTRATION.md, per R4’s header) |
no N stated in this doc (the pilot’s own deviation note records “N = 6, not 20”) | in-house deliberation | The ablation has three listed runs. R4 (S4e, in-house arms and coding): status as in register; its own §6 verdict line: “worth nothing as inference and a lot as calibration.” R5 (S4g, external coders, pre-anchoring instrument): “No null armed; Theory C is NOT retired; the death-condition stays armed-not-firing,” with its own corridor sentence: “under the preserved 1.5× margin, the observed 1.45× would MEET condition 1’s threshold and the null would stand armed pending cross-model confirmation — the ratified 1.2× line, not the data, is what keeps this read C-favorable,” and its own broken-series rule: the 1.45× “may never be pooled with, trended against, or averaged into post-anchoring codings.” R6 (S4h, post-anchoring instrument): “FINAL: INDETERMINATE — INSTRUMENT UNRELIABLE… No null posts; no rung moves; nothing at run level is read”; its own §3: the powered coder’s 1.15× is “inside the null margin under BOTH lines … permutation p = 0.24 … it arms nothing, posts nothing, and may never be cited as ‘the null fired.’” The external red-team and fork-integration tests: no listed result file reports a run. |
| §4 | The baseline problem (Q-012): the panel and the single author “are the same underlying model wearing different prompts,” so the ablation tests the narrower question (“does a disagreement-structured process beat an unstructured one, holding the underlying reasoner fixed?”), not the grand claim; the uncontaminated baseline arrives only with real external forks | Q-012; the two-surfaces plan | no N stated | in-house deliberation | R4 touches this directly — its own C-run-1: the pilot “sharpens Q-012 from ‘the baseline is entangled’ to ‘the ablation is self-administered’”; R4 status as in register. R5 and R6 used external coders (cross-model coding) with the caveat carried in both status lines verbatim: “Cross-model-confirmed, n=1 of a kind (L-013).” The grand claim (human plurality): no listed result file reports a run; R4 §1 states “The grand claim — human plurality beats individual genius — is not tested (Q-012).” |
| §5 | Four falsification conditions: (1) ablation null; (2) catches collapse onto the author (provenance); (3) forking accumulates without integrating; (4) external red-team finds a large missed volume | none beyond internal cross-refs | no N stated | in-house deliberation | Condition (1): R5 — “No null armed … the death-condition stays armed-not-firing” (its body: “Condition 1 (ON < 1.2× OFF) not met — observed 1.45×”); R6 — “No null posts; no rung moves” (its body: the run is “INDETERMINATE — INSTRUMENT UNRELIABLE,” and the powered coder’s within-construct ratio “may never be cited as ‘the null fired’”). Condition (2): R5’s body — “Condition 2 (zero structure-provenance S3) not met — 26 struct S3”; R6’s body — “Condition 2 (provenance collapse) is not met in any coder (GPT-5 struct-S3 = 32 of 60).” Conditions (3) and (4): no listed result file reports a run. |
| §6 | First study: the §3 ablation, pre-registered, “runnable immediately,” no data pipeline or external contributors needed; “reported with the baseline caveat (§4) front and centre” | the pre-registration; §4 | no N stated in this doc | in-house deliberation | Run three times: R4, R5, R6 — statuses as in the register (R4: “Directional only; NOT an inferential result”; R5: “No null armed; Theory C is NOT retired; the death-condition stays armed-not-firing”; R6: “FINAL: INDETERMINATE — INSTRUMENT UNRELIABLE…”). Instrument acceptance for the coding apparatus: R7 (“FINAL: INDETERMINATE — instrument unreliable (the sealed reliability floor fired). No rung fires”) and R8 (“SEVERITY ANCHORING: PASSED EVERY GATE. CONVERGENCE INSTRUMENT: mechanically reliable, but … instrument-acceptance FAILURE for the placebo-reading purpose… NO RUNG FIRES IN EITHER DIRECTION”). |
| §7 | The meaning half is “not operationalized, and cannot be”; “a method is not a meaning” (Nietzsche’s cut); measuring it would be a category error; D-001, D-004, Q-002 remain open | Nietzsche’s cut; D-001, D-004, Q-002 | no N stated | in-house deliberation | No listed result file reports a run (the doc states this half is outside measurement by design). |
| §8 | Goal S5 “moves from an aspiration to a pre-registered ablation with coded outcomes”; C-006 “is no longer only a standing warning — catch-provenance is the number that would expose it”; Q-012 is named onto the frontier | METRICS.md; C-006; Q-012 |
no N stated | in-house deliberation | The provenance number has been produced in two listed runs: R5 body — “26 struct S3”; R6 body — “GPT-5 struct-S3 = 32 of 60” (statuses as in the register; R6: “nothing at run level is read”). |
| §9 | Four self-stated weaknesses: the baseline problem bounds the study; self-serving convenience (“C must clear a higher bar” — Luhmann’s warning, THE_LIVING_DOCUMENT.md §6 capture mode); metric gaming (“the guard is a pre-registered severity-weighted taxonomy”); the meaning half untouched |
Q-012; Luhmann’s warning; THE_LIVING_DOCUMENT.md §6; the 1/3/9 taxonomy |
no N stated | in-house deliberation | On the severity-taxonomy guard: R8 — “SEVERITY ANCHORING: PASSED EVERY GATE” (its body: “Severity now travels across model families,” with Campbell’s broken-series rule restated there); R7 — “FINAL: INDETERMINATE — instrument unreliable (the sealed reliability floor fired)”; R6 — its own §2: “Severity-banding was repaired; catch-DETECTION was never anchored.” The other three statements: no listed result file reports a run against them. |
Author: Hulki Okan Tabak — with Claude · License: CC BY-SA 4.0.