THEORY B, OPERATIONALIZED — the optimization ecology as a measurable claim
Version: 1.0 · Status: Ratified baseline (author, Session 4) · still living · Last updated: Session 4
Rung labels — APPLIED (Session 4k, the author’s ratifying pass; C-034 discharged, C-033 divergence closed). Under the adapted evidence ladder (
docs/PRESSURE_TESTS_CAUSAL_HYPOTHESIS.md§9): Theory B’s anchor mechanisms (D2→S2, the attention economy as optimization’s direct product; the documented de-optimization instances) aremechanism · pre-registered · frozen-unrun— the O→P pipeline is built and frozen with no data (C-017), so it produced no result and fabricated nothing. The headline O→P claim itself istheory · untested. The one reading on record — the H3 documented-case cross-check by four project-blind LLMs — isrun: SPLIT(Facebook 2018 followed the engagement incentive unanimously; YouTube 2019 mixed), read at its true n=1 (one corpus, one model family of coders = pseudo-replication, L-013), never as “4/4”. Nothing in this document isrun:survived.
The second Phase-1 work product: turning Theory B (the Optimization Ecology / Autopoietic Capture, CANDIDATE_THEORIES.md) from a mechanism-story into a stated, falsifiable claim — and, unlike Theory A, one that can be tested largely with data that already exists. B’s central advantage over A was always tractability: engagement-optimization is already measured, by the companies doing it and by researchers studying them. That makes B the natural place to run the project’s first empirical study, even though A is the flagship theory. This document specifies the measurable variable, the two discriminating tests that give B its integrity (selection-not-design; and B-vs-A), the first study, and the conditions that would prove it wrong. The honest posture: B’s greatest danger is sliding into a conspiracy theory, so its operationalization lives or dies on a test that distinguishes selected-to-capture from designed-to-capture.
1. The claim, restated as a measurable proposition
Informal (Theory B): self-reproducing optimizers — markets, media, technical systems — capture human drives (attention, dopamine, outrage, desire) as fuel and increasingly run without human ends, by selection without a designer.
Measurable restatement:
Let O = the optimization intensity with which a system is tuned to capture a human drive, and P = the capture/pathology that results (compulsive use, affective polarization, structural spread of unreality, well-being decrements). Theory B predicts that O positively predicts P across systems and over time; that systems deliberately re-optimized for a human end show reduced P (a natural experiment already partly running); and — the integrity condition — that outcomes track a selection pressure rather than any designer’s intent.
Three things make this a claim rather than a mood: O must be measured independently of P (not “it’s addictive, therefore it was optimized”); the re-optimization prediction must be checkable against real de-optimization events; and the selection-not-design test (§5) must be able to come out either way.
2. The key measurable: optimization intensity (O)
B’s advantage is that O is, in part, directly observable — it is what growth and recommendation teams literally do. Candidate proxies (faster/harder optimization → higher O):
| Facet | Candidate proxy | Direction |
|---|---|---|
| Objective function | is the system optimized for engagement/time-on-platform vs. a human-centered metric (declared or inferred) | engagement-objective ↑ |
| Personalization depth | recommender aggressiveness; how tightly the feed is fitted per user | deeper ↑ |
| Iteration speed | rate of A/B testing and model re-optimization (how fast selection runs) | faster ↑ |
| Capture tightness | conversion of the captured drive into system output (revenue per user-hour; attention→ad yield) | tighter ↑ |
| Autonomy from human ends | share of ranking decided by learned objective vs. user-chosen controls (chronological, opt-outs) | more autonomous ↑ |
The capture/pathology index (P): a pre-registered composite — compulsive-use metrics (session length, return frequency, self-reported loss of control), affective-polarization measures, the spread-rate differential of false vs. true claims, and use-associated well-being decrements.
3. The core prediction
- Cross-system. Across platforms/products, higher O predicts higher P, controlling for audience and content type.
- Longitudinal / natural experiment. When a system is re-optimized toward a human end — ships a “time-well-spent” objective, offers a chronological feed, is subjected to a chronological-feed or design-code mandate — P falls; when it re-optimizes harder for engagement, P rises. Several such events have already happened, which means part of this test can be run retrospectively rather than waiting.
4. The discriminator vs Theory A (making “A is the condition, B is the mechanism” testable)
A and B are said to be complementary — A the condition, B what rushes in. That relationship is only meaningful if they make different predictions somewhere. They do:
A predicts pathology tracks the rate-gap (broad, cross-domain, wherever change outruns adaptation). B predicts pathology tracks optimization intensity (specific to systems being tuned to capture a drive). The discriminating cases:
- Wide gap, low O: a slow-moving, hard-to-keep-up-with domain that nobody is optimizing to capture you (e.g. a bureaucratic or infrastructural domain). A predicts pathology; B predicts less.
- Narrow gap, high O: an intensely optimized attention product in an otherwise stable domain. B predicts pathology; A predicts less.
Running both O and a rate-gap measure in the same model, on domains chosen to separate them, tells you whether the symptoms are driven by the gap, by optimization, or (most likely) by both in different mixtures per symptom. This turns the tidy “condition vs mechanism” slogan into an empirical decomposition — exactly the “draw the arrows” discipline (L-005).
5. The integrity test: selection, not design (B’s make-or-break)
B’s authors named its deepest risk: it is narratively one step from a paranoid worldview (“everything is engineered to exploit us”). The theory’s integrity depends entirely on the claim that the outcome is selected, not planned — and that claim must be empirically distinguishable, or B is untestable and dangerous. The test:
Find cases where a designer’s intent and the selection pressure point in different directions, and see which the outcome follows.
- Strong evidence for B (selection): systems whose designers intended a benign or human-centered outcome but which drifted toward capture anyway — intent said A, the market selected B, the outcome followed B. (This pattern is common: teams that set out to “connect people” and produced outrage machines.)
- Evidence against B, toward conspiracy: outcomes that track a designer’s stated plan better than any selection pressure — where you can find the plan and the outcome follows the plan, not the market. That would make it a (weaker, different) intentional-harm claim.
- Bounding the strong version (agency is not erased): if deliberate de-optimization by designers reliably reverses P, that demonstrates humans can steer these systems — refuting the fatalistic “humans are just inputs” reading while leaving the selection mechanism intact.
This is the cleanest falsifiable discrimination in the whole project: it converts B’s worst liability (conspiracy-shape) into a test with a pre-committed verdict.
6. Falsification conditions (stated in advance)
Theory B is falsified if:
- P rises independent of O — pathologies grow where optimization intensity is low, so O is not the operative variable.
- De-optimization does nothing — systems re-optimized toward human ends show no reduction in P (the natural experiment comes back null).
- Intent beats selection — outcomes track designers’ stated plans better than selection pressures, collapsing B into a conspiracy claim it explicitly disavows.
- O is not measurable non-circularly — every attempt to measure optimization intensity smuggles in the pathology it is meant to predict (O and P cannot be separated), making the claim untestable.
7. The first study (pre-registered, runnable, cheapest of the three)
- Units: a set of comparable digital platforms/products (or product-versions over time).
- O: engagement-vs-human objective (declared/inferred) + personalization depth + presence/absence of user-chosen ranking controls (normalized composite).
- P: compulsive-use composite + affective-polarization + false-vs-true spread differential.
- Design: (a) cross-system correlation of O with P; (b) the natural experiment — before/after P around real de-optimization events (a shipped “time-well-spent” objective; a chronological-feed mandate); (c) the selection-vs-design case analysis (§5) on a curated set of documented cases.
- Pre-registration: proxy list, the P composite, and the case-selection rule fixed before outcomes are examined.
Because much of the input already exists (platform research, well-being studies, documented de-optimization events), B is the cheapest and fastest of the three studies to run — which is why, despite A being the flagship theory, B is the recommended first empirical test.
8. What this changes
- Falsifiability coverage. D2 (optimization) and S2 (attention economy) move from “partial” to a stated, testable design; the selection-not-design test gives B something A lacks — a crisp verdict on its own integrity. (
METRICS.md§4.) - The A/B relationship becomes empirical. “Condition vs mechanism” is now a decomposition to run, not a slogan (§4).
- The project gets a runnable first study. B is the on-ramp to Phase-3 data work — the concrete thing heavy Code would compute first.
9. Known weaknesses of the operationalization itself (skeptical close)
- O is hard to measure from outside. Optimization intensity lives inside proprietary systems; external proxies (personalization depth, ranking controls) are imperfect shadows of the real objective function. The defense is triangulation and transparency about proxy limits.
- The conspiracy line is psychologically slippery. Even with the §5 test, readers will collapse “selected to capture” into “designed to addict us.” The operationalization reduces but cannot eliminate this; the discipline is to state the verdict the test would return, every time.
- Functionalism risk (shared with A). “It exists because it was selected to capture” can back-rationalize anything; the guard is the specific measured pressure and the natural experiment, not the story.
- Publication and access bias. The documented cases and datasets skew toward large Western platforms — the D-005 vantage problem again. Generalization is a separate question.
As with Theory A: none of these is fatal, and all are named so the next contributor knows where to push. A theory you can attack precisely is worth more than a mechanism you can only admire.
Author: Hulki Okan Tabak — with Claude · License: CC BY-SA 4.0