Evaluation

Pressure-test it. Here is what it claims, and what the claims are worth.

Every panel that shaped LEXIS — the premise challenge, the ten-seat council, the playtest, the audits — is the same model that built it. Same-model agreement can find real defects and converge on a defensible shape, but it cannot certify quality; it measures internal coherence. So this page separates the evidence that comes from outside the model (arithmetic, live measurement) from the verdicts that come from within it, and it names what the process cannot settle.

Ship state: provisional — pending the author's full ratification. The one gate this series has shown can kill a concept — the author's own look — has now been exercised once: his first real use found that the piece's meaning never reached the eye (the words lived only in the screen-reader channel) and that mobile stayed silent. That verdict was critical, and it was answered — the visible answer line under the field and the mobile-audio hardening exist because of it. Full ratification of the whole remains open.

What it claims, and the grade of the evidence

Model-external (measured, not asserted by the builder):

Simulated (same-model personas — a defect-finding pass, not proof of quality): eight first-time personas passed every pre-registered gate bar (a hypothesis within twenty seconds, intentional composition, resolved surprise, learning something true, no dark-pattern trap), with the two kill-criteria clear. Useful for finding friction; it does not certify the piece is good.

The strongest criticism we could not patch

LEXIS models the rule-inducing kind of curiosity — the kind most native to a language model — and the very properties that make it excellent (determinism, honesty, fairness, combinatorial depth) are what give its curiosity a short half-life and keep it from ever touching wonder. Once you have induced the rules, it deepens in combination, not in mystery. It is one genuinely good mechanic wrapped in a competent system, with an unusually honest process around it. Whether that clears the bar for a real answer to curiosity, or only for a well-made honest small thing, is precisely what a same-model apparatus cannot decide.

Honest limits, named as bets

Deliberately out of scope (chosen, not skipped): wonder, aimless sensory delight, morbid curiosity, and curiosity about other people. LEXIS is one register, and says so.

Check the arithmetic yourself

A number a reader cannot re-run is an appeal to authority. The button below enumerates the state space on your own device, using the same shipped code that runs the field — no network, no telemetry, arithmetic in the open.

Two honesty notes on interaction

Every visitor can read the line as a timeline: resting on a glyph in the line (pointer or keyboard focus) shows the state the line meant up to that point, and on touch the first tap on a line glyph previews that same prefix — the second tap removes it. And the human-observation protocol this piece committed to (docs/OBSERVE.md in the workshop repo — a handful of genuine first-timers, watched quietly) is written and pending; until those sittings happen, the first-contact claims here rest on simulated visitors plus one real first look — the author's, whose verdict ("the meaning does not appear anywhere") produced the visible answer line under the field — and should be read that way.

If you are evaluating it

  1. Is the loop genuinely curiosity, or competent puzzle-decipherment wearing the word?
  2. Does "inexhaustible, not secret" hold for you on a real session, or does it bottom out?
  3. Is making a gap honest a real answer to the clickbait critique, or a relocation?
  4. Cold-start: did you spontaneously try an operator, or plateau at stacking marks?
  5. Grade the object apart from the process: good, or honest-and-ordinary? Say which.

← Back to the field  ·  Colophon  ·  Concepts