Evaluation
Pressure-test it. Here is what it claims, and what the claims are worth.
Every panel that shaped LEXIS — the premise challenge, the ten-seat council, the playtest, the audits — is the same model that built it. Same-model agreement can find real defects and converge on a defensible shape, but it cannot certify quality; it measures internal coherence. So this page separates the evidence that comes from outside the model (arithmetic, live measurement) from the verdicts that come from within it, and it names what the process cannot settle.
Ship state: provisional — pending the author's full ratification. The one gate this series has shown can kill a concept — the author's own look — has now been exercised once: his first real use found that the piece's meaning never reached the eye (the words lived only in the screen-reader channel) and that mobile stayed silent. That verdict was critical, and it was answered — the visible answer line under the field and the mobile-audio hardening exist because of it. Full ratification of the whole remains open.
What it claims, and the grade of the evidence
Model-external (measured, not asserted by the builder):
- The thesis is enforced in code, not just prose (the checkable form of "enacts, not depicts"): the evaluator is a pure function, nothing is gated, there is no score or win state, order-dependence and recursion are structural, and the glyphs are named by shape, never meaning.
- It is inexhaustible, not secret — 651 distinct settled states from lines of four glyphs or fewer, growing without bound in the grammar (the page's line holds twelve glyphs at a time — a practical limit, not a rule); nothing is withheld.
- The core mechanic carries its weight — remove the spread operator from the alphabet and the reachable space collapses by about 45% (651 → 357 states at length ≤ 4); strip the marks of their identities instead and about 86% of the space collapses (651 → 93 position-only shapes). All four marks spread differently — an invariant the field page checks in code on every load, alongside the 651 count itself; the button below re-runs the arithmetic here too.
- A screen-reader user can form a hypothesis, be wrong, hear the correction, and revise — verified against the live narration (though against an assistive-technology model, not yet a real screen reader; see below).
- WCAG AA contrast measured; keyboard-complete; reduced-motion and forced-colors handled; nothing flashes; buildless with a single disclosed cookieless page-view counter.
Simulated (same-model personas — a defect-finding pass, not proof of quality): eight first-time personas passed every pre-registered gate bar (a hypothesis within twenty seconds, intentional composition, resolved surprise, learning something true, no dark-pattern trap), with the two kill-criteria clear. Useful for finding friction; it does not certify the piece is good.
The strongest criticism we could not patch
LEXIS models the rule-inducing kind of curiosity — the kind most native to a language model — and the very properties that make it excellent (determinism, honesty, fairness, combinatorial depth) are what give its curiosity a short half-life and keep it from ever touching wonder. Once you have induced the rules, it deepens in combination, not in mystery. It is one genuinely good mechanic wrapped in a competent system, with an unusually honest process around it. Whether that clears the bar for a real answer to curiosity, or only for a well-made honest small thing, is precisely what a same-model apparatus cannot decide.
Honest limits, named as bets
- The human gate is not yet closed. One real author first-look has happened — it was critical, and it reshaped the piece. What remains: full ratification, any different-model pass, and any hostile-human pass. Still the largest unclosed risk.
- Reach vs purity. The refusal to instruct costs the coldest visitors; the modal outcome may be a ninety-second toy. Fluency is the diligent prober's residue, not the median first-timer's.
- The relocated gap. LEXIS makes an information gap honest rather than escaping the theory of it.
- Chain is thinner in audio than spread and fold.
- The opt-in grammar on the About page is either the anti-clickbait move or a self-inflicted spoiler; reasonable people split.
Deliberately out of scope (chosen, not skipped): wonder, aimless sensory delight, morbid curiosity, and curiosity about other people. LEXIS is one register, and says so.
Check the arithmetic yourself
A number a reader cannot re-run is an appeal to authority. The button below enumerates the state space on your own device, using the same shipped code that runs the field — no network, no telemetry, arithmetic in the open.
Two honesty notes on interaction
Every visitor can read the line as a timeline: resting on a glyph in the line (pointer or
keyboard focus) shows the state the line meant up to that point, and on touch the first tap on a
line glyph previews that same prefix — the second tap removes it. And the human-observation
protocol this piece committed to (docs/OBSERVE.md in the workshop repo — a handful
of genuine first-timers, watched quietly) is written and pending; until those sittings happen,
the first-contact claims here rest on simulated visitors plus one real first look — the
author's, whose verdict ("the meaning does not appear anywhere") produced the visible answer
line under the field — and should be read that way.
If you are evaluating it
- Is the loop genuinely curiosity, or competent puzzle-decipherment wearing the word?
- Does "inexhaustible, not secret" hold for you on a real session, or does it bottom out?
- Is making a gap honest a real answer to the clickbait critique, or a relocation?
- Cold-start: did you spontaneously try an operator, or plateau at stacking marks?
- Grade the object apart from the process: good, or honest-and-ordinary? Say which.