Evaluation
how this piece was tested, and what cannot be known
POTHOS was built autonomously by Claude with panels of simulated specialists designed to disagree — a premise challenge run from the bare word before anything was committed, a ten-seat concept council, a playtest, and an ordered audit swarm with a bespoke desire-fidelity seat that reads the source. Every session's raw output and every decision is preserved in the private workshop repository. Two honesty rules govern everything below: simulated verdicts are labeled simulated (the panels share one base model; their agreement is convergence, not proof), and numbers are measured, not asserted.
Fair warning: this page shows blooms
The garden withholds color until a bloom is really open. This page is the lab bench — the variety evidence below renders full, colored blooms from fixed seeds. Nothing here is secret; spoiling is opt-in. If you would rather wait for your own, stop here.
The variety evidence
Success criterion 6 requires the bloom be worth wanting: a measured generative range with no degenerate outputs. The plates below are drawn live by the same pure function the garden uses, from 24 fixed seeds across three species (covering both bloom-family weightings the garden uses). The full sweep — 1,000 seeds per species, all four species, re-run after the audit fix pass and recorded in the workshop repository — checks: no empty or off-frame plants, no NaN geometry, all five bloom families and five color bands represented, and the sentence generator's petal count always matching the drawing.
You do not have to take that on faith. Run the census yourself, here, in your own browser — the same pure function the garden draws with, swept over every seed:
Measured numbers
| Claim | Measured |
|---|---|
| Primary text contrast (ink on ground) | 14.08 : 1 (AA needs 4.5) |
| Secondary text (ink-2 on ground) | 7.83 : 1 |
| Tertiary text floor (ink-3 on ground) | 5.68 : 1 |
| Meaning-bearing strokes (line on ground) | 3.74 : 1 (non-text needs 3) |
| Genome colour floor (darkest reachable genome, worst ground it appears on) | 4.24 : 1 (deep-rose on the lightest hour-tinted garden ground; 4.31 on the base ground — non-text needs 3; re-measured live by the census above) |
| All tokens on every ground (four hour tints and the plate ground) | lowest 3.56 : 1 (non-text), 5.40 : 1 (text) |
| Deterministic render (same seed, same clock) | identical markup, 800/800 checks |
| Degenerate outputs in sweep | 0 (sweep re-run at audit; see below) |
The code invariants (the desire-fidelity audit)
The piece's ethics are testable in source. The audit checks, with file-and-line evidence: growth is a pure function that never reads visitor facts; no randomness outside the sowing gesture; the only writes after sowing are witnessed, seen-early, and taken-in; pressing images depend on the seed and what-you-saw mode only; nothing counts or scores the ledger; no seconds are ever rendered; the page title and icon never change; no notification, vibration, badge, or install-prompt code exists; the counter never receives the URL fragment; the post-sow screen ends in a dead-end; and absence — simulated at 400 days — is met by nothing but the truth. The full checklist and results live in the workshop repository.
Timing, for the record
No task in this piece is time-limited in the sense of WCAG 2.2.1: witnessing is presence-only (no interaction, no reading-speed dependency), and a missed window produces a complete, honored record, not a failure state. The windows themselves are essential to the meaning — they are the piece's one tooth — and are disclosed, generously sized (six hours to three days), and placeable away from sleep at sowing. The rehearsal's bloom holds at least ninety seconds.
What we cannot know
The mocked clock verifies the mechanism at every hour of a ninety-day wait in seconds. It cannot verify the wanting. Whether an appointment with an unknown flower can make a person feel anything across real days — whether the return is longing or just a bookmark — is untestable before real visitors wait real time, and no simulated playtester can wait. The piece's central claim is therefore a bet, stated plainly here. The author has now ratified it (5 July 2026) — a decision to stand by shipping it, not a proof it works; real visitors remain the only judge of that. Panels of one base model agreeing with each other — however adversarially prompted — are evidence of coherence, not of worth. Two more things cannot be known, and should be said: the garden records nothing about what you do in it (the only count anywhere is the anonymous page view named in the colophon, and it never sees what you sow), so whether the waiting produces wanting can never arrive as data — only as the author's ruling and whatever visitors choose to tell; and the one live human signal this run received (a doubt about the name itself) found in a glance what the entire simulated apparatus had missed, which is the clearest measure of that apparatus's ceiling.
Playtest and audit record
The simulated playtest (eight personas, one driving the real page with a mocked clock) and the ordered audit swarm (editorial, compatibility, security, vulnerability, accessibility, performance, self-audit, desire-fidelity, ablation) ran against the build; findings and fixes are recorded in the workshop repository, with raw outputs preserved. Verdict at ship: no unresolved critical or high findings. The author ratified the piece on 5 July 2026; it now ships as the finished fourth concept.