Vitrine IV

What the ratio is predicted by

Every seated piece was counted on both sides. The measurement is old; what is new is that it now has a predictor, and the predictor is not in the Turkish.

The criterion

The project’s own published counting criterion, printed so the figures can be re-run by another hand:

tokens matching [\w’'-]+ over the body, front matter excluded, H1 title line removed; list numerals count, the title does not

English, over the sixty
50,080 words
Turkish, over the sixty
35,351 words
The book’s ratio
0.706

The question was malformed

For most of the stage a piece was convicted when its ratio left the band of its family — documents in one band, narration in another. Two explanations for the bands were tested and both failed. Marker tokens that cross at one by construction, such as field numerals and build directives, move the spread of the family that carries most of them from 0.061 to 0.058: they explain the form before story one exactly and the book not at all. Length gives r = −0.26, too weak to carry a band.

What holds is the English’s own function-word density. English the, a, of, to, is, and have no Turkish word: they are absorbed into suffixes or they vanish. So a piece dense in them must compress, and a piece of bare labels cannot.

Sixty pieces: the Turkish word count against the English, plotted against the density of the EnglishA scatter of sixty points falling from upper left to lower right, with the least-squares line through them and two points ringed below it.0.200.250.300.350.400.450.600.650.700.750.800.85DS-042DS-043DS-026English function-word densityTurkish words ÷ English words
Sixty pieces. The line is the least squares of the ratio on the density: ratio = 0.927 − 0.598 × density, r = −0.718, residual standard deviation 0.034. The two ringed points are the only two that fall further than two standard deviations from their own prediction: DS-042 (−0.096) and DS-043 (−0.091) and DS-026 (0.069).
The least squares
ratio = 0.927 − 0.598 × density
Correlation, over 60 pieces
r = −0.718
Within a single family of 7 pieces
r = −0.874 — stronger inside the family than across all of them, which is what shows the family is not the variable
Residual standard deviation
0.034
Beyond two standard deviations of their own prediction
3 of 60 — DS-042 (−0.096), DS-043 (−0.091), DS-026 (0.069)

The sixty pieces plotted are the 59 seated pieces that carry a ratio and the form before story one. DS-038 carries none, by the law in vitrine I: it is the same text on both sides, so its ratio is one by construction and measures nothing. The returned field is too short to measure and is excluded by the instrument’s own floor.

And every conviction before this one measured the wrong thing

The register’s own row, on the band that had been convicting pieces for the whole stage:

the families were a PROXY and the bands convict a piece for the DENSITY OF ITS ENGLISH. Two hypotheses tested and both failed — incompressible marker tokens (explains the PROLOGUE exactly and the book not at all) and length (r = −0.27). What explains it is the English’s own function-word density: r = −0.732 over 60 pieces, and −0.872 WITHIN family D — the within-family correlation being STRONGER than the across-family one is the proof that the family is not the variable

docs/VERIFIER-LOG.md, row V-80. The instrument is tools/vh_tr_ratios.py, and the only figure a pass may convict on is the residual.

Two figures for the length hypothesis stand on this page and they differ in the second decimal: −0.26 above, derived at this build over the 60 pieces plotted, and −0.27 inside the row, derived at the filing over the set that filing used. Neither is adjusted to agree with the other and neither is dropped. Both scopes are stated instead, which is the only thing that makes two numbers for one test readable rather than a contradiction — and which is the subject of vitrine VII.


Previous · III The pavilion V · Next

Copyright © 2026 Hulki Okan Tabak. All rights reserved.