Every figure, checked
Every number published here is measured on text the lexicon never trained on, recomputed independently from its source measurement, and held to thresholds the tooling enforces automatically. The figures are reproducible on request, and we are glad to walk through the methodology in detail.
Whole documents, not truncated samples. The adversarial set covers the cases that break tokenizers quietly: zero-width-joiner emoji, right-to-left scripts, combining marks, raw control bytes, mixed line endings, and private-use codepoints — the places where a collision corrupts text without raising an error anywhere downstream.
The question worth asking of any tokenizer result, answered with a measurement rather than a reassurance.
Holding out by content hash still draws both halves from one pool of tools, so we went further: the agentic corpora were re-partitioned by tool. A portion of the tool vocabulary was reserved out of training entirely, and the lexicon was then scored only on records using tools it had never encountered.
68–81% of the gain is structural and carries across to an entirely new tool set. We quote the range rather than a single figure because it varies by corpus, and we lead with the floor.
Where a lexicon meets traffic in an unfamiliar format the ratio is lower, and that lower figure is the one we put in front of you. A tuned result is never quoted for untuned traffic.
Gains vary by domain, so every engagement opens with a measurement of the customer's own corpus. Nothing is committed before that figure exists.
Training size and source travel with every sector result, so the basis for a number is always visible alongside it rather than a page away.
A tuned vocabulary changes what a request costs and who can read it. We describe it as exactly that, and size every claim to what has been measured.