Oculi·Dei
Oculi Dei LLC — Michigan

Every figure, checked

Verification

Every number published here is measured on text the lexicon never trained on, recomputed independently from its source measurement, and held to thresholds the tooling enforces automatically. The figures are reproducible on request, and we are glad to walk through the methodology in detail.

Lossless at scale

573,218
Documents round-tripped
904 MB
Total text verified
0
Round-trip failures
0
Adversarial failures

Whole documents, not truncated samples. The adversarial set covers the cases that break tokenizers quietly: zero-width-joiner emoji, right-to-left scripts, combining marks, raw control bytes, mixed line endings, and private-use codepoints — the places where a collision corrupts text without raising an error anywhere downstream.

Measured on held-out traffic

Split by content hashEvery sector figure is measured on text the lexicon never trained on. Training and evaluation sets are disjoint by construction.
Matched-train, all elevenEach sector's lexicon is trained on 50 MB or more of that sector's traffic, so the sectors are directly comparable to one another.
Independently recomputedEach published figure is recalculated from its source measurement by a separate validator — ratio, tokens per terabyte, percentage and dollars — and both must agree before it ships.
Quoted against the hardest baselineFigures are published against the most token-efficient tokenizer we test. Measured across six production tokenizers from four families, the same lexicon performs better against every alternative.
Thresholds enforced by toolingMinimum corpus size and maximum plausible ratio are applied mechanically when the record is generated, not left to judgment.

The gain survives tools the lexicon has never seen

The question worth asking of any tokenizer result, answered with a measurement rather than a reassurance.

Holding out by content hash still draws both halves from one pool of tools, so we went further: the agentic corpora were re-partitioned by tool. A portion of the tool vocabulary was reserved out of training entirely, and the lexicon was then scored only on records using tools it had never encountered.

68–81% of the gain is structural and carries across to an entirely new tool set. We quote the range rather than a single figure because it varies by corpus, and we lead with the floor.

What we quote, and what we don't

THE CONSERVATIVE CASE

We lead with the floor, not the ceiling

Where a lexicon meets traffic in an unfamiliar format the ratio is lower, and that lower figure is the one we put in front of you. A tuned result is never quoted for untuned traffic.

YOUR OWN CORPUS

The number you get is measured on your traffic

Gains vary by domain, so every engagement opens with a measurement of the customer's own corpus. Nothing is committed before that figure exists.

PROVENANCE

Each figure names the corpus behind it

Training size and source travel with every sector result, so the basis for a number is always visible alongside it rather than a page away.

SCOPE

Compression and privacy, stated as such

A tuned vocabulary changes what a request costs and who can read it. We describe it as exactly that, and size every claim to what has been measured.