Oculi·Dei
Oculi Dei LLC — Michigan

The first engagement

The Token Audit

We build a vocabulary on your corpus and measure what it does to your token counts — on held-out text, against the tokenizer you run today. You receive a reproducible figure for your own traffic, the cost it removes at your own rate, and the context headroom it returns.

50–100 MB
Representative sample required, by domain — a single export, not a data program
≥97.6%
Of achievable compression captured at that sample size, measured across three domains
6
Production tokenizers benchmarked against, spanning four model families
Yours outright
A quantified baseline of what your traffic costs you today

What the audit establishes

Four figures, each specific to your corpus and each reconstructible from the measurement that produced it.

01

Your compression ratio

A lexicon trained on your corpus and scored on a held-out partition it never saw, against your current tokenizer. A deterministic count over fixed text — the same inputs return the same number on re-run, which is what separates a measurement from a projection.

02

The cost it removes

Your ratio applied to your throughput at your own price per million tokens. Across published sectors that lands between $226k and $466k per terabyte at $3.00/M; your figure is computed from your volume, with the arithmetic shown so you can substitute your own rate.

03

Your context headroom

A window is counted in tokens, not in meaning. At the measured agentic ratio a 128K window carries 373K tokens-worth of your material, and 1M becomes 2.92M. This is the half of the return that is never billed and rarely priced.

04

Where the gain concentrates

A ranked decomposition across the traffic classes in your corpus, so deployment begins where the return per unit of work is highest. Agentic and structured traffic consistently lead; document-style prose follows.

What we need from you

Established empirically rather than assumed, by training at increasing corpus sizes against one fixed holdout.

Sample size is the question every engagement opens with, so we measured it. Lexicons were trained on disjoint pools from 1 MB to 400 MB and scored against a single held-out partition, making corpus size the only variable. The answer is domain-dependent, and we size the request accordingly.

Share of achievable compression captured at each sample size, measured per domain.
Traffic class25 MB50 MB100 MBWe ask for
Technical & research prose98.4%99.4%50 MB
Agentic & tool traffic95.0%97.9%100%100 MB
Source code93.5%97.6%100%100 MB

Homogeneous prose saturates early: beyond 50 MB an eightfold increase in corpus moved the ratio by 0.04%. Agentic and code traffic carry more structural variety and continue to reward larger samples, so we ask for 100 MB where the corpus supports it. Breadth across your traffic classes is what matters — a representative slice, not volume for its own sake.

Benchmarked against what you actually run

Your ratio is computed against your own tokenizer, not a published figure borrowed from another vendor's stack.

We measure against six production tokenizers spanning four model families. Across all of them the spread is only 7%, which means tokenizer choice moves the result far less than domain composition does — and our published figures are quoted against o200k_base, the most token-efficient of the six. Whatever you run, the comparison you receive is computed on your stack.

How the engagement runs

1 — SampleYou provide a representative slice of the text your models process, sized per the table above. Where data handling is constrained, the transfer is scoped to your requirements before anything moves.
2 — PartitionA portion is reserved out of training entirely. Every figure you receive is measured on text the lexicon has never encountered, split by content hash so the two sets are disjoint by construction.
3 — Build and measureA lexicon is trained on your corpus and scored against your current tokenizer, with the round trip verified byte-exact on your own documents — including control bytes, mixed encodings and unusual scripts.
4 — DeliverYour ratio, its value at your rate, your context headroom, and the ranked decomposition of where the return concentrates — with the method behind each figure.

Scope, timeline and fee are set against the size and composition of your corpus. Tell us what you run and we will come back with both.

What you are paying for

A lexicon is built and measured on your own corpus before any figure exists — engineering against real data, held to the same standard as every number we publish. What you acquire is a quantified account of what your traffic costs you today, on your infrastructure, at your rate.

The deliverable is yours outright. It stands as the baseline every subsequent decision is measured against, and it is reproducible — the same corpus and the same method return the same number, by anyone, at any time.

Next step

Tell us what you run

Domain, approximate volume, and whether you are self-hosted or on an API. That is enough for us to scope the audit.

Request an audit →