The first engagement
We build a vocabulary on your corpus and measure what it does to your token counts — on held-out text, against the tokenizer you run today. You receive a reproducible figure for your own traffic, the cost it removes at your own rate, and the context headroom it returns.
Four figures, each specific to your corpus and each reconstructible from the measurement that produced it.
A lexicon trained on your corpus and scored on a held-out partition it never saw, against your current tokenizer. A deterministic count over fixed text — the same inputs return the same number on re-run, which is what separates a measurement from a projection.
Your ratio applied to your throughput at your own price per million tokens. Across published sectors that lands between $226k and $466k per terabyte at $3.00/M; your figure is computed from your volume, with the arithmetic shown so you can substitute your own rate.
A window is counted in tokens, not in meaning. At the measured agentic ratio a 128K window carries 373K tokens-worth of your material, and 1M becomes 2.92M. This is the half of the return that is never billed and rarely priced.
A ranked decomposition across the traffic classes in your corpus, so deployment begins where the return per unit of work is highest. Agentic and structured traffic consistently lead; document-style prose follows.
Established empirically rather than assumed, by training at increasing corpus sizes against one fixed holdout.
Sample size is the question every engagement opens with, so we measured it. Lexicons were trained on disjoint pools from 1 MB to 400 MB and scored against a single held-out partition, making corpus size the only variable. The answer is domain-dependent, and we size the request accordingly.
| Traffic class | 25 MB | 50 MB | 100 MB | We ask for |
|---|---|---|---|---|
| Technical & research prose | 98.4% | 99.4% | — | 50 MB |
| Agentic & tool traffic | 95.0% | 97.9% | 100% | 100 MB |
| Source code | 93.5% | 97.6% | 100% | 100 MB |
Homogeneous prose saturates early: beyond 50 MB an eightfold increase in corpus moved the ratio by 0.04%. Agentic and code traffic carry more structural variety and continue to reward larger samples, so we ask for 100 MB where the corpus supports it. Breadth across your traffic classes is what matters — a representative slice, not volume for its own sake.
Your ratio is computed against your own tokenizer, not a published figure borrowed from another vendor's stack.
We measure against six production tokenizers spanning four model families. Across all of them the spread is only 7%, which means tokenizer choice moves the result far less than domain composition does — and our published figures are quoted against o200k_base, the most token-efficient of the six. Whatever you run, the comparison you receive is computed on your stack.
Scope, timeline and fee are set against the size and composition of your corpus. Tell us what you run and we will come back with both.
A lexicon is built and measured on your own corpus before any figure exists — engineering against real data, held to the same standard as every number we publish. What you acquire is a quantified account of what your traffic costs you today, on your infrastructure, at your rate.
The deliverable is yours outright. It stands as the baseline every subsequent decision is measured against, and it is reproducible — the same corpus and the same method return the same number, by anyone, at any time.
Next step
Domain, approximate volume, and whether you are self-hosted or on an API. That is enough for us to scope the audit.