Private vocabulary for trained models
Every language model's vocabulary is frozen at pretraining, so every deployment pays token overhead on text that vocabulary never saw. Oculi Dei retrofits a private, domain-tuned vocabulary onto a model that is already trained — cutting token counts losslessly, without a pretraining run.
The part nobody prices
The same window, holding more of your content. You are not billed for the difference and you do not change models to get it.
A context window is measured in tokens, not in meaning. Encode the same material in 2.92× fewer tokens and the window you already pay for carries 2.92× more of it — the whole codebase instead of a slice, the full trace instead of the tail. Cost falls and capability rises together, which is why this lands on both sides of a buying conversation.
Shown at the measured 2.916× agentic ratio. A general-purpose corpus gains 1.527×; a lexicon tuned to one customer's own traffic reaches 3.14×.
Where to look
The first engagement: a lexicon tuned to your corpus, scored on held-out text, so you get your reduction, its dollar value at your rate, and your context headroom before committing to anything.
Token reduction per sector on held-out traffic, the value that represents per terabyte, and the context-window expansion that comes with it.
What a domain-tuned vocabulary is, how it is retrofitted onto a model that is already trained, and what is prior art versus what is ours.
Lossless round-trip at scale, held-out by content hash, and how much of the gain carries across to tools the lexicon has never seen.
Tell us the domain, the rough volume, and whether you are self-hosted or on an API. That is enough to say what is measurable about it.