Oculi·Dei
Oculi Dei LLC — Michigan

Private vocabulary for trained models

Your model,
speaking your language.

Every language model's vocabulary is frozen at pretraining, so every deployment pays token overhead on text that vocabulary never saw. Oculi Dei retrofits a private, domain-tuned vocabulary onto a model that is already trained — cutting token counts losslessly, without a pretraining run.

2.92×
Fewer tokens on agentic tool traffic, against GPT-4o's tokenizer — the hardest baseline we test
573,218
Documents round-tripped losslessly — 904 MB, zero failures
92.05%
Vocabulary in active use, against o200k's 72.60%
3 / 3
Architectures passed, including a sparse mixture-of-experts

The part nobody prices

Fewer tokens is also a bigger context window

The same window, holding more of your content. You are not billed for the difference and you do not change models to get it.

A context window is measured in tokens, not in meaning. Encode the same material in 2.92× fewer tokens and the window you already pay for carries 2.92× more of it — the whole codebase instead of a slice, the full trace instead of the tail. Cost falls and capability rises together, which is why this lands on both sides of a buying conversation.

Shown at the measured 2.916× agentic ratio. A general-purpose corpus gains 1.527×; a lexicon tuned to one customer's own traffic reaches 3.14×.

Where to look

The detail, in full

TOKEN AUDIT — START HERE

Measure your own traffic first

The first engagement: a lexicon tuned to your corpus, scored on held-out text, so you get your reduction, its dollar value at your rate, and your context headroom before committing to anything.

Request an audit →

RESULTS

Eleven sectors, measured and priced

Token reduction per sector on held-out traffic, the value that represents per terabyte, and the context-window expansion that comes with it.

See the results →

METHOD

The lexicon and the transplant

What a domain-tuned vocabulary is, how it is retrofitted onto a model that is already trained, and what is prior art versus what is ours.

How it works →

VERIFICATION

How every figure is checked

Lossless round-trip at scale, held-out by content hash, and how much of the gain carries across to tools the lexicon has never seen.

See the checks →

CONTACT

Operators, researchers, investors

Tell us the domain, the rough volume, and whether you are self-hosted or on an API. That is enough to say what is measurable about it.

Get in touch →