Introducing LiTiL’s models for legal search
LiTiL released three specialized models for legal query-to-passage retrieval.
We've released three LiTiL models for legal search: LiTiL Octen Law 8B, LiTiL Embed 0.6B, and LiTiL ColBERT 300M. They find passages in contracts, cases, statutes, and policy libraries that are relevant to a legal question. All three are available on Hugging Face for teams building their own search and RAG systems.
LiTiL Octen Law 8B took the top spot on MTEB Law, with LiTiL Embed 0.6B in fourth. LiTiL ColBERT 300M placed second on NanoLaw. These rankings are as of September 14, 2026.
What LiTiL built
An intelligence layer needs a good foundation. For legal work, that starts with getting the right clause, case, or statute in front of the person, agent, or model doing the work. More room in a prompt helps you fit documents. You still have to select the useful passages.
We focused on that selection step. We trained the models on legal questions paired with relevant passages, using two approaches to search: whole-passage embeddings and ColBERT's token-level matching. The smaller embedding model also learned from a larger teacher model. Our hypothesis was that training on legal examples would help the models distinguish a useful passage from one that merely looks related.
Three models, three deployment choices
LiTiL Octen Law 8B is the larger dense encoder. It turns a question and a passage into vectors, then ranks the closest matches. Use it when you want the higher-capacity dense option for a substantial legal corpus.
LiTiL Embed 0.6B is the smaller dense encoder. It produces 1,024-dimensional vectors and uses Octen-Embedding-8B as a teacher during training. It is a drop-in option for a local embedding index when model size matters.
LiTiL ColBERT 300M uses token-level late interaction. With PyLate, it scores a question against the words inside each passage instead of reducing the whole passage to one vector. It uses a different kind of index from dense search.
The smaller models need roughly 1.2 GB and 0.6 GB just to hold their weights at 16-bit precision. That is the memory floor; running them also needs room for computation and the search index.
| Model | Weight memory at 16-bit precision | Apple silicon Mac |
|---|---|---|
| LiTiL Embed 0.6B | About 1.2 GB, plus runtime memory | An 8 GB Mac is a reasonable starting point for short passages and small batches. |
| LiTiL ColBERT 300M | About 0.6 GB, plus runtime memory | An 8 GB Mac has room for the weights and a light workload; the index and PyLate backend need testing. |
| LiTiL Octen Law 8B | About 15.2 GB, plus runtime memory | The 16-bit weights alone exceed an 8 GB Mac's memory. Start testing with 32 GB unified memory or a 24 GB GPU. |
These are calculated weight sizes, with device suggestions for testing. We have not measured minimum runtime memory on these machines. FP32 weights take twice as much memory as 16-bit weights. Apple silicon shares memory with macOS and other applications; keep passages short and batches small on an 8 GB machine. Sentence Transformers supports CPU and Apple MPS execution. Reduced precision and the chosen backend should be checked against the model's published results.
Where they fit
In retrieval-augmented generation (RAG), a search step supplies documents to the model writing the answer. These models handle that search. Split documents into passages, keep their source locations, and index them. When a question comes in, retrieve the relevant passages and send them with their locations to the next step.
For a law firm, that could mean searching a deal archive for relevant clauses. An in-house team could retrieve contract language for a playbook review. A product team could use the same search step to supply sources to an assistant. Running the model and index locally lets the team process search queries and documents within its own environment. The application still manages access permissions and where those passages go next.
The results
Our LiTiL ColBERT 300M model is roughly 25× smaller by parameter count than NVIDIA's Nemotron-3-Embed-8B-BF16, and placed second behind it on NanoLaw. The leaderboard lists 312M parameters for ours and 7.95B for NVIDIA's. NVIDIA led on Borda score, 96.67 to 95.67; LiTiL's mean nDCG@10 was 75.01 versus 71.67. Borda combines task ranks, so the overall winner can differ from the model with the highest mean score.
LiTiL Embed 0.6B placed fourth on MTEB Law with roughly one-thirteenth the parameters of the models ranked above it.
That makes the 0.6B model worth testing for a legal RAG application with a conventional vector index and a limited memory budget. ColBERT offers token-level matching in a similarly small encoder, useful for ranking candidate passages or building search with PyLate. Its index stores multiple vectors per passage, so the smaller model does not automatically mean a smaller index or faster search.
See how the models ranked on MTEB Law and NanoLaw as of September 14, 2026. Follow either benchmark link for the latest results.
| Benchmark | LiTiL model | Rank | Mean nDCG@10 |
|---|---|---|---|
| MTEB(Law, v1) | LiTiL Octen Law 8B | #1 | 76.77 |
| MTEB(Law, v1) | LiTiL Embed 0.6B | #4 | 73.30 |
| NanoLaw | LiTiL ColBERT 300M | #2 | 75.01 |

The MTEB Law leaderboard covers eight legal retrieval tasks, including case law, statutes, contract questions, and legal question answering.

The current NanoLaw view ranks models across four tasks: NanoGerDaLIRSmall, NanoLeCaRDv2, NanoLegalBenchConsumerContractsQA, and NanoLegalQuAD.
nDCG@10 rewards relevant passages near the top of the first ten results. The means shown here use a 0 to 100 scale. Both boards use Borda ranking, which combines a model's place across tasks. Read each board on its own task set.
The scores measure retrieval on the evaluated tasks. A team should still test its own documents and questions before relying on the result in a workflow.
An open intelligence layer for legal work
Retrieval is one component in the open intelligence layer LiTiL is building for legal work. The goal is a set of specialist models that teams can inspect, test, and arrange around their own data and workflows. Other LiTiL releases handle intake routing, sensitive-information marking, contract terms, and review playbooks.
The Open Intelligence Models collection brings those releases together. Each model has a defined job. Teams decide how the pieces fit their own stack.
Foundations and credits
The two dense models build on Octen's work, which builds on Qwen3 Embedding. The 0.6B model also uses Octen-Embedding-8B as its teacher. LiTiL ColBERT builds on LightOn's mLateOn and, upstream, mmBERT. We credit those teams.
Training sources named across the model cards include Hanno Labs and Clause Logic, their German Wikipedia retrieval pairs, CUAD, GerDaLIR, LeCaRDv2, LegalPincite, STARD, ContractNLI, CaseHOLD, GermanQuAD, MIRACL, and Natural Questions. We credit those projects and their authors; each model card identifies its named training sources.