Back to Blog

Give every contract clause a stable name: LiTiL Clause Classifier

A contract classifier that assigns one category to a provision and returns scores across the available taxonomy.

Contract language varies. The work still has to land in a stable category before a system can search it, compare it, apply a playbook, or report on it.

LiTiL Clause Classifier assigns one category from a 100-label contract taxonomy to an individual provision. It also returns a score for every category, which gives an application more to work with than a label alone.

The model is a complete ModernBERT checkpoint. Load it directly from the LiTiL Labs repository, pass one provision, and get the selected label plus its score distribution. There is no adapter to attach and no separate base-model download.

Where it fits

Clause classification belongs after document parsing and clause segmentation. Keep the source location attached to every provision, then classify each provision independently.

The selected category can drive several parts of a legal intelligence stack:

  • Write normalized labels into a clause index so differently worded provisions appear in the same search.
  • Route confidentiality, assignment, governing-law, or change-of-control language to the matching playbook.
  • Build agreement-level coverage views from the clauses actually present.
  • Send low-scoring or closely split predictions to a reviewer.

That last use matters. The score distribution has not been calibrated as a probability, but it still helps rank alternatives and identify uncertain classifications. A production interface can show the top few labels beside the clause rather than pretending every argmax deserves equal confidence.

Multi-topic language still receives one category. If a long paragraph combines several provisions, better segmentation will usually help more than trying to make the classifier solve the document structure too.

What was trained

LiTiL Clause Classifier starts from answerdotai/ModernBERT-large. It was fine-tuned on 60,000 public LEDGAR contract provisions distributed through the LexGLUE dataset. The model reads up to 512 tokens and applies softmax across 100 mutually exclusive labels.

The ordered label map ships with the model and stays aligned with the classification head. That sounds like a small implementation detail until a label file drifts and every output silently points to the wrong category. Keep the released configuration and labels together.

No private client or user data was used in the reviewed post-training set.

What the measurements say

The selected checkpoint was evaluated on the 10,000-provision public LEDGAR test set. It reached 88.82% accuracy and 83.17% macro F1 across the 100 labels.

Accuracy answers how often the top label was correct. Macro F1 gives each category equal weight, so the common categories cannot completely hide weak performance on smaller ones. Both numbers are useful for a taxonomy this wide.

A separate CPU interface check produced the expected category on four synthetic provisions covering confidentiality, governing law, severability, and counterparts. That check is operational proof that the packaged model, tokenizer, and label mapping work together. The 10,000-provision test result carries the real aggregate measurement.

What I would build with it

Start with clause search. Segment a contract, classify the provisions, and let a reviewer filter all governing-law or assignment clauses across a document set. Then connect the same label to downstream playbooks and specialist models.

For example, a provision labeled Governing Laws can go to jurisdiction extraction. A confidentiality provision can go to a deviation checker. The classifier decides the next useful question. It does not have to answer every question itself.

That division of work is the point of the open intelligence layer. A small, measured model creates structure. Other components can then retrieve precedent, apply policy, or ask for human review using the same category.

The complete checkpoint, tokenizer, label map, and offline runner are available under Apache 2.0 in the LiTiL Clause Classifier repository.