LiTiL ClauseTagger 4B: give contract passages a common vocabulary
A multi-label contract model that gives passages a consistent taxonomy for indexing and downstream routing.
Contract language is messy. The same provision can appear under different headings, use unfamiliar wording, or do more than one thing at once. A contract intelligence system still needs stable labels if it is going to index, route, and review those passages consistently.
LiTiL ClauseTagger 4B takes one contract passage and assigns one or more labels from the included contract taxonomy. It returns a short JSON object containing sorted category IDs. An empty list means the passage does not directly support a category.
This is a multi-label model. A sentence that gives both the effective date and the expiration date can carry both labels. That is useful when splitting the language into smaller fragments would lose context or create duplicate source records.
Where it fits
The model expects a passage that has already been found. Put it after a parser, clause segmenter, or retrieval layer. Keep the source document and location attached to the passage, then store the returned labels on the same record.
Those labels can support:
- clause-level search and filtering;
- category-specific review queues;
- analytics across a contract repository;
- selection of the right extraction question; and
- selection of the right playbook rule.
For example, a renewal label can trigger an extraction step for the renewal term and notice deadline. Another category label can route a passage to the relevant company rule. The tag is a routing primitive. The source text remains the evidence.
Input and output
ClauseTagger is a PEFT LoRA adapter for Qwen/Qwen3.5-4B. The system prompt contains the contract taxonomy. The user message starts with Contract segment:, includes the passage, and asks for exactly one JSON object.
{"class_ids":[...]}The application should validate four things before accepting the result: one class_ids key, an array of integers, no duplicate IDs, and IDs that appear in the included catalog. The published runtime uses greedy decoding, disables thinking, and allows up to 64 generated tokens.
The model card includes a saved passage with both an effective date and an expiration date. ClauseTagger returns the two matching IDs. The unchanged base chose a different pair on that case.
What the measurements support
The published evaluation uses 500 attorney-annotated positive passages from 70 CUAD documents. The base and adapter received the same prompt and used the same renderer, decoding, and parser.
Macro F1 increased from 0.6195 for the base to 0.7243 for ClauseTagger. Micro F1 increased from 0.7160 to 0.8193. Exact label-set match rose from 63.6% to 70.8%, and both models returned valid JSON on every case. The paired bootstrap interval for the macro-F1 improvement ran from +0.0711 through +0.1431.
These measurements start from passages that were already selected and annotated. In an end-to-end contract system, retrieval and segmentation still determine whether the right language reaches the tagger. The useful deployment test therefore has two parts: passage recall upstream and label quality here.
A practical implementation
Build the first version around one downstream decision. If the goal is renewal review, retrieve likely renewal passages, run ClauseTagger, and validate the returned IDs. Send accepted renewal labels to an extractor that captures the term and notice period. Record any passage the model leaves unlabeled or sends to an unexpected category.
The adapter is 342,027,648 bytes. The public card recommends 12 to 16 GB of available accelerator or unified memory as a practical starting point for short passages with the 4B BF16 base. It also identifies the exact Qwen base revision and the disabled-thinking renderer used for evaluation.
Everything needed to load the adapter and map IDs to labels is in the repository: LiTiL ClauseTagger 4B on Hugging Face.