LiTiL Contract Extractor 1.7B: capture the term and keep the source
A contract extraction model that returns the language answering a specific question or a stable absence value.
A clause label tells you what kind of language you found. Most contract workflows still need the exact answer: the effective date, governing law, party name, renewal term, insurance requirement, or another value from the agreement.
LiTiL Contract Extractor 1.7B handles that second step. Give it contract text, a category, and one question. It returns the relevant language inside an <answer> element. If the term is absent, it returns <answer>NOT_PRESENT</answer>.
That output is easy to parse and easy to retain with its source. A contract system can write the extracted value to a structured record while preserving the original span and clause location for review.
Where it fits
The extractor belongs after document parsing and retrieval. For a long agreement, the application should first find the passage most likely to contain the answer. Then it sends one category-specific question with that text.
<context>
CONTRACT TEXT OR RETRIEVED PASSAGE
</context>
Category: Effective Date
Question: What is the effective date of this agreement?The model returns either a supporting span or NOT_PRESENT. The parser should require one <answer> element. Preserve the returned text before applying date, currency, or entity normalization, then store the normalized value beside the original span.
This separation helps when the next system needs evidence. A renewal dashboard may display a normalized date, but a reviewer can still open the clause that produced it. A playbook model may use the extracted term as deal context, but the application retains the contract language behind that fact.
What the measurements support
The public evaluation contains 2,091 contract-question pairs from 51 contracts. The 459 training contracts and 51 test contracts are disjoint.
LiTiL Contract Extractor reached 74.46% normalized exact match, compared with 70.78% for the matched Qwen base. Token F1 increased from 0.7306 to 0.7690; character Jaccard increased from 0.7460 to 0.7927. The largest saved gains were on document name, agreement date, effective date, parties, and insurance questions.
The aggregate contains two different behaviors: copying the right span when a term is present and returning NOT_PRESENT when it is absent. A useful application test should report those behaviors separately. It should also measure the retrieval step because the extractor cannot return language it never receives.
A practical implementation
Start with the fields the business already tracks. For each field, define the contract category, canonical question, output parser, and normalization rule. Retrieve a focused passage, call the model once, and attach the result to the source location.
A basic pipeline can look like this:
- Segment the agreement and keep page or character offsets.
- Use a tagger, classifier, or retrieval query to find likely passages.
- Ask one extraction question at a time.
- Parse exactly one
<answer>block. - Convert
NOT_PRESENTto a structured null. - Normalize the value without discarding the returned span.
- Route uncertain or conflicting fields to review.
Contract Extractor is a PEFT LoRA adapter for Qwen/Qwen3-1.7B. The adapter is about 266 MiB. The public card suggests 6 to 8 GB of accelerator memory as a practical starting point for short or retrieved passages with the BF16 base. Training used a 2,048-token envelope, so focused passages are the best match for the published setup.
The repository includes the adapter, prompt format, parser, loading example, and exact measured results: LiTiL Contract Extractor 1.7B on Hugging Face.