Back to Blog

LiTiL Legal Intake 4B: bounded decisions at the front of a legal workflow

A bounded-output model for intake decisions, routing codes, issue codes, citations, and short extracted values.

Legal intake breaks when the answer cannot be consumed by the system that asked the question. A paragraph may sound reasonable and still fail to select a queue, fill a field, or tell the workflow whether it has enough information to continue.

LiTiL Legal Intake 4B is trained around bounded answers. The application supplies the source text, the specific task, and the output choices it will accept. The model returns one short action, route, issue code, section identifier, or extracted value.

That makes it useful near the front of a legal intelligence stack. It can check whether the request contains the needed information, select the right queue, identify a supplied issue or topic code, cite a section, or return a short value. The workflow validates the response against the allowed vocabulary before it changes state.

What bounded intake looks like

For a workflow disposition, the application can allow only proceed, insufficient evidence, counsel review required, privacy review required, or escalate: prompt injection. For a clause-presence task, it can require yes or no. Routing requests name their allowed queues in the prompt. Issue spotting and topic classification work the same way with application-defined codes.

The output vocabulary is part of the interface. If a route should be one of six queue names, put those six names in the request and reject any other completion. This keeps the model's job narrow and gives the workflow a deterministic boundary.

The same pattern works for short extraction. If the intake record needs a particular value, the prompt defines the task and expected shape. For evidence citation, the model can return a lowercase section identifier such as s1. Store that identifier beside the result so a reviewer can open the supporting text.

Where it fits

Use Legal Intake after the incoming request and its attachments have been converted to text. A typical flow looks like this:

  1. Normalize the request and identify its source documents.
  2. Supply one bounded intake task and its allowed answers.
  3. Run the model with greedy decoding and a short generation limit.
  4. Validate the answer against the supplied vocabulary.
  5. Store the accepted answer with the matter ID and source text.
  6. Select the next queue, form, specialist model, or review step.

The model is broader than the separate LiTiL Legal Request Router 1.5B. The router is for top-level application domains. Legal Intake can handle missing-information decisions, workflow dispositions, issue codes, topic codes, source citations, and short values once the workflow knows the task it wants answered.

What the measurements support

The public card reports a 511-case panel covering several trained task formats. Legal Intake scored 94.91% across the panel versus 75.93% for the matched Qwen base. On 151 real contract-clause tasks, it scored 95.36% versus 91.39%. Routing improved from 42.50% to 82.50%, and OCR-noisy inputs improved from 82.50% to 95.00%.

The adapter scored 40/40 on the included prompt-injection dispositions, 40/40 on legal-boundary dispositions, and 20/20 on privacy-review dispositions. It scored 37/40 when the correct answer was insufficient evidence. Those results support the published bounded-output tasks. An application should still test its own source formats, queue vocabulary, and escalation rules.

Training and runtime

Legal Intake is a PEFT LoRA adapter for Qwen/Qwen3-4B-Instruct-2507. It was trained on 1,730 unique cleanroom train and development rows from public LegalBench and LEDGAR tasks plus synthetic intake, extraction, missing-information, privacy, and embedded-instruction examples. The reviewed post-training set contained no private material.

The 4B base is about 8 GB in BF16 weights, and the adapter adds about 126 MB. The card suggests roughly 10 to 12 GB of accelerator memory for short prompts. The retained long-context panel used much longer prompts and more memory, so a production workflow should set its context window from the actual intake documents it expects.

The public repository includes the adapter, tested base revision, exact system prompt, output contracts, and a loading example: LiTiL Legal Intake 4B on Hugging Face.