Back to Blog

A privacy checkpoint for legal documents: LiTiL Legal PII Finder

A legal-document privacy checkpoint that returns labeled sensitive spans with source offsets.

Legal text usually has to leave its original file before it becomes useful. It gets extracted, indexed, embedded, searched, shared, and sometimes sent to another service. Names, email addresses, phone numbers, account numbers, addresses, private URLs, and sensitive dates can travel with it.

LiTiL Legal PII Finder gives that pipeline a clear privacy checkpoint. It reads legal document text and returns the exact spans that match its seven-label policy. Each result includes the label, character offsets, text, and a model score. Your application gets something it can act on, not a general warning that a document might contain sensitive information.

Where it fits

The model belongs after OCR or text extraction and before search, retrieval, sharing, or outside processing.

A practical flow looks like this:

  1. Extract text while preserving document and page offsets.
  2. Run LiTiL Legal PII Finder against the seven published labels at the recommended 0.90 threshold.
  3. Merge overlaps created by chunking.
  4. Add deterministic detectors for fields your policy must catch outside the model's label set.
  5. Send the spans to a redaction view, replace them with typed placeholders, or queue the document for privacy review.

Keep the original offsets and the source of every match. That lets a reviewer see exactly what the model found, and it lets the team measure model spans separately from rule-based matches. If the same document later enters a retrieval system, the redaction record can travel with the processed copy.

This gives an intelligence stack a narrow job, a typed output, and a clear place to add policy. The result can be inspected without asking another language model to explain what it meant.

What was trained

This release is a complete fine-tuned GLiNER2 checkpoint based on fastino/gliner2-privacy-filter-PII-multi. You do not need to load a separate base and attach an adapter.

The post-training mix contained 2,000 rows: 1,660 generated examples and 340 negatives derived from public sources. The reviewed post-training inputs contained no private client or user data.

The policy covers:

  • account_number
  • address
  • email
  • person
  • phone_number
  • private_url
  • sensitive_date

That list is part of the model contract. An SSN, credential, or jurisdiction-specific identifier outside those labels needs a separate detector or an expanded, tested label policy.

What the measurements say

On 1,046 deduplicated synthetic-form rows, the selected checkpoint reached exact-span F1 of 0.9668. A separate six-case comparison, using the same text, labels, and threshold for both models, produced F1 of 0.9333 for the LiTiL model and 0.8571 for the matched base.

The larger exact-span result is the better aggregate measurement. The six-case comparison is useful because it checks the same public interface against the starting model, but six cases are still six cases.

Exact-span evaluation matters here. Finding the right name but returning the wrong boundaries creates bad redactions and broken placeholders. The measurement checks whether the predicted label and character span match the expected result.

What to build with it

The obvious use is assisted redaction. Show each suggested span in the document, let a reviewer accept or reject it, and save the action. The same output can support a masking service before retrieval, a policy check before external processing, or a privacy queue for documents with a high concentration of sensitive fields.

Start with the published threshold, run a representative sample from the target workflow, and measure recall by label. Some teams care far more about missing an account number than overflagging a date. That decision belongs in the application policy, where it can be tested and changed.

The model, tokenizer, runnable example, and exact seven-label contract are available under Apache 2.0 in the LiTiL Legal PII Finder repository.