A Correct Answer Can Still Fail the Assignment
Legal work ends in something another person has to review, use, send, sign, or file. That changes how agents should be built and measured.
A model can answer the legal question and still fail the assignment.
The memo may be incomplete. The redline may be unusable. The diligence table may be three documents short. The conclusion may be right but unsupported by a citation anyone can check. Every sentence can look reasonable on its own while the work product fails review.
In short-form legal benchmarks, the answer is the unit of measurement. A task may ask whether the system identified an issue, selected a label, retrieved a source, or produced an expected string. Those tests measure specific capabilities. Completing an assignment asks more of the system.
An assignment has a record, instructions, constraints, tools, and a deliverable. It has things the agent must notice and things it must ignore. It has a stopping point. It often has more than one defensible answer, but not an unlimited number of acceptable work products.
The work ends in an object someone else has to use.
The deliverable changes the system
Consider contract review. Finding a change-of-control provision is one capability. Reviewing the agreement is a larger job.
The agent may need to identify which party it represents, read the agreement and the relevant deal context, compare the language against a versioned playbook, propose an edit, explain the issue, keep the issue list consistent with the redline, and preserve the document well enough for another person to continue working in it.
A system can get the central legal issue right and still fail almost anywhere along that path.
Once the target is a deliverable, the agent needs more than a prompt. It needs a place to work: the matter files, the available tools, the current state, the actions it is allowed to take, and a place to put the result. It also needs a definition of completion that reaches past surface plausibility.
That is what an environment contributes.
ASSIGNMENT
instructions + record + constraints
↓
ENVIRONMENT
state + tools + permitted actions
↓
TRAJECTORY
what the agent read, changed, checked, and produced
↓
REVIEW
artifact checks + rubric + expert judgmentEach line introduces choices about what the system sees, does, records, and rewards.
The environment has to decide what part of the work to preserve. A closed-universe matter makes the source set clear and the experiment reproducible. It also removes some of what makes practice difficult: incomplete instructions, late-arriving documents, local conventions, permissions, and facts that refuse to stay still.
More realism is not automatically better. If every task becomes an unbounded simulation of a law firm, it becomes expensive to run and difficult to grade. The useful question is narrower: which parts of the assignment have to survive the abstraction for the result to mean anything?
Benchmarks and workplaces answer different questions
LegalBench made “legal reasoning” more specific by collecting expert-built tasks across distinct forms of reasoning. Its authors also describe clear limits: the benchmark excludes long-document work, subjective and ambiguous tasks, most jurisdictions outside U.S. federal law, and multilingual work. Those boundaries define what the benchmark can support.
Pile of Law serves a different purpose. It made a large legal and administrative corpus available for training and research. A corpus supplies material. A benchmark supplies a measurement. Neither, by itself, gives an agent a workplace.
Harvey’s Legal Agent Benchmark is a useful public example of the next step. Its tasks place an agent inside a bounded matter, give it files and tools, require a deliverable, and score the result against granular criteria. Harvey’s published work also shows the tradeoff: synthetic matters improve control and scale, but they are still synthetic and carry imperfections of their own.
Changing the unit from an answer to an assignment changes what has to be built and what failure looks like.
Completion is its own capability
A correct conclusion can still arrive in an incomplete assignment.
Review also has to ask:
- Did the agent cover the required issues?
- Did it use the right sources and apply them to the right proposition?
- Did it preserve the requested format?
- Do the work products agree with one another?
- Can another person inspect what changed and why?
- Did the agent stop with the assignment complete, or merely with an answer in hand?
Some of those questions can be checked mechanically. Some need rubrics. Some need expert review, and experts will not always agree. The evaluation should record that disagreement instead of hiding it inside one score.
The trajectory matters too. A clean final document can conceal a brittle process. The agent may have ignored half the record, fabricated a source, or arrived at the right result by accident. Conversely, a surprising answer may be a defensible alternative rather than an error. Expert review helps distinguish a model failure from a missing tool, a bad instruction, or a task that was never well specified.
That feedback can improve the agent, the environment, the verifier, or the training data. It should not all be flattened into “the model was wrong.”
Not every workflow needs post-training
A strong base model with retrieval, tools, and a good in-context playbook may be enough for some work. Post-training adds cost and creates new ways to overfit the system to its own measurements.
The case for an environment is strongest when an organization needs the behavior to repeat across many matters, wants to test tool use under controlled conditions, or needs improvement to survive a change in the base model. The environment records what the agent is asked to do, what it can use, what it must produce, and how people decide whether the result is usable.
Can the system finish the assignment?