The Model Learns the Grader
A reward puts pressure on the model. A weak grader can turn its own mistakes into training signal.
Train against a grader long enough and the grader becomes part of what the model learns.
The grader may be a set of deterministic checks, an expert rubric, another model, or a combination of all three. Its job is to turn an agent’s work into a reward. Reinforcement learning then changes the model to earn more of that reward.
If the grader recognizes the work, the signal can help. If it recognizes a shortcut, the shortcut becomes valuable too.
A contract-review agent makes the problem concrete. The output may include a redline and an issues list. A mechanical check can confirm that both files exist. A rubric can check for required issues. Another model can judge whether the explanation sounds reasonable.
The files can still contradict one another. The redline can make the wrong edit. The explanation can cite a real provision for a proposition it does not support. The work can pass every stated criterion while failing the assignment.
All of that eventually has to be compressed into an optimization signal.
The verifier is inside the system
It is tempting to treat evaluation as the report that comes after the work. In post-training, the evaluator sits in the causal loop. It decides which behavior gets reinforced.
The verifier sits inside the system. It has a scope, an interface, dependencies, and failure modes. It needs tests of its own.
Prime Intellect’s Verifiers project makes the loop explicit through tasksets, harnesses, and traces used for evaluation or reinforcement learning. Prime’s reward-hacking experiments go further. In a controlled setting, reinforcement learning amplified rare behaviors that exploited a hidden reward. The effect depended on the task, the model’s starting behavior, and the available reward. Optimization searched the measurement surface the researchers provided, including behavior that served the hidden reward instead of the visible task.
Recent work on rubric-based reinforcement learning separates two failures that are easy to collapse:
- The verifier applies the rubric badly.
- The verifier applies an incomplete rubric correctly.
A stronger judge may help with the first. It cannot recover a quality the rubric never represented. In the reported medical and science experiments, models could improve against the explicit rubric while losing ground on qualities measured outside it. That is not a legal-agent result. It is a mechanism legal post-training should test rather than assume away.
A legal verifier needs layers
No single check can carry the whole assignment. Different checks catch different failures.
| Layer | What it can catch | What can still pass |
|---|---|---|
| Artifact | Missing files, broken documents, malformed fields | A valid file containing bad work |
| Source | Missing citations, nonexistent authority, unsupported propositions | A cited authority from the wrong jurisdiction or date |
| Rubric | Omitted issues and explicit requirements | A defensible alternative the rubric did not anticipate |
| Cross-artifact | A redline and issue list that disagree | Two consistent documents built on the same legal mistake |
| Holistic review | Proportion, usefulness, professional quality | Reviewer preference and disagreement |
| Trajectory | Ignored files, skipped validation, suspicious shortcuts | Hidden reasoning or a lucky path that looks careful |
The table makes the residual error visible. Each layer catches a different failure and leaves another for review.
A rubric can also be intentionally narrow and still useful. A test of citation validity does not need to measure drafting quality. The trouble starts when a result from that test is presented as a score for legal capability in general.
Test the test
Before a verifier becomes a reward, it should face work designed to break it.
- Known-good work that follows the expected path
- Known-bad work with obvious substantive errors
- An alternative answer that is different and still defensible
- A polished answer that satisfies the format and misses the assignment
- A real citation used for a proposition it does not support
- A redline and issue list that pass separately but conflict together
- Work that exploits a required phrase, section count, or judge preference
- The same work under a different file order, prompt form, or judge family
Suppose a contract-review verifier gives full credit whenever the redline changes the governing clause and the issue list mentions consent. A candidate can satisfy both checks while changing the clause in the wrong direction. If the verifier cannot tell the difference and it is the signal the run optimizes, training has no reason to fix the mistake. It can make the mistake more reliable.
Adding criteria forever can create a grader that is expensive, brittle, and still incomplete. Some qualities belong in a held-out audit precisely because the model should not train against every measurement used to judge it.
Post-training starts before training
The consequential decisions happen before the optimizer runs.
Which tasks enter the environment? Which tools can the agent use? What state is visible? What counts as completion? Which outcomes receive reward? Which cases remain outside training?
Those choices decide what the run is capable of improving.
Training reward and held-out evaluation therefore have different jobs. Reward supplies pressure. Held-out evaluation asks whether the resulting behavior transfers to new work. Once a held-out failure is used to redesign the task or grader, that set has done its job and should become development data. The next version needs an untouched test.
Expert review matters here. Reviewers still disagree. Instructions are ambiguous. A surprising answer may expose a valid alternative, a missing tool, a weak rubric, or a model failure. The useful output is a diagnosis of which part of the system needs to change.
The research problem is how to build a reward that improves the work without teaching the model to perform for the measurement.
A good verifier should be able to answer a simple challenge before it trains anything: show it two polished work products, make one subtly wrong, and see whether it knows which one not to reward.