A seven-model study of governed knowledge in Harvey LAB-derived legal agent handoffs.
When one agent hands work to the next, a summary is not enough.
The receiving agent—or the human reviewing its work—still needs to answer three questions:
- What may I rely on now?
- Why may I rely on it?
- Under whose authority?
Memory can retrieve what happened. A trace can show what the system did. A wiki can organize what people wrote. None of those systems, by themselves, determines which conclusions are current, in scope, supported by evidence, and authorized for reuse.
That is the problem Proofpress is built to address: governed knowledge for agents.
The experiment
We tested whether governed knowledge changes what happens after an agent handoff.
The frozen study covers seven models, three multi-step legal task episodes, and 126 valid paired runs. The tasks were composed by Proofpress from version-pinned public Harvey LAB Contracts materials. They are not official Harvey benchmark tasks or leaderboard scores.
Each paired run held the model, task, evaluator, tools, and operating limits fixed. What changed was the handoff condition:
- Ordinary handoff: the receiving agent received a readable summary of prior work.
- Proofpress handoff: the receiving agent received selected conclusions bound to evidence, scope, version, lifecycle, policy, and review authority.
The study evaluated Claude Opus 4.8, Qwen 3.8 27B, GLM 5.2, GPT-5.6 Sol, Muse Spark 1.1, DeepSeek V4 Flash 0731, and Inkling.
What we observed

Across the complete frozen panel, ordinary handoffs passed 10,654 of 11,928 matched rubric criteria (89.3%). Proofpress handoffs passed 11,141 of 11,928 (93.4%).
That is an observed difference of 487 additional matched criteria, or 4.1 percentage points. The Proofpress condition had higher rubric completion for every model in the panel, although the size of the difference varied by model and task.
We also ran 63 controlled stress pairs containing stale, superseded, unsupported, or out-of-scope state. Ordinary handoffs produced 8 observed unsafe-propagation events. Proofpress handoffs produced 0.
These are descriptive frozen-panel results. They do not establish population-level causality or statistical significance.
Why the handoff changed
The intervention was not simply a longer prompt or a second memory store.
Proofpress treated prior conclusions as candidate knowledge rather than automatically reusable context. Evidence and deterministic integrity checks were evaluated first. Policy could recommend admission, rejection, or escalation. Only the configured review authority could authorize a conclusion for reuse. Rejected, unresolved, expired, superseded, or actor-mismatched conclusions were excluded from the governed context delivered to the next agent.
The separation matters:
- Retrieval finds potentially relevant information.
- Evidence shows where a conclusion came from.
- Policy evaluates whether stated conditions are satisfied.
- Authorized review determines whether the conclusion may be relied on.
The ledger, receipts, and verification checks are mechanisms. The product outcome is governed knowledge that a receiving agent or human can evaluate before continuing the work.
What this result does—and does not—say
The study provides a bounded product-mechanism signal: in these frozen Harvey LAB-derived handoff episodes, governed knowledge coincided with higher rubric completion and no observed unsafe propagation.
It does not establish:
- an official Harvey leaderboard result;
- a general improvement in legal intelligence;
- a population-level causal effect across agent workflows;
- statistical significance;
- that every workflow should adopt the same policy or review design.
The result is deliberately narrower. It shows that the state passed between agents can be treated as a governed object—and that this treatment was associated with materially different downstream behavior in the frozen panel.
Read the technical record
- Formal technical report
- Public results and claim boundaries
- Study design
- Content-addressed final receipt manifest
- Per-model and per-task results
The result was frozen on August 26, 2026. Additional models or scenarios require a new preregistered study rather than modifying this result after outcomes are known.
The broader question
Agents are becoming better at remembering, searching, and acting. The next infrastructure problem is deciding what should survive from one execution context to another.
For every important handoff, the receiving agent or human should be able to ask:
What may I rely on now? Why, and under whose authority?
Proofpress is building the governed knowledge layer that makes that question answerable.
