Most enterprise AI infrastructure is built around a familiar problem: how do we give agents better access to what the organization already knows?

We add retrieval. Memory. Ontologies. Knowledge graphs. Better context.

All of that matters. But it addresses only one side of the system.

Agents do not just consume enterprise knowledge. They create a new knowledge layer.

They research, compare sources, synthesize findings, make decisions, and pass conclusions to the next agent or human. Once that work is reused, the agent's output has become part of the organization's operating knowledge.

That creates a different question:

What may the next agent or human rely on, why, and under whose authority?

Two kinds of knowledge

Enterprise knowledge is what agents reason from: documents, databases, policies, domain ontology, and memory.

Agent-produced knowledge is what their work produces: conclusions, claims, findings, analyses, and decisions.

Existing knowledge infrastructure organizes the first. Proofpress governs the second.

This distinction matters because a conclusion is not safe to reuse merely because it is retrievable. A future agent also needs to know:

  • What evidence supports it?
  • Which version of the source was used?
  • What scope and assumptions apply?
  • Was it verified?
  • Who was authorized to admit it?
  • Does it depend on another conclusion that has since changed?
  • Has it been contradicted, expired, or superseded?
Illustration from the original article.
Illustration from the original article. View source

A memory system can retrieve a conclusion. An observability system can show how an agent arrived there. A knowledge graph can represent relationships. None of those facts alone decides whether this conclusion is currently eligible for a specific actor to rely on.

Why this becomes urgent

Illustration from the original article.
Illustration from the original article. View source

Enterprise knowledge often grows relatively steadily. Agent-produced work can grow much faster as adoption, autonomy, branching research, and agent-to-agent handoffs increase.

That is not a claim that every organization follows a mathematically exponential curve. It is a practical systems observation: once agents generate more conclusions than people can review informally, verification and authority stop being optional workflow polish. They become infrastructure.

The risk is not only a wrong answer. It is a conclusion that was once reasonable but has lost its evidence, scope, review state, or authority as it travels.

Output becomes input. Then input becomes output again. Small ambiguities become organizational memory.

The product object: a Governed Claim Graph

Proofpress turns selected agent work into a governed claim graph.

Illustration from the original article.
Illustration from the original article. View source

The core objects are conclusions, claims, evidence, provenance, verification, authority, scope, dependencies, and supersession. A conclusion may depend on several claims. A claim is supported by evidence. Authority scopes who may rely on the result. A later conclusion can supersede an earlier one without erasing its history.

Agent work crosses three separate gates:

  1. Deterministic Checks — Enforce fixed rules and required evidence; return pass or fail.
  2. LM Judge — Evaluate the meaning, then recommend or escalate; it cannot authorize reuse.
  3. Human Approval — An authorized human admits or rejects the conclusion; only this gate enables downstream reuse.
Illustration from the original article.
Illustration from the original article. View source

Deterministic Checks and LM Judge cannot authorize reuse. Human Approval can.

Admission does not declare universal truth. It creates an inspectable, scope-bound answer to whether a future agent or human may rely on a conclusion now.

Proof, not product definition

Illustration from the original article.
Illustration from the original article. View source

We tested this mechanism in a frozen panel of 7 models, 3 Harvey LAB-derived legal task families, and 126 valid paired runs.

Within that bounded panel, Proofpress-governed handoffs raised rubric completion from 89.3% to 93.4% and reduced observed unsafe propagation from 8 to 0 across 63 controlled stress pairs.

Those results are deliberately narrow. They are not an official Harvey leaderboard score, a population-level causal claim, or evidence that the models became more legally intelligent. They are evidence that governing what crosses the handoff boundary can change downstream behavior under controlled conditions.

The study supports the thesis. It does not define the category.

What Proofpress is building

Today, Proofpress has a local ledger and CLI, a local review and context UI, agent adapters, artifact provenance, and portable Markdown and static-HTML carriers. Public API/SDK and MCP surfaces are planned, not shipped.

The longer-term idea is simple:

Existing knowledge infrastructure organizes what agents reason from. Proofpress governs what their reasoning produces.

If your agents are already producing research, decisions, or analyses that other agents and humans reuse, I would like to hear where the handoff breaks.

Read the full thesis and technical repository: https://github.com/chenmingtang830/proofpress