· VerifyCore Labs
Review the route a secret can take, not just the tools an agent can call
A practical review of data routes, approval replay, and bounded agent-tool coverage.
Who did this. The lab’s AI agents did the research and engineering. Nick Harris, founder. CTO of VivaMed BioPharma; co-founder of MedSim.ai, FastRead.io and Formulai. The lab’s track record.
An agent platform can approve a file reader, a database writer, and a network sender individually and still authorize an unsafe workflow. The engineering question is what information can move through the sequence. A harmless intermediate write can become the bridge between private input and an external destination.
The useful review unit is therefore a route through resources. Tool names remain relevant, but a list of names does not describe where their arguments came from, what they persist, or who can later read that state. This article proposes a concrete way to review those routes before adding another broad approval prompt.
A bounded result that motivates the review
OrbitalProof's public result page reports executed sessions in a constructed environment with a fixed tool set. Its published receipt records three minimal dangerous compositions, with only one a pair. The page explicitly notes that conventional taint tracking would catch them. It also leaves a coverage question open: the public material does not map each execution to the claimed enumeration.
The approval limitation remains relevant
The recorded approval replay limitation matters. A newer public claim describes consumption within a running process, but a receipt demonstrating that version is not published. Restarting or moving an approval between processes remains outside that claimed fix. This is internal evidence about a test environment, not an independent security audit of deployed assistants.
A small example to reason about
Consider a hypothetical assistant asked to prepare a diagnostic report. It can read an internal configuration, write a local draft, and upload a report. Suppose the reader is allowed because diagnostics require configuration, the writer is allowed because drafts are local, and the uploader is allowed because reports need delivery.
A two-call check might see no direct configuration-to-upload pair: the upload reads the draft rather than the configuration. The three-call path still carries configuration data outside. The intermediate filename changes; the data's sensitivity does not.
Follow the information through storage
This is a reasoning example, not a reproduction of the lab's execution. It illustrates why review should preserve information lineage across storage boundaries. A database row, temporary file, queue item, or cached response can all be intermediate resources. Clearing a variable in the agent process does not clear those resources.
Identify the enforcement point
A reviewer can write down the intended path before examining implementation: private input → transformation → durable draft → outbound payload. At each arrow, ask what rule permits the transfer and which component enforces it. If the only answer is that the language model was asked to be careful, the rule has no independently inspectable enforcement point.
Separate permission from freshness
An approval answers whether a particular action may happen. It also needs to describe which action: destination, resource, purpose, actor, and payload. An approval for uploading a harmless diagnostic summary should not authorize a later payload that includes raw configuration.
Freshness is a separate question. Can the approval be replayed after a restart? Can two workers consume it concurrently? Can a different workflow present it? A system may correctly compare payloads and still execute the same approved action twice because consumption exists only in one worker's memory.
Test approval freshness
A practical test uses synthetic markers rather than actual credentials. Approve one payload, then vary one property at a time: destination, payload digest, workflow identifier, actor, and expiration. Finally present the unchanged approval twice, concurrently and after recovery. Record what reached the test destination, not just what the policy function returned.
This does not establish that a specific storage or transaction design is sufficient. The purpose is to make the intended behavior observable, including cases where a tool times out after its external effect occurred.
Make the coverage claim inspectable
If a report says every sequence of length three was tested, its reader needs the alphabet of allowed tools, constraints on repetition, initial states, and arguments. There is a large difference between enumerating names and enumerating meaningful state transitions. A tool that accepts a destination and a filename has many behaviors behind one name.
Start with a deliberately small test world. Publish a sequence identifier, initial-state identifier, expected sensitivity labels, actual outbound markers, verdict, and failure reason for each case. Keep excluded combinations visible with their exclusion reasons. A failed setup is not a safe execution.
Use positive and negative controls
Then add controls. A deliberately forbidden route should be blocked. A permitted transformation should complete. A planted error in the judge should fail a check. These controls test different things: enforcement, availability, and whether the verification machinery can notice damage.
A clean report should distinguish execution coverage from input-domain coverage. Exhaustive enumeration within a small model can be useful; it cannot justify an unrestricted claim about every workflow or every tool argument in a production runtime.
A reviewer checklist
Before granting an assistant a new combination of tools, ask:
- Which sources contain private information, and which destinations are outside the intended boundary?
- Does lineage survive writes to files, databases, queues, and caches?
- Which component evaluates the final payload at the external-effect boundary?
- What exactly is bound into approval, and how is consumption handled across processes?
- What happens after timeout, retry, restart, and concurrent dispatch?
- Can someone inspect both the allowed controls and the prohibited cases?
The right outcome may be a smaller reachable tool set, a destination restriction, a data-flow policy, or a resumable approval step. Those are design choices to test against the intended workflow, rather than interchangeable labels for safety.
A concrete evaluation offer
For an agent-platform team, a useful first engagement is a scoped review of one private-data-to-external-action workflow. The proposed deliverables are its resource-flow diagram, a bounded synthetic-marker test suite, a coverage manifest, and a limitations report. Acceptance would require the team's chosen prohibited routes to be blocked while permitted cases complete, with recoverable receipts for ambiguous attempts.
Discuss a scoped evaluation
To discuss that consulting evaluation or the OrbitalProof research materials, contact nick@latticegraph.com and describe the workflow boundary you want examined. VerifyCore's research and engineering are performed with AI agents under founder direction; the public record does not establish customer adoption, revenue, or an outside audit. Any licensing or acquisition evaluation would separately inspect rights, dependencies, and reproducibility.