OrbitalProof · AI agents
Tool combinations that leak a secret
A run the lab reports as exhaustive, in a test world the lab built, over every sequence of up to three AI-assistant tool calls, finding the smallest combinations that leak a stored secret.
Who did this. The lab’s AI agents did the research and engineering. Nick Harris, founder. CTO of VivaMed BioPharma; co-founder of MedSim.ai, FastRead.io and Formulai. The lab’s track record.
What we showed
covered, by the lab’s account, every sequence of up to three AI-assistant tool calls in the lab’s test world. Three smallest combinations leak the stored secret and only one of them is a pair, so approving tools two at a time misses the other two. This site does not show how the number of runs maps onto the number of sequences.
- Limit.
- One test world the lab built. A standard technique, tracking where the secret’s data flows, would also catch all three. The run-time check that enforces the result remembers an approval was used only while the program runs; after a restart it counts as unused.
The problem
A platform that approves an assistant’s tools one at a time, or two at a time, can let a combination of three tools leak a stored secret. Ordinary tracking of where a secret’s data flows (taint analysis) already closes that gap; what this adds is a run the lab reports as executing every combination of up to three calls (this site does not show how its runs map onto those combinations).
What it means for a buyer
If you decide which tools an AI assistant may combine: approving tools one or two at a time can miss a combination of three that leaks a stored secret. This is an executed record, which the lab reports as exhaustive, of which combinations are dangerous in one test world, run on real files and a real database.
Who we expect would buy
Teams we expect would care (no customer or pilot yet): AI agent-platform and runtime-security teams that decide which tool combinations an assistant may use.
Why now
Two published write-ups describe dangerous combinations of AI-assistant tools: Simon Willison’s “lethal trifecta” (2025-06-16), which says an agent that combines private data, untrusted content and outside communication can be tricked into sending the private data to an attacker, and Invariant Labs’ analysis of “toxic flows” in agent systems (2025-07-29).
Why you can trust the check
The lab’s check re-runs the analysis, and it failed as it should on a deliberately broken input. The tools, the test world, the secret and the judge are all the lab’s own, so a buyer would want the same run on their own assistant’s tools.
No outside firm has audited it. How this result’s check works, step by step.
What this does not show yet
- It is one constructed test world: ten tools, one stored secret, outlets the lab designed, and a verdict that depends on three of the five trust assumptions the lab declared, meaning assumptions about what the test may take as trusted. This site does not list the five.
- A standard technique, tracking where the secret’s data flows (taint analysis), would also catch all three combinations. What this adds is an executed record of which combinations leak.
- The run-time check that enforces the result lets an approval be used only once, but it remembers that only while the program runs. In the lab’s test, one approval used for three calls of the network-send tool let exactly one copy of the secret reach the outlet, so the once-only rule held while the program ran. But nothing stores that the approval was used, so an approval loaded again from disk, or shown to another process, counts as unused and works again.
- The result is about the lab’s ten tools; it is not shown for the tools of real assistants.
Prior work
Named in the lab’s prior-art search for this result, and credited here.
- The lethal trifecta for AI agents: private data, untrusted content, and external communication, Simon Willison, simonwillison.net, 2025-06-16
- Invariant Labs Exposes Novel Prompt Injection Attack Vulnerabilities, "Toxic Flows," in Agentic Systems & MCP Servers, Luca Beurer-Kellner, Marco Milanta, Marc Fischer, Invariant Labs blog, 2025-07-29
- On lightweight mobile phone application certification, William Enck, Machigar Ongtang, Patrick McDaniel, ACM CCS 2009 (Penn State research portal record)
The exact wording, for a technical reader
The lab’s own sentences and figures for this result, word for word, its limits in full: Tool combinations that leak a secret, exact wording.
Check it yourself
- This result’s file: every sentence and figure on this page that is the lab’s own, copied from its current record at the commit the file names.
- The lab’s result file, copied from its codebase at the commit it names.
- OrbitalProof, the company that carries this result.
Related results across the group
- This result on orbitalproof.com, OrbitalProof’s own site.
- Crash recovery for AI assistants’ multi-step changes (OrbitalProof, ai agents)
- An Ultra Ethernet recovery design, proved in a model (OrbitalProof, ai-cluster networking)
- An automated attack search against quantum-safe Wi-Fi sign-in (OrbitalProof, wi-fi security)
- A memory ceiling for post-quantum Wi-Fi, proved in a model (OrbitalProof, wi-fi security)