OrbitalProof · AI agents
Crash recovery for AI assistants’ multi-step changes
A transaction layer meant to finish or undo an AI assistant’s multi-step change after a crash, for plans it accepts before they run. At every crash point the lab chose, it recovered the change whole or not at all; killed at random moments, some recoveries were not, and the lab has not fixed that yet.
Who did this. The lab’s AI agents did the research and engineering. Nick Harris, founder. CTO of VivaMed BioPharma; co-founder of MedSim.ai, FastRead.io and Formulai. The lab’s track record.
What we showed
The lab killed the program mid-change at each of them, and each time a fresh program recovered the change whole or not at all from the log on disk.
- Limit.
- 2,893 further kills at seeded random moments found 31 recoveries that were not atomic, some undoing work already committed (recorded in the lab’s claim; the run record is not published). The lab has not fixed them yet. Power cuts are untested.
The problem
An assistant whose program dies halfway through a multi-step change can leave files and database rows half-updated, with nothing to finish or undo the change afterwards. Databases solve this with a standard design, a log written to disk before the changes; the lab applies it to an assistant’s tool calls.
What it means for a buyer
If your assistants change files and databases, a change that dies halfway can leave them half-updated, with nothing to finish or undo it. This layer writes each change to a log on disk first and recovers from it after a crash, for plans it accepts before they run. At the moments the lab chose, recovery was all-or-nothing every time. At random moments it was not always, and the lab publishes their count (the run record is not published), which is what a buyer needs to know before relying on it.
Who we expect would buy
Teams we expect would care (no customer or pilot yet): AI agent-platform teams that let assistants make changes to files and databases.
Why now
Other researchers have published transaction layers for AI agents’ tool calls: SagaLLM (arXiv 2503.11951, first submitted 2025-03-15) and Cordon (arXiv 2606.17573, submitted 2026-06-16). That published work shows the problem is being worked on now.
Why you can trust the check
The lab showed that its crash check can fail: when a separate program changed a number the check reads, the check caught it. The test world, the scenarios and the crash points are the lab’s own, and no outside party has graded the layer, so a buyer would want to run it against their own tools.
No outside firm has audited it. How this result’s check works, step by step.
What this does not show yet
- It covers only plans the layer accepts before they run. A plan is accepted only if it has at most one action that cannot be undone (such as sending a message), that action comes last, and it carries a replay key (a label that makes repeating the action harmless). Other plans are refused or left for a human operator.
- At the moments the lab chose, every recovery was all-or-nothing. Killed at random moments instead, some recoveries were not, and some undid work that had already been committed. This is recorded in the lab’s claim; the run record is not published. The lab has not fixed these failures yet.
- The tests killed the program; they did not cut the power, so whether the log’s writes reach the disk in the right order is untested.
- The random kills sample the ways a crash can happen (they reached 85 distinct crash states) rather than covering them all.
- The world tested is small: a set of files and one database table, with scenarios and a grader the lab wrote.
- The method is standard in databases, and other researchers have published transactions for AI agents’ tool calls already: SagaLLM and Cordon.
Prior work
Named in the lab’s prior-art search for this result, and credited here.
What is new, by the lab’s own account: “what may be new is only its application, with admission rules, to an assistant's tool calls”. The lab’s own prior-art search, in its words, is on the exact-wording page.
- How SQLite Is Tested (section Crash Testing), SQLite developers, sqlite.org
- All File Systems Are Not Created Equal: On the Complexity of Crafting Crash-Consistent Applications, Thanumalayan Sankaranarayana Pillai, Vijay Chidambaram, Ramnatthan Alagappan, Samer Al-Kiswany, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau, OSDI 2014
- SagaLLM: Context Management, Validation, and Transaction Guarantees for Multi-Agent LLM Planning, Edward Y. Chang, Longling Geng, arXiv 2503.11951, first submitted 2025-03-15 (PVLDB 18)
- Cordon: Semantic Transactions for Tool-Using LLM Agents, Zheng Chen, Hanqing Liu, Duling Xu, Dong Dong, Jialin Li, Bangzheng Pu, Jidong Zhai, arXiv 2606.17573, submitted 2026-06-16
The exact wording, for a technical reader
The lab’s own sentences and figures for this result, word for word, its limits in full: Crash recovery for AI assistants’ multi-step changes, exact wording.
Check it yourself
- This result’s file: every sentence and figure on this page that is the lab’s own, copied from its current record at the commit the file names.
- The lab’s earlier result file: older than the current record above.
- OrbitalProof, the company that carries this result.
Related results across the group
- This result on orbitalproof.com, OrbitalProof’s own site.
- An Ultra Ethernet recovery design, proved in a model (OrbitalProof, ai-cluster networking)
- Tool combinations that leak a secret (OrbitalProof, ai agents)
- An automated attack search against quantum-safe Wi-Fi sign-in (OrbitalProof, wi-fi security)
- A memory ceiling for post-quantum Wi-Fi, proved in a model (OrbitalProof, wi-fi security)