{
 "_doc": "One result as verifycorelabs.com prints it: the headline, the problem and the buyer cut (never reworded) from the lab's plain-English account of the result, the one figure a buyer is shown and its limit, and the lab's current claim and limits verbatim, copied at the commit named in source. A field is left out, never reworded, when it uses a word this site keeps for its own process or names a file this site does not serve (not_reprinted).",
 "slug": "txn-crash-recovery",
 "company": "OrbitalProof",
 "order": 1,
 "rank": 4,
 "served_receipt": "wireless-reliability/results/txn-crash-recovery.json",
 "lab_file_older": true,
 "headline": [
  "When an AI assistant carries out a multi-step change to files and one database table, the lab's transaction layer is meant to either complete all of it or undo all of it — but only for plans it admits in advance"
 ],
 "qualification": [
  "which may contain at most one action that cannot be undone (such as sending a message), placed last and carrying a replay key, other plans being refused or left for a human operator",
  "killed at 29 moments the authors chose, the program each time recovered a clean all-or-nothing state from its on-disk log in a fresh process, while a deliberately broken version did not",
  "but killed 2,893 more times at random moments it recovered 31 times to a state that was neither, some of them undoing work already committed"
 ],
 "why": [
  "An assistant whose program dies halfway through a multi-step change can leave files and database rows half-updated, with nothing to finish or undo the change afterwards"
 ],
 "buyer": [
  "AI agent-platform teams that let assistants make changes to files and databases"
 ],
 "figure": "atomic recovery recorded at every one of those 29 points",
 "figure_from": "claim",
 "limit": "The 2,893 random kills are process kills, not power cuts, so the fsync ordering the journal relies on is still untested",
 "limit_from": "limits",
 "plain": "atomic recovery recorded at every one of those 29 points",
 "limfig": "2,893 further kills at seeded random moments found **31 recoveries that were not atomic**",
 "stronger": "2,893 further kills at seeded random moments found **31 recoveries that were not atomic**",
 "quotes": {
  "declared": "atomic recovery recorded at every one of those 29 points and a negative control that is genuinely non-atomic",
  "novelty": "what may be new is only its application, with admission rules, to an assistant's tool calls",
  "priorart": "So this is INCREMENTAL, a real test of one more system, done to less than the published standard"
 },
 "claim": "Twenty-nine real processes killed with `os._exit(9)` at declared interruption points, each recovered in a **fresh interpreter** from the on-disk journal, with atomic recovery recorded at every one of those 29 points and a negative control that is genuinely non-atomic; 2,893 further kills at seeded random moments found **31 recoveries that were not atomic**, some undoing committed work.",
 "limits": "The 2,893 random kills are process kills, not power cuts, so the fsync ordering the journal relies on is still untested; they sample the interleaving space (85 distinct crash states) rather than exhaust it; and the three windows they found are recorded, not fixed.",
 "not_reprinted": [],
 "prior_work": [
  {
   "title": "How SQLite Is Tested (section Crash Testing)",
   "by": "SQLite developers, sqlite.org",
   "url": "https://www.sqlite.org/testing.html"
  },
  {
   "title": "All File Systems Are Not Created Equal: On the Complexity of Crafting Crash-Consistent Applications",
   "by": "Thanumalayan Sankaranarayana Pillai, Vijay Chidambaram, Ramnatthan Alagappan, Samer Al-Kiswany, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau, OSDI 2014",
   "url": "https://usenix.org/node/186195"
  },
  {
   "title": "SagaLLM: Context Management, Validation, and Transaction Guarantees for Multi-Agent LLM Planning",
   "by": "Edward Y. Chang, Longling Geng, arXiv 2503.11951, first submitted 2025-03-15 (PVLDB 18)",
   "url": "https://arxiv.org/abs/2503.11951"
  },
  {
   "title": "Cordon: Semantic Transactions for Tool-Using LLM Agents",
   "by": "Zheng Chen, Hanqing Liu, Duling Xu, Dong Dong, Jialin Li, Bangzheng Pu, Jidong Zhai, arXiv 2606.17573, submitted 2026-06-16",
   "url": "https://arxiv.org/abs/2606.17573"
  }
 ],
 "source": {
  "claims_file": "registry/CLAIMS_CURRENT.jsonl",
  "claims_sha256": "81776df15a1498abec7228ea179c952d01c54e031ffd4144498a30660963074b",
  "ranking_file": "views/VALUE_RANKING.md",
  "ranking_sha256": "2e980e26178b34dc49ec1d365f966be292f57d7c636143fc62c12feab72ae470",
  "dossier_sha256": "43b5336b3d004c109164221a25838394c400057618f44dfe4ee0f0423f0ec9dc",
  "commit": "25d755d2f06518b632c3658ae83f20312daac53e"
 }
}
