AxiomLimit · AI-cluster networking
Fixed-size memory for reordering AI network data
In the lab’s own network simulator, a network-card design that keeps track of out-of-order data in a small, fixed amount of memory, compared on the same simulated grid with the lab’s model of STrack, a published fixed-size design, and with a second bounded-state design in the lab’s simulator.
Who did this. The lab’s AI agents did the research and engineering. Nick Harris, founder. CTO of VivaMed BioPharma; co-founder of MedSim.ai, FastRead.io and Formulai. The lab’s track record.
What we showed
The lab ran its own network simulator over a grid of 64 simulated cases. In it, the lab’s model of STrack, a published fixed-size design by Le, Pan and Newman, needs this many times the reorder memory that the lab’s design does, on average (a geometric mean). The lab’s own check confirms the figure only to two decimal places, so read it as about twice. In bytes the gap is often small, and in some cases STrack needs less.
- Limit.
- Against a second bounded-state design in the lab’s simulator, which the lab calls CTS, the lab’s record gives 1.443x, CTS better in 31 of 64 simulated cases, a ratio of CTS’s memory to the lab design’s whose averaging it does not state; gaps are often tens of bytes, so the result is mixed; STrack needs less in some cells. A network simulation, not silicon. Much larger savings in the lab’s files compare against cards that buffer all out-of-order data, which is not the fair comparison.
The problem
AI clusters spread traffic over many network paths, so data reaches the receiving network card out of order. A card that holds it in a buffer until it can be put back in order needs on-chip memory that grows with the data in flight.
What it means for a buyer
If you build network cards or switch chips for AI clusters: traffic spread over many paths arrives out of order, and a card that buffers it needs memory that grows with the data in flight. This design keeps a small note of fixed size instead. Measured against the lab’s model of a published fixed-size design (STrack) on the same simulated grid it uses less memory on average, but often by only tens of bytes.
Who we expect would buy
Teams we expect would care (no customer or pilot yet): makers of network-interface cards and switch chips for AI data-centre Ethernet, and large cloud operators that design their own cluster networks.
Why now
The Ultra Ethernet Consortium published version 1.0.2 of its specification, dated 2026-01-28. Its architecture paper (Hoefler and others, arXiv 2508.08906, 2025-08-12) says that unordered transport modes may deliver data to memory out of order. STrack, a published transport for AI clusters (arXiv 2407.15266, 2024), also allows out-of-order delivery over many paths.
Why you can trust the check
Every number comes from the lab’s own network simulator, and nobody outside the lab has checked it. The lab’s check recomputes the larger headline ratio from the saved simulation results (the figure against STrack is checked to two decimal places), so it would catch a mis-copied figure but not a wrong simulator; a buyer would want to re-run the simulation, or measure on hardware.
No outside firm has audited it. How this result’s check works, step by step.
What this does not show yet
- It is a network simulation, not silicon, and every number comes from the lab’s own simulator, including its model of each design it is compared with.
- The fixed-size designs it is measured against are the lab’s model of STrack, a published design by Le, Pan, Newman and others, and a second bounded-state design in the lab’s simulator: the lab’s model of the receiver-driven clear-to-send (CTS) admission Meta describes (Gangidi and Zeng, SIGCOMM 2024, section 5.2.2). No public Meta source states that design’s reorder memory. Against both, the memory saving is small, often tens of bytes, and in some cells they hold less memory: the second design nearly half the time.
- The lab’s files also show much larger savings. Those compare against cards that buffer all out-of-order data, which is not the fair comparison, or are separate measurements that must not be added together. They are in the exact wording.
- STrack was added to the lab’s adversarial test runs against the design only late in the work; the details are on the evidence page.
Prior work
Named in the lab’s prior-art search for this result, and credited here.
- RFC 5041: Direct Data Placement over Reliable Transports, H. Shah, J. Pinkerton, R. Recio, P. Culley, IETF RFC, 2007
- Multi-Path Transport for RDMA in Datacenters (MP-RDMA), Yuanwei Lu, Guo Chen, Bojie Li, Kun Tan, Yongqiang Xiong, Peng Cheng, Jiansong Zhang, Enhong Chen, Thomas Moscibroda, USENIX NSDI, 2018
- Revisiting Network Support for RDMA (IRN), Radhika Mittal, Alexander Shpiner, Aurojit Panda, Eitan Zahavi, Arvind Krishnamurthy, Sylvia Ratnasamy, Scott Shenker, arXiv (its Comments field: "Extended version of the paper appearing in ACM SIGCOMM 2018"), 2018
- STrack: A Reliable Multipath Transport for AI/ML Clusters, Yanfang Le, Rong Pan, Peter Newman, Jeremias Blendin, Abdul Kabbani, Vipin Jain, Raghava Sivaramu, Francis Matus, arXiv, 2024
The exact wording, for a technical reader
The lab’s own sentences and figures for this result, word for word, its limits in plain words where the lab’s text cannot be reprinted: Fixed-size memory for reordering AI network data, exact wording.
Check it yourself
- This result’s file: every sentence and figure on this page that is the lab’s own, copied from its current record at the commit the file names.
- The lab’s result file, copied from its codebase at the commit it names.
- AxiomLimit, the company that carries this result.
Related results across the group
- This result on axiomlimit.com, AxiomLimit’s own site.
- Where deletion receipts can be fooled (AxiomLimit, ai inference)
- The least bookkeeping a shared AI cache needs, in a model (AxiomLimit, ai inference)