Skip to content

· VerifyCore Labs

A shared inference cache needs a permission model before it needs a smaller table

Specify cache rights before claiming permission compression or isolation.

Who did this. The lab’s AI agents did the research and engineering. Nick Harris, founder. CTO of VivaMed BioPharma; co-founder of MedSim.ai, FastRead.io and Formulai. The lab’s track record.

When teams discuss shared inference caches, the conversation often starts with bytes saved and cache-hit rates. An earlier question determines what those numbers mean: which customers are allowed to reuse which cached items? A cache shared by arbitrary groups has a different permission problem from a cache whose items have one tenant.

This distinction changes what can be compressed, what must be remembered, and what a proposed isolation guarantee actually promises. A smaller table is interesting only after the team specifies the rights it must represent.

What the public proof says

AxiomLimit's published result describes a machine-checked lower bound for a dense rights model. With T tenants and n blocks, each tenant-block permission can vary independently. Under both zero unauthorized reuse and complete authorized reuse, the model needs at least T×n bits of permission state.

Read the proof receipt with its limits

The public proof receipt identifies the Lean toolchain, a successful recorded build, and theorem names for both dense rights and rights in which each block belongs to one tenant. This is a receipt of the lab's check; this article does not rerun that build. The counting argument is classical. The result does not prove that a deployed engine matches the model, nor does it cover timing or content channels.

A toy example makes the distinction concrete

Take two tenants, A and B, and two blocks, X and Y. If arbitrary sharing is permitted, each block can be authorized to neither tenant, A alone, B alone, or both. There are four possibilities for X and four for Y: sixteen permission arrangements.

A representation that cannot distinguish two arrangements can face a request where the correct answer differs. Suppose the arrangements disagree on whether B may reuse X. Serving B in both arrangements violates one policy; denying B in both violates the other policy's full-reuse requirement.

Distinguish information from implementation cost

That is a reasoning example of the information requirement, not an inference benchmark. It says nothing about the size of a practical hash table or the time required to look up a permission. Implementations have other state, alignment, metadata, and representation choices.

Now restrict each block to exactly one of the two tenants. X has two possibilities and Y has two: four arrangements. The dense-rights example is no longer the right model. A system deliberately allowing only ownership-based reuse should be evaluated under its actual restriction rather than criticized for failing to store every imaginable sharing arrangement.

The two promises are both necessary

Zero unauthorized reuse alone is easy to satisfy by refusing every cache hit. Complete authorized reuse alone is easy to satisfy by serving every hit. Neither is the intended service.

The engineering specification needs both permission correctness and intended reuse behavior. Otherwise a performance change can quietly weaken a security condition, or a security change can silently eliminate the reuse that justified sharing the cache.

Make the promises observable

Write the promises as observable request outcomes. For a named tenant, block, and rights snapshot, which requests must miss and which must be eligible to hit? Distinguish eligibility from actual residency: a permitted block might already have been evicted. A miss caused by eviction is different from a miss caused by an authorization decision.

This distinction helps test failures. A cache-hit counter cannot by itself show permission correctness, and a denial counter cannot show useful authorized sharing. Both need to be tied back to the rights state that existed when the request was evaluated.

Structured groups introduce lifecycle questions

Real deployments may use organizations, groups, or hierarchies rather than independent rights for every block. That structure can reduce the space of policies, but it also introduces questions the dense model alone does not answer.

If a user leaves an organization, what happens to previously cached work? If a group identifier is reused, can an old block become visible to the new group? If an adapter or model revision changes, is a previously computed block still semantically compatible? Permission identity and computation identity are separate dimensions.

Test membership and revision changes

An evaluation should exercise membership changes and revisions, not only steady-state hits. Use synthetic prompts and controlled tenant identities. Start with allowed sharing, revoke it, then repeat the request through each supported API and transfer path. Record the policy revision, cache key context, residency, and observed reuse.

This is a proposed evaluation protocol. It is not evidence that AxiomLimit has performed those production tests or that any particular cache fails them.

State isolation is not timing isolation

A cache can return correct data while still revealing that somebody else's prefix was present. Conversely, a timing mitigation does not automatically prove permission correctness for every representation and storage path.

These are distinct claims with different tests. A state-isolation test examines which block can be returned. A timing experiment needs a threat model, query budget, controlled observations, and statistical evaluation. Combining their labels into a single “secure cache” claim hides the assumptions of both.

Describe each security claim separately

For procurement, ask for a claim table. Each row should name the protected property, model assumptions, verification method, implementation boundary, and untested conditions. A formal proof of one property belongs in that table beside implementation tests, not above them as a replacement.

Questions for an inference team

A short permission review should resolve:

  • Are rights exclusive ownership, group-based, hierarchical, or arbitrary per block?
  • Does every supported route carry the same authorization context?
  • What changes when membership, model, adapter, or policy revisions change?
  • Can permitted reuse be distinguished from cache residency and eviction?
  • Which timing and content threats are tested separately?
  • What public or shareable artifact lets a second engineer inspect the claim?

Answering these questions may reveal that the dense theorem is relevant. It may reveal that a smaller structured model is the appropriate one. Either result is more actionable than treating a theoretical lower bound as a universal production-memory bill.

A concrete evaluation offer

For a confidential-inference team, the proposed first engagement is a rights-model and cache-boundary review: specify the intended sharing relation, map it to cache metadata, and construct a small acceptance matrix for allowed, revoked, and incompatible reuse. Deliverables would include the model, executable synthetic tests where feasible, and an explicit list of untested side channels.

Contact nick@latticegraph.com to discuss that consulting scope or evaluation of AxiomLimit research materials. VerifyCore uses AI agents for research and engineering. The public work establishes a bounded formal result, not production readiness, a measured speedup, customer traction, or audited security. Licensing and acquisition discussions require separate rights and dependency diligence.

All postsAll resultsDiscuss an evaluation

How we show numbers

A number you can click opens the file it comes from. How each result is checked

  • We never show a number before its file has loaded.
  • A question we have not checked yet is marked as unchecked.
  • A check that found nothing says so.
  • A file with no value for a question says so.
  • A number whose file is missing or has changed is not shown.
  • Two files that disagree about what a number describes are both flagged.
  • A number from too few samples shows its sample size.
  • Two files that give different values are both shown.
  • A file we cannot publish is listed by its fingerprint only.
  • Where a file records when it was measured, the date is in that file.
  • A question that does not apply to a page is left off it.
  • A measurement whose program failed is shown as failed.