Five Approaches Making LLM Inference Cryptographically Verifiable

Five Approaches Making LLM Inference Cryptographically Verifiable

When a language model API returns tokens, those tokens do not identify the computation that produced them. A provider could run cheaper weights, apply heavier quantization than advertised, or return cached responses, and a customer using standard APIs would have no mechanism to detect any of it. Sampled Layerwise Proofs, submitted September 23, 2026, makes the scale stakes concrete: Youki Lim and Sam Yong at TrueOpen sealed a complete Llama-2-70B inference trace on a 2 TB CPU host, proved five of 163 chunks, produced a 4.34 MiB proof in 1,259 seconds, and verified it in 46.3 seconds without the weights. That is not production-ready serving. It is the first measured demonstration that cryptographic verification of a 70-billion-parameter inference trace is physically possible.

That result joins four other systems, each approaching the same trust problem from a different direction: hardware enclaves, delegation of linear algebra, tensor-native commitments, and layerwise zero-knowledge circuits. Together they map a space whose perimeter is becoming clearer even if the center is not yet reached. This article examines what each system proves, what it does not prove, and what the gap between the two tells engineering teams evaluating outsourced AI.

Why This Comparison Is Timely

The problem is not hypothetical. NanoZK’s introduction cites empirical audits that found discrepancies between advertised and deployed model capabilities. The paper notes that users paying for frontier models, at enterprise contracts reaching millions annually, receive outputs through opaque APIs with no verification mechanism. TensorCommitments frames the same concern from a multi-agent angle: when LLMs execute tool calls and coordinate workflows, a single silently corrupted inference can propagate through downstream systems acting on those tokens as ground truth. As LLM usage scales to persistent, tool-using, agentic systems, the gap between the impact of a faulty inference and the ability to cheaply verify it is widening.

The September 23 SLP paper is significant because it demonstrates end-to-end verification at 70B scale for the first time. All prior published demonstrations of cryptographic LLM inference verification operated at GPT-2 scale or below. The 70B result is slow and covers only sampled chunks, but it shifts the discussion from “theoretically possible” to “demonstrably achievable at cost.”

The Shift: Provider Assertion to Independently Checkable Execution Evidence

Current cloud inference works entirely on provider assertion. The customer sends a prompt, the API returns tokens, and the bill arrives. The provider controls the weights, the hardware, the precision settings, and the runtime. There is no mechanism for the customer to verify any of those parameters independently.

The five systems in this comparison replace that with an additional artifact: execution evidence that an independent verifier can check without re-running the inference and without trusting the provider. The evidence varies in strength, cost, and what exactly it covers. VeriAttn delegates computation to a GPU that cannot be trusted but places a small trusted hardware enclave in the verification path. Maverick uses an information-theoretically sound mathematical delegation protocol with no hardware trust assumption. TensorCommitments binds inference to a cryptographic tag. NanoZK generates a formal zero-knowledge proof per layer. SLP commits to the full trace then lets the verifier sample which parts to audit. Each of these is a different answer to the question of what “verifiable” means.

VeriAttn — TEE as Root of Trust, GPU as Untrusted Compute

VeriAttn, published June 15, 2026, takes the approach of hardware-assisted verification. The baseline method it improves, called TEE-shielded DNN partitioning (TSDP), uses an Intel TDX Trusted Execution Environment to compute non-linear operations and verify the results of linear operations offloaded to an untrusted GPU. The problem with applying that baseline directly to Transformer LLMs is that attention produces large intermediate states that must cross the TEE-GPU boundary repeatedly, and SoftMax over long sequences is expensive inside a TEE.

VeriAttn’s core contribution is offloading both linear and non-linear attention computations to the GPU while keeping only lightweight verification and pre/post-processing inside the TEE. This removes the major data movement bottleneck. For prefill, a two-level pipeline overlaps TEE memory copies, data transfers, GPU computation, and in-TEE verification across attention head blocks and exponentiation row tiles. For decoding, when the KV cache grows beyond GPU memory, VeriAttn partitions attention: the GPU handles KV entries that fit in its memory, while the TEE locally processes the remaining non-resident KV blocks rather than moving them across the boundary for every decode step.

The paper evaluates the system on an Intel TDX platform across four Transformer models ranging from 3B to 14B parameters (LLaMA, Qwen, Phi architectures). At a prompt length of 6,000 tokens, VeriAttn improves full-model time to first token by 2.60 to 3.38× over the TSDP baseline and by 3.14 to 5.19× over running everything inside the TEE. For long-context decoding with 10,000 output tokens, it improves time per output token by 3.86 to 5.42× over TSDP. The paper also benchmarks VeriAttn against zkLLM, a cryptographic proof-based system: zkLLM achieves negligible cryptographic soundness error but incurs two orders-of-magnitude longer TTFT and TPOT. That gap makes the hardware-versus-cryptography trade-off concrete.

The limitation worth stating directly: VeriAttn’s security guarantee is only as strong as the TEE hardware. Intel TDX and AMD SEV protect memory and execution state from compromised privileged software, but they are not formally cryptographically sound in the same sense that a zero-knowledge proof is. Side-channel attacks on TEE platforms have been demonstrated in research. An adversary who can exploit a TEE vulnerability breaks VeriAttn’s guarantees entirely, whereas a cryptographic soundness error bounds the probability of an undetected cheat even against computationally unbounded adversaries.

Maverick — Information-Theoretic Verification With No Server Overhead

Maverick, published September 9, 2026, approaches the problem from a different direction: a mathematical delegation protocol for matrix-vector multiplication. Matrix-vector products are the dominant operation in LLM inference, accounting for the vast majority of compute. If a client can efficiently verify that a server computed the right matrix-vector product without re-doing the multiplication, it can verify LLM inference. Maverick presents what the paper describes as the first information-theoretically sound verification protocol for matrix-vector multiplication delegation with transparent preprocessing, efficient batch verification, and virtually no server overhead.

The protocol works by combining sparse challenge vectors with error-correcting codes. Rather than sampling a uniform random challenge vector and computing a full matrix product to check the server’s result (which is as expensive as the original computation), the client samples a sparse challenge vector and uses a linear code to spread any nonzero error across enough positions that the sparse check is likely to intersect it. Preprocessing computes a matrix transformation that is independent of client inputs, so the verifier can use it to check any number of subsequent computations. For input privacy, LPN-based pseudorandom masking hides the client’s prompt from the server without introducing server overhead: the mask is generated from sparse random vectors and a preprocessed matrix product, allowing the client to remove it efficiently without recomputing the full product.

The end-to-end Maverick prototype evaluates on Qwen3-4B (an open-source model whose weights are public). With one client thread and a CPU server using up to 128 threads, Maverick achieves throughput gains over local inference of up to 17× when privacy masks are generated online, 45× when they are precomputed, and 44× in a verification-only configuration without privacy. The client online work at matrix dimension n=2^14 completes in 9.33 ms versus 116.00 ms for a local multiplication, a 12.44× client-side speedup. The paper compares the vMVMD protocol against two baselines: Dumas-Zucca (specialized for matrix-vector verification) and Sum-Check instantiated with BaseFold (the approach used by DeepProve and similar ZK systems). Across evaluated dimensions up to n=2^12, Maverick’s implementation is up to 34.8× faster than Dumas-Zucca and up to 194.8× faster than Sum-Check+BaseFold, including the cost of input privacy in that figure.

The key limitation is scope. Maverick outsources linear operations to the server and runs nonlinear operations locally on the client. In the Qwen3-4B evaluation, nonlinear operations account for less than 1% of total inference time. But this design requires the client to handle nonlinear computations, which may not be practical for very resource-constrained clients. It also requires that model weights be public, since the preprocessing involves the weight matrices. Maverick is most directly useful for open-source model inference outsourcing, not for proprietary model weight protection.

TensorCommitments — Lightweight Binding With Tensor-Native Trees

TensorCommitments, published February 13, 2026, targets a different point in the cost curve: very low overhead for the prover, no GPU required for the verifier. The paper’s starting observation is that existing cryptographic verification schemes encode model execution into arithmetic circuits or constraint systems, producing proof overhead that is impractical for large models and hard to amortize in interactive workloads. TensorCommitments instead uses a commitment scheme: an irreversible tag that binds the inference to its internal states, detectable when tampered with, without requiring the verifier to re-execute inference.

The key technical insight is that transformer parameters are naturally tensors, and committing to multivariate polynomials respecting those tensor axes is substantially cheaper than committing to flat vectors using univariate polynomials. The paper demonstrates that moving from univariate to bivariate interpolation reduces polynomial evaluation runtime from 4.1 seconds to 0.125 seconds for a fixed grid of 2^12 samples, a 30-fold speedup. This makes the commitment cost decrease as model dimensionality increases, the opposite of the scaling behavior that makes monolithic ZK systems impractical at LLM scale. TensorCommitments organizes these multivariate commitments into Terkle Trees, a tensor-native authentication structure that tracks evolving hidden states with a single root.

The evaluation on LLaMA2 reports 0.97% added prover time and 0.12% added verifier time over plain inference. The paper measures robustness to tailored LLM attacks, defined as adversarial output perturbations that a verifier must detect, and shows improvements of up to 48% over the best prior work that requires a verifier GPU. The verifier checks cryptographic pairings rather than re-running any inference, enabling lightweight verification on a CPU. The paper also introduces a layer selection algorithm based on spectral heavy-tail scores, so when full verification is too expensive, the verifier can challenge the most important layers first.

TensorCommitments does not produce a zero-knowledge proof. The commitment scheme detects tampering probabilistically based on which layers the verifier challenges, not deterministically across all operations. A sophisticated adversary who knows the challenge selection algorithm could concentrate tampering in layers unlikely to be challenged. The paper addresses this with its layer selection methodology, but the security is weaker than a cryptographic soundness bound over all operations.

NanoZK — Layerwise Zero-Knowledge Proofs at GPT-2 Scale

NanoZK, published March 2026, takes the full zero-knowledge proof approach but structures it to avoid the scalability barrier that makes monolithic ZK systems impractical for large models. The core observation is that transformer inference decomposes naturally into independent layer computations: each layer’s output depends only on the previous layer, not on global state shared across non-adjacent layers. This means the monolithic proof can be replaced by L independent layer proofs connected by cryptographic commitments that ensure each layer’s input matches the previous layer’s output.

The peak memory reduction is immediate. A monolithic proof requires memory proportional to the sum of all layer constraint sizes. NanoZK requires memory proportional to the largest single layer, a factor of L improvement that makes GPT-2 scale tractable without the hours of proving time that prior full-circuit approaches required. The framework also enables parallel proving: since layer proofs are independent, they can be generated concurrently. Each layer generates a constant-size proof regardless of model width, achieving 6.9 KB proofs for layers ranging from 64 to 768 dimensions.

Non-arithmetic operations present a separate challenge. Softmax, GELU, and LayerNorm cannot be directly represented in arithmetic circuits. NanoZK develops lookup table approximations with 16-bit precision that, importantly, introduce zero measurable perplexity change on standard benchmarks. This means the proved computation matches the original model exactly rather than approximating it. For resource-constrained scenarios, Fisher information provides a principled measure of layer importance: verifying only high-Fisher layers captures 65 to 86% of model sensitivity with 50% of proving cost, compared to 51 to 79% coverage from random layer selection. The soundness error using Halo2 IPA is below 10^-37 for 32-layer models.

The scale boundary is NanoZK’s main limitation. The paper evaluates on GPT-2, GPT-2-Medium, and TinyLLaMA-1.1B. GPT-2 scale transformer block proofs take 43 seconds with 23 ms verification time. NanoZK achieves 52× speedup over EZKL at these sizes, reaching 228× for larger models where EZKL encounters memory pressure. What the paper does not demonstrate is how the 43-second proving time scales beyond 1.1B parameters, or what the practical overhead looks like for models in the 7B to 70B range where most production deployments now operate. That gap between the 1.1B scale demonstrated and 7B+ production scale is significant.

SLP — Sampled Proofs That Commit the Whole Trace Before Any Challenge

Sampled Layerwise Proofs (SLP), from Youki Lim and Sam Yong at TrueOpen (September 22, 2026), is the freshest entry and the one that scales farthest, reaching Llama-2-70B at the cost of a 1,259-second proof. Its design separates the commitment phase from the proving phase. The system first seals the boundary activations of every chunk of the inference trace, absorbing all commitments into the transcript before any challenge is derived. Only after all boundaries are committed does it determine which chunks to prove. The verifier always receives proofs for the chunks adjoining the input and the output (the endpoint anchors), plus a verifier-selected subset of interior chunks.

This seal-then-sample structure has a specific property: the prover commits to all boundaries before knowing which interior chunks will be audited. This prevents after-the-fact manipulation of a specific boundary to cover a tampered chunk. The paper is careful to distinguish commitment coverage from arithmetic coverage: a complete manifest prevents a sealed boundary from being changed, but it does not establish that an unproved interior transition is correct. In the 70B configuration, three random chunks plus two anchors are proved out of 163 chunks, meaning a single fixed invalid chunk is covered with probability 3/161. That detection probability grows with the audit budget.

Proof cost in SLP is dominated by model weights, not by tokens. On TinyLlama, doubling the context from 8 to 16 tokens changed the proving time from 94.8 seconds to 99.0 seconds. This observation motivates batch proving: SLP packs multiple concurrent requests into one trace under a block-diagonal causal mask, decoding all slots in lockstep and binding each request’s prompt and answer to its position. Twelve packed requests produce a single proof in 181.9 seconds, 6.5× less proving time than twelve separate proofs at the measured single-proof cost. A simulated service proves twelve requests at 30.6 seconds per request with 0.6 seconds of verification each.

The engineering challenges at 70B scale are substantial. With eight bytes per integer weight and 32 bytes per field element, 70 billion parameters require roughly 560 GB and 2.24 TB respectively, and holding both representations would require approximately 2.8 TB. SLP spills quantized weight tensors to disk and commits weight polynomials in bounded batches. The recorded 70B run peaked at about 387 GB resident memory with a 60 GB working set. The fixed-point canonical model used for proving achieves 84.8 to 84.9% argmax agreement with the floating-point reference over 334,705 WikiText-2 test positions, a fidelity gap that must be reported as a property of the registered model version rather than concealed. A timing attack on the challenge derivation exists: if the Fiat-Shamir challenge comes only from a worker-chosen manifest, the worker can grind attempts at 12.5 ms each to avoid unfavorable chunk selections. SLP addresses this by requiring an externally ordered challenge source.

How They Compare

System Trust basis Coverage Model scale demonstrated Prover overhead Verification time Weight privacy Prompt privacy
VeriAttn Intel TDX hardware enclave Full (TEE verifies all operations) 3B-14B (LLaMA, Qwen, Phi) Not reported; 2.60-5.42× faster than baselines Milliseconds (online) Partial (TEE protects, GPU is untrusted) Yes (TEE-protected)
Maverick Information-theoretic (mathematical) Linear ops verified; nonlinear run locally Qwen3-4B (end-to-end) Virtually zero server overhead 9.33 ms online client work at n=2^14 No (weights must be public) Yes (LPN masking)
TensorCommitments Cryptographic commitment (multivariate) Probabilistic; verifier challenges layers LLaMA2 0.97% prover, 0.12% verifier overhead Lightweight; no GPU needed Yes Not addressed
NanoZK Zero-knowledge proofs (Halo2 IPA) Full per-layer; Fisher-guided partial option TinyLLaMA-1.1B 43 s proof per transformer block at GPT-2 scale 23 ms Yes Yes
SLP Polynomial commitments (sampled) Anchors + verifier-selected chunks Llama-2-70B (sampled) 1,259 s for 5/163 chunks at 70B 46.3 s at 70B without weights Yes Not addressed

What This Category Reveals

The five approaches are ordered from the most production-viable today (hardware TEE) to the highest cryptographic ambition and largest demonstrated model scale (SLP). That ordering is not a simple progression: each approach makes different trade-offs between hardware assumptions, proof strength, prover cost, verifier cost, and what exactly is verified.

The critical design dimension separating these systems is what the verifier trusts. VeriAttn trusts a hardware manufacturer’s TEE security guarantees. Maverick trusts mathematics only, assuming the server runs the specified linear operations and the client handles the nonlinear residue. TensorCommitments and SLP trust that polynomial commitments bind their inputs, but not that all interior transitions are correct unless specifically proved. NanoZK provides the strongest formal guarantee per layer, with soundness error below 10^-37, but only at GPT-2 scale. No single system in this list provides both production-latency serving and cryptographically complete verification of a frontier-scale model with weight privacy, prompt privacy, and no hardware trust assumptions. That combination remains undemonstrated.

The open question the category has not answered is the fidelity gap. Every cryptographic system ultimately proves a fixed-point or quantized representation of a model, not the floating-point checkpoint that customers think they are paying for. SLP’s 84.8 to 84.9% argmax agreement figure makes this concrete: the proved computation and the production computation are related, but not identical. Until verification systems can close the gap between the provable fixed-point model and the production floating-point model without requiring the customer to re-run inference, “the provider ran the claimed model” and “the provider ran a model that produces identical outputs to the claimed model” remain two different statements.

Limitations and Open Questions

Scale is the dominant limitation across the cryptographic approaches. NanoZK’s 43-second proof at GPT-2 scale (~117M parameters) gives no direct evidence of what proving time would be for a 7B model, let alone a 70B. SLP’s 1,259-second 70B proof covers only 5 of 163 chunks and requires disk-backed weights and 387 GB peak memory. Neither is compatible with the latency expectations of production serving.

The fidelity problem is underreported. Proving that a fixed-point computation occurred is not the same as proving that a floating-point API call occurred. Every system that converts a model to fixed-point or integer representation for proving purposes creates a gap between the proved computation and the one the user thinks they paid for. SLP addresses this honestly, measuring and reporting 84.8 to 84.9% argmax agreement and treating fidelity as a separate property of a registered model version. Other systems should report this figure explicitly before claiming to prove LLM inference in a meaningful commercial sense.

VeriAttn’s TEE-based approach sidesteps the fidelity problem but introduces hardware supply chain trust. If the TEE has a vulnerability, the security guarantee disappears. The paper compares favorably to zkLLM (two orders of magnitude faster), but the comparison is between a hardware guarantee and a cryptographic guarantee, which are not equivalent in what they protect against.

Prompt privacy is handled differently across systems. Maverick provides it explicitly through LPN masking. VeriAttn provides it through TEE memory protection. NanoZK and TensorCommitments do not directly address prompt privacy in the evaluated configurations. Engineering teams handling sensitive user inputs must check this property specifically for any system they evaluate.

What This Means for Engineering Teams

Understanding what the actual risks of LLM deployment are in your context is the right starting point. For most production applications today, provider trust is not the most acute risk: misaligned model behavior, prompt injection, and hallucination are more operationally impactful than a provider secretly substituting cheaper weights. Verifiable inference becomes more important in regulated settings (healthcare, legal, finance), multi-agent pipelines where one inference feeds the next, and high-value inference contracts where billing is based on model version and precision.

For teams evaluating these systems now: VeriAttn is the only approach in this list that operates at serving latencies on real hardware for models up to 14B. Its limitation is the hardware trust assumption, which matters differently depending on whether the threat model includes the hardware vendor. Teams building on AI infrastructure that runs on Intel TDX-capable cloud instances (available on Azure Confidential Computing and similar offerings) can evaluate VeriAttn today without waiting for cryptographic proof systems to close the scale gap.

Maverick is the strongest near-term candidate for open-source model verification without hardware assumptions. Its requirement that model weights be public limits it to open-source deployments, but that covers a substantial fraction of internal enterprise serving today. Teams already familiar with how LLM inference works mechanically will find the matrix-vector delegation design straightforward to reason about, because it maps cleanly onto the structure of the linear layers that dominate inference compute.

For teams designing new compliance frameworks or governance requirements around LLM-as-a-service procurement, these systems indicate the right shape of what is possible to demand: commitment to a model identifier and version, sampling-based audit rights over inference traces, and eventually cryptographic proof of model identity. The right governance requirement is not “prove every inference” but “commit before execution and allow audit on demand,” which SLP’s seal-then-sample design already implements for non-production audit workloads.

Key Takeaways

  • SLP (TrueOpen, September 22, 2026) achieved the first published end-to-end cryptographic verification of a 70B inference trace, sealing 163 chunks and proving 5 in 1,259 seconds with 84.8 to 84.9% argmax agreement to the floating-point reference over 334,705 WikiText-2 positions.
  • VeriAttn (June 2026) operates at serving latencies on real Intel TDX hardware, improving TTFT by 2.60 to 3.38× over standard TEE-shielded DNN partitioning at 6,000-token prompts, while running two orders of magnitude faster than cryptographic proof systems at the cost of hardware trust assumptions.
  • Maverick (September 9, 2026) eliminates server overhead entirely by delegating only matrix-vector multiplication and running nonlinear operations locally, achieving 12.44× client-side speedup at n=2^14 and up to 194.8× speedup over Sum-Check+BaseFold verification on Qwen3-4B.
  • TensorCommitments adds only 0.97% prover overhead and 0.12% verifier overhead to LLaMA2 inference by committing to tensor-shaped polynomials in a Terkle Tree structure, improving attack robustness by up to 48% over prior work while requiring no verifier GPU.
  • NanoZK’s layerwise decomposition achieves 43-second proofs with 6.9KB proof size and 23ms verification at GPT-2 scale, with soundness error below 10^-37, but the gap from 1.1B to 7B+ production scale is not yet characterized.
  • No system in this list simultaneously provides production-latency serving, cryptographically complete coverage, frontier-scale (70B+) model verification, weight privacy, and prompt privacy without hardware trust assumptions. That combination is the open design problem for the category.

Work With Origins AI

Origins AI builds production AI infrastructure for engineering teams. If you are designing governance frameworks for LLM-as-a-service procurement or need to evaluate verifiable inference for regulated applications, talk to our team.