Five Systems, Two Phases: How LLM Inference Is Being Split Apart
Two papers published within a day of each other in September 2026 pushed a well-known idea to a new level. Disaggregated Quantization (September 22) trains a separate NVFP4 checkpoint specifically for prefill, leaving the decode checkpoint entirely unchanged. KITE (September 23) goes further: it redesigns the model itself so that newly added parameters never participate in building the KV cache at all. The underlying premise of both papers is the same premise that motivated Splitwise at ISCA 2024, SPAD’s proposed silicon, and HMA-Serve’s cross-vendor cluster: prefill and decode are not the same computation, and treating them as identical is leaving performance and money on the table.
If both phases must run the same model checkpoint on the same hardware, the serving fleet is a compromise for both. If they can differ, each can be optimized on its own terms. This article traces five systems across that spectrum, from Splitwise’s foundational scheduler to KITE’s architectural separation inside the model itself, and explains what each separation costs and what it buys.
Why This Comparison Is Timely
Prefill-decode disaggregation, routing the two phases to separate GPU pools, is already standard infrastructure. NVIDIA Dynamo documents it as an ordinary serving configuration, and AWS documentation now treats it as a baseline option for high-traffic workloads. What changed in September 2026 is that the separation moved inside the model and the checkpoint. Disaggregated Quantization demonstrates that a prefill checkpoint can be trained to use NVFP4 compute formats while the decode checkpoint uses a compact weight-only encoding, and the two can serve a single request without accuracy loss. KITE demonstrates that model capacity added during scaling does not need to touch KV construction at all, keeping the KV-producing subsystem small as the model grows.
Those two results make the question sharper. If the phases can run on different silicon, train under different objectives, use different numerical representations, and correspond to different towers in the model architecture, what exactly does it mean for a serving fleet to run “a model”?
The Shift: One Homogeneous Executable to Two Specialized Products
Classic LLM serving treats a request as a sequence of operations on a single model checkpoint running on homogeneous hardware. Every layer, every linear operation, every attention head uses the same weights, the same numerical format, and the same accelerator type. Prefill and decode are two phases in a single pipeline, not two independent systems.
The systems in this comparison move away from that picture at different abstraction levels. Splitwise separates the machines while keeping the model identical. HMA-Serve pairs different hardware vendors and applies different numerical formats to each phase. SPAD designs different chips for each machine from the ground up. Disaggregated Quantization trains different weight representations for each phase. KITE changes the model architecture so that KV construction and token generation happen in structurally separate towers. The common thread across all five is that prefill and decode are becoming distinct products that happen to be joined at the KV cache boundary.
Splitwise — Separate Machines, Identical Model
Splitwise, published at ISCA 2024 by Pratyush Patel, Esha Choukse, and colleagues at Microsoft Research, established the core argument for phase separation with production numbers. LLM inference has two phases with measurably different hardware profiles: a compute-intensive prompt computation phase and a memory-intensive token generation phase. Splitwise routes them to separate machines and transfers request state between them using high-speed back-plane interconnects available in modern GPU clusters.
The results justify the transfer overhead. Splitwise clusters achieve up to 1.4× higher throughput at 20% lower cost than co-located baselines. Under the same power and cost budgets, the advantage grows to 2.35× more throughput. The key observation is that token generation does not need the compute capability of the latest GPUs. Running it on hardware matched to its memory-bandwidth profile is more efficient than running it on the same premium silicon that handles prefill.
Splitwise does not change the model. Both machines hold identical weights and run identical code. What it changes is the scheduling policy and the physical machine assignment. That simplicity is both its strength and its ceiling: the weights themselves remain a compromise for both phases, and the hardware on each side is still a general-purpose GPU. Splitwise is the baseline that every more aggressive separation in this list tries to exceed.
HMA-Serve — Cross-Vendor Hardware With Phase-Wise Quantization
HMA-Serve, published in June 2026, asks a follow-on question: if prefill and decode run on different machines anyway, those machines can come from different vendors with different silicon architectures. The paper pairs Tenstorrent Blackhole p150b GDDR accelerators for prefill with NVIDIA HBM-based GPUs for decode, and introduces three coordinated mechanisms to make cross-vendor serving work without manual format translation.
The motivation is concrete. A Tenstorrent Blackhole p150 provides 664 TFLOPS of BFP8 compute at approximately $1,300, an order of magnitude cheaper than a comparable NVIDIA A100 while offering similar low-precision throughput. HBM sits almost entirely idle during compute-bound prefill; HMA-Serve measures an A100 wasting over 97% of its HBM bandwidth during 4K-token prefill. High-bandwidth memory is the dominant cost component of modern accelerators, accounting for 18% of A100 manufacturing cost and rising to 45% on the B200. That expensive bandwidth is useful during decode, not prefill.
Cross-vendor disaggregation breaks two assumptions that single-vendor deployments take for granted: a KV format both ends consume natively, and a shared software stack. HMA-Serve addresses both through three mechanisms that work together. Phase-wise quantization runs prefill in Tenstorrent’s vendor-native BFP8 and decode in BF16, avoiding accuracy loss on the memory-bound generation side while maximizing compute throughput on prefill. Compute-transfer pipelining exposes per-layer completion events from the prefill runtime and overlaps each layer’s KV egress, device-to-host DMA plus RDMA, with later-layer prefill computation, keeping transfer off the critical path for time-to-first-token. Deferred dequantization ships raw quantized bytes verbatim across the network, halving wire traffic, and reconstructs them lazily inside a fused decode-side kernel using the GPU’s integer ALU rather than the tensor cores that decode already saturates.
Across four Qwen3 models (4B to 32B) and three production traces, HMA-Serve delivers up to 3.2× higher goodput than memory-homogeneous serving methods and 4.8× higher goodput-per-dollar, with no measurable loss on generation-quality benchmarks. Unlike SPAD, HMA-Serve runs on real deployed hardware, not simulated silicon.
SPAD — Purpose-Built Silicon for Each Phase
SPAD, from Hengrui Zhang, Pratyush Patel, August Ning, and David Wentzlaff at Princeton and the University of Washington, asks what happens when hardware is designed from scratch for one phase rather than adapted for it. The paper’s starting point is that current datacenter GPU design philosophy, which maximizes both compute capacity and memory bandwidth on every chip, creates systematic underutilization in disaggregated deployments.
The paper quantifies that underutilization with a targeted experiment. Reducing the HBM bandwidth of a modeled H100 by 40% increases prefill latency by only 17%. Reducing compute capacity by half increases decode latency by only 22%. Both phases are tolerant of the expensive resource they do not use heavily. SPAD calls this a “less-is-more” philosophy: right-size each chip for its actual workload rather than over-provisioning both.
SPAD proposes two chip designs built around those tolerances. Prefill Chips carry larger systolic arrays for parallel matrix computation and use cost-effective GDDR memory, since prefill’s compute-bound nature leaves expensive HBM bandwidth underutilized. Decode Chips retain high memory bandwidth but reduce compute capacity, matching the memory-bound, low-arithmetic-intensity profile of sequential token generation. Simulations against modeled H100s show Prefill Chips delivering 8% higher prefill performance at 52% lower hardware cost. Decode Chips achieve 97% of decode performance at 28% lower TDP.
End-to-end cluster simulations on production chatbot and code generation traces show hardware cost reductions of 19% to 41% and TDP reductions of 2% to 17% compared to homogeneous H100 baselines at the same performance level. When models or workloads change, SPAD chips can be reallocated to run either phase; those simulations still show 11% to 43% lower hardware costs than equivalent H100 clusters, which the authors describe as demonstrating the longevity of the design.
The critical caveat is that SPAD remains simulation-based throughout. The chips are not fabricated silicon. All cost projections and performance figures come from LLMCompass simulations of proposed designs. The gap between modeled and manufactured cost for a new chip family is historically large, and HBM pricing, packaging yield, and memory controller design all affect the real-world cost-performance ratio in ways that simulations typically underestimate.
Disaggregated Quantization — Different Weights for Each Phase
Disaggregated Quantization (DQ), submitted September 22, 2026, moves the separation from hardware into the model checkpoint. The observation driving the work is that quantization formats that help one phase actively hurt the other. Quantizing activations alongside weights enables hardware-native matrix multiply, which accelerates compute-bound prefill. Applying the same activation quantization to decode, where weight loading dominates rather than compute, degrades accuracy without improving speed. The right numerical format for each phase is different, and standard quantization pipelines do not account for this.
The paper measures how asymmetric this sensitivity is. On decode-heavy benchmarks, quantizing decode alone incurs 2 to 4× the accuracy loss of quantizing prefill alone across most tested models, reaching 7× on Gemma3-1B. DQ introduces three complementary schemes to exploit that asymmetry. Format disaggregation keeps shared model weights but disables activation quantization specifically on decode, combining NVFP4 prefill compute with weight-only NVFP4A16 decode. It recovers most of the decode-side accuracy loss while adding no weight storage overhead and no prefill cost.
Full disaggregation goes further and trains separate prefill weights using QADD, Quantization-aware Distillation with Disaggregation. QADD uses the SFT label mask, normally used only for loss masking, to route prompt tokens through the prefill pathway and response tokens through the decode pathway in a single forward-backward pass. Both pathways optimize toward the same response objective. The result is a native NVFP4 prefill checkpoint and a separate low-bit decode checkpoint. For 1-bit GGUF compression on Qwen3.8-27B, training an NVFP4 prefiller improves accuracy by 32.5 points on MMLU-Pro and 35.3 points on MMMU-Pro without modifying the decode checkpoint at all.
The third scheme, Offloaded Disaggregated Prefill (ODP), addresses the device-memory cost of holding two checkpoints on one accelerator. ODP streams prefill weights from SSD block by block, reusing device memory buffers as the context propagates through the network. Once a prefill transformer block has produced its outputs, its weights are no longer needed until the next request, so the same buffer space can hold the next block. Prefill compute grows with context length while SSD loading cost is fixed per block, so relative loading overhead decreases as prompts grow. On Qwen3.8-27B, compute overtakes SSD loading around 8K context for all tested model sizes. The llama.cpp implementation delivers a 1.78× time-to-first-token speedup over the weight-only baseline at 8K context. The paper validates the full DQ framework under disaggregated serving in vLLM and reports results via post-training quantization on models up to 2.8T parameters.
KITE / Step Scale Transformer — Architecture Where Added Capacity Skips KV Construction
KITE, KV-Invariant Transformer Expansion, submitted September 23, 2026, works at the model architecture level. Its core idea is that when a Transformer is scaled up, newly added parameters do not need to participate in building the KV cache. If the KV-producing portion of the model stays constant while additional capacity is added outside that region, the prefill cost of the larger model stays equal to the prefill cost of the smaller model, but model quality improves.
The paper introduces the Step Scale Transformer (SST) as a concrete instantiation. SST is a two-tower decoder. The first tower, the Prefiller, processes the prompt and produces layer-wise KV. The second tower, the Decoder, reads those KVs to predict the next token. During bulk prefill, only the Prefiller runs. During generation, both towers run per decode step, but the KV computation belongs entirely to the first tower. The Prefiller is trained first as a standard model; the Decoder is added and both towers continue training jointly, a form of upcycling that avoids running the full two-tower compute from token zero.
The motivation comes from measured production traffic. The paper analyzes OpenRouter requests for six top models over August 22 to September 20, 2026. For Claude Fable 5, uncached input tokens exceed output tokens by 16.29× on average, and uncached-input charges account for 76.7% of total charges. For GPT-6 Astra, uncached input exceeds output by 14.64× with a 74.6% charge share. Making prefill architecturally cheaper directly reduces serving cost for the workloads that dominate real production traffic.
The experiments compare SST against same-compute baselines. A 67B MoE SST with 2.15B active body parameters per decode token reaches an EMA-200 training loss of 1.5900, which is 0.0106 lower than a 47B classic Transformer and 0.0021 lower than a 63B classic Transformer trained on comparable cumulative compute. With illustrative cost weights of 75% for prefill and 25% for decode, SST’s analytical inference-cost proxy is 6.7% lower than the 47B baseline and 31.6% lower than the 63B baseline. On downstream tasks, SST scores higher than both baselines on OpenBookQA, MMLU, GSM8K, MATH, HumanEval, MBPP, and BBH.
The limitation worth stating directly: KITE requires training. An organization running an existing deployed model cannot adopt KITE without committing to a new training run structured around the two-tower paradigm. The paper notes that SST is probably not the optimal KITE instantiation and is a proof of concept of the paradigm. MoE and hybrid attention are orthogonal to SST and can be combined with it, but the interactions between those techniques at scale under production serving conditions remain unexplored.
How They Compare
| System | Depth of separation | Different weights per phase | Different hardware | Training required | Key measured result | Readiness |
|---|---|---|---|---|---|---|
| Splitwise | Scheduling only | No | Optional | No | 1.4× throughput, 20% lower cost | Published ISCA 2024; clusters characterized |
| HMA-Serve | Cross-vendor hardware + phase quantization | No (same weights, different formats) | Yes (Tenstorrent GDDR + NVIDIA HBM) | No | 3.2× goodput; 4.8× goodput-per-dollar | Deployed cluster on real hardware |
| SPAD | Purpose-designed silicon | No | Yes (GDDR Prefill Chip + HBM Decode Chip) | No | 19-41% cluster cost reduction vs H100 | Simulation only; no fabricated silicon |
| Disaggregated Quantization | Separate checkpoints and numerical formats | Yes (NVFP4 prefill + weight-only decode) | Optional | Yes (QADD fine-tuning) | 1.78× TTFT speedup at 8K; +32.5 MMLU-Pro pts | Demonstrated in llama.cpp and vLLM |
| KITE / SST | Model architecture (separate towers) | Yes (structurally distinct Prefiller and Decoder) | Optional | Yes (new training paradigm) | 31.6% lower inference cost vs same-quality 63B baseline | Research prototype; requires full training |
What This Category Reveals
The five systems are ordered by how deep the separation goes: scheduling, then hardware type, then silicon design, then checkpoint, then model architecture. At each level, the separation becomes harder to undo. A different scheduling policy can be reverted with a config change. A different checkpoint can be deprecated. A different model architecture requires a training decision made before deployment. The ordering makes the trade-off visible: more separation means more optimization potential and more operational commitment.
The critical ingredient separating useful systems from theoretical ones is an honest accounting of what the split introduces. Every phase separation adds KV transfer overhead, storage overhead for separate configurations, and operational complexity in the form of two pipeline stages to monitor, scale, and upgrade independently. Systems that show net improvements despite those costs are demonstrating that the phase mismatch was large enough to justify the overhead. HMA-Serve showing 4.8× goodput-per-dollar on production Qwen3 traces with real Tenstorrent hardware is a different class of claim from SPAD’s simulations. DQ showing 1.78× TTFT in llama.cpp at 8K context is a real measurement. When evaluating entries in this category, the first question to ask is whether the result comes from real hardware and real serving software, not a simulator.
The open question the category has not answered is how to make phase-specific optimization dynamic. All five systems commit to a hardware and checkpoint configuration at deployment time. If a workload shifts from prompt-heavy to generation-heavy, or the ratio of input to output tokens changes, a statically disaggregated cluster has no efficient way to rebalance. SPAD gestures at this by designing chips that can run either phase. Production systems will need autoscaling that adjusts the ratio of prefill capacity to decode capacity as traffic evolves, not just the ability to statically provision the right ratio at deployment.
Limitations and Open Questions
Scale is the clearest limitation across the board. Most results are on models from 4B to 67B parameters under controlled traffic conditions. Production serving at frontier scale introduces additional variables: model parallelism across many GPUs, KV cache migration costs that grow with context, and power draw profiles that differ from single-chip measurements. DQ validates shared-weight format disaggregation via post-training quantization on models up to 2.8T parameters, which is an important step, but full disaggregation with separately trained prefill weights at that scale is not yet demonstrated.
SPAD’s results are simulation-based throughout. The claim that Prefill Chips deliver 8% higher performance at 52% lower cost than modeled H100s depends on chip cost projections, GDDR pricing assumptions, and a performance simulator rather than fabricated silicon. The gap between simulated and real hardware cost for a new chip family is historically large.
DQ’s Offloaded Disaggregated Prefill scheme, which streams prefill weights from SSD, works for dense models but the paper explicitly states the approach does not transfer to MoE models. The ratio of compute cost to loading cost grows with the fraction of active parameters in an MoE, making SSD-streamed prefill unworkable at practical context lengths for sparse models, which is where many frontier deployments now sit.
KITE requires a new training paradigm. An existing deployed model cannot adopt it without a full training run with the two-tower architecture. The paper acknowledges SST is a first instantiation, not the optimal one, and the interactions between the KITE paradigm and standard MoE scaling at frontier model sizes remain open.
What This Means for Engineering Teams
Teams evaluating prefill-decode disaggregation for production serving should separate the question of whether to disaggregate from the question of how deep to go. Disaggregated scheduling (Splitwise-style) is available in frameworks like NVIDIA Dynamo today and is a low-risk starting point. It does not change the model and can be reversed. Building an understanding of LLM inference mechanics, specifically how prefill and decode differ in hardware demand, makes the cost-benefit calculation for any of the above approaches concrete before committing to infrastructure changes.
Cross-hardware disaggregation, pairing GDDR accelerators with HBM GPUs as HMA-Serve demonstrates, is operationally viable today and the cost-per-dollar case is strong for prefill-heavy workloads. The challenge is software: cross-vendor stacks require careful format translation at the KV boundary. Teams serving workloads where uncached input tokens dominate should run the math on what GDDR-based prefill hardware costs at their input token volume. The KITE paper’s OpenRouter data showing 70-77% of charges attributable to uncached input across the largest production models is a concrete benchmark to check against your own serving logs.
Disaggregated Quantization is the most immediately applicable for teams already using quantized serving. Adding a prefill-specific checkpoint requires a QADD fine-tuning run, but the result can attach to existing weight-only GGUF decoders without modifying them. The 1.78× TTFT gain at 8K context matters for any application processing long prompts: document analysis, code review, retrieval-augmented generation over large contexts. Teams already doing model compression work should evaluate whether phase-specific quantization formats improve accuracy at their target bit-width before committing to a uniform quantization scheme across both phases.
KITE’s architectural separation is a longer-horizon decision, relevant for organizations training new model families where inference cost is a first-class constraint from the start. For teams building on AI infrastructure that expects high input-to-output token ratios, it is the direction that eliminates the phase mismatch at the root rather than compensating for it at the serving layer.
Key Takeaways
- Splitwise (ISCA 2024, Microsoft Research) demonstrated that routing prefill and decode to separate machines achieves up to 1.4× throughput at 20% lower cost, establishing the baseline for all subsequent phase-splitting work without changing the model itself.
- HMA-Serve paired Tenstorrent Blackhole p150b GDDR accelerators (664 TFLOPS BFP8 at approximately $1,300) with NVIDIA HBM GPUs for decode on a real deployed cluster, delivering 3.2× goodput and 4.8× goodput-per-dollar across Qwen3 4B to 32B under production traces.
- SPAD proposes GDDR-based Prefill Chips and high-bandwidth Decode Chips that simulations show reducing cluster cost by 19% to 41% compared to equivalent H100 deployments, though no silicon is fabricated and results are from LLMCompass simulation.
- Disaggregated Quantization (September 22, 2026) trains a separate NVFP4 prefill checkpoint alongside a 1-bit GGUF decode checkpoint for Qwen3.8-27B, improving MMLU-Pro accuracy by 32.5 points and delivering a 1.78× TTFT speedup at 8K context via SSD-streamed prefill weights, with no changes to the decode checkpoint.
- KITE’s Step Scale Transformer (September 23, 2026) places newly added model capacity outside the KV-producing region, reducing estimated inference cost by 31.6% compared to a same-quality 63B classic Transformer while achieving lower training loss and higher scores on seven downstream benchmarks at 67B MoE scale.
- The deepest separation, different model architecture for each phase, is now demonstrated at 67B-MoE scale; whether it transfers to frontier-scale dense models under production serving conditions is the next open question for the category.
Work With Origins AI
Origins AI builds production AI infrastructure for engineering teams. If your serving costs are dominated by prefill-heavy workloads and you want to evaluate phase-specific optimization strategies, talk to our team.

