Skip to main content

NanoFlow: Towards Optimal Large Language Model Serving Throughput

How NanoFlow raises LLM serving throughput by overlapping compute, memory and network work inside a single GPU: it shows serving is compute-bound, splits each batch into nano-batches, and shares SMs between concurrent kernels, reaching 1.91x the throughput of TensorRT-LLM.

TL;DR

  • LLM serving is usually described as memory-bound. NanoFlow's cost model shows that with large batches, which grouped-query attention makes possible, end-to-end serving is compute-bound for most models and workloads.
  • Serving engines run each operation to completion before the next starts. Each operation keeps its own bottleneck resource about 80% busy, but compute, memory and network take turns, so overall compute utilization is only about 40%.
  • NanoFlow splits each batch into nano-batches and runs copies of each operation on them, so a GEMM for one nano-batch overlaps decode attention or a network collective for another, inside one GPU.
  • An auto-search picks the number and size of nano-batches, their order, and how many SMs each concurrent kernel gets. On LLaMA-2 70B this gives 1.91x the throughput of TensorRT-LLM and 68.5% of the theoretical optimum.

Serving is compute-bound

The paper models one serving iteration (every request in the batch passing through every layer) from three sides:

  • Memory: the GPU streams its whole memory, the weights plus the KV cache, once per iteration.
  • Compute: each dense GEMM costs about 2 ร— batch ร— parameters FLOPs.
  • Network: tensor-parallel collectives move activations between GPUs.

Memory time stays fixed once the KV cache fills the GPU, while compute and network time grow with the dense batch. Grouped-query attention shares each KV head across several query heads, so the same memory holds many more requests. LLaMA-2 70B on 8 A100s reaches a dense batch of about 2048 tokens, and at that size compute takes about 2.5 times as long as memory.

In the compute-bound regime, the best possible throughput depends only on the GPUs' compute and the model size: compute divided by two FLOPs per parameter per token. For LLaMA-2 70B on 8 A100s, the paper puts that optimum at 1857 tokens/s per GPU. Offline, vLLM, DeepSpeed-FastGen and TensorRT-LLM reach only 22.0%, 22.9% and 37.8% of it.

Overlapping work with nano-batches

The gap comes from running operations one after another. During decode attention the tensor cores wait on memory, and during an all-reduce both wait on the interconnect. NanoFlow cuts the batch into nano-batches with no data dependencies between them, so while one nano-batch is in a memory-bound or network-bound operation, another can use the compute units.

Splitting has a cost: smaller GEMMs are less efficient and the weights are read once per nano-batch. The paper measures a 13.2% slowdown from nano-batching alone, and overlap more than recovers it. In the automatically generated LLaMA-2 70B pipeline, the start of each layer (KQV generation, where compute, memory and network all overlap) uses four nano-operations, and the rest of the layer uses two, over token ranges 0โ€“768 and 768โ€“2048.

Sharing SMs between concurrent kernels

Overlapping kernels do not get separate hardware. They split the GPU's streaming multiprocessors, and a GEMM that gives up SMs slows down. The key measurement is how much speed the other kernel gets back for the SMs it takes. Memory-bound and network-bound kernels saturate their bandwidth with a fraction of the SMs, so the trade is favourable: in the LLaMA-2 70B pipeline, decode attention takes 40% of the GPU's resources and still runs at 80% of its full speed.

NanoFlow profiles these trade-offs once per kernel pair, turns them into a table of resource share versus performance, and feeds it to a two-stage mixed-integer linear program. The first stage chooses the pipeline structure (number, size and order of nano-operations), and the second assigns each kernel its resource share. A practical pipeline is found in about 10 minutes.

The runtime around the pipeline

  • Asynchronous scheduling. Forming the next batch (retiring finished requests, admitting new ones, updating the PagedAttention page table) runs on the CPU while the GPU executes the current iteration, instead of stalling the GPU between iterations.
  • Memory-aware admission. New prefill requests are admitted only if predicted peak memory use stays within the GPU; if memory still runs out, a request is offloaded to the CPU and reloaded later without recomputation.
  • KV-cache offloading. For multi-round conversations, finished KV caches are copied to host memory and SSD while the GPU computes, so a follow-up turn does not recompute its history. This costs about 3% throughput from kernel interference.

Results

On 8 A100 80GB GPUs serving LLaMA-2 70B with request lengths drawn from ShareGPT, LMSYS-Chat-1M and Splitwise, NanoFlow delivers on average 4.18x the throughput of vLLM, 3.45x that of DeepSpeed-FastGen and 1.91x that of TensorRT-LLM. At its best it reaches 68.5% of the optimal throughput. At low request rates its latency is similar to TensorRT-LLM, and it sustains up to 1.64x higher request rates within the latency target. The same approach carries to LLaMA-3 8B and 70B, Qwen2 72B, DeepSeek 67B and Mixtral 8x7B, reaching 50% to 72% of optimal.

The ablation separates the pieces. Nano-batching alone costs 13.2%. Overlapping network-bound kernels with compute gives 1.07x over the non-overlapped baseline on a prefill-only workload, and overlapping both network-bound and memory-bound kernels gives 1.17x on a decode-heavy one.

Why it mattered

NanoFlow argues that once large batches make serving compute-bound, the goal is to keep the tensor cores busy rather than to minimize memory traffic, and that the remaining waste is the time GPUs spend idle between heterogeneous operations. Its answer, overlapping those operations inside one device on an automatically searched schedule, complements the memory-side techniques: it runs on a PagedAttention-managed KV cache, and it targets idle time that faster individual kernels do not remove.

If you found this paper review helpful, consider sharing it with others.

Mastodon