While model training commands immense capital expenditure, operational artificial intelligence costs are overwhelmingly dominated by real-time inference serving. When millions of concurrent users interact with conversational AI agents, autonomous coding assistants, and enterprise search platforms, inference engines must process unpredictable request arrival rates, dynamic prompt lengths, and streaming token generation under strict Service Level Objectives (SLOs). Naive inference architectures suffer from catastrophic memory fragmentation and memory-bound latency stalls, leaving expensive GPU tensor cores operating at single-digit utilization.
The modern inference serving revolution is anchored by breakthrough systems software: PagedAttention in vLLM and kernel-fused graph execution in NVIDIA TensorRT-LLM. By treating GPU High-Bandwidth Memory like virtual operating system memory pages, implementing continuous request batching, and quantizing key-value caches, modern serving engines deliver 10x to 25x higher request throughput on identical physical hardware. This architectural deep dive examines the memory bottlenecks of autoregressive inference, PagedAttention virtual memory mechanics, dynamic continuous batching, and KV-cache compression strategies.
Table of Contents
- 1. The Mechanics of Autoregressive Decoding: Prefill vs Decode Phases
- 2. The KV-Cache Bottleneck: Why Memory Capacity Limits Concurrency
- 3. PagedAttention: Virtual Memory and Paging for GPU Tensors
- 4. Continuous Dynamic Batching vs Static Batching
- 5. Serving Engine Architectural Matrix: vLLM vs TensorRT-LLM vs TGI
- 6. KV-Cache Optimization: FP8, INT4 Quantization, and Multi-Query Attention
- 7. Speculative Decoding: Accelerating Generation via Small Draft Models
- 8. Frequently Asked Questions
1. The Mechanics of Autoregressive Decoding: Prefill vs Decode Phases
Transformer inference consists of two fundamentally distinct operational phases with opposing computational characteristics:
- Prefill Phase (Prompt Processing): The inference engine ingests the user’s input prompt (e.g., 2,000 tokens) all at once. Because all prompt tokens are known simultaneously, matrix multiplications are dense:
O(N × D). This phase is compute-bound, achieving high Tensor Core arithmetic utilization (high TFLOPS). - Decode Phase (Token Generation): Tokens are generated sequentially, one token at a time. The model generates token N, appends it to context, and feeds it back to predict token N+1. Because each step processes a batch of single vectors (batch size × 1), the computation cannot saturate thousands of parallel arithmetic units. The decode phase is strictly memory-bandwidth bound; performance is dictated by how fast model weights and historical context can be read from High-Bandwidth Memory (HBM).
2. The KV-Cache Bottleneck: Why Memory Capacity Limits Concurrency
During the attention computation, each new token must compute dot-product attention against all preceding tokens in the sequence. To avoid recomputing Key (K) and Value (V) projections for every past token at every single step, inference engines cache previous keys and values in GPU memory: the KV-Cache.
Calculating KV-Cache Memory Consumption:
The memory footprint of the KV-Cache per active request scales dynamically with context length:
KV-Cache Bytes = 2 × 2 × n_{layers} × n_{heads} × d_{head} × s_{seq} × BytesPerPrecision
For a standard 70-billion parameter model (80 layers, 64 heads, head dimension 128) operating in 16-bit precision (2 bytes per element):
- Each token consumes approximately 2.6 Megabytes of memory across the network layers.
- A single user request with a 4,096-token context consumes 10.7 Gigabytes of GPU memory purely for its KV-cache.
- A request with a 32,768-token context consumes 85.5 Gigabytes.
Before PagedAttention, legacy serving systems allocated contiguous blocks of GPU memory based on the maximum potential request length (e.g., reserving 85GB upfront for a request that might only generate 50 tokens). This caused up to 60% to 80% of GPU memory to sit empty as internal and external fragmentation, capping request concurrency to small numbers.
3. PagedAttention: Virtual Memory and Paging for GPU Tensors
Developed by Woosuk Kwon et al. at UC Berkeley in 2023, PagedAttention solved the KV-cache fragmentation crisis by applying classical operating system virtual memory paging to GPU memory management.
How PagedAttention Functions:
- Physical Block Pool: GPU memory is divided into fixed-size physical blocks (e.g., blocks holding keys and values for 16 tokens).
- Logical-to-Physical Block Table: Each incoming request maintains an OS-style page table mapping logical token sequences to non-contiguous physical memory blocks scattered across HBM.
- On-Demand Dynamic Allocation: As a request generates new tokens, the engine allocates individual blocks only when previous blocks fill. Near-zero memory is pre-allocated or wasted.
- Copy-on-Write Memory Sharing: In parallel sampling (generating multiple responses for the same prompt) or beam search, multiple requests share identical physical prompt blocks. If a request branches, the engine applies Copy-on-Write (CoW), copying only the diverged block. This reduces prompt memory overhead by over 50%.
4. Continuous Dynamic Batching vs Static Batching
In classical web serving, requests are batched together: Request A and Request B execute in parallel. In conversational AI, Request A might require 20 tokens (finishing in 500ms), while Request B requires 800 tokens (finishing in 20 seconds).
Under Static Batching, the engine must either wait for all requests in a batch to finish before returning responses, or pad completed sequences with dummy tokens, wasting immense compute.
Continuous Batching (Iteration-Level Batching) operates at the individual iteration level:
- As soon as Request A outputs an
<EOS>(end-of-sequence) token, it is immediately evicted from the batch, and its response is streamed back to the client. - A new incoming Request C from the waiting queue is inserted into the batch on the very next forward pass, without resetting other active generations.
- GPU tensor cores remain 100% saturated with zero idle cycles between varying request lengths.
5. Serving Engine Architectural Matrix: vLLM vs TensorRT-LLM vs TGI
| Evaluation Metric | vLLM (UC Berkeley / Open Source) | NVIDIA TensorRT-LLM | Hugging Face TGI |
|---|---|---|---|
| Core Strength | PagedAttention, ease of deployment, broad open weights | Peak throughput on NVIDIA GPUs, deep C++ kernel fusion | Production reliability, Hugging Face ecosystem integration |
| Hardware Portability | Universal (NVIDIA, AMD ROCm, Intel Gaudi, AWS Neuron) | Proprietary (NVIDIA hardware exclusively) | Broad (NVIDIA, AMD, AWS Inferentia) |
| Compilation Latency | Instant startup (Python / PyTorch / CUDA kernels) | Ahead-of-time compilation required (TRT engines) | Fast startup (Rust / Python hybrid) |
| Peak Throughput (QPS) | Very High (Exceptional batch concurrency) | Maximum (Custom fused GEMM + in-flight batching) | High |
6. KV-Cache Optimization: FP8, INT4 Quantization, and Multi-Query Attention
To scale concurrent user capacity beyond physical memory limits, modern architectures employ algorithmic and data-type optimizations:
- KV-Cache Quantization (FP8 / INT4): Compressing keys and values from 16-bit to 8-bit (FP8) cuts memory usage by 50% with negligible loss in perplexity. Advanced schemes compress to 4-bit (INT4), quadrupling concurrent request capacity on a single GPU.
- Multi-Query Attention (MQA) & Grouped-Query Attention (GQA): Standard Multi-Head Attention maintains independent key and value heads for each query head. Grouped-Query Attention (used in LLaMA 2/3 and Mistral) shares a single key-value head across 8 query heads, shrinking KV-cache footprint by 87.5% in hardware.
- Cross-Attention Prefill/Decode Disaggregation: Routing prefill requests to compute-dense GPU nodes (like B200) while streaming decode requests to memory-dense nodes across high-speed networks, isolating the opposing latency requirements of both phases.
7. Speculative Decoding: Accelerating Generation via Small Draft Models
Because the decode phase is memory-bound (reading 140GB of weights per single token generated), token generation is latency-constrained by memory bandwidth, typically limited to 30 to 50 tokens per second.
Speculative Decoding circumvents this bottleneck by coupling a small, ultra-fast “draft model” (e.g., a 1B model generating 250 tokens/sec) with the large primary “target model” (e.g., 70B):
- The small draft model rapidly guesses K candidate tokens sequentially in memory (e.g., predicting 5 candidate tokens).
- The primary target model evaluates all 5 candidate tokens simultaneously in a single compute-bound forward prefill pass.
- If the target model accepts 4 of the 5 tokens, the system outputs 4 tokens in the time it would normally take to generate a single token, delivering a 2x to 3x wall-clock latency speedup with mathematically identical output distributions.
8. Frequently Asked Questions
What is Time To First Token (TTFT) vs Time Per Output Token (TPOT)?
TTFT measures the prefill latency: how long the user waits before the first response token appears. TPOT measures the decode latency: the time required to generate each subsequent token (streaming cadence). Optimizing TTFT requires raw compute FLOPS, whereas optimizing TPOT requires memory bandwidth.
Does PagedAttention alter the mathematical accuracy of model outputs?
No. PagedAttention is strictly a systems memory management innovation. The attention dot-product mathematics are 100% bit-exact and identical to standard un-paged attention mechanisms.
Why is vLLM preferred over TensorRT-LLM for many startup deployments?
vLLM is written in Python and PyTorch with instant startup times, supporting virtually all open-source models immediately upon release on Hugging Face. TensorRT-LLM requires building compiled C++ engine binaries for specific GPU architectures, which introduces operational friction during rapid prototyping.
Architectural Takeaway
Mastering AI inference serving requires moving beyond basic model hosting. By combining PagedAttention virtual memory, continuous iteration-level batching, Grouped-Query Attention, and speculative decoding, infrastructure engineers transform memory-starved GPU clusters into hyper-efficient real-time token factories capable of serving global enterprise traffic at minimum cost.