← Back to list
AI & Data
#LLM 추론#KV 캐시#연속 배칭#PagedAttention#추측 디코딩#추론 서빙#TTFT#ITL
Last updated · 2026-09-30

Large Language Model Inference Optimization: KV Cache, Continuous Batching, and Speculative Decoding

1. Overview

Definition: Large language model (LLM) inference optimization is the design of model, runtime, and system techniques that improve response latency, throughput, memory use, and cost per token while preserving quality and safety when serving a trained model.

LLM inference processes an input prompt and then repeatedly generates the next token based on the preceding context. A user sees one answer, but the server receives jobs with different prompt lengths, output lengths, concurrency, and model sizes. Adding GPUs alone therefore does not ensure efficient resource use and a consistent response experience.

The key-value cache is essential to autoregressive generation because it stores previously computed context states instead of recalculating them. However, the cache occupies GPU memory as it grows and creates pressure across long contexts and concurrent users. Cache efficiency, request scheduling, parallel execution, and quality loss must therefore be optimized together.

The objective cannot be reduced to throughput alone. Interactive systems care about time to first token and inter-token latency, while offline summarization may prioritize total throughput and cost. Response quality, reliability, fairness, and security must also be expressed in the service SLO.

In practice, low-precision quantization and model compression, KV-cache management, batching and scheduling, parallel execution, and request routing are combined in one serving architecture. Their effectiveness depends on hardware, model structure, prompt distribution, output length, and concurrency, so representative workload benchmarks are indispensable.

2. Inference architecture and bottlenecks

A. End-to-end request path

In the architecture below, the gateway checks authentication, quotas, and policies before routing a request to a suitable model engine. The engine schedules queued work and returns generated tokens as a stream or a completed response.

flowchart LR
  U[User request] --> G[API gateway]
  G --> P[Policy quota routing]
  P --> Q[Request queue and scheduler]
  Q --> E[LLM inference engine]
  E --> W[Model weights]
  E --> K[KV cache pool]
  E --> R[Streaming response]
  R --> G
  G --> U
  M[Metrics logs traces] -.observability.-> G
  M -.observability.-> Q
  M -.observability.-> E

The gateway is a policy boundary separate from model execution. It validates token limits, maximum context length, tenant priorities, and privacy controls, then selects an engine that matches the request's model, region, and hardware requirements. Putting all of these responsibilities inside the inference engine can couple policy changes to runtime deployments.

The scheduler decides what runs next based on current GPU memory and the progress of active requests. Static batching collects requests for a fixed batch, but when output lengths vary, short requests may wait for the longest request to finish. Continuous batching rebuilds the batch at token-generation steps by removing completed requests and adding queued ones.

The inference engine manages not only model weights but also temporary activations, the KV cache, and kernel buffers. Memory allocation and kernel scheduling interact: reserving too much cache can prevent new requests from being admitted, while reserving too little can reduce throughput.

B. Prefill and decode

sequenceDiagram
  participant C as Client
  participant S as Scheduler
  participant M as Model engine
  participant K as KV cache
  C->>S: Submit prompt
  S->>M: Assign prefill work
  M->>M: Compute attention over input tokens
  M->>K: Store layer Key Value states
  M-->>C: Stream first token TTFT
  loop Generate output tokens
    S->>M: Execute decode step
    M->>K: Read past Key Value
    M->>M: Compute next token
    M-->>C: Deliver token ITL
  end
  M->>K: Reclaim cache or apply retention policy

Prefill processes the prompt tokens and builds KV states for each layer. Long inputs require substantial computation and activation memory, so GPU compute utilization may be high. A long prompt also creates a large cache and competes for resources with other users' decode work.

Decode is the autoregressive stage that computes one next token at a time, including the previously generated tokens. Each step performs attention over the context and repeatedly uses model weights, so memory bandwidth or sequential per-token latency can often become the bottleneck rather than arithmetic throughput. The actual bottleneck varies with workload and model and must be measured.

Time to first token (TTFT) is strongly affected by input queueing and prefill. Inter-token latency (ITL or TPOT) is sensitive to decode speed and scheduling. Measuring the two separately distinguishes TTFT degradation from long prompts from ITL degradation caused by slow token generation.

3. KV cache and memory management

A. Cache principle and size

Transformers produce Key and Value representations for attention at each token. Reusing earlier K and V means that their projections do not need to be recomputed when generating a new token. The system must still read earlier K and V and calculate attention with the new Query, so caching does not eliminate all computation.

For a model using full attention, the cache memory for one request can be estimated as follows:

KV-cache bytes ≈ 2 × number of layers (L) × number of KV heads (Hₖᵥ) × head dimension (d) × token count (T) × bytes per element (b) × concurrent requests (N)

The factor 2 accounts for storing both Key and Value. Grouped-query attention (GQA) and multi-query attention (MQA) can reduce the cache because they use fewer KV heads than Query heads. Longer contexts, higher concurrency, and higher precision increase cache capacity requirements.

For example, a hypothetical model with 32 layers, 8 KV heads, a head dimension of 128, an 8,192-token context, and FP16 (2 bytes per element) requires approximately 1 GiB of cache for one request. The actual value is an estimate that varies with attention structure, windows, cache precision, and runtime implementation.

This is separate from weight memory. Large models may use most GPU memory for weights; once cache and activations are included, request concurrency is constrained. Serving feasibility must therefore not be judged by whether model weights alone fit.

B. Paged cache and sharing

Allocating per-request cache in contiguous memory can create external fragmentation when request lengths vary; reserving a maximum length in advance wastes unused space. A paged approach divides the cache into fixed-size blocks and links only the blocks needed as a request generates tokens. The key idea is to separate logical token order from physical memory blocks using a mapping table.

PagedAttention-style approaches are associated with this block-oriented KV memory layout. Smaller blocks can reduce wasted space but may increase block-management and address-translation overhead. Larger blocks may simplify management, while changing internal waste at the end of a sequence and the granularity of sharing.

When many requests share a system prompt or repeated RAG instructions, sharing the cache for a common prefix can reduce duplicate prefill. Cache sharing requires a key design that verifies not only content identity but also matching tenant, authorization, and model version.

An eviction policy decides which blocks to release when GPU memory is scarce. It may use recency (LRU), priority, reuse likelihood, or retention time. Keeping one user's data longer to improve hit rates may conflict with security and deletion policies, so cache behavior must be defined together with operational policy.

C. Compression, quantization, and offloading

KV-cache quantization lowers the bit width used to represent cached states and can reduce memory use. Lower-precision representations such as FP8 may allow more tokens or requests, but quantization scales and value ranges can affect output quality or stability. Weight quantization and cache quantization should therefore have separate quality-validation criteria.

Offloading moves less frequently used cache blocks to CPU memory or another memory tier. It frees GPU capacity but consumes transfer latency, CPU memory, and PCIe or network bandwidth. If the system cannot predict when an offloaded block will be needed again, retrieving it may be slower than recomputing it.

Models with sliding-window attention may reclaim tokens earlier when some layers do not need an unlimited context. Hybrid attention or stateful architectures can have different cache lifetimes and pool sizes across layers. Uniformly shrinking the cache without understanding model-specific layouts can break accuracy or execution compatibility.

4. Batching and scheduling

A. Static and continuous batching

Static batching runs a fixed group of requests. It is simple to implement and operate for offline jobs with similarly sized inputs, but if request generation lengths vary, the whole batch is often bound by the longest request. Empty slots may remain unavailable until the batch finishes.

Continuous or in-flight batching removes completed requests and admits queued ones at generation-step boundaries. It separates request progress and can reduce GPU idle time, increasing throughput for continuously arriving requests. However, scheduler admission and fairness become important, and long prefill work must not indefinitely delay decode requests.

Continuous batching does not automatically reduce latency. Adding requests may improve throughput while making each request wait longer for a generation step. Measure queueing delay and p95/p99 TTFT and ITL rather than relying on average throughput alone.

B. Chunked prefill and scheduling policy

Chunked prefill divides a long prompt into token chunks and processes them incrementally. Prefill work can be interleaved with decode requests, reducing the time that one long prompt monopolizes other users' token generation. Chunk transitions and scheduling add overhead, and total prefill completion time may change.

A scheduler may consider token budgets, free cache blocks, request priority, deadlines, and estimated output length. A throughput-first policy may serve short requests quickly but starve long ones. Fairness policies can apply tenant weights or maximum wait times to reduce service imbalance.

Increasing the maximum batch-token budget and maximum active sequence count can raise parallelism but rapidly consume cache and activation memory. Excessive admission of small requests can repeatedly reject later long prompts. Admission control should respond proactively to queue delay and memory headroom, not only after errors occur.

5. Model execution optimizations

A. Quantization and kernel optimization

Weight quantization stores and computes model weights at lower precision than FP16 or BF16, reducing capacity and bandwidth pressure. Four-bit or eight-bit formats do not guarantee the same quality across all models and tasks, so domain-specific evaluation sets should test accuracy, hallucinations, and safety changes.

Fused kernels and optimized attention kernels combine operations and reduce memory round trips for intermediate data. Their effect depends on GPU architecture, input length, batch shape, and library version. When changing a serving engine, also evaluate operator coverage, debuggability, operational maturity, and vendor dependence.

B. Parallelism and model placement

Data parallelism places model replicas on multiple GPUs to handle different requests and is useful for scaling throughput. Each replica needs another copy of the weights, so it is difficult to apply when the model does not fit on one GPU. More replicas can also reduce per-GPU cache headroom, which affects long-context capacity.

Tensor parallelism partitions layer matrix operations across GPUs and can fit a larger model, but increases inter-GPU communication. Adding GPUs may split computation while communication and synchronization increase latency. Pipeline parallelism places groups of layers on different GPUs but requires consideration of pipeline bubbles and request scheduling.

Choose parallelism based not only on parameter count but also on target latency, network topology, batch size, and model structure. Mixture-of-Experts (MoE) models may add expert parallelism and token routing. When combining strategies, check that fault isolation and deployment complexity do not exceed the value of performance tuning.

C. Speculative decoding

Speculative decoding uses a smaller draft model or auxiliary predictor to propose several candidate next tokens, then verifies them together with the larger target model. If candidates are accepted, multiple tokens may be committed in one verification step, reducing the number of sequential decode rounds.

Algorithms can accept or correct proposals while preserving the target model's distribution, but a poor draft match adds verification overhead. Measure draft-generation cost, acceptance rate, batch behavior, and memory use. Do not assume speculative decoding provides the same acceleration for every request.

6. Comparison and selection

Technique Primary bottleneck Expected effect Cost or caution
KV cache Repeated K/V computation for past tokens Less redundant decode work More cache memory
Paged cache Fragmentation and fixed-reservation waste Dynamic block allocation and sharing More mapping and eviction complexity
Continuous batching Empty slots and waiting in static batches Better GPU utilization and throughput Queue fairness and tail latency
Chunked prefill Long inputs blocking decode Scheduling between prefill and decode Chunk-transition overhead
Cache quantization KV memory capacity More concurrency or context capacity Precision and quality validation
Speculative decoding Sequential per-token generation delay Faster commitment of accepted token groups Draft cost and acceptance-rate dependence
Weight quantization Weight memory and bandwidth Better model fit and execution efficiency Possible task-specific quality loss

These techniques address different bottlenecks and are complementary rather than direct substitutes. Continuous batching improves execution-slot use while a paged cache improves KV-memory placement, so both can be applied together. However, increasing the maximum batch size first may worsen cache pressure and reverse the benefit.

Select techniques in stages using the service profile. Short-prompt, short-response workloads may benefit from a smaller model or more replicas, while long-context RAG services may depend on cache capacity and reusable prefixes. Interactive services may prioritize TTFT and ITL SLOs, while offline inference may prioritize tokens per hour and cost per token.

7. Industry use cases

A. Customer-support conversation service

A customer-support system applies a common system prompt and policy instructions to many requests. Safely reusing that prefix can reduce duplicate prefill, while customer-specific conversation history must use cache keys that prevent cross-tenant mixing.

During an evening traffic spike, continuous batching and queue limits can be applied, with weighted scheduling to prevent high-priority support work from being crowded out by general questions. Observability should include p95 TTFT, p95 ITL, response success rate, queue-overflow rejection rate, and GPU memory headroom.

B. Internal RAG knowledge search

Internal knowledge search uses the same system instructions and retrieval-output format for each query, while retrieved document length varies. Prefix caching and chunked prefill can reduce duplicate work and limit resource monopolization by long contexts.

In organizations with sensitive documents, cache reuse should be scoped by organization and authorization, and cache invalidation should be considered when employees leave, permissions change, or documents are deleted. Access control and information separation take precedence over cache hit rate; prompt and result logs should follow data minimization.

C. Code-generation assistant

Code generation mixes long file context with short iterative completion requests. Short suggestions are sensitive to first-token and inter-token latency, while long analysis requests care more about prefill cost and cache capacity. Separate queues or model-engine pools by request type can isolate conflicting SLOs.

Speculative decoding with a draft model can be tested in code domains where draft and target token distributions align. Before deployment, compare code correctness, insecure-code generation, acceptance rate, mean latency, and inference cost against a baseline using the same input distribution.

8. Performance measurement and operations

Benchmark with concurrent requests, input and output length distributions, and arrival rates rather than one isolated average request. If production logs cannot be reused, generate representative workloads with personal information removed while preserving prompt- and response-length distributions.

Core measures include TTFT, ITL/TPOT, total request latency, input and output token throughput, queue delay, success, timeout and rejection rates, GPU utilization, and weight and KV-cache memory. Average throughput during low usage can hide SLO violations under traffic spikes.

Cost measures include not only GPU time but also model size, idle replicas, CPU offloading, network transfer, retries, and post-processing caused by quality degradation. If lower cost per token reduces accuracy or safety and increases human rework, total cost of ownership may increase.

An operational sequence should include baseline measurement, applying one optimization at a time, quality regression testing, replaying load, canary deployment, comparing SLOs, and gradual rollout. Record runtime, kernel, and tokenizer versions together, and check whether gains are limited to a specific GPU or input length.

9. Advanced topic: separating prefill and decode, reusing cache

Because prefill and decode have different resource profiles, large systems may consider separating them into distinct work pools or hardware. The prefill pool can be optimized for long inputs, while the decode pool retains cache and streams tokens.

Separation enables independent scaling of each stage but requires transferring KV state between machines. If the data volume is large, network bandwidth and copy latency may outweigh the benefits, while routing and failure recovery become more complex. Compare end to end, including cache-transfer time, prefill and decode queue delay, and network contention.

Combining prefix caching with external routing can direct requests to an engine that already holds a prefix. This is a problem of balancing cache affinity with workload distribution across the cluster. If too many requests target one engine, hotspots form; balance reuse benefits against per-engine queue length.

Cross-request cache reuse is a security boundary. Design strong hash-based cache keys, tenant and user isolation, retention periods for sensitive prompts, and eviction and deletion evidence. Include cache side channels and prompt-inference risks in the threat model. Actual features vary by product and version, so verify official documentation against the deployed release.

10. Considerations and implications

  1. Set objectives from the SLO: Interactive services must optimize tail TTFT and ITL, error rate, and fairness rather than throughput alone. Define acceptable latency and maximum queue wait by business priority.

  2. Manage the combined memory budget: Manage model weights, activations, KV cache, and runtime buffers within one GPU-memory budget. Bound context length and concurrency together, and monitor cache use and safety headroom before OOM.

  3. Validate quality and safety regressions: Pair low-precision weights or KV cache and speculative decoding with representative task evaluation and safety testing. Check for degradation in specific languages, high-risk work, or long contexts, not only average scores.

  4. Isolate tenants and data lifecycles: Constrain prefix and cache sharing to authorization boundaries and retention policy. Document how long sensitive data remains in cache, how deletion requests are applied, and how logs differ from cache.

  5. Control fairness and load: Protect high-priority latency without starving ordinary requests through maximum wait times, weights, and admission control. When demand exceeds capacity, provide clear retry guidance and fallback responses.

  6. Optimize end-to-end cost: Do not judge improvement solely by GPU cost per token. Compare unit-work cost including quality, retries, cache transfer, idle resources, and operational complexity.

  7. Manage runtime and vendor changes: Cache layouts, scheduling features, and quantization formats vary by runtime version. Before changing engines or GPU generations, revalidate compatibility, performance, and recovery procedures with standard workloads.

An information-management professional should go beyond model choice and GPU sizing to connect request profiles, cache budgets, batching policies, and information-security controls into one capacity plan derived from service SLOs. This turns short-term performance gains into sustainable service quality and cost efficiency.

References


In one line: LLM inference optimization jointly designs KV caching, batching, and model execution for SLOs, quality, security, and cost, then validates the design under representative load.