← Back to list
AI & Data
#양자화#모델 경량화#PTQ/QAT#INT8/INT4#LLM 추론 최적화
Last updated · 2026-10-11

Model Quantization

1. Overview

Model Quantization is a model compression technique that converts a neural network's weights and activations from high-precision floating point such as FP32/FP16 into low-bit integers such as INT8/INT4 (or low-precision floating point such as FP8/FP4), reducing the model's memory footprint, computation, inference latency, and power consumption while minimizing the loss of accuracy.

The parameter scale of deep learning models, and large language models (LLMs) in particular, has exploded from hundreds of millions to hundreds of billions, and performance has improved in proportion to scale. This growth in scale has delivered richer expressive power and higher task performance, but it has simultaneously driven up both the cost of "training the model" and the cost of "actually serving it for inference." Inference in particular recurs throughout the service's lifetime, so in cumulative terms it often becomes a greater burden than training.

Yet this growth leads directly to an explosion in deployment cost. Storing a single parameter as a 32-bit floating-point number requires 4 bytes each, so a 70B-parameter model occupies roughly 140GB of memory even in FP16 (2 bytes). This is a scale that can only barely be loaded by ganging together several expensive data-center GPUs, and it is impossible to deploy as-is in memory- and power-constrained environments such as smartphones, automotive ECUs, and industrial edge devices. Quantization emerged precisely to resolve this fundamental dilemma that "performance comes from large models, but deployment must be small."

The idea of quantization itself is not new. Its roots lie in the quantization concept of sampling an analog signal into digital form in signal processing, and in the embedded domain lightweight inference has long been implemented with fixed-point arithmetic. However, as model scale exploded and inference cost came to dominate service operating expenses, quantization was elevated from an "optional optimization" to a "precondition of deployment." Today's commercial LLM serving stacks almost invariably assume low-precision formats by default, and quantization quality translates directly into service unit-cost competitiveness.

The core principle by which quantization works lies in the empirical fact that the weight distribution of a trained neural network is generally concentrated in a narrow interval near zero and does not demand extreme numerical precision. In other words, of the roughly 4.2 billion levels of resolution that 32 bits express, what is actually needed is far fewer levels, so even approximating with a well-designed integer grid preserves most of the model's judgment tendencies. In practice, changing weights from FP32 to INT8 cuts storage to one quarter, improves throughput by 2–4x on hardware that supports integer arithmetic, and yields an even greater perceived speedup in LLM inference, which is dominated by memory-bandwidth bottlenecks. For this reason, quantization has become the de facto standard compression method for on-device AI, sLLM (lightweight language model) serving, and real-time inference services.

2. Basic Principles and Mathematical Structure of Quantization

flowchart LR
  R["Real-valued weights (FP32 distribution)"] --> SC["Compute scale and zero-point"]
  SC --> Q["Quantize: q = round(r/scale) + zero_point"]
  Q --> QI["Integer representation (stored as INT8/INT4)"]
  QI --> DQ["Dequantize: r' = scale * (q - zero_point)"]
  DQ --> O["Reconstructed real value (approximate, with error)"]
  O --> ERR["Quantization error"]

The mathematical essence of quantization is mapping a continuous real-valued interval onto a finite number of integer grid points. Deciding how wide a range of the real distribution to place on the grid is called range setting (clipping/calibration); if the range is set too wide the resolution of common values degrades (clipping error is small but rounding error is large), and if it is set too narrow extreme values are cut off (rounding error is small but clipping error is large). Range setting is therefore an optimization problem of finding the point that minimizes the sum of the two errors, and instead of the simple approach of using the maximum and minimum values directly, techniques that set the threshold by distribution percentiles or by minimizing KL divergence are widely used. When converting a real value r into an integer q, two parameters are used: scale and zero-point. Scale is the resolution that determines how many integer steps one unit of real value corresponds to, and zero-point is the reference point that determines which integer the real value 0 corresponds to. Quantization is computed as q = round(r / scale) + zero_point, and dequantization as r' = scale × (q - zero_point). In this process a quantization error, the difference between the original r and the reconstructed r', inevitably arises during rounding, and how small this error can be suppressed is the central task of every quantization technique.

This range setting is especially sensitive for activations. Weights are fixed once training is complete, so the range need only be set once, but activations have a distribution that changes with every inference depending on the input, so a "typical range" must be estimated in advance using representative calibration data. If the calibration data is far removed from actual production inputs, values that fall outside the range at inference time become frequent and quality degrades, so composing the calibration set becomes a practical pressure point that governs PTQ quality.

Depending on how the zero-point is handled, quantization is divided into symmetric and asymmetric (affine) quantization. Symmetric quantization fixes the zero-point at 0 and assumes the value range is symmetric about 0, so computation is simple and fast, but it wastes representation range when the distribution is skewed to one side. Asymmetric quantization leaves the zero-point free, so it captures skewed distributions (e.g., activations with no negatives such as ReLU outputs) more precisely but adds a zero-point correction operation. In general, a hybrid strategy of applying symmetric quantization to weights that are distributed around 0 and asymmetric quantization to activations that have skewed distributions is widely used.

Depending on whether the grid spacing is kept constant, quantization is also distinguished into uniform and non-uniform quantization. Uniform quantization keeps the real-valued spacing between integer steps all equal, so it pairs well with hardware integer arithmetic and is the most widely used. Non-uniform quantization, by contrast, places the grid densely near the crowded region around 0 and sparsely in the sparse outer region, obtaining smaller error at the same bit width. NF4 (NormalFloat4), adopted by QLoRA, is a representative example of non-uniform quantization in which the grid is designed so that the probability mass of each interval is equal under the assumption that weights follow a normal distribution, minimizing information loss even under the extreme condition of 4 bits.

In addition, precision varies greatly depending on the unit over which the scale is shared. Per-tensor quantization uses one scale for an entire layer, so it is lightweight, but if the value range of a particular channel is unusually wide the overall resolution is dragged toward that channel and error increases. Per-channel quantization, by contrast, assigns a separate scale to each output channel and absorbs the per-channel distribution differences, greatly raising accuracy in weight quantization. For example, the combination of quantizing the weights of convolution/linear layers per-channel and activations per-tensor is widely adopted as a balance point between accuracy and efficiency.

Axis of distinction Type Characteristics Main target
Zero-point handling Symmetric / Asymmetric Symmetric is simple and fast; asymmetric is precise for skewed distributions Weights / Activations
Scale sharing per-tensor / per-channel Per-channel is precise, per-tensor is lightweight Activations / Weights
Precision INT8·INT4·FP16·BF16·FP8·FP4 Lower bits are lighter and faster, with increased error Chosen per serving requirements

3. Types of Quantization — PTQ and QAT

flowchart TB
  M["Pre-trained model (FP32/FP16)"] --> B{"Retraining feasibility and precision target"}
  B -->|"Fast, without retraining"| PTQ["PTQ: Post-Training Quantization"]
  B -->|"Accuracy first, low bits"| QAT["QAT: Quantization-Aware Training"]
  PTQ --> CAL["Estimate range with small calibration data"]
  CAL --> PQ["Fix weight/activation range, then convert to integers"]
  QAT --> FQ["Insert fake quant during training"]
  FQ --> RT["Learn and correct even quantization error via backprop"]
  PQ --> DEP["Deploy lightweight model"]
  RT --> DEP

Two broad streams exist depending on when quantization is applied. PTQ (Post-Training Quantization) applies quantization after the fact to an already-trained model. It requires no full retraining and estimates each layer's value range using only hundreds to thousands of calibration data points to compute scale and zero-point. Because it is very low-cost and fast, it forms the mainstream of current LLM deployment. However, since it cannot correct error through retraining, accuracy degradation can become pronounced in aggressive low bits of 4 bits or below. PTQ is further subdivided into a dynamic method that computes the activation range in real time at inference and a static method that fixes the range in advance at the calibration stage. The dynamic method is easy to implement and accurate but has overhead at inference, while the static method is faster with no extra computation but is sensitive to the representativeness of the calibration data.

QAT (Quantization-Aware Training) inserts fake quantization nodes that simulate quantization into the training process itself, so that the model "knows in advance" it will be represented in low bits and adapts its parameters accordingly. In the forward pass it uses values that have passed through quantization and dequantization, but because the rounding operation is non-differentiable, in the backward pass gradients are approximated and passed through with the STE (Straight-Through Estimator). This way the model reflects quantization error in the loss function and corrects it on its own, so it greatly reduces the loss relative to PTQ, especially for extremely low bits such as INT4 or for accuracy-sensitive tasks. The downside is that it requires the cost of full (or substantial) retraining and a training-data pipeline, so the typical trade-off between cost and accuracy holds.

In practice, a staged strategy of "first try PTQ, and switch to QAT if the target accuracy is not met" is common. In addition, weight-only quantization, which lowers only the weights to low bits while keeping activations at relatively higher bits, is widely used in LLM serving, because LLM inference is dominated by the bandwidth of reading weights from memory rather than by the amount of computation. In other words, lowering just the weights to INT4 greatly eases the memory load and bandwidth burden, improving perceived performance.

An actual quantization project generally proceeds through the following steps.

  1. Define the goal: First fix the deployment hardware, the acceptable accuracy floor (SLA), and the memory/latency budget.
  2. Select the technique: Choose a primary candidate between PTQ (dynamic/static) and QAT based on retraining feasibility and the target bits.
  3. Calibrate/train: PTQ computes the range with representative calibration data, while QAT inserts fake quantization and retrains.
  4. Evaluate/tune: Measure accuracy and latency on an evaluation set, and raise the bits of sensitive layers via mixed precision.
  5. Format/engine conversion: Export to the format required by the target inference engine (such as GGUF) and verify kernel compatibility.
  6. Deploy/monitor: Continuously observe quality degradation due to input-distribution shift in production and set a recalibration cycle.

As a middle ground between the two methods, mixed-precision quantization, which exploits the fact that each layer has a different sensitivity, is treated as important. Rather than lowering every layer equally to 4 bits, keeping sensitive layers that strongly affect output error (e.g., some projection layers of attention, or the first and last layers) at 8 bits while lowering the rest to 4 bits lets you lower the overall average bits while avoiding accuracy collapse. Here, each layer's sensitivity is measured in advance by the loss increase under quantization or by second-derivative (Hessian) based metrics to assign bits differentially, and this sensitivity analysis becomes a key tool for safely pushing aggressive low-bit quantization.

4. LLM-Specific Quantization Techniques and the Outlier Problem

The fundamental reason LLM quantization is trickier than that of ordinary neural networks is the outliers in activations. In certain channels (feature dimensions) of a transformer, activations tens to hundreds of times larger than other channels appear systematically, and in per-tensor quantization matching the range to these enormous values smears the resolution of the vast majority of remaining values and causes error to explode. The interesting point is that these outliers are not random but are concentrated repeatedly and structurally in a particular small set of dimensions, and this paradoxically becomes a clue to the solution that "if you handle only the outliers separately, the rest can be safely lowered to low bits." To solve this phenomenon, several dedicated techniques have been developed as follows.

flowchart TB
  LLM["Large language model (weights and activations)"] --> OUT["Activation outlier problem"]
  OUT --> M1["LLM.int8(): outliers in FP16, rest in INT8 mixed precision"]
  OUT --> M2["SmoothQuant: migrate activation difficulty to weights"]
  OUT --> M3["GPTQ: error-compensating quantization with second-order (Hessian) info"]
  OUT --> M4["AWQ: preserve important channels by activation (per-channel scale)"]
  M3 --> FMT["Package into storage format (such as GGUF), load into inference engine"]
  M4 --> FMT
  M1 --> FMT
  M2 --> FMT

LLM.int8() is a mixed-precision decomposition that computes separately in FP16 only the small number of outlier dimensions among activations that exceed a threshold, and processes most of the rest in INT8, almost eliminating accuracy loss while cutting memory to about half. SmoothQuant migrates the "difficulty" of hard-to-quantize activations, through a mathematically equivalent transformation, toward the relatively easy-to-quantize weights, stably lowering both activations and weights to INT8.

In the weight-only low-bit field, GPTQ and AWQ are de facto standards. GPTQ quantizes weights layer by layer, but uses second-derivative information (a Hessian approximation) to update the still-unquantized remaining weights so as to compensate for the error that arises when quantizing one block, suppressing performance degradation even at 3–4 bits. AWQ (Activation-aware Weight Quantization) starts from the observation that "not all weights are equally important," and applies per-channel scaling that selectively protects a small number of high-importance channels based on activation magnitude. Weights quantized this way are often packaged into a storage format such as GGUF (the llama.cpp family) and loaded into various inference engines. Meanwhile, QLoRA is a technique that freezes the weights to 4-bit NF4 (NormalFloat4) and trains only a small number of low-rank adapters (LoRA), enabling fine-tuning of models with tens of billions of parameters on a single GPU, and is regarded as a representative case combining quantization and efficient fine-tuning.

Another key device these LLM quantization techniques commonly adopt is group-wise quantization. Binding an entire channel with one scale still leaves large error, while assigning a scale to every value hurts storage efficiency, so weights are usually divided into small groups of 32, 64, or 128 and each group is given a separate scale (and zero-point). The smaller the group size the higher the precision, but the overhead bits for storing scales increase, so, as in GGUF's various "k-quant" variants, multiple profiles with different group-size and bit combinations are provided to let users weigh their memory budget against quality. In the end, the actual quality of LLM quantization is determined not only by "how many bits" but also by "at what granularity, and with what outlier handling, it was quantized."

5. Comparison and Real-World Applications

Comparing the characteristics by precision makes the trade-offs of the choice clear. INT8 cuts memory to one quarter versus FP32 while keeping accuracy loss generally under 1%, so it passes as a "safe default." INT4, when combined with GPTQ/AWQ, keeps loss on standard benchmarks at roughly the 1–3%p level, and is evaluated as the practical sweet spot for LLM serving where the memory bottleneck is severe. The reason is that halving the bits exponentially reduces the representable steps and so increases error (INT8 has 256 steps, INT4 has 16 steps), but the redundancy of LLM weights is large so performance holds up to a certain level.

Precision Bits Memory vs FP32 Characteristics/use Accuracy loss (trend)
FP32 32 1x (baseline) Training default, high precision None (baseline)
FP16/BF16 16 1/2 Training/inference mixed precision; BF16 has wide exponent Negligible
FP8 8 1/4 Native support on latest GPUs, close to BF16 Small
INT8 8 1/4 General inference standard, broad hardware support Generally under 1%
INT4 4 1/8 LLM weight-only (GPTQ/AWQ) About 1–3%p
FP4 (NVFP4, etc.) 4 1/8 Low-precision inference on Blackwell-class GPUs Slight, case-dependent

As a concrete case, a 7B-parameter-class LLM occupies about 13–14GB in FP16, which is tight for consumer GPUs, but quantizing to INT4 cuts it to about 3.5–4GB, making it runnable even on ordinary laptops and edge devices. A 70B-class model likewise needs about 140GB in BF16, requiring multiple GPUs, but lowering it to INT4 brings it to around 35–40GB, widening the serving range to a single high-end GPU. In the mobile domain, the Qualcomm Hexagon NPU, Apple Neural Engine, and others accelerate INT8 arithmetic in hardware, so on-device speech recognition, translation, and camera AI operate in real time. In the data center, the NVIDIA Hopper (H100) generation began supporting FP8 natively in the tensor cores, and this maintains quality close to BF16 while raising throughput and power efficiency, becoming a key means of lowering the unit cost of large-scale inference services.

A subtle point to note here is that weight-only INT4 quantization is closer to "reducing the amount of data read from memory" than to "making the computation itself fast with integers." Many LLM inference kernels read in weights stored as INT4 and then dequantize to FP16 just before the operation to perform the multiplication. That is, storage and transfer are in 4 bits but the actual computation is in high precision, so the benefit is more pronounced in small-batch, low-latency serving where memory bandwidth is the bottleneck than in large-batch situations dominated by computation. Conversely, the path of also lowering activations to INT8 to accelerate the integer matrix multiplication itself yields greater effect in compute-intensive workloads. Therefore, "which quantization is faster" varies by workload characteristics, and this must be considered in design.

Quantization has long been deployed in practice in the vision and recommendation fields as well. The MobileNet family used for mobile object recognition, under INT8 quantization, suppresses accuracy drop to around 1%p while greatly reducing inference latency and model size, enabling real-time processing on battery-powered devices. In large-scale recommendation systems too, because embedding tables occupy most of memory, cases of quantizing embeddings to INT8/INT4 to cut serving cost have become commonplace. As this shows, the point where quantization's benefit appears most dramatically is commonly a serving environment where "memory/bandwidth is the bottleneck and many inference requests recur."

6. Deep Dive — Recent Trends in Low-Precision Formats and Hardware Co-evolution

The latest trends in quantization can be summarized as "from integer (INT) to low-precision floating point (FP8/FP4)" and "co-evolution of software techniques and hardware instructions." Past quantization relied mainly on INT8 integer arithmetic, but recently, as tensor cores directly support low-precision floating point such as FP8 of the NVIDIA Hopper generation and FP4 (NVFP4·MXFP4) of the Blackwell generation, it is evolving toward securing a wide dynamic range by exploiting the exponent even at the same bit width. According to industry reports, FP4 inference on the Blackwell family shows several-fold throughput improvement over the previous FP8 generation and a drastic reduction in per-token power consumption, while on some large reasoning models the drop in benchmark scores such as MMLU was reported to stay at about the 0.1%p level (though such figures vary by vendor, specific model, and serving stack, so they cannot be asserted as a general performance improvement).

Along with this trend, the standardization and ecosystem formation of low-precision formats is also advancing. In the past, each vendor had a different bit allocation (exponent/mantissa) for FP8, making interoperability difficult, but recently, through industry consortia, the MX (Microscaling) family of FP8/FP4 formats has been organized, and major inference stacks such as Hugging Face, vLLM, and TensorRT-LLM are converging toward being able to commonly consume GPTQ/AWQ/FP8 checkpoints. This shows that quantization is becoming a standard layer for model exchange and deployment rather than a trick tied to a particular library.

On the research frontier, extremely low-bit binary/ternary quantization is also active. The BitNet family, which represents weights with the three values {-1, 0, +1}, is an attempt to maximize energy efficiency by replacing multiplication with addition, and there are cases reporting performance close to FP16 at a certain scale, though it is still too early to regard as generalized. An important implication from the perspective of an Information Management Professional Engineer is that quantization no longer stays a software post-processing technique but has expanded into a system-level problem in which chip design, instruction sets, compilers, and inference engines are designed together. Therefore, a model lightweighting strategy must become an integrated decision that, beyond the choice of a particular quantization algorithm, also considers which low-precision format the target hardware accelerates natively.

7. Considerations and Implications (Professional Engineer Perspective)

  • Quantitative management of the accuracy–efficiency trade-off: Quantization is not a free lunch but a deal that gives up some accuracy to gain efficiency. Therefore, before adoption one must measure per-bit performance on an evaluation set based on representative tasks and real data, first define the accuracy floor (SLA) the service permits, and then make a data-driven decision to choose the lowest bits within that range. A staged approach of "verified INT8 → QAT/mixed precision if it falls short" is safer than "INT4 right away."

  • Verification of hardware/inference-engine suitability: No matter how good a quantization format is, if the target hardware cannot accelerate that operation, the theoretical gain does not translate into real performance. One must check in advance the precision/instructions supported by the deployment target (data-center GPU, mobile NPU, edge MCU) and the kernel support range of the inference engine (TensorRT, ONNX Runtime, llama.cpp, etc.) and work backward to the format and technique.

  • Calibration data and reproducibility/governance: Since PTQ quality is greatly swayed by the representativeness of calibration data, composing a calibration set that reflects the actual production distribution and managing its provenance/licensing are important. In addition, quantized models should be included in MLOps governance that records version, hash, and evaluation metrics together, so that performance changes due to lightweighting are traceable and auditable.

  • Mutually complementary combination of lightweighting techniques: Quantization is not in an exclusive relationship with knowledge distillation, pruning, or low-rank decomposition but combines orthogonally and complementarily. For example, a multi-stage pipeline that takes a student model made small by distillation, then quantizes it to INT4 and fine-tunes with QLoRA, yields optimal efficiency in practice. A Professional Engineer must be able to design not a single technique but a lightweighting strategy portfolio that synthesizes required performance, cost, and deployment environment.

  • Risk and security considerations: Aggressive low-bit quantization may keep average accuracy but make predictions unstable on a small number of sensitive inputs, or worsen fairness/safety metrics. In high-risk domains such as healthcare, finance, and autonomous driving in particular, the evaluation scope must include edge cases and adversarial robustness before and after quantization.

  • Total cost of ownership (TCO) and sustainability outlook: The ultimate value of quantization lies not in mere model shrinkage but in the reduction of inference unit cost and energy consumption. As generative AI services become large-scale and always-on, inference cost is becoming the dominant item of operating expense, and low-precision arithmetic lowers per-token power consumption to a fraction, reducing even carbon emissions and cooling cost. Therefore, a Professional Engineer must position quantization as part of a green IT and sustainability strategy beyond a performance optimization technique, and be able to explain the effect of the lightweighting investment from a TCO perspective across the model lifecycle.

References


In one line: Model quantization is a core lightweighting technique that converts weights and activations into low-bit integers and low-precision floating point to reduce memory, computation, and power, and it must be designed and validated from the accuracy–efficiency trade-off perspective across PTQ/QAT, LLM-specific techniques such as GPTQ/AWQ/QLoRA, and the hardware co-evolution heading toward FP8/FP4.