MachinoAI explainer / LLM Quantization
LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
An outlier-aware INT8 matrix multiplication method that enables large Transformer inference with near-full-precision accuracy.
Tim Dettmers, Mike Lewis, Younes Belkada, Luke ZettlemoyerAug 15, 2022NeurIPS 202212 min read
FOUNDATIONALLLM QuantizationAdvanced
01
Abstract
- The paper introduces LLM.int8(), an INT8 matrix-multiplication procedure for Transformer language models at billion-parameter scale.
- It targets feed-forward and attention-projection matrix multiplications.
- The method cuts inference memory roughly in half while preserving full-precision performance.
- A 175B-parameter 16/32-bit checkpoint can be converted to INT8 and used directly without performance degradation.
- The method combines vector-wise quantization with mixed-precision decomposition.
- More than 99.9% of values are multiplied in 8-bit.
- The paper demonstrates inference up to 175B parameters.
- OPT-175B and BLOOM-176B are highlighted as large-scale examples.
02
Introduction
- For models at and beyond 6.7B parameters, feed-forward and attention projection layers account for about 95% of parameters.
- Those layers account for roughly 65–85% of computation.
- Earlier 8-bit Transformer quantization methods could reduce memory but degrade performance.
- Prior studies largely focused on models below 350M parameters.
- The paper identifies degradation-free multi-billion-parameter quantization as an open challenge.
- The central difficulty is preserving precision in the presence of emergent outlier features.
- The paper studies why conventional quantization fails as model size increases.
03
Problem
- The problem is to perform Transformer matrix multiplication in INT8 without meaningful accuracy degradation.
- A single tensor-wide scale can let one extreme value reduce precision for many ordinary values.
- More localized scaling is needed to reduce quantization error.
- Vector-wise quantization alone becomes insufficient when systematic feature outliers emerge.
- These outliers are sparse but disproportionately important to model behavior.
- The method must preserve high precision for those dimensions while keeping most computation in INT8.
- The solution must remain memory efficient for very large models.
04
Background
- LLM.int8() has two principal components: vector-wise quantization and mixed-precision decomposition.
- Vector-wise quantization assigns separate normalization constants to input rows and weight columns.
- INT8 matrix multiplication is accumulated in INT32.
- The INT32 output is dequantized using the corresponding row/column scale factors.
- Mixed-precision decomposition isolates outlier feature dimensions.
- Outlier dimensions are multiplied in FP16.
- Non-outlier dimensions are multiplied in INT8.
- The two outputs are accumulated in FP16.
05
Methodology
- The method starts from FP16 inputs and FP16 weights.
- Row-wise and column-wise absolute maxima provide vector-wise scaling constants.
- Quantized values are mapped into the INT8 range.
- INT8 multiplication produces INT32 accumulations.
- Dequantization uses the outer product of input-row and weight-column normalization constants.
- Outlier dimensions are detected using a hidden-state magnitude threshold.
- The paper finds alpha = 6.0 sufficient to reduce degradation close to zero.
- Outlier and regular matrix products are combined into FP16 outputs.
Figure 1
Figure 2
Figure 3
Figure 4
06
Architecture
- Quantization is applied primarily to feed-forward and attention-projection linear layers.
- Matrix multiplication is treated as a sequence of independent inner products.
- Vector-wise scaling uses one constant per input row and one per weight column.
- Dequantization is equivalent to applying an outer product of those constants.
- Outlier feature dimensions are selected across the hidden dimension.
- The same feature dimensions can contain outliers across many sequence positions.
- The implementation was released through bitsandbytes and integrated with Hugging Face Transformers.
07
Dataset
- Language-model experiments use dense autoregressive Transformers from 125M to 13B parameters.
- The models were pretrained in fairseq.
- Training corpora include Books, English Wikipedia, CC-News, OpenWebText, CC-Stories, and English CC100.
- C4 validation perplexity is used to measure quantization degradation.
- C4 is a subset of Common Crawl.
- NVIDIA A40 GPUs were used for the language-model evaluation.
- OPT models are used for zero-shot downstream evaluation.
- Zero-shot evaluation uses the EleutherAI language-model evaluation harness.
08
Training
- Baselines include tensor-wise absmax quantization.
- Zeropoint quantization is evaluated.
- Row-wise quantization is evaluated.
- Vector-wise quantization is evaluated.
- Vector-wise quantization plus decomposition is evaluated as LLM.int8().
- FP32/FP16 baselines provide the full-precision reference.
- The comparisons are designed to expose scaling failures rather than only single-model accuracy.
09
Experiments
- On C4, LLM.int8() maintains a favorable scaling trend from 125M through 13B parameters.
- At 13B, FP32 perplexity is 12.45 and LLM.int8() absmax is 12.45.
- LLM.int8() zeropoint at 13B also reaches 12.45 perplexity.
- Plain INT8 absmax reaches 19.08 perplexity at 13B.
- Plain row-wise INT8 reaches 16.49 perplexity at 13B.
- Vector-wise INT8 reaches 16.48 perplexity at 13B.
- LLM.int8() restores approximately full-precision perplexity.
- It maintains full 16-bit zero-shot performance across OPT models up to 175B.
10
Baselines
- Conventional absmax, row-wise, and zeropoint quantization show worsening scaling behavior.
- Zeropoint helps with asymmetric distributions but does not solve the large-model outlier problem.
- Vector-wise quantization improves over global scaling but remains insufficient at larger scales.
- Adding mixed-precision decomposition restores lost precision.
- The experiments show that outlier handling is a critical design issue.
- For BLOOM-176B, the reported model-memory reduction is 1.96×.
- The method therefore preserves quality while substantially reducing parameter memory.
11
Results
- Large-magnitude features emerge systematically as Transformers scale.
- Features can reach magnitudes up to about 20× larger than other dimensions.
- Outliers first appear in roughly 25% of layers at smaller scales.
- Around 6.7B parameters, a phase shift occurs.
- At that scale, all Transformer layers are affected.
- Approximately 75% of sequence dimensions can be affected by extreme-magnitude features.
- Around 150,000 outliers can occur per sequence at 6.7B scale.
- Those outliers are concentrated in only about six feature dimensions.
12
Ablation
- Outlier features represent only about 0.1% of input features but strongly affect model behavior.
- Zeroing the identified outlier dimensions can reduce top-1 attention softmax probability mass by more than 20%.
- Removing them can increase validation perplexity by roughly 600–1000%.
- Removing an equal number of random features reduces probability mass by at most about 0.3%.
- Random feature removal increases perplexity by only about 0.1% in the reported comparison.
- The contrast shows that the outliers are systematic and important rather than ordinary numerical noise.
- Preserving those dimensions in FP16 enables aggressive quantization elsewhere.
13
Limitations
Error Analysis
- Outlier feature dimensions contain at least one hidden-state value above the selected threshold.
- The reported threshold alpha = 6.0 works effectively across the studied Transformers.
- For models up to 13B, the number of outlier dimensions is no larger than seven.
- The additional high-precision computation consumes about 0.1% extra memory.
- More than 99.9% of values remain in INT8 multiplication.
- This sparse FP16 path preserves the memory benefit of quantization.
- The decomposition operates at matrix-multiplication time rather than retraining the model.
- This enables immediate post-training conversion.
14
Conclusion
- LLM.int8() can introduce inference overhead for models below about 6.7B parameters.
- Smaller models commonly fit on available GPUs, reducing the practical need for quantization.
- For large matrix multiplications comparable to those in 175B models, the paper reports about 2× speedup in relevant measurements.
- End-to-end inference runtime is maintained for large models such as BLOOM-176B in the reported experiments.
- Memory reduction is the primary objective rather than universal latency improvement.
- The software was open-sourced.
- Hugging Face Transformers integration made the method broadly usable for linear layers.
15
References
- The paper establishes degradation-free INT8 inference at multi-billion and 175B scale.
- It identifies systematic activation outliers as the key reason naive INT8 methods fail.
- Vector-wise quantization provides finer-grained scaling for ordinary values.
- Mixed-precision decomposition preserves high precision for sparse important feature dimensions.
- The result retains roughly half the memory of FP16 representation for the quantized matrix components.
- The work demonstrates that successful LLM quantization requires understanding model structure.
- It became foundational for later outlier-aware and mixed-precision LLM quantization research.
- The implementation connected the research result directly to practical LLM deployment.
Continue reading