Research

MachinoAI explainer / LLM Quantization

LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale

An outlier-aware INT8 matrix multiplication method that enables large Transformer inference with near-full-precision accuracy.

Tim Dettmers, Mike Lewis, Younes Belkada, Luke ZettlemoyerAug 15, 2022NeurIPS 202212 min read
FOUNDATIONALLLM QuantizationAdvanced

01

Abstract

  • The paper introduces LLM.int8(), an INT8 matrix-multiplication procedure for Transformer language models at billion-parameter scale.
  • It targets feed-forward and attention-projection matrix multiplications.
  • The method cuts inference memory roughly in half while preserving full-precision performance.
  • A 175B-parameter 16/32-bit checkpoint can be converted to INT8 and used directly without performance degradation.
  • The method combines vector-wise quantization with mixed-precision decomposition.
  • More than 99.9% of values are multiplied in 8-bit.
  • The paper demonstrates inference up to 175B parameters.
  • OPT-175B and BLOOM-176B are highlighted as large-scale examples.

02

Introduction

  • For models at and beyond 6.7B parameters, feed-forward and attention projection layers account for about 95% of parameters.
  • Those layers account for roughly 65–85% of computation.
  • Earlier 8-bit Transformer quantization methods could reduce memory but degrade performance.
  • Prior studies largely focused on models below 350M parameters.
  • The paper identifies degradation-free multi-billion-parameter quantization as an open challenge.
  • The central difficulty is preserving precision in the presence of emergent outlier features.
  • The paper studies why conventional quantization fails as model size increases.

03

Problem

  • The problem is to perform Transformer matrix multiplication in INT8 without meaningful accuracy degradation.
  • A single tensor-wide scale can let one extreme value reduce precision for many ordinary values.
  • More localized scaling is needed to reduce quantization error.
  • Vector-wise quantization alone becomes insufficient when systematic feature outliers emerge.
  • These outliers are sparse but disproportionately important to model behavior.
  • The method must preserve high precision for those dimensions while keeping most computation in INT8.
  • The solution must remain memory efficient for very large models.

04

Background

  • LLM.int8() has two principal components: vector-wise quantization and mixed-precision decomposition.
  • Vector-wise quantization assigns separate normalization constants to input rows and weight columns.
  • INT8 matrix multiplication is accumulated in INT32.
  • The INT32 output is dequantized using the corresponding row/column scale factors.
  • Mixed-precision decomposition isolates outlier feature dimensions.
  • Outlier dimensions are multiplied in FP16.
  • Non-outlier dimensions are multiplied in INT8.
  • The two outputs are accumulated in FP16.

05

Methodology

  • The method starts from FP16 inputs and FP16 weights.
  • Row-wise and column-wise absolute maxima provide vector-wise scaling constants.
  • Quantized values are mapped into the INT8 range.
  • INT8 multiplication produces INT32 accumulations.
  • Dequantization uses the outer product of input-row and weight-column normalization constants.
  • Outlier dimensions are detected using a hidden-state magnitude threshold.
  • The paper finds alpha = 6.0 sufficient to reduce degradation close to zero.
  • Outlier and regular matrix products are combined into FP16 outputs.
Figure 1
Figure 2
Figure 3
Figure 4

06

Architecture

  • Quantization is applied primarily to feed-forward and attention-projection linear layers.
  • Matrix multiplication is treated as a sequence of independent inner products.
  • Vector-wise scaling uses one constant per input row and one per weight column.
  • Dequantization is equivalent to applying an outer product of those constants.
  • Outlier feature dimensions are selected across the hidden dimension.
  • The same feature dimensions can contain outliers across many sequence positions.
  • The implementation was released through bitsandbytes and integrated with Hugging Face Transformers.

07

Dataset

  • Language-model experiments use dense autoregressive Transformers from 125M to 13B parameters.
  • The models were pretrained in fairseq.
  • Training corpora include Books, English Wikipedia, CC-News, OpenWebText, CC-Stories, and English CC100.
  • C4 validation perplexity is used to measure quantization degradation.
  • C4 is a subset of Common Crawl.
  • NVIDIA A40 GPUs were used for the language-model evaluation.
  • OPT models are used for zero-shot downstream evaluation.
  • Zero-shot evaluation uses the EleutherAI language-model evaluation harness.

08

Training

  • Baselines include tensor-wise absmax quantization.
  • Zeropoint quantization is evaluated.
  • Row-wise quantization is evaluated.
  • Vector-wise quantization is evaluated.
  • Vector-wise quantization plus decomposition is evaluated as LLM.int8().
  • FP32/FP16 baselines provide the full-precision reference.
  • The comparisons are designed to expose scaling failures rather than only single-model accuracy.

09

Experiments

  • On C4, LLM.int8() maintains a favorable scaling trend from 125M through 13B parameters.
  • At 13B, FP32 perplexity is 12.45 and LLM.int8() absmax is 12.45.
  • LLM.int8() zeropoint at 13B also reaches 12.45 perplexity.
  • Plain INT8 absmax reaches 19.08 perplexity at 13B.
  • Plain row-wise INT8 reaches 16.49 perplexity at 13B.
  • Vector-wise INT8 reaches 16.48 perplexity at 13B.
  • LLM.int8() restores approximately full-precision perplexity.
  • It maintains full 16-bit zero-shot performance across OPT models up to 175B.

10

Baselines

  • Conventional absmax, row-wise, and zeropoint quantization show worsening scaling behavior.
  • Zeropoint helps with asymmetric distributions but does not solve the large-model outlier problem.
  • Vector-wise quantization improves over global scaling but remains insufficient at larger scales.
  • Adding mixed-precision decomposition restores lost precision.
  • The experiments show that outlier handling is a critical design issue.
  • For BLOOM-176B, the reported model-memory reduction is 1.96×.
  • The method therefore preserves quality while substantially reducing parameter memory.

11

Results

  • Large-magnitude features emerge systematically as Transformers scale.
  • Features can reach magnitudes up to about 20× larger than other dimensions.
  • Outliers first appear in roughly 25% of layers at smaller scales.
  • Around 6.7B parameters, a phase shift occurs.
  • At that scale, all Transformer layers are affected.
  • Approximately 75% of sequence dimensions can be affected by extreme-magnitude features.
  • Around 150,000 outliers can occur per sequence at 6.7B scale.
  • Those outliers are concentrated in only about six feature dimensions.

12

Ablation

  • Outlier features represent only about 0.1% of input features but strongly affect model behavior.
  • Zeroing the identified outlier dimensions can reduce top-1 attention softmax probability mass by more than 20%.
  • Removing them can increase validation perplexity by roughly 600–1000%.
  • Removing an equal number of random features reduces probability mass by at most about 0.3%.
  • Random feature removal increases perplexity by only about 0.1% in the reported comparison.
  • The contrast shows that the outliers are systematic and important rather than ordinary numerical noise.
  • Preserving those dimensions in FP16 enables aggressive quantization elsewhere.

13

Limitations

Error Analysis

  • Outlier feature dimensions contain at least one hidden-state value above the selected threshold.
  • The reported threshold alpha = 6.0 works effectively across the studied Transformers.
  • For models up to 13B, the number of outlier dimensions is no larger than seven.
  • The additional high-precision computation consumes about 0.1% extra memory.
  • More than 99.9% of values remain in INT8 multiplication.
  • This sparse FP16 path preserves the memory benefit of quantization.
  • The decomposition operates at matrix-multiplication time rather than retraining the model.
  • This enables immediate post-training conversion.

14

Conclusion

  • LLM.int8() can introduce inference overhead for models below about 6.7B parameters.
  • Smaller models commonly fit on available GPUs, reducing the practical need for quantization.
  • For large matrix multiplications comparable to those in 175B models, the paper reports about 2× speedup in relevant measurements.
  • End-to-end inference runtime is maintained for large models such as BLOOM-176B in the reported experiments.
  • Memory reduction is the primary objective rather than universal latency improvement.
  • The software was open-sourced.
  • Hugging Face Transformers integration made the method broadly usable for linear layers.

15

References

  • The paper establishes degradation-free INT8 inference at multi-billion and 175B scale.
  • It identifies systematic activation outliers as the key reason naive INT8 methods fail.
  • Vector-wise quantization provides finer-grained scaling for ordinary values.
  • Mixed-precision decomposition preserves high precision for sparse important feature dimensions.
  • The result retains roughly half the memory of FP16 representation for the quantized matrix components.
  • The work demonstrates that successful LLM quantization requires understanding model structure.
  • It became foundational for later outlier-aware and mixed-precision LLM quantization research.
  • The implementation connected the research result directly to practical LLM deployment.

Continue reading

Related research

Browse all research