MachinoAI explainer / LLM Quantization
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
A one-shot, approximate second-order post-training quantization method that compresses 175B-scale language models to 3–4 bits with minimal accuracy loss.
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, Dan AlistarhOct 31, 2022ICLR 202315 min read
FOUNDATIONALLLM QuantizationAdvanced
01
Abstract
- GPTQ is a one-shot post-training weight quantization method based on approximate second-order information.
- It targets OPT and BLOOM at hundreds-of-billions parameter scale.
- OPT-175B and BLOOM-176B can be quantized to 3–4 bits in roughly four GPU hours.
- The reported 3–4 bit models retain near-baseline language-modeling accuracy.
- GPTQ enables OPT-175B generative inference on a single 80GB NVIDIA A100 at 3-bit precision.
- The paper also demonstrates useful 2-bit and ternary regimes.
- Custom kernels provide about 3.25× A100 and 4.5× A6000 speedups for 3-bit OPT-175B generation.
02
Introduction
- GPT3-175B contains about 175 billion parameters and occupies 326GB in FP16.
- FP16 inference therefore exceeds a single high-end GPU's capacity.
- Full retraining or fine-tuning is prohibitively expensive at this scale.
- Post-training quantization compresses pretrained models without retraining.
- Earlier accurate PTQ methods were difficult to scale to billions of parameters.
- ZeroQuant, LLM.int8(), and nuQmm primarily relied on round-to-nearest weight quantization.
- Round-to-nearest works for mild compression but fails at aggressive 3–4 bit precision.
- GPTQ asks whether one-shot PTQ can remain accurate at much higher compression.
- The paper targets OPT-175B and BLOOM-176B, the largest publicly available models at the time.
03
Problem
- The goal is to compress massive pretrained language models without retraining them.
- Independent round-to-nearest decisions ignore correlations induced by layer inputs.
- Second-order error compensation can improve accuracy but traditionally scales poorly.
- OBQ has cubic input-dimension dependence for a layer matrix.
- A scalable method must preserve second-order compensation while reducing computation.
- Low-arithmetic-intensity updates can become GPU memory-bandwidth bottlenecks.
- Numerical instability becomes increasingly important as model size grows.
- The quantizer must also support practical low-bit inference.
04
Background
- GPTQ formulates quantization layer-by-layer as reconstruction of full-precision layer outputs.
- For weights W and calibration inputs X, it minimizes ||WX − W-hat X||².
- The quantization grid is fixed before optimization while weights can move through error compensation.
- Optimal Brain Quantization (OBQ) is GPTQ's direct precursor.
- OBQ quantizes one weight at a time and updates remaining weights to compensate its error.
- OBQ uses H_F = 2 X_F X_F^T and its inverse Hessian.
- OBQ updates its inverse Hessian through Gaussian-elimination-style row/column removal.
- OBQ has O(d_row·d_col³) runtime and becomes impractical for giant layers.
- GPTQ retains the reconstruction objective but redesigns execution for LLM scale.
05
Methodology
- GPTQ observes that exact greedy ordering provides little benefit on large heavily parameterized layers.
- It quantizes all rows in the same fixed column order.
- All rows share the same inverse-Hessian state because the Hessian depends on layer inputs.
- Hessian updates therefore occur once per column rather than once per weight.
- Runtime changes from O(d_row·d_col³) to O(max{d_row·d_col²,d_col³}).
- Lazy batch updates process B=128 consecutive columns at a time.
- Local updates stay within the current 128-column block before global updates.
- Lazy batching addresses GPU memory-throughput pressure and gives about an order-of-magnitude practical speedup.
- A Cholesky reformulation provides numerically stable Hessian-inverse information.
- The implementation uses damping equal to 1% of the average Hessian diagonal.
- The full algorithm quantizes columns, records errors, updates the active block, then propagates errors to remaining columns.
Figure 1
Figure 2
Figure 3
Figure 4
06
Architecture
- GPTQ is a layer-wise weight-only quantization pipeline, not a new Transformer architecture.
- Calibration examples are propagated through the model to obtain layer inputs.
- One Transformer block is loaded into GPU memory at a time.
- Each reported Transformer block contains 6 layers in the quantization procedure.
- Layer Hessians are accumulated from calibration activations.
- Low-bit weights retain scale/zero-point metadata.
- After a block is quantized, its outputs become inputs for the next block's quantization.
- Inference dynamically dequantizes weights for matrix-vector products.
- Activations remain full precision; GPTQ does not require activation quantization.
- Custom quantized-matrix/full-precision-vector kernels target autoregressive memory bandwidth.
07
Dataset
- Calibration uses the C4 dataset.
- The calibration set contains 128 randomly sampled segments.
- Each segment contains 2048 tokens.
- The calibration text comes from randomly crawled websites and represents generic text.
- No task-specific calibration data is used for the main GPTQ procedure.
- WikiText2 is a primary language-modeling evaluation benchmark.
- Penn Treebank and C4 are also used for perplexity.
- C4 evaluation is not fully zero-shot because calibration samples come from C4 training data.
- LAMBADA, ARC Easy/Challenge, and PIQA are used for zero-shot evaluation.
08
Training
- GPTQ does not retrain or fine-tune the foundation model.
- Quantization is performed after pretrained weights are available.
- The method is post-training quantization rather than quantization-aware training.
- The implementation is written in PyTorch.
- Hugging Face integrations are used for BLOOM and OPT.
- All reported model quantization experiments, including 175B models, use one NVIDIA A100 80GB GPU.
- Calibration requires only 128×2048-token C4 segments.
- One Transformer block is quantized and then used to produce inputs for the next block.
- The main setup uses uniform per-row asymmetric quantization on a min-max grid.
- Group-wise quantization can be added without changing the core error-compensation procedure.
09
Experiments
- Experiments compare GPTQ against accurate but expensive PTQ methods on smaller models.
- Runtime scaling is measured on large OPT and BLOOM models.
- Complete OPT and BLOOM families are evaluated at 3-bit and 4-bit precision.
- WikiText2, PTB, and C4 are used for language-modeling perplexity.
- Zero-shot evaluations include LAMBADA, ARC, and PIQA.
- Small vision experiments use ResNet18 and ResNet50.
- Smaller language-model comparisons include BERT-base and OPT-125M.
- The largest experiments focus on OPT-175B and BLOOM-176B.
- Grouped quantization and extreme 2-bit/ternary settings are also tested.
- Runtime and generation latency are measured on A100 and A6000 GPUs.
10
Baselines
- The primary large-model baseline is round-to-nearest (RTN) using the same asymmetric per-row grid.
- RTN represents the simple weight quantization strategy used by earlier large-LLM PTQ systems.
- AdaRound is used in small vision-model comparisons.
- AdaQuant is included as a faster prior PTQ baseline.
- BRECQ is included as a reconstruction-based PTQ baseline.
- OBQ is GPTQ's direct second-order predecessor.
- ZeroQuant-LKD is referenced for large-model runtime comparison.
- LLM.int8() is discussed as a prior 8-bit large-model quantizer.
- GPTQ is competitive with accurate PTQ methods on small models while being much faster.
- The key large-model comparison is GPTQ versus RTN because prior sophisticated PTQ methods did not scale adequately.
11
Results
- GPTQ quantizes OPT-175B in 4.2 hours on one A100 and BLOOM-176B in 3.8 hours.
- OPT-13B, 30B, and 66B require 20.9 minutes, 44.9 minutes, and 1.6 hours.
- BLOOM-1.7B, 3B, and 7.1B require 2.9, 5.2, and 10.0 minutes.
- On OPT-175B WikiText2, FP16 perplexity is 8.34, GPTQ-4bit is 8.37, and RTN-4bit is 10.54.
- On OPT-175B WikiText2, GPTQ-3bit reaches 8.68 while RTN-3bit is about 7300.
- On BLOOM-176B WikiText2, GPTQ-4bit reaches 8.21 versus 8.37 RTN and 8.11 FP16.
- On BLOOM-176B WikiText2, GPTQ-3bit reaches 8.64 while RTN-3bit is about 571.
- GPTQ-4bit on OPT-175B is only 0.03 perplexity above FP16 on WikiText2.
- At 3-bit, GPTQ typically loses only about 0.3–0.6 perplexity points on the largest models.
- Group size 1024 adds about 0.02 bits and improves perplexity by about 0.2 on average.
- Group size 128 adds about 0.15 bits and improves perplexity by another roughly 0.1.
- 3-bit OPT-175B uses about 63GB including FP16 embeddings/output layer, plus about 9GB for 2048-token KV history.
- It fits on one 80GB A100 versus five A100s for FP16 and three for LLM.int8().
- Average per-token latency drops from 230ms FP16 on five A100s to 71ms on one A100 at 3-bit.
- On two A6000 GPUs, latency drops from 589ms FP16 on eight GPUs to 130ms at 3-bit.
- The reported speedups are about 3.24× on A100 and 4.53× on A6000.
12
Ablation
- GPTQ is compared with AdaRound, AdaQuant, BRECQ, and OBQ on ResNet18 and ResNet50.
- On ResNet18 at 4-bit, GPTQ reaches 69.37% versus 69.34% AdaRound and 69.56% OBQ.
- On ResNet18 at 3-bit, GPTQ reaches 67.88%, above AdaQuant's 59.21%.
- On ResNet50 at 4-bit, GPTQ reaches 75.71%, close to BRECQ 75.88% and OBQ 75.72%.
- On ResNet50 at 3-bit, GPTQ reaches 74.87% versus AdaQuant 64.98%.
- On BERT-base and OPT-125M, GPTQ and greedy OBQ are similar at 4-bit and GPTQ is slightly better at 3-bit.
- Fixed arbitrary ordering loses little versus greedy ordering on large layers.
- Lazy batch updates provide about an order-of-magnitude practical speedup on very large models.
- Cholesky reformulation addresses failures caused by indefinite inverse-Hessian updates.
- Group sizes 1024 and 128 improve 3-bit large-model accuracy with small metadata overhead.
- Group size 32 at around 2.6 bits gives a 0.6–0.7 perplexity increase on the largest models.
13
Limitations
Error Analysis
- The dominant RTN failure at 3-bit is catastrophic perplexity degradation.
- GPTQ reduces this by compensating each quantization error through correlated remaining weights.
- The benefit of second-order information is strongest under aggressive quantization.
- Repeated inverse-Hessian updates can accumulate numerical errors.
- The inverse Hessian can become indefinite on large models.
- Indefinite Hessians can drive remaining weights in incorrect directions and produce arbitrarily bad layers.
- Damping is adequate for smaller models but less robust for the largest models.
- OPT-66B is an exception to the general trend that larger models are easier to quantize.
- The authors associate the OPT-66B anomaly with many dead units in early layers.
- At 4-bit, RTN can degrade substantially while GPTQ stays close to FP16.
- At 3-bit, RTN frequently collapses to extremely large perplexity values.
- Quantization difficulty depends on model scale, architecture, and granularity, not bitwidth alone.
14
Conclusion
- GPTQ demonstrates accurate one-shot quantization at hundreds-of-billions-parameter scale.
- Approximate second-order error compensation becomes scalable through fixed ordering, batching, and Cholesky reformulation.
- 3-bit and 4-bit weights preserve language-model quality far better than RTN.
- OPT-175B can run on a single 80GB A100 at 3 bits.
- Custom kernels turn memory savings into practical generation-speed improvements.
- Group-wise quantization extends the accuracy/compression trade-off toward lower average bitwidth.
- Around 2.2-bit and 2.6-bit configurations remain viable in the reported experiments.
- Ternary quantization reaches 9.20 WikiText2 perplexity on OPT-175B with group size 8.
- GPTQ speedups primarily come from reduced memory movement, not fewer arithmetic operations.
- Activation quantization is outside the main scope and remains a future direction.
15
References
- Frantar et al., GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers, arXiv:2210.17323, ICLR 2023.
- Frantar et al., Optimal Brain Compression: A Framework for Accurate Post-Training Quantization and Pruning, 2022.
- Yao et al., ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers, 2022.
- Dettmers et al., LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale, 2022.
- Park et al., nuQmm: Quantized Matmul for Efficient Inference of Large-Scale Generative Language Models, 2022.
- Nagel et al., Up or Down? Adaptive Rounding for Post-Training Quantization, ICML 2020.
- Li et al., BRECQ: Pushing the Limit of Post-Training Quantization by Block Reconstruction, ICLR 2021.
- Hubara et al., Accurate Post Training Quantization with Small Calibration Sets, 2021.
- Zhang et al., OPT: Open Pre-trained Transformer Language Models, 2022.
- Scao et al., BLOOM and the BigScience language-model effort, 2022.
Continue reading