Research

MachinoAI explainer / LLM Quantization

GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

A one-shot, approximate second-order post-training quantization method that compresses 175B-scale language models to 3–4 bits with minimal accuracy loss.

Elias Frantar, Saleh Ashkboos, Torsten Hoefler, Dan AlistarhOct 31, 2022ICLR 202315 min read
FOUNDATIONALLLM QuantizationAdvanced

01

Abstract

  • GPTQ is a one-shot post-training weight quantization method based on approximate second-order information.
  • It targets OPT and BLOOM at hundreds-of-billions parameter scale.
  • OPT-175B and BLOOM-176B can be quantized to 3–4 bits in roughly four GPU hours.
  • The reported 3–4 bit models retain near-baseline language-modeling accuracy.
  • GPTQ enables OPT-175B generative inference on a single 80GB NVIDIA A100 at 3-bit precision.
  • The paper also demonstrates useful 2-bit and ternary regimes.
  • Custom kernels provide about 3.25× A100 and 4.5× A6000 speedups for 3-bit OPT-175B generation.

02

Introduction

  • GPT3-175B contains about 175 billion parameters and occupies 326GB in FP16.
  • FP16 inference therefore exceeds a single high-end GPU's capacity.
  • Full retraining or fine-tuning is prohibitively expensive at this scale.
  • Post-training quantization compresses pretrained models without retraining.
  • Earlier accurate PTQ methods were difficult to scale to billions of parameters.
  • ZeroQuant, LLM.int8(), and nuQmm primarily relied on round-to-nearest weight quantization.
  • Round-to-nearest works for mild compression but fails at aggressive 3–4 bit precision.
  • GPTQ asks whether one-shot PTQ can remain accurate at much higher compression.
  • The paper targets OPT-175B and BLOOM-176B, the largest publicly available models at the time.

03

Problem

  • The goal is to compress massive pretrained language models without retraining them.
  • Independent round-to-nearest decisions ignore correlations induced by layer inputs.
  • Second-order error compensation can improve accuracy but traditionally scales poorly.
  • OBQ has cubic input-dimension dependence for a layer matrix.
  • A scalable method must preserve second-order compensation while reducing computation.
  • Low-arithmetic-intensity updates can become GPU memory-bandwidth bottlenecks.
  • Numerical instability becomes increasingly important as model size grows.
  • The quantizer must also support practical low-bit inference.

04

Background

  • GPTQ formulates quantization layer-by-layer as reconstruction of full-precision layer outputs.
  • For weights W and calibration inputs X, it minimizes ||WX − W-hat X||².
  • The quantization grid is fixed before optimization while weights can move through error compensation.
  • Optimal Brain Quantization (OBQ) is GPTQ's direct precursor.
  • OBQ quantizes one weight at a time and updates remaining weights to compensate its error.
  • OBQ uses H_F = 2 X_F X_F^T and its inverse Hessian.
  • OBQ updates its inverse Hessian through Gaussian-elimination-style row/column removal.
  • OBQ has O(d_row·d_col³) runtime and becomes impractical for giant layers.
  • GPTQ retains the reconstruction objective but redesigns execution for LLM scale.

05

Methodology

  • GPTQ observes that exact greedy ordering provides little benefit on large heavily parameterized layers.
  • It quantizes all rows in the same fixed column order.
  • All rows share the same inverse-Hessian state because the Hessian depends on layer inputs.
  • Hessian updates therefore occur once per column rather than once per weight.
  • Runtime changes from O(d_row·d_col³) to O(max{d_row·d_col²,d_col³}).
  • Lazy batch updates process B=128 consecutive columns at a time.
  • Local updates stay within the current 128-column block before global updates.
  • Lazy batching addresses GPU memory-throughput pressure and gives about an order-of-magnitude practical speedup.
  • A Cholesky reformulation provides numerically stable Hessian-inverse information.
  • The implementation uses damping equal to 1% of the average Hessian diagonal.
  • The full algorithm quantizes columns, records errors, updates the active block, then propagates errors to remaining columns.
Figure 1
Figure 2
Figure 3
Figure 4

06

Architecture

  • GPTQ is a layer-wise weight-only quantization pipeline, not a new Transformer architecture.
  • Calibration examples are propagated through the model to obtain layer inputs.
  • One Transformer block is loaded into GPU memory at a time.
  • Each reported Transformer block contains 6 layers in the quantization procedure.
  • Layer Hessians are accumulated from calibration activations.
  • Low-bit weights retain scale/zero-point metadata.
  • After a block is quantized, its outputs become inputs for the next block's quantization.
  • Inference dynamically dequantizes weights for matrix-vector products.
  • Activations remain full precision; GPTQ does not require activation quantization.
  • Custom quantized-matrix/full-precision-vector kernels target autoregressive memory bandwidth.

07

Dataset

  • Calibration uses the C4 dataset.
  • The calibration set contains 128 randomly sampled segments.
  • Each segment contains 2048 tokens.
  • The calibration text comes from randomly crawled websites and represents generic text.
  • No task-specific calibration data is used for the main GPTQ procedure.
  • WikiText2 is a primary language-modeling evaluation benchmark.
  • Penn Treebank and C4 are also used for perplexity.
  • C4 evaluation is not fully zero-shot because calibration samples come from C4 training data.
  • LAMBADA, ARC Easy/Challenge, and PIQA are used for zero-shot evaluation.

08

Training

  • GPTQ does not retrain or fine-tune the foundation model.
  • Quantization is performed after pretrained weights are available.
  • The method is post-training quantization rather than quantization-aware training.
  • The implementation is written in PyTorch.
  • Hugging Face integrations are used for BLOOM and OPT.
  • All reported model quantization experiments, including 175B models, use one NVIDIA A100 80GB GPU.
  • Calibration requires only 128×2048-token C4 segments.
  • One Transformer block is quantized and then used to produce inputs for the next block.
  • The main setup uses uniform per-row asymmetric quantization on a min-max grid.
  • Group-wise quantization can be added without changing the core error-compensation procedure.

09

Experiments

  • Experiments compare GPTQ against accurate but expensive PTQ methods on smaller models.
  • Runtime scaling is measured on large OPT and BLOOM models.
  • Complete OPT and BLOOM families are evaluated at 3-bit and 4-bit precision.
  • WikiText2, PTB, and C4 are used for language-modeling perplexity.
  • Zero-shot evaluations include LAMBADA, ARC, and PIQA.
  • Small vision experiments use ResNet18 and ResNet50.
  • Smaller language-model comparisons include BERT-base and OPT-125M.
  • The largest experiments focus on OPT-175B and BLOOM-176B.
  • Grouped quantization and extreme 2-bit/ternary settings are also tested.
  • Runtime and generation latency are measured on A100 and A6000 GPUs.

10

Baselines

  • The primary large-model baseline is round-to-nearest (RTN) using the same asymmetric per-row grid.
  • RTN represents the simple weight quantization strategy used by earlier large-LLM PTQ systems.
  • AdaRound is used in small vision-model comparisons.
  • AdaQuant is included as a faster prior PTQ baseline.
  • BRECQ is included as a reconstruction-based PTQ baseline.
  • OBQ is GPTQ's direct second-order predecessor.
  • ZeroQuant-LKD is referenced for large-model runtime comparison.
  • LLM.int8() is discussed as a prior 8-bit large-model quantizer.
  • GPTQ is competitive with accurate PTQ methods on small models while being much faster.
  • The key large-model comparison is GPTQ versus RTN because prior sophisticated PTQ methods did not scale adequately.

11

Results

  • GPTQ quantizes OPT-175B in 4.2 hours on one A100 and BLOOM-176B in 3.8 hours.
  • OPT-13B, 30B, and 66B require 20.9 minutes, 44.9 minutes, and 1.6 hours.
  • BLOOM-1.7B, 3B, and 7.1B require 2.9, 5.2, and 10.0 minutes.
  • On OPT-175B WikiText2, FP16 perplexity is 8.34, GPTQ-4bit is 8.37, and RTN-4bit is 10.54.
  • On OPT-175B WikiText2, GPTQ-3bit reaches 8.68 while RTN-3bit is about 7300.
  • On BLOOM-176B WikiText2, GPTQ-4bit reaches 8.21 versus 8.37 RTN and 8.11 FP16.
  • On BLOOM-176B WikiText2, GPTQ-3bit reaches 8.64 while RTN-3bit is about 571.
  • GPTQ-4bit on OPT-175B is only 0.03 perplexity above FP16 on WikiText2.
  • At 3-bit, GPTQ typically loses only about 0.3–0.6 perplexity points on the largest models.
  • Group size 1024 adds about 0.02 bits and improves perplexity by about 0.2 on average.
  • Group size 128 adds about 0.15 bits and improves perplexity by another roughly 0.1.
  • 3-bit OPT-175B uses about 63GB including FP16 embeddings/output layer, plus about 9GB for 2048-token KV history.
  • It fits on one 80GB A100 versus five A100s for FP16 and three for LLM.int8().
  • Average per-token latency drops from 230ms FP16 on five A100s to 71ms on one A100 at 3-bit.
  • On two A6000 GPUs, latency drops from 589ms FP16 on eight GPUs to 130ms at 3-bit.
  • The reported speedups are about 3.24× on A100 and 4.53× on A6000.

12

Ablation

  • GPTQ is compared with AdaRound, AdaQuant, BRECQ, and OBQ on ResNet18 and ResNet50.
  • On ResNet18 at 4-bit, GPTQ reaches 69.37% versus 69.34% AdaRound and 69.56% OBQ.
  • On ResNet18 at 3-bit, GPTQ reaches 67.88%, above AdaQuant's 59.21%.
  • On ResNet50 at 4-bit, GPTQ reaches 75.71%, close to BRECQ 75.88% and OBQ 75.72%.
  • On ResNet50 at 3-bit, GPTQ reaches 74.87% versus AdaQuant 64.98%.
  • On BERT-base and OPT-125M, GPTQ and greedy OBQ are similar at 4-bit and GPTQ is slightly better at 3-bit.
  • Fixed arbitrary ordering loses little versus greedy ordering on large layers.
  • Lazy batch updates provide about an order-of-magnitude practical speedup on very large models.
  • Cholesky reformulation addresses failures caused by indefinite inverse-Hessian updates.
  • Group sizes 1024 and 128 improve 3-bit large-model accuracy with small metadata overhead.
  • Group size 32 at around 2.6 bits gives a 0.6–0.7 perplexity increase on the largest models.

13

Limitations

Error Analysis

  • The dominant RTN failure at 3-bit is catastrophic perplexity degradation.
  • GPTQ reduces this by compensating each quantization error through correlated remaining weights.
  • The benefit of second-order information is strongest under aggressive quantization.
  • Repeated inverse-Hessian updates can accumulate numerical errors.
  • The inverse Hessian can become indefinite on large models.
  • Indefinite Hessians can drive remaining weights in incorrect directions and produce arbitrarily bad layers.
  • Damping is adequate for smaller models but less robust for the largest models.
  • OPT-66B is an exception to the general trend that larger models are easier to quantize.
  • The authors associate the OPT-66B anomaly with many dead units in early layers.
  • At 4-bit, RTN can degrade substantially while GPTQ stays close to FP16.
  • At 3-bit, RTN frequently collapses to extremely large perplexity values.
  • Quantization difficulty depends on model scale, architecture, and granularity, not bitwidth alone.

14

Conclusion

  • GPTQ demonstrates accurate one-shot quantization at hundreds-of-billions-parameter scale.
  • Approximate second-order error compensation becomes scalable through fixed ordering, batching, and Cholesky reformulation.
  • 3-bit and 4-bit weights preserve language-model quality far better than RTN.
  • OPT-175B can run on a single 80GB A100 at 3 bits.
  • Custom kernels turn memory savings into practical generation-speed improvements.
  • Group-wise quantization extends the accuracy/compression trade-off toward lower average bitwidth.
  • Around 2.2-bit and 2.6-bit configurations remain viable in the reported experiments.
  • Ternary quantization reaches 9.20 WikiText2 perplexity on OPT-175B with group size 8.
  • GPTQ speedups primarily come from reduced memory movement, not fewer arithmetic operations.
  • Activation quantization is outside the main scope and remains a future direction.

15

References

  • Frantar et al., GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers, arXiv:2210.17323, ICLR 2023.
  • Frantar et al., Optimal Brain Compression: A Framework for Accurate Post-Training Quantization and Pruning, 2022.
  • Yao et al., ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers, 2022.
  • Dettmers et al., LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale, 2022.
  • Park et al., nuQmm: Quantized Matmul for Efficient Inference of Large-Scale Generative Language Models, 2022.
  • Nagel et al., Up or Down? Adaptive Rounding for Post-Training Quantization, ICML 2020.
  • Li et al., BRECQ: Pushing the Limit of Post-Training Quantization by Block Reconstruction, ICLR 2021.
  • Hubara et al., Accurate Post Training Quantization with Small Calibration Sets, 2021.
  • Zhang et al., OPT: Open Pre-trained Transformer Language Models, 2022.
  • Scao et al., BLOOM and the BigScience language-model effort, 2022.

Continue reading

Related research

Browse all research