Paper page - GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Elias Frantar ,

Published on Oct 31, 2022

Authors:

,

Abstract

GPTQ, a novel one-shot weight quantization method using approximate second-order information, significantly reduces the computational and storage requirements of large GPT models while maintaining accuracy and enabling execution on a single GPU.

Generative Pre-trained Transformer

models, known as

GPT

or

OPT

, set themselves apart through breakthrough performance across complex language modelling tasks, but also by their extremely high computational and storage costs. Specifically, due to their massive size, even

inference

for large, highly-accurate

GPT

models may require multiple performant

GPU

s, which limits the usability of such models. While there is emerging work on relieving this pressure via

model compression

, the applicability and performance of existing compression techniques is limited by the scale and complexity of

GPT

models. In this paper, we address this challenge, and propose

GPT

Q, a new one-shot weight quantization method based on

approximate second-order information

, that is both highly-accurate and highly-efficient. Specifically,

GPT

Q can quantize

GPT

models with 175 billion parameters in approximately four

GPU

hours, reducing the bitwidth down to 3 or 4 bits per weight, with negligible accuracy degradation relative to the uncompressed baseline. Our method more than doubles the compression gains relative to previously-proposed

one-shot quantization

methods, preserving accuracy, allowing us for the first time to execute an 175 billion-parameter model inside a single

GPU

for generative

inference

. Moreover, we also show that our method can still provide reasonable accuracy in the extreme quantization regime, in which weights are quantized to 2-bit or even ternary quantization levels. We show experimentally that these improvements can be leveraged for

end-to-end inference speedups

over

FP16

, of around 3.25x when using high-end

GPU

s (NVIDIA A100) and 4.5x when using more cost-effective ones (NVIDIA A6000). The implementation is available at https://github.com/IST-DASLab/

gpt

q.

View arXiv page

View PDF

Add to collection

Get this paper in your agent:

hf papers read 2210.17323

Don't have the latest CLI?curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 234

RedHatAI/Meta-Llama-3.1-70B-Instruct-quantized.w4a16 Text Generation • 71B • Updated Feb 12, 2025 • 7.17k • 33

BelleGroup/BELLE-7B-gptq Text Generation • Updated Apr 20, 2023 • 30 • 26

astronomer/Llama-3-8B-Instruct-GPTQ-4-Bit Text Generation • 8B • Updated Apr 22, 2024 • 83 • 26

astronomer/Llama-3-8B-Instruct-GPTQ-8-Bit Text Generation • 8B • Updated Apr 22, 2024 • 80 • 25

Browse 234 models citing this paper

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2210.17323 in a dataset README.md to link it from this page.

Spaces citing this paper 8

Collections including this paper 12

Browse 12 collections that include this paper