Articles

Quantization: how much precision a model can lose

Fewer bits per weight save memory; what that costs in quality must be measured.

Reading: 6 minAutomation & AI

Article cover: Quantization: how much precision a model can lose

Quantization stores each weight in fewer bits. The vendor documentation states the trade in one sentence: it “lowers the memory requirements of loading and using a model by storing the weights in a lower precision while trying to preserve as much accuracy as possible”. Trying is the operative word. Memory is predictable arithmetic; quality is an empirical question that has to be answered on your task, with your data, before the smaller file is treated as the same model.

What actually changes

A transformer’s weights are floating-point numbers. Storing them at half precision costs two bytes each; at 8 bits, one; at 4 bits, about half a byte. On that arithmetic alone, a 7-billion-parameter model moves from roughly 14 GB to 7 GB to about 3.5 GB. The published method tables state the precision each family targets — for instance 4 and 8 bits for bitsandbytes, 4 for AWQ, 1 and 2 for AQLM, 8 for EETQ — and mark, per method, whether it works on the fly or needs CPU or a specific accelerator.

Two consequences follow, and only the first is guaranteed:

  • the file and the resident weights get smaller, which is what makes a model fit at all;
  • speed may or may not improve. A lower-precision format only becomes a faster computation if the hardware and the kernels actually implement it, which is why the same method is marked as supported on one backend and not on another.

How the methods differ

8-bit integer arithmetic with explicit outlier handling. LLM.int8() “cut the memory needed for inference by half while retaining full precision performance” by quantizing most values vector-wise and isolating the outlier dimensions into a 16-bit matrix multiplication, “while still more than 99.9% of values are multiplied in 8-bit”. The outlier layer is the whole point: naive quantization fails on models because a small number of feature dimensions dominate the computation.

One-shot weight quantization with calibration. GPTQ is “a new one-shot weight quantization method based on approximate second-order information, that is both highly-accurate and highly-efficient”: it “can quantize GPT models with 175 billion parameters in approximately four GPU hours, reducing the bitwidth down to 3 or 4 bits per weight, with negligible accuracy degradation relative to the uncompressed baseline”. Its authors report end-to-end inference speedups “of around 3.25x when using high-end GPUs (NVIDIA A100)”. Note what that sentence is: a result obtained on named hardware, by the method’s authors, on their benchmarks.

Activation-aware weight quantization. AWQ starts from the finding that “not all weights in an LLM are equally important. Protecting only 1% salient weights can greatly reduce quantization error”, and chooses which weights to protect from the activation distribution rather than from the weights themselves. It “does not rely on any backpropagation or reconstruction, so it generalizes to different domains and modalities without overfitting the calibration set” — an explicit answer to the risk that a calibrated method is tuned to the data used for calibration.

4-bit with adapters. QLoRA “reduces memory usage enough to finetune a 65B parameter model on a single 48GB GPU while preserving full 16-bit finetuning task performance” — a training result rather than an inference one, and the origin of the 4-bit normal-float type that later served inference as well.

The documentation sorts the families accordingly: “Some methods require calibration for greater accuracy and extreme compression (1-2 bits), while other methods work out of the box with on-the-fly quantization.”

Where the loss appears

Quantization does not remove information uniformly, and the damage is not visible in a sample of fluent text:

  • Precision is lost first where it is least redundant — exact recall of a rare fact, faithful copying of an identifier, strict format compliance, arithmetic. A degraded model still writes well; it fails on the things you were going to check with a script.
  • The compression ratio predicts size, not quality. The same nominal 4 bits describes several different procedures, and the method’s own benchmark is not your workload.
  • A quantized model is a different model. It shares a name and a shape with its full-precision origin, not its behaviour. Treating the two as interchangeable is the same mistake as treating two versions of a model as interchangeable.

What this means operationally

Choose the smallest file that passes the checks you actually run, not the smallest file that exists. The measurement — a fixed task set, the same prompts, side-by-side outputs, a decision rule — belongs to Progetti; the knowledge claim here is narrower: that quality is traded away, that the trade is task-dependent, and that nobody can tell you in advance where your model will break.

Level and prerequisites. L2 — operational: enough to interpret a quantization label, to know what to expect from it, and to know that the expectation must be tested. It is not a conversion guide and prescribes no method. Prerequisites: the L1 sheets on tokens and on the model/runtime boundary, and the sheet in this batch on model formats.

Where to go next

References

  • Hugging Face — Quantization overview — the definition of quantization, the calibration/on-the-fly distinction, and the per-method precision and backend table.
  • Hugging Face — Bitsandbytes documentation — LLM.int8() and QLoRA in the library’s own terms, and the hardware minimums per method.
  • Dettmers et al. — LLM.int8() — 8-bit inference with outlier isolation.
  • Frantar et al. — GPTQ — one-shot 3–4 bit weight quantization.
  • Lin et al. — AWQ — activation-aware protection of salient weight channels.
  • Dettmers et al. — QLoRA — 4-bit quantization with adapters (training result).