Skip to main content

Quantization

Updated 3 min read

Share this page

Send the link, quote the definition with a link back, or show it as a card on your own site.

https://softwaredictionary.org/terms/quantization

In short

Quantization is a technique that shrinks an AI model by storing its parameters in fewer bits, such as 8 or 4 instead of 16, so inference is faster and cheaper.

What is quantization in AI?

Quantization reduces the precision of the numbers inside a machine learning model. Models are usually trained with parameters stored as 16- or 32-bit floating-point numbers, and quantization converts them to smaller formats such as 8-bit or 4-bit integers. The model keeps the same architecture and behaves almost the same, but it takes a fraction of the memory and often runs faster.

The basic method maps a range of values onto a small set of levels. For 8-bit quantization, each group of weights gets a scale factor, and every weight is divided by that scale and rounded to one of 256 whole numbers; at run time the numbers are multiplied back by the scale. Post-training quantization applies this to a finished model, while quantization-aware training simulates the rounding during training so the model learns to tolerate it. The payoff is large: a 70-billion-parameter model needs about 140 GB at 16 bits but roughly 35 GB at 4 bits.

An analogy is saving a photo with fewer colors or rounding prices to the nearest dollar: you lose a little detail, but the result is much smaller and still does the job. Quantization is what makes it practical to run capable models on a single GPU, a laptop, or a phone, and it lets servers handle more requests with the same hardware. The trade-off is some loss of accuracy, which is usually small at 8 bits and more noticeable at 4 bits and below.

Quantization is sometimes confused with compressing a file, but a quantized model is not unzipped before use; it runs directly on the lower-precision numbers, and the lost precision is gone for good. It is also different from distillation, which trains a new, smaller model to imitate a larger one, and from pruning, which removes weights entirely. These techniques are often combined, and despite both turning things into numbers, quantization has nothing to do with tokenization.

Key takeaways

  • Quantization stores model parameters with fewer bits, such as 8 or 4.
  • It cuts memory use sharply and often speeds up inference.
  • A scale factor maps the original values to a small range of integers and back.
  • Accuracy loss is usually small at 8 bits and larger at very low bit widths.
  • Distillation and pruning are different ways to make models smaller.

Example

Quantizing weights to 8-bit integers and backpython
# Quantize a few floating-point weights to 8-bit integers and back
weights = [0.4213, -1.27, 0.0318, 0.8871, -0.5096]

scale = max(abs(w) for w in weights) / 127       # map the largest value to 127
quantized = [round(w / scale) for w in weights]  # small integers: 1 byte each
restored = [q * scale for q in quantized]        # what the model computes with

print(quantized)                        # [42, -127, 3, 89, -51]
print([round(r, 3) for r in restored])  # [0.42, -1.27, 0.03, 0.89, -0.51]
# Close to the originals, but the small rounding errors are permanent

Readers ask

Does quantization reduce model accuracy?

Usually a little. At 8 bits the difference is often hard to notice, while 4-bit and lower can cause measurable drops, especially on complex reasoning, so quantized models should be tested on your own tasks.

What does 4-bit quantization mean?

It means each parameter is stored in 4 bits, which allows only 16 distinct values per group of weights. It uses about a quarter of the memory of 16-bit weights, which lets much larger models fit on consumer hardware.

What is the difference between quantization and distillation?

Quantization keeps the same model but stores its numbers with less precision. Distillation trains a separate, smaller student model to reproduce the behavior of a larger teacher model.

See also

Spotted a mistake or something missing on this page?Suggest an edit

Read a random page
Open today's review
Switch to the dark theme
Read this page in Türkçe

More

Settings