AI Model Quantization

 

What is AI Model Quantization?

AI model quantization is the process of reducing the precision of the numbers (weights and activations) that make up a neural network.

Normally, machine learning models are trained using 32-bit floating-point precision (FP32). Quantization maps these high-precision numbers to lower-bit representations (such as 16-bit, 8-bit, 4-bit, or even 2-bit integers or floats).

Why is it used?

  • Reduced Memory Footprint: A model's size shrinks almost proportionally to the drop in bit-width, allowing large language models (LLMs) to run on consumer hardware or edge devices.

  • Faster Inference: Lower-bit operations require less memory bandwidth and can be processed faster by modern hardware accelerators (like GPUs, TPUs, and NPUs).

  • Lower Energy Consumption: Moving less data around reduces power usage, which is crucial for mobile and embedded devices.

Breakdown of Bit Levels: 32-bit down to 2-bit

1. 32-bit Quantization (FP32 / INT32)

  • Precision: Full precision (the baseline for training). Each parameter takes 4 bytes of memory.

  • Use Case: Model training and scenarios where maximum accuracy is required, and memory/compute constraints do not exist.

  • Trade-off: High memory usage and slower inference speeds.

2. 16-bit Quantization (FP16 / BF16)

  • Precision: Half precision. Each parameter takes 2 bytes of memory.

  • Use Case: Standard for training modern LLMs and running inference on modern GPUs without noticeable loss in accuracy.

  • Trade-off: Cuts memory usage in half compared to FP32 with virtually zero degradation in model performance.

3. 8-bit Quantization (INT8)

  • Precision: Standard quantization level where weights are compressed to 1 byte.

  • Use Case: Highly popular for deploying models in production. Techniques like PTQ (Post-Training Quantization) and QAT (Quantization-Aware Training) make 8-bit models run efficiently on standard hardware.

  • Trade-off: Minimal to negligible accuracy loss while cutting memory requirements by 75% compared to FP32.

4. 4-bit Quantization (INT4 / NF4)

  • Precision: Extremely compressed, averaging 4 bits per parameter (often using formats like NormalFloat4 introduced by QLoRA).

  • Use Case: Running massive LLMs (like 70B+ parameter models) on local consumer GPUs (e.g., running a model that normally needs 140GB of VRAM on a single desktop GPU).

  • Trade-off: A slight degradation in perplexity/accuracy, though advanced algorithms minimize this impact heavily.

5. 2-bit Quantization (INT2)

  • Precision: Ultra-low precision, packing multiple weights into a single byte.

  • Use Case: Highly constrained edge devices, microcontrollers, or extreme compression research where fitting a model into minimal memory is paramount.

  • Trade-off: Noticeable drop in model intelligence and accuracy unless specialized, highly sophisticated quantization-aware training methods are applied.

Summary Comparison Table

Bit-WidthMemory per ParameterRelative Size (vs FP32)Typical Accuracy ImpactPrimary Use Case
32-bit (FP32)4 bytes100%Baseline (None)Training
16-bit (FP16/BF16)2 bytes50%NegligibleStandard Inference & Training
8-bit (INT8)1 byte25%Very LowEfficient Production Deployment
4-bit (INT4/NF4)0.5 bytes12.5%Low to ModerateRunning Large LLMs locally
2-bit (INT2)0.25 bytes6.25%Moderate to HighExtreme edge computing
sources:
https://www.cloudflare.com/learning/ai/what-is-quantization/
https://medium.com/@isanghao/what-is-quantization-and-why-it-matters-for-inference-c62135f7cfa7
https://huggingface.co/docs/transformers/main_classes/quantization
https://medium.com/@isanghao/optimizing-llm-inference-with-dynamic-quantization-056026701667
https://medium.com/@isanghao/io-bound-or-compute-bound-in-ai-c9c541cd6696
https://developer.nvidia.com/blog/model-quantization-concepts-methods-and-why-it-matters/
https://www.ibm.com/think/topics/quantization
https://developers.google.com/edge/litert/conversion/tensorflow/quantization/post_training_quantization
https://www.medoid.ai/blog/a-hands-on-walkthrough-on-model-quantization/