AI Model Quantization
What is AI Model Quantization?
AI model quantization is the process of reducing the precision of the numbers (weights and activations) that make up a neural network.
Normally, machine learning models are trained using 32-bit floating-point precision (FP32). Quantization maps these high-precision numbers to lower-bit representations (such as 16-bit, 8-bit, 4-bit, or even 2-bit integers or floats).
Why is it used?
Reduced Memory Footprint: A model's size shrinks almost proportionally to the drop in bit-width, allowing large language models (LLMs) to run on consumer hardware or edge devices.
Faster Inference: Lower-bit operations require less memory bandwidth and can be processed faster by modern hardware accelerators (like GPUs, TPUs, and NPUs).
Lower Energy Consumption: Moving less data around reduces power usage, which is crucial for mobile and embedded devices.
Breakdown of Bit Levels: 32-bit down to 2-bit
1. 32-bit Quantization (FP32 / INT32)
Precision: Full precision (the baseline for training). Each parameter takes 4 bytes of memory.
Use Case: Model training and scenarios where maximum accuracy is required, and memory/compute constraints do not exist.
Trade-off: High memory usage and slower inference speeds.
2. 16-bit Quantization (FP16 / BF16)
Precision: Half precision. Each parameter takes 2 bytes of memory.
Use Case: Standard for training modern LLMs and running inference on modern GPUs without noticeable loss in accuracy.
Trade-off: Cuts memory usage in half compared to FP32 with virtually zero degradation in model performance.
3. 8-bit Quantization (INT8)
Precision: Standard quantization level where weights are compressed to 1 byte.
Use Case: Highly popular for deploying models in production. Techniques like PTQ (Post-Training Quantization) and QAT (Quantization-Aware Training) make 8-bit models run efficiently on standard hardware.
Trade-off: Minimal to negligible accuracy loss while cutting memory requirements by 75% compared to FP32.
4. 4-bit Quantization (INT4 / NF4)
Precision: Extremely compressed, averaging 4 bits per parameter (often using formats like NormalFloat4 introduced by QLoRA).
Use Case: Running massive LLMs (like 70B+ parameter models) on local consumer GPUs (e.g., running a model that normally needs 140GB of VRAM on a single desktop GPU).
Trade-off: A slight degradation in perplexity/accuracy, though advanced algorithms minimize this impact heavily.
5. 2-bit Quantization (INT2)
Precision: Ultra-low precision, packing multiple weights into a single byte.
Use Case: Highly constrained edge devices, microcontrollers, or extreme compression research where fitting a model into minimal memory is paramount.
Trade-off: Noticeable drop in model intelligence and accuracy unless specialized, highly sophisticated quantization-aware training methods are applied.
Summary Comparison Table
| Bit-Width | Memory per Parameter | Relative Size (vs FP32) | Typical Accuracy Impact | Primary Use Case |
| 32-bit (FP32) | 4 bytes | 100% | Baseline (None) | Training |
| 16-bit (FP16/BF16) | 2 bytes | 50% | Negligible | Standard Inference & Training |
| 8-bit (INT8) | 1 byte | 25% | Very Low | Efficient Production Deployment |
| 4-bit (INT4/NF4) | 0.5 bytes | 12.5% | Low to Moderate | Running Large LLMs locally |
| 2-bit (INT2) | 0.25 bytes | 6.25% | Moderate to High | Extreme edge computing |