What is quantization?
Quantization stores model numbers with less numerical precision so they take less memory. The useful question is how much memory you save while keeping the answers good enough for your task.
On this page
What changes?
A model's weights are the numbers learned during training. When the model runs, it uses them to calculate its next output. Quantizing weights means replacing their values with approximations that need fewer bits to store.
Think of rounding a measurement to fewer decimal places: you keep a useful approximation but lose some detail. Quantization uses a defined numeric format and scaling rules, rather than simply deleting decimal digits. Some methods also quantize intermediate calculations, called activations.
What do BF16 and NVFP4 mean?
BF16 is a floating-point format that uses 16 bits per value. NVFP4 uses 4-bit floating-point values plus scaling metadata that helps represent their range. These are storage and calculation formats, not quality scores.
Our published experiment used NVFP4 for selected parts of the model while other weights remained in BF16. A chart label saying NVFP4 does not mean every weight, calculation or cache uses that format. The experiment write-up describes exactly what changed.
Format references: PyTorch numeric types and NVIDIA's NVFP4 documentation.
Why it saves memory
Using fewer bits per weight reduces the space needed for the weights. This can make a model fit in the memory of a GPU that could not hold the original version.
The model file is only part of the memory budget. Running it also needs temporary working memory and space for active requests. The quantized format itself may need extra scale values or keep sensitive parts at higher precision. A smaller download does not tell you how many simultaneous users the server can handle.
Less memory does not guarantee the same answers or more speed
Rounding changes the calculations, and those changes can affect the output. A model may still do well on general questions while getting worse at your code, Hebrew text or required output format. Compare the original and quantized versions on representative examples.
Speed depends on the hardware, the software that runs the model, and the request. Moving fewer bytes can help, but converting the compressed values for calculation also takes work. Measure response time on the actual deployment, including long inputs and concurrent requests.
The KV cache is separate from the weights
During text generation, many models keep intermediate results from attention, the mechanism that relates tokens to other tokens in the input. These saved results belong to tokens the model has already processed. This is the KV cache. Reusing it avoids repeating those calculations at each next token.
It uses memory too. Its size depends on the model, the amount of text being processed and the number of active requests. Quantizing the weights does not automatically quantize the KV cache. Quantizing that cache is a separate choice with its own quality and speed checks.
What to check before choosing
- Task quality: does the model still meet your acceptance criteria on examples it was not tuned on?
- Memory under load: does it fit with the input lengths and simultaneous requests you need?
- Response time: how long until useful text arrives, and until the full answer is ready?
- Deployment support: can your chosen software run this exact format efficiently on your hardware?
The right choice is the version that meets your quality and operating requirements. The smallest file is not automatically the best deployment.