The usual answer to post training quantization vs quantization aware training is a two line table: PTQ is fast, QAT is accurate. Both lines are true and neither helps you decide. What decides it is where the accuracy goes when you drop to integers, how much of that loss your system can absorb, and what each method does to the evidence you already hold about the model.
PTQ keeps the weights you trained and only chooses how to round them. QAT goes back into training and lets the weights move. For a model that ranks search results, that difference is an engineering cost. For a model that closes a loop on a physical process, it decides whether your validation record survives the compression.
The short answer
Run post training quantization first. The Qualcomm white paper on neural network quantization states that in most cases PTQ is sufficient for 8 bit quantization with close to floating point accuracy, and it needs no retraining and no labelled data. Measure the result where your system is fragile, not on the aggregate score. Repair the calibration if the loss sits in the tails. Move to quantization aware training only when the loss that remains is larger than what your tolerance budget and your safety layer can absorb, or when you need to go below 8 bits.
That ordering is not a compromise between the two. It is how the measured evidence falls.
What each method actually does
Both methods end in the same place: a model whose weights, and usually activations, are stored and computed as low precision integers. Going from 32 to 8 bits cuts the memory for stored tensors by a factor of 4, per the same white paper. The difference is how each method gets there.
Post training quantization: keep the weights, choose the scales
PTQ takes a trained network and maps each tensor onto an integer grid. The work is choosing the scale factor, and often a zero point, for each tensor or each channel. For activations, that means running a small calibration set through the model and recording the ranges it produces. No gradients, no labels, no training loop. It is close to a push button operation, which is why every deployment toolchain ships it.
The weakness is that the rounding is imposed after the fact. The model never had a chance to adapt to it, so any layer that is sensitive to precision takes the error at full strength.
Quantization aware training: let the weights learn the rounding
QAT inserts fake quantization nodes into the forward pass, so the network computes with rounded values during training while gradients flow through the rounding via the straight through estimator. As the practical QAT vs PTQ comparison on apxml puts it, the optimiser directly accounts for the error introduced by the integer mapping, and the advantage is largest at very low bit widths such as INT4. The technique was introduced by Jacob et al. in 2018, according to a 2025 survey of QAT applications.
In practice QAT is a fine tuning phase that starts from the trained model rather than a full retrain. It still needs your training pipeline, your labelled data and a training budget.
What the measurements say
The cleanest head to head on the SERP is not a benchmark paper. It is a six month production A/B of PTQ and QAT on an edge vision model: a defect classifier on a steel rolling line, running on a Jetson Orin Nano with a 14 ms budget for the classification path.
| Variant | Latency | mAP | Activation memory |
|---|---|---|---|
| FP16 baseline | 17.3 ms | 92.4 | 71 MB |
| INT8 PTQ | 8.9 ms | 88.1 | 38 MB |
| INT8 QAT | 9.2 ms | 91.2 | 38 MB |
Three things are worth reading off that table. The unquantized model misses the deadline, so quantization is not optional here. The two integer models are almost identical at inference, same memory and a 0.3 ms difference in latency, so the cost of QAT is paid entirely before deployment. And the aggregate gap of about 3 mAP points understates the problem: under PTQ, the rare defect classes lost between 5 and 9 mAP points each. The failure lived in the tails, in the 1 to 2 percent of samples that paid for the project.
The same pattern appears at lower precision and on harder tasks, only larger. NVIDIA reports that quantization aware distillation recovers 4 to 22 percent accuracy over PTQ on Math-500 and AIME 2024 for Llama Nemotron Super. A systematic study of low bit QAT for reasoning models finds PTQ produces large accuracy drops on reasoning tasks in low bit settings, and its QAT workflow surpasses GPTQ by 44.53 percent on MATH-500 for Qwen3-0.6B.
Read together, the evidence is consistent. At INT8, on the aggregate metric, the gap between the two methods is small. At 4 bits and below, and in the rare regimes of the input distribution, it is not.
Repair the calibration before you reach for QAT
The most useful part of the Orin Nano write up is what the team tried before QAT. Their first calibration set was drawn in proportion to the training data, which meant 92 percent clean background. The activation ranges were tuned to the background, and the hot pixels of rare defects were clipped into it.
| Approach | Rare class mAP | Engineer weeks |
|---|---|---|
| PTQ, default calibration | 79.2 | 0.5 |
| PTQ, defects oversampled 10x | 84.6 | 1.0 |
| PTQ, 99.99 percentile + rebalanced | 86.1 | 1.5 |
| QAT, 10 epochs | 89.4 | 2.5 |
Most of the distance between default PTQ and QAT was closed by changing what the model saw during calibration, not by changing the method. Switching weights from per tensor to per channel scales was worth another 1.5 mAP on rare classes at essentially no inference cost on that hardware.
The practical lesson: a disappointing PTQ result is first a question about your calibration set. Build it to cover the operating envelope, oversample the regimes that matter, use per channel weight scales, and try percentile clipping before you conclude that the method has failed.
What QAT costs
When QAT is the answer, it is worth knowing what you are signing up for. In the same deployment, turning on fake quantization stretched the training cycle from 2 hours to 8, and a custom group norm layer broke quantized export entirely, forcing a fall back to batch norm on the deployment branch. The ACM survey lists the general limitations: training instability, degraded performance at extremely low bit widths, poor adaptability across architectures and high computational cost.
Toolchains add their own friction. A systematic review of edge AI deployment on embedded systems identifies variability in hardware toolchains and limited support for ultra low precision inference as persistent gaps, and proposes a five stage workflow in which requirement definition and hardware alignment are explicit stages alongside optimisation. If the operators you need do not quantize on your target, QAT cannot rescue them.
None of this argues against QAT. It argues against starting with it.
The question the SERP skips: what happens to your evidence
Every comparison on this SERP scores the two methods on accuracy and cost. Neither metric captures what matters most once a model acts on something physical.
PTQ does not change the trained weights. It adds a bounded, characterisable rounding error on top of a model you already validated. You can measure that error directly against the float model on the same test campaign, and the argument that the model is fit for purpose can be carried forward and rechecked rather than rebuilt.
QAT changes the weights. The quantized model is the product of a new training run, which makes it a new artefact. Whatever evidence you held about the float model is now evidence about a different function, and the validation has to be run again on the thing that will actually ship.
In our own control work, documented across the OptiVX and IntelliBot programmes, end to end edge latency came down from 1.2 seconds to 285 ms, the safety layer answers in under 2 ms, growth model accuracy holds at R-squared above 0.95, and validation showed zero safety violations across 400 simulated years. Those numbers are only worth something because they describe the model that runs at the plant. Any compression step that silently changes that model is a step that invalidates them.
This is also why the safety layer matters to the choice. When a separate guarantee, such as safe reinforcement learning with control barrier functions, bounds what the controller is allowed to do, a small quantization error in the learned policy is absorbed by a mechanism that does not depend on the policy being exact. That is the budget PTQ has to fit inside. If the residual error exceeds it, QAT is justified, and so is the cost of rerunning the sim to real transfer evidence on the new weights.
A decision rule you can run
- Fix the target first. Hardware, integer format and latency deadline, before any model is compressed. The deadline tells you whether you need 8 bits or fewer; the toolchain tells you which operators survive.
- Run PTQ with a calibration set built for your tails. Cover the operating envelope, oversample rare regimes, use per channel weight scales.
- Evaluate where the system is fragile. Rare classes, boundary conditions, transients. An aggregate score is where PTQ failures hide.
- Compare the residual loss with your tolerance. If it fits inside the accuracy budget and the margin your safety layer provides, ship PTQ. You keep the trained weights and most of your evidence.
- Otherwise, run QAT starting from the PTQ model. The reasoning QAT study finds that PTQ is a strong initialisation for QAT, improving accuracy while reducing training cost, and that matching the calibration domain to the training domain speeds convergence. Budget for revalidation, because you are now shipping new weights.
- If even the float model misses the deadline, the problem is the architecture. No rounding scheme fixes that. The tool is distillation into a smaller model, which we compare in knowledge distillation vs quantization.
Size and speed gains are real at either end of this procedure. A comparative PTQ and QAT analysis across 10M to 1B parameter models reports up to 2.4x throughput for INT8 and 3x for INT4 on edge devices. The rule above is about keeping those gains without losing what you cannot afford to lose. For where quantization sits next to pruning and distillation in a full pipeline, see our guide to model compression for edge AI, and for the training mechanics in more depth, our explainer on what quantization aware training is.
FAQ
Is quantization aware training always more accurate than post training quantization?
Not by a margin that always matters. At 8 bits, well calibrated PTQ often lands close to floating point accuracy, and QAT's lead there is often only a few points. The gap widens at 4 bits and below, and on rare inputs that a calibration set underrepresents.
Does a QAT model run slower than a PTQ model?
No. At the same integer format both produce the same kind of model, with the same memory footprint and near identical latency. QAT costs training time and pipeline work, not inference speed.
Do I need labelled data for post training quantization?
No. PTQ needs only a small, representative calibration set without labels. QAT is a training procedure and needs labelled data and access to your training pipeline.
Can I use PTQ and QAT together?
Yes, and it is the better order. Quantize with PTQ, then use that model as the starting point for QAT. The fine tuning starts closer to a good integer solution, so it converges faster and usually ends more accurate than QAT from the float weights.

