Guides/Techniques

INT8 Quantization for Embedded AI: A Practical Guide

How int8 quantization works under the hood — scales, zero-points, QAT vs PTQ, per-channel vs per-tensor — and how to debug accuracy loss on the target.

10 min read·

Int8 quantization is the single most consequential step in embedded AI deployment — it's what turns a float model into something a microcontroller executes natively. It's also where accuracy silently leaks if you're not careful.

This guide covers the mechanics (scales, zero-points, per-tensor vs per-channel), the PTQ versus QAT decision, and how to chase down accuracy loss when it happens.

The mechanics: scales and zero-points

Quantization maps a float range to 256 integer levels using two numbers per tensor: a scale (step size) and a zero-point (the integer that represents 0.0). A float value becomes q = round(f / scale) + zero_point. Every layer's weights and activations get their own scale and zero-point, and the runtime requantizes between layers to keep arithmetic in integers.

The art is choosing ranges. Too tight a range clips real values; too loose wastes precision on values that never occur. Calibration — running representative data through the model to observe activation ranges — is what sets those bounds honestly.

PTQ: fast, cheap, usually good enough to start

Post-training quantization takes a finished float model plus a small calibration set — a few hundred representative inputs — and produces an int8 model without retraining. For well-conditioned small networks it's often within a point of float accuracy, which makes it the right first move.

The two knobs that matter: calibration data quality (it must look like real field inputs, not random noise — ranges are set from what the model actually sees) and granularity. Per-channel quantization for weights (each output channel gets its own scale) buys back significant accuracy over per-tensor at zero inference cost on most targets.

QAT: recovering the accuracy PTQ lost

Quantization-aware training inserts fake-quantization nodes into the model during training, so the forward pass experiences int8 rounding while gradients flow through in float. The network learns weights that are robust to quantization noise instead of discovering it at deploy time.

QAT typically costs a few epochs of fine-tuning from the trained float checkpoint, not a full retrain. It's standard practice for anything shipping to int8 at scale — the difference between 'within a fraction of a point' and 'why did we lose 4%'.

  • PTQ — minutes, calibration set only, good enough to validate size
  • QAT — fine-tune with fake-quant nodes, recovers most/all lost accuracy
  • Per-channel weights — near-free accuracy win on supported targets
  • Fake-quant — training sees int8 noise, gradients stay float

Debugging accuracy loss on target

When the int8 model underperforms, isolate which layer hurts. Run the quantized model layer-by-layer against the float reference on the same inputs and find where outputs diverge — usually a layer with wide activation range or tiny weights that get crushed to zero.

Common fixes: improve calibration data to cover the missing range, move sensitive layers to per-channel or higher precision (some runtimes allow selective float), or adjust the architecture — a batchnorm-folded, properly scaled model quantizes far better than one with pathological activation ranges.

Frequently Asked Questions

What is a calibration dataset and how big should it be?

Calibration data is a small sample of real inputs — typically 100–1000 examples — run through the model to observe actual activation ranges so quantization scales can be set correctly. It must be representative of production inputs: calibrating on dev-kit data while shipping to different sensors produces systematically wrong ranges and degraded accuracy.

Why did my accuracy collapse after quantization?

The usual suspects: calibration data that doesn't represent real inputs, a layer with extreme activation range dominating the scale (fix with per-channel or range clamping), weights crushed to zero in a sensitive layer, or an operator the runtime quantizes differently than expected. Layer-wise comparison against the float model pinpoints the culprit quickly.

Is int8 always the right choice for MCUs?

It's the sweet spot: most MCU kernels and NPUs are optimized for int8 MACs. Float16 halves memory over float32 but needs FPU support and still costs more than int8; int4 and lower exist but with accuracy and toolchain caveats. For a handful of ultra-sensitive layers, mixed-precision (keeping a layer float) is a legitimate escape hatch.

Does quantization affect inference speed or just size?

Both, and speed is often the bigger win. Int8 MACs run on the integer pipeline — far faster than software-emulated float on M-class cores, and it's what NPU accelerators execute natively. A quantized model is both ~4x smaller in storage and typically several times faster per inference than its float version.

Building something that should run AI on-device?

Edgehound designs, compresses, and deploys TinyML models on microcontrollers — from feasibility audit to field-ready firmware.

Related guides