A model trained on a workstation is a draft, not a product. Before it ships to a microcontroller it goes through compression — the set of techniques that cut a network from megabytes to kilobytes while holding onto the accuracy that makes it worth shipping.
Three techniques do almost all the work: quantization, pruning, and distillation. Here's what each one actually does, and how to sequence them.
Quantization: float to int8
Quantization converts 32-bit floating-point weights and activations to 8-bit integers — a 4x memory reduction and, more importantly, a shift from slow float math to the integer MAC units microcontrollers are built for. Post-training quantization (PTQ) converts a finished model in minutes; quantization-aware training (QAT) simulates int8 rounding during training so the network learns to be robust to it.
In practice: PTQ to check the size fits, QAT to recover the accuracy PTQ cost you. Well-chosen architectures lose under a percentage point to int8; poorly conditioned ones lose several. For the tightest targets, int8 isn't the floor — sub-byte quantization (int4 weights) exists, at real accuracy cost.
Pruning: deleting what doesn't matter
Networks trained on desktop GPUs end up with far more weights than they use. Pruning removes them: unstructured pruning zeroes individual weights (compressing storage but not speed, since sparse math still costs cycles), while structured pruning removes entire channels or filters — shrinking both size and latency on real hardware.
Structured pruning is what matters for MCUs. Cutting 30–50% of filters in a modest CNN often costs little accuracy after fine-tuning, and it directly shrinks both the weight file and the activation buffers. The catch: pruning interacts badly with later-stage fusion in some toolchains, so it's done before final quantization.
Knowledge distillation: small model learns from big
Distillation trains a small 'student' network to mimic a large 'teacher' — not just on ground-truth labels, but on the teacher's full output distributions, which carry richer signal about class similarity. A student with a tenth of the parameters can land within a point or two of the teacher on focused tasks.
This is the technique that lets you train big and deploy small honestly: build an unconstrained model for accuracy, then distill it into an architecture shaped for the device. It pairs naturally with quantization — distill first, then QAT the student.
- Quantization — 4x smaller, integer-native math (int8, sometimes int4)
- Structured pruning — remove channels, shrink weights + activations
- Distillation — small model trained on a big model's outputs
- Sequence: distill → prune → QAT
Architecture choices that help compression
Some architectures compress gracefully and some fight you. Depthwise-separable convolutions (MobileNet-style) start small and quantize well. Large fully-connected layers are pure weight bulk — usually replaceable with a bottleneck or global pooling. Attention mechanisms are activation-heavy at MCU scale.
The practical rule: pick an architecture family known to survive int8 (small CNNs, depthwise nets, compact LSTMs for sequences), validate PTQ size early, and spend your accuracy budget on QAT and distillation rather than exotic architectures.
Measuring what compression cost you
Compression success is three numbers, not one: model size (flash), peak activation memory (SRAM), and accuracy on a held-out field dataset — not the training split. A quantized model that aces the lab set but drifts on real device data tells you the problem was in your data pipeline, not the compression.
Benchmark latency on the actual target too. Two models with identical parameter counts can differ 5x in inference time depending on operator support in your runtime — a layer the NPU can't execute falls back to CPU and eats the power budget you saved.
Frequently Asked Questions
How much accuracy do I lose going to int8?
With post-training quantization alone, expect anything from negligible loss to a few percentage points depending on the architecture and how well conditioned it is. With quantization-aware training, most well-designed small networks recover to within a fraction of a point of the float model — often indistinguishable in field metrics.
What's the difference between structured and unstructured pruning?
Unstructured pruning removes individual weights, creating sparse matrices — it compresses storage but doesn't speed up inference on standard MCU runtimes, which can't exploit sparsity efficiently. Structured pruning removes whole channels, filters, or heads, producing a genuinely smaller dense network that runs faster and uses less memory everywhere.
Can I use int4 or even smaller quantization?
Yes — sub-byte quantization (int4 weights, sometimes binary/ternary) is an active area and supported by some toolchains and NPUs. The accuracy cost grows steeply below 8 bits for most tasks, and not all runtimes execute sub-byte ops efficiently. It's worth exploring when int8 almost fits, not as a first resort.
Do I compress before or after collecting field data?
Compress iteratively throughout. Validate size with PTQ early (so you're never surprised), then do final QAT and pruning once field data and thresholds are settled. Re-running QAT on a mature dataset is cheap; discovering your model doesn't fit after data collection ends is not.
Building something that should run AI on-device?
Edgehound designs, compresses, and deploys TinyML models on microcontrollers — from feasibility audit to field-ready firmware.