TensorFlow Lite Micro is the runtime that executes your quantized model on bare silicon. It takes a flatbuffer graph and a small kernel library and runs inference with no OS dependency — the same piece of code whether your target is a Cortex-M4, a RISC-V core, or a DSP.
This guide walks the pipeline end to end: train, convert, quantize, compile, integrate.
Train and export the model
Start in Keras with an architecture sized for the target — small CNNs with depthwise-separable layers are the standard choice. The critical design rule: only use operators TFLM supports, because unsupported ops either fail conversion or drag in code size you can't afford.
Export via the TFLite converter with quantization enabled: a full integer model with int8 weights and int8 activations is the target artifact. The output is a .tflite flatbuffer file — typically tens to a few hundred kilobytes, ready to embed.
From .tflite to C
The flatbuffer ships inside firmware as a byte array — generated with xxd -i or the toolchain's equivalent — and compiled into flash. The model becomes just another constant in the image, versioned with the firmware or loaded separately for OTA model updates.
At runtime you hand the buffer to the interpreter: set up a tensor arena (a single static buffer that holds all activations), resolve the operators your model uses, allocate tensors, and call Invoke(). A keyword-spotting model with a 30 KB arena allocates in microseconds and runs per audio frame.
The tensor arena and memory planning
TFLM's entire activation memory lives in one contiguous arena you declare — no heap, no malloc. Sizing it is empirically driven: the interpreter tells you the minimum required at allocation time, and you pad for safety. This design is why TFLM fits on RAM-constrained devices; it's also why arena size becomes a first-class build artifact.
Operator resolution matters for code size too. An AllOpsResolver links every kernel; a MicroMutableOpResolver links only the ops your model uses — often cutting tens of KB of flash. Always use the mutable resolver in production firmware.
- Declare one static arena sized for peak activations
- Use MicroMutableOpResolver — link only needed kernels
- Keep the model buffer const — it stays in flash, not RAM
- Invoke() per window — the pipeline drives the cadence
Kernel libraries: CMSIS-NN and NPUs
TFLM's reference kernels are portable C — correct but not fast. On Arm, kernels backed by CMSIS-NN exploit SIMD (MVE/Helium on M55, DSP extensions on M4/M7) for several-times-faster conv and dense ops. Most vendor SDKs ship TFLM builds already wired to optimized kernels.
On NPU parts the flow changes: Arm's Vela compiler takes the .tflite graph and produces an Ethos-U command stream, executed through an EthosU driver instead of CPU kernels. Ops the NPU doesn't support fall back to CPU — which is why checking the compiled graph's op partitioning matters for both latency and power.
Alternatives and when to use them
TFLM isn't the only runtime. Vendor frameworks — ST's X-CUBE-AI, NXP eIQ, Renesas DRP-AI — compile models to target-specific code and sometimes outperform TFLM on their silicon. ONNX Runtime Micro, microTVM (Apache TVM), and interpreter-free codegen approaches each have niches.
The pragmatic default remains TFLM: broad op support, battle-tested on Cortex-M, huge community, and the same graph artifact works from dev board to NPU compiler. Switch when a vendor toolchain demonstrably beats it on your specific part — not before.
Frequently Asked Questions
How much overhead does TFLM itself add?
The interpreter core is roughly 20–40 KB of code with a lean op set, plus the kernels your model links. With a MicroMutableOpResolver and CMSIS-NN kernels, total TFLM footprint typically lands under 100 KB of flash — acceptable on 256 KB+ devices, tight but workable below that.
Can I use PyTorch models with TFLM?
Not directly — TFLM consumes TFLite flatbuffers. The standard path is PyTorch → ONNX → TensorFlow → TFLite, using tools like onnx2tf. It works but adds conversion risk; if deployment on TFLM is the plan, training in Keras/TF from the start eliminates a whole class of export bugs.
What operators does TFLM not support?
TFLM supports a subset of TFLite ops focused on common inference patterns: conv, depthwise conv, fully connected, pooling, activations, some LSTM/GRU, and basic elementwise ops. Dynamic shapes, control flow, string ops, and many modern ops are absent. Design the model against the supported-op list early — conversion surprises at the end are expensive.
How do I update just the model without reflashing firmware?
Store the .tflite blob in a dedicated flash partition with a header (version, checksum), and have the bootloader or app pass its address to the interpreter. OTA then delivers a new blob to that partition — same mechanism as firmware updates, smaller payload. Versioning and rollback logic lives in your update manager, not TFLM.
Building something that should run AI on-device?
Edgehound designs, compresses, and deploys TinyML models on microcontrollers — from feasibility audit to field-ready firmware.