Guides/Fundamentals

AI on Microcontrollers: How Neural Networks Run on Cortex-M and RISC-V

Inside MCU machine learning: memory constraints, int8 arithmetic, inference runtimes like TFLM and CMSIS-NN, vendor NPUs, and what kinds of models actually fit.

11 min read·

Running a neural network on a microcontroller sounds paradoxical until you look at the numbers: modern MCUs execute billions of int8 multiply-accumulates per second, and a well-designed model needs far less than that per inference. The real constraint isn't compute — it's memory.

This guide covers the hardware landscape, the runtime software stack, and the arithmetic realities that shape MCU-scale AI.

The hardware landscape

The mainstream target is Arm Cortex-M — M4 and M7 for most TinyML work, M33 and M55 with Helium vector extensions for heavier models. RISC-V cores are growing quickly, and DSP cores (CEVA, Cadence Tensilica) remain strong for audio. A newer tier adds dedicated neural accelerators: Arm Ethos-U55/U65, and vendor NPUs inside MCUs from Renesas, NXP, and STMicroelectronics.

What defines each tier is memory: a typical TinyML device has 256 KB–1 MB of flash and 64–512 KB of SRAM. Weights live in flash; activations live in SRAM — and activation buffers are almost always the binding constraint on model size.

  • Cortex-M4/M7 — the workhorse, MVE-free baseline
  • Cortex-M33/M55 — Helium SIMD for faster int8 kernels
  • Ethos-U55/U65 — dedicated NPU, 10–100x inference speedup
  • RISC-V + DSP cores — vendor alternatives for audio/vision

Why int8 is the native tongue

Most MCUs lack serious floating-point throughput, and float32 weights quadruple memory for no gain you'd notice after quantization-aware training. The industry standard is int8: weights and activations quantized to 8-bit integers, with the heavy math done by integer MAC units — exactly what these chips are built for.

Quantization-aware training (QAT) simulates int8 rounding during training so the model learns to compensate. Post-training quantization (PTQ) is faster but costs accuracy; the standard workflow is PTQ to validate the size, then QAT to recover the accuracy.

The runtime stack

TensorFlow Lite Micro (now evolving under LiteRT Micro) is the dominant interpreter — it executes a flatbuffer model graph with a small kernel library. On Arm, kernels are backed by CMSIS-NN, a hand-optimized SIMD library that gets the most out of M-series cores. Vendors ship alternatives: ST's X-CUBE-AI, NXP's eIQ, Microchip's MPLAB ML plugin.

NPU-equipped parts use a compiler instead — Arm's Vela compiler retargets a TFLM graph into Ethos-U command streams, offloading convolutions entirely. The firmware difference matters: interpreter-based runtimes cost CPU cycles per layer; NPU compilation moves the math to dedicated silicon.

What actually fits

Rule-of-thumb budgets: a keyword-spotting CNN runs in 30–80 KB; an IMU activity classifier in 20–60 KB; an anomaly-detection autoencoder or classifier in 50–150 KB; a small vision model (person detection at 96×96) in 250–350 KB. Models beyond ~400 KB push past most single-MCU devices into MPU territory.

The architectures that fit share traits: depthwise-separable convolutions, small input resolutions, few classes, and integer-friendly ops. Architectures that don't fit — transformers with large attention, high-res vision, generative models — belong on bigger silicon.

Frequently Asked Questions

What's the smallest chip that can run AI?

Practical TinyML starts around a Cortex-M0+/M3 with 64 KB flash and 8 KB RAM — enough for a tiny keyword or simple classifier. Below that you're in DSP-assembly territory. The comfortable sweet spot is 256 KB flash / 64 KB SRAM, which covers most sensor-analytics models with room for firmware.

How fast is inference on an MCU?

A compact CNN inference takes roughly 1–50 ms on a Cortex-M4/M7 at 80–480 MHz, depending on model size. Ethos-U accelerators cut that by an order of magnitude or more. For context: wake-word pipelines must run per ~20 ms audio frame, and anomaly detectors often need just one inference per second — both easily met.

Do I need an RTOS to run ML on an MCU?

No — plenty of TinyML products run bare-metal with a simple super-loop and interrupts. An RTOS (Zephyr, FreeRTOS) helps when inference shares the chip with other real-time duties like sensor sampling, radio stacks, or UI, since it gives you task prioritization and cleaner power management.

Can I train on the microcontroller itself?

Training on-MCU is an active research area (on-device fine-tuning, federated learning at the edge) but production TinyML is almost always trained off-device. The standard pattern is train in the cloud or on a workstation, quantize and compile for the target, then ship updates over the air.

Building something that should run AI on-device?

Edgehound designs, compresses, and deploys TinyML models on microcontrollers — from feasibility audit to field-ready firmware.

Related guides