TinyML is the discipline of running machine-learning inference on microcontrollers — the cheap, milliwatt-class chips inside sensors, wearables, appliances, and industrial equipment. No cloud connection, no GPU, no fan: just a model compressed to fit in kilobytes of flash, making decisions at the sensor.
This guide explains what TinyML actually is, the hardware and software stack behind it, where it shines, and where it doesn't.
The definition, in plain terms
Machine learning usually means big models on big computers. TinyML inverts that: it targets devices with a few hundred kilobytes of RAM, clock speeds under a few hundred megahertz, and power budgets measured in milliwatts — sometimes microwatts. The canonical targets are Arm Cortex-M and RISC-V microcontrollers, DSPs, and small neural accelerators.
A TinyML model is not a smaller version of a server model copied over. It is a model designed, trained, quantized, and compiled from the start to fit the device's memory and energy envelope. That constraint-first mindset is what separates TinyML engineering from ordinary ML deployment.
Why run ML on a microcontroller at all?
Three reasons dominate. Latency: an inference that happens on the chip answers in single-digit milliseconds, not after a network round trip. Privacy: raw audio, images, or biosignals never leave the device. Cost: a $3 MCU running a 200 KB model has no per-inference cloud bill, ever.
There's a fourth reason that matters in the field: connectivity is unreliable or absent. Factory floors shield RF, agricultural sensors sit kilometers from the nearest tower, and battery devices can't afford a radio that stays awake. TinyML makes the device useful regardless.
- Millisecond latency — decisions at the sensor
- Privacy by default — raw data never leaves the device
- Works offline — no connectivity dependency
- No cloud bill — inference costs milliwatts, not dollars
The TinyML stack
A typical pipeline starts in a familiar framework — TensorFlow or PyTorch — where a small network (usually well under a million parameters) is trained and quantized to int8. The quantized graph is then compiled for the target: TensorFlow Lite Micro is the most common runtime, with CMSIS-NN on Arm and vendor NPUs on silicon that has them.
Below the runtime sits firmware: sensor drivers, a DSP or feature-extraction stage, the inference call itself, and the application logic that acts on the result. TinyML success is as much firmware engineering as it is data science — the model is useless if the input pipeline feeds it garbage.
Where TinyML shows up today
Wake-word detection is the most famous example — always-on listening at microwatt budgets. Beyond that: predictive-maintenance sensors that classify vibration signatures on motors, wearables that detect falls and arrhythmias, smart meters that spot anomalies in current draw, and pet trackers that distinguish scratching from sleeping.
The common thread is sensor data with a learnable pattern and a product that benefits from immediate, private, low-power answers. If that describes your device, TinyML is probably the right architecture — and a feasibility audit is the right first step.
Frequently Asked Questions
How much memory does a TinyML model need?
TinyML models typically range from tens of kilobytes to a few hundred kilobytes. A wake-word model might run in 30–60 KB of flash and 8–16 KB of RAM; a compact anomaly-detection or classification model in 100–250 KB. The constraint is usually RAM for activations, not flash for weights — which is why model architecture and input resolution are co-designed with the hardware.
Can a microcontroller really run a neural network?
Yes — routinely. A Cortex-M4 at 100+ MHz performs an int8 inference on a compact CNN or MLP in single-digit to tens of milliseconds. Vendor NPUs like Arm's Ethos-U push that to microseconds per layer. The limitation isn't whether it runs, it's whether the model fits the memory and energy budget while staying accurate.
What can't TinyML do?
TinyML handles focused, single-task inference: keyword spotting, anomaly detection, classification of IMU/audio/vibration signals, small image models. It does not run large language models, high-resolution video analytics, or anything requiring more than roughly a megabyte of model parameters. Those belong on larger edge processors or in the cloud.
What hardware do I need to start with TinyML?
Development boards like the Arduino Nano 33 BLE Sense, STM32 discovery kits, or Nordic nRF52/nRF53 boards are common starting points — most cost under $50 and include sensors. For production, the choice is driven by memory, power budget, sensor interfaces, and unit cost, not dev convenience.
Building something that should run AI on-device?
Edgehound designs, compresses, and deploys TinyML models on microcontrollers — from feasibility audit to field-ready firmware.