Ask someone to picture on-device AI and they picture the neural network. Ask an embedded AI engineer and they picture the pipeline: sampling, conditioning, feature extraction, inference, post-processing — of which the model is maybe a fifth of the code and none of the interesting bugs.
This guide walks the full path from analog signal to device action, stage by stage.
Sampling and signal conditioning
Everything starts with the sensor interface — I2S for microphones, SPI/I2C for IMUs, ADC reads for current and analog sensors. Sample rates are chosen for the phenomenon, not the marketing: 16 kHz for speech, a few kHz for vibration harmonics, tens of hertz for slow environmental drift.
Conditioning happens next: DC removal, anti-aliasing, gain normalization, sometimes bandpass filtering to the frequency band that carries your signal. Garbage in here is permanent — no downstream stage recovers a clipped or aliased signal.
Windows and feature extraction
Models rarely see raw samples. The stream is cut into windows — say, 1-second audio frames or 2-second IMU segments — and each window is converted to features. For audio that's typically MFCCs or a mel spectrogram; for IMU, spectral features or statistics per axis; for current draw, RMS, crest factor, and harmonic content.
Feature extraction is where domain knowledge earns its salary. A well-chosen feature set lets a small model learn fast and run cheap; a lazy raw-signal approach demands a bigger model to learn the filtering itself. Fixed-point DSP on Cortex-M handles all of this in real time — CMSIS-DSP is your friend.
Inference: the small, fast middle
The quantized model consumes features and emits a distribution or score — one per class, or an anomaly score. On a Cortex-M this is typically a few milliseconds of int8 math; on an NPU-equipped part, microseconds per layer. The inference call itself is the most boring stage of the pipeline, which is exactly what you want.
What matters around the call is scheduling: inference must run on a cadence (per window, per second, on wake) without starving sensor ISRs or the power manager. This is why inference sits in its own task or is carefully placed in the main loop.
Post-processing and decision logic
Raw model outputs aren't decisions — a 0.7 wake-word score on one frame isn't a wake event. Post-processing applies smoothing (score averaging, debounce), thresholds tuned for precision/recall trade-offs, and hysteresis so outputs don't flicker. Stateful logic — 'three detections in five seconds' — lives here too.
Then the action: wake the host, raise a GPIO, queue an event for the cloud, actuate a relay. The event format matters as much as the detection — what does upstream software need to see? Designing the pipeline ends at the interface contract, not at the model's softmax.
- Smoothing: rolling average or majority vote over windows
- Thresholds: tuned to the product's precision/recall trade-off
- Hysteresis & debounce: no flickering outputs
- Events: compact derived data upstream, not raw streams
Frequently Asked Questions
What sampling rate should I use for my sensor?
Sample at least twice the highest frequency component of interest (Nyquist), plus margin for your anti-alias filter — but no more. Oversampling costs power, memory, and DSP cycles at every downstream stage. Vibration monitoring typically needs 1–10 kHz, speech 16 kHz, human motion 50–200 Hz, and slow environmental drift under 1 Hz.
Should I feed raw samples or extracted features to the model?
Extracted features, in almost every MCU case. Handcrafted features (MFCC, spectral bands, statistics) let a much smaller model reach the same accuracy, run faster, and train on less data. End-to-end raw-input learning pays off mainly when you have massive labeled datasets and compute headroom — neither of which describes a typical TinyML product.
How do I tune the detection threshold?
From a validation set of real field data — plot precision/recall or ROC across thresholds, then pick the operating point matching the product's cost of false alarms versus misses. Tune once in the lab, then re-tune with early field data. The right number is a product decision, not a model property.
Why is most of the engineering not the model?
Because correctness is a system property. The model sees what the pipeline feeds it: if windowing is off-by-one, features drift between firmware versions, or thresholds never got field-tuned, the best network in the world misbehaves. Roughly 80% of embedded AI effort is data plumbing, DSP, scheduling, and validation — the model is the easy fifth.
Building something that should run AI on-device?
Edgehound designs, compresses, and deploys TinyML models on microcontrollers — from feasibility audit to field-ready firmware.