Wake-word detection — 'Hey device' — is TinyML's flagship workload: an always-on microphone pipeline that listens at microwatts and wakes the rest of the product only when the magic words are spoken.
This guide covers the pipeline architecture, model design, and the false-wake tuning that decides whether users love or hate the feature.
The two-stage pipeline
Production wake-word systems are staged. A cheap always-on stage — sometimes a hardware voice-activity detector, sometimes a tiny model — runs continuously, gating the power-hungrier keyword model. The second stage only fires when the first suspects speech, keeping average power near the cheap stage's draw.
The audio pipeline itself is standard: 16 kHz sampling, 20–40 ms frames, MFCC or log-mel features per frame, and a small temporal model (CNN, DS-CNN, or compact GRU) that sees a rolling 1-second window. Each stage adds power budget — the design job is making sure the expensive parts run rarely.
Choosing and training the word
Not all phrases are equally learnable. Distinctive multi-syllable words with strong phoneme diversity ('hey silicondog') separate cleanly from background speech; short or common words ('on', 'dog') collide with everyday audio constantly. Pick the wake phrase for discriminability before branding commits to it.
Training data is the usual bottleneck. You need thousands of utterances across voices, accents, distances, and noise conditions — which is why teams augment recordings with noise mixing, room-simulation (convolution with impulse responses), and text-to-speech variety. Negative data matters just as much: hours of TV, speech, and household audio teach the model what the wake word is not.
- Long, phoneme-diverse phrases detect better
- Augment: noise mixing, room impulse responses, TTS variety
- Negative examples (speech, TV, household noise) are half the dataset
- Field recordings beat studio recordings
False wakes vs missed wakes
Every wake-word product lives on a trade-off curve: false wakes (fires when nobody said it) versus false rejects (ignores the real word). Users tolerate occasional misses; they don't tolerate a device that lights up constantly during TV shows. Tune the threshold for precision first, then measure recall in field conditions.
Post-processing does heavy lifting: score smoothing across consecutive frames, a refractory period after each wake, and secondary confirmation (a slightly larger model or a confidence margin) all cut false wakes dramatically at minimal compute cost.
Power budget in practice
The math that matters: mic + ADC + DSP frontend draws the baseline; each inference costs energy at the MCU's active current. A model invoked once per 30 ms frame at a few milliamps lands in the 1–5 mW range — good for mains-powered devices, marginal for coin cells.
Battery products go further: hardware VAD in the tens-of-microwatts tier, duty-cycled DSP, and inference only when speech energy is detected. This is the architecture inside always-listening wearables and battery doorbells — the always-on budget lives in microwatts, and the model is the exception, not the rule.
Frequently Asked Questions
How small can a wake-word model be?
Production keyword-spotting models run 30–80 KB int8 — a small DS-CNN or CNN over MFCC features fits comfortably in a Cortex-M4's memory. Research models dip below 20 KB at some accuracy cost. The model is rarely the constraint; the always-on audio frontend's power draw usually is.
Can I have multiple wake words or custom phrases?
Yes — multi-keyword models classify a handful of phrases plus 'unknown' in one pass, and it's the standard approach for products with several commands. Custom wake words per user require retraining or few-shot adaptation; that's why most products fix the phrase and customize what happens after the wake.
How do I test wake-word performance before shipping?
Build a test set of field-style audio: target phrases across voices/distances/noise, plus hours of negative audio (speech, TV, music) for false-wake measurement. Measure per-hour false wakes and recall at distance — these field-style metrics predict user experience far better than held-out clean audio.
Is wake-word detection the same as speech recognition?
No — wake word (keyword spotting) detects one or a few fixed phrases and runs entirely on-MCU at microwatts. Full speech recognition transcribes arbitrary speech and needs far bigger models — typically on-device NPUs or the cloud. Products usually combine them: the MCU wake word triggers a connection to a bigger ASR system.
Building something that should run AI on-device?
Edgehound designs, compresses, and deploys TinyML models on microcontrollers — from feasibility audit to field-ready firmware.