Guides/Applications

Edge AI for Consumer Devices: Voice, Gesture & Context on a Coin Cell

What actually runs on-device in consumer products: wake words, gesture controls, context awareness for wearables, hearables, and smart home — at power budgets that fit a coin cell.

10 min read·

The consumer device is the harshest product environment AI ships into: battery budgets in microwatts, users who judge in seconds, and privacy expectations that grow every year. The winners run intelligence on-device by default.

This guide is the consumer-device map: what runs, at what power, and which guides go deeper on each workload.

Voice: the always-on problem

Every voice product starts with the same architecture — a hardware or TinyML wake stage running at microwatts, gating the expensive inference. Keyword spotting on a Cortex-M keeps a hearable or smart speaker listening all day on a coin cell; the second-stage command models run only on trigger.

The consumer-specific challenge is acoustics diversity: the same wake phrase must fire across voices, accents, bathroom echo, and TV background — and false-wake during a movie kills reviews faster than any missing feature. Field-style negative data is the discipline.

Gesture and context: the IMU everywhere

The accelerometer is the cheapest sensor in the BOM and the most ML-friendly: tap/double-tap controls on hearables, wrist gestures on wearables, gesture-free fall detection, head-gesture nod/shake detection on earbuds. These run on tiny IMU classifiers — 20-60 KB, inference in milliseconds, power in microamps.

Context models go further: 'in ear vs in case', 'walking vs transport', 'indoor vs outdoor' — classifiers that gate power-hungry features (GPS, ANC profiles, always-on listening) on the inferred state. Context is invisible when it works; users notice only the battery life.

  • Wake words — microwatt always-on stage, MCU keyword spotting
  • Gestures — tap/nod/wrist controls from IMU at 20–60 KB
  • Context — in-ear, in-transit, indoor inference gates features
  • Presence — mmWave/audio for smart home, privacy-first

Presence and event detection at home

The smart-home tier adds presence sensing (mmWave micro-motion, audio presence models) and audio event detection (glass break, alarm beep, baby cry) — all on-device so nothing recorded leaves the room. The privacy story is the product feature: 'nothing leaves your home' is the pitch that sells.

Battery-powered nodes (door, window, leak sensors) push the duty-cycle discipline hardest — always-on sensing is impossible at coin-cell budgets, so cheap hardware triggers gate the model. The guide on power budgeting is the engineering core.

The consumer deployment reality

Consumer fleets are the largest in embedded AI — hundreds of thousands to millions of units — which makes the OTA model update loop existential: models must improve in the field from anonymized fleet signals, because the launch model is never the best one the product ships.

Cost discipline is the other consumer reality: every BOM cent multiplies by fleet size, which is why the 'slightly bigger MCU' debate matters — and why TinyML models engineered to run on the existing chip often beat the NPU upgrade on unit economics alone.

Where it shows up

Hearable voice control

Always-on keyword spotting plus on-ear tap/gesture detection on the earbud MCU — no phone, no cloud, coin-cell battery.

Smart-home presence

mmWave + audio models detect stillness and count people for thermostats and lighting — private by default.

Wearable context

Activity + in-transit + indoor classifiers gate GPS, sensors, and features — the battery life users notice.

Frequently Asked Questions

What's the smallest always-on AI feature that ships?

Hardware wake-word triggers and tap detection run in the tens-of-microwatts tier — often inside the sensor or codec itself, waking the MCU only on trigger. Above that, a TinyML keyword spotter on a Cortex-M at a few mW is the standard always-listening budget.

How do consumer devices handle the model update problem?

The same OTA pattern as industrial — versioned model blobs, staged rollout, health-gated progression — but at consumer scale the fleet-data flywheel is the real advantage: millions of devices reporting compact event features make the next model version dramatically better.

Why do consumer products choose edge over cloud?

Privacy (home audio/video never leaves), latency (wake-word response can't wait on a round trip), cost (no per-device cloud inference at million-unit scale), and reliability (the feature works during outages). On-device is the default; cloud is the enhancement layer.

Can consumer AI run without a dedicated NPU?

Usually yes — the bulk of shipped consumer TinyML (wake words, gesture, context, presence) runs on standard Cortex-M/RISC-V with DSP-optimized kernels. NPUs appear in flagship tiers (phones, premium hearables) for heavier models; the entry tier is still an efficient MCU.

Building something that should run AI on-device?

Edgehound designs, compresses, and deploys TinyML models on microcontrollers — from feasibility audit to field-ready firmware.

Related guides