Guides/Deployment

Gesture & Activity Recognition from IMU Data on Microcontrollers

Turning accelerometer and gyroscope data into gestures and activities: windowing, features, model architectures, and the classic pitfalls of IMU-based ML.

9 min read·

The IMU — accelerometer plus gyroscope — is the most ML-friendly sensor in embedded: cheap, low-power, and rich with learnable patterns. Wrist gestures, falls, gait, tool use, machine motion — all visible in a few axes of sampled motion.

This guide covers the standard IMU-ML pipeline and the pitfalls that sink first attempts.

Windows are the unit of meaning

A single IMU sample is meaningless; a gesture is a pattern over a window. Standard practice: 1–4 second windows at 50–200 Hz, with 50% overlap between windows for smooth detection. Window length trades responsiveness against context — long enough to contain the gesture, short enough to feel instant.

Sampling rate follows the physics. Human motion lives under ~20 Hz; 50 Hz captures it comfortably. Machine vibration and impact events need kilohertz. Oversampling wastes power and model input size for no accuracy gain.

Features or raw input?

Two families of approach. Feature-based: compute per-window statistics (mean, variance, energy, dominant frequency, spectral entropy) and feed a small dense network — tiny model, trains on hundreds of examples, highly interpretable. Raw-input: feed the window itself to a small CNN that learns the features — more data-hungry but captures subtle morphology.

The pragmatic route is hybrid: start with statistical features and a small classifier (it's 90% of the value in days), then move to raw-input models when the feature approach plateaus. Many shipped products never need the second step.

  • Statistical features + dense net — smallest, trains on little data
  • Raw windows + small CNN — captures shape, needs more data
  • Hybrid — features for screening, CNN for hard classes

Orientation invariance and mounting

The classic IMU pitfall: the model memorizes orientation. A wrist gesture model trained on data where the watch faced a particular way fails when users wear it differently. Fixes: augment training data with rotations, feed orientation-invariant features (magnitude, per-axis statistics), or normalize orientation at the pipeline level.

Mounting consistency matters more than the fix, though. If sensor position varies freely — pocket vs hand vs wrist — either constrain the mounting in the product design or collect data across every expected placement. Models can't learn invariance to variations they've never seen.

Classes, context, and rejection

Define an 'other' class or a confidence floor from the start. A model forced to choose among N gestures will confidently misclassify random motion; a rejection path ('none of the above') is what separates demos from products.

Activity recognition adds a temporal layer: single windows give noisy labels ('walking', 'walking', 'sitting'?). Smoothing across windows — majority vote, HMM-style state logic — turns jittery per-window guesses into stable activity segments users can trust.

Frequently Asked Questions

How much training data does IMU gesture recognition need?

Feature-based approaches work with tens of examples per class; raw-input CNNs want hundreds to thousands per class, across users. Whatever the approach, collect data from many different people — IMU patterns are highly user-specific and models trained on three volunteers fail on the fourth.

Accelerometer only, or accel + gyro?

Gyroscope adds rotational information that disambiguates many gestures (twists vs shakes) at a modest power cost — the IMU usually has both anyway. For battery-critical always-on use, accel-only at low rate with the gyro gated on is a valid staged architecture.

Can the model detect falls on a wearable?

Fall detection is a flagship IMU application — high-g impact signature plus post-impact stillness, detectable on a small model. The challenge is the data: real falls are rare and dangerous to collect, so training uses stunt/synthetic data plus anomaly-detection strategies, and field tuning against false alarms from impacts like dropping the device.

How do I handle left-handed users or different placements?

Either normalize at the input (mirror augmentation, canonicalize the dominant axis) or collect training data across every supported placement. Multi-placement products often ship per-placement models selected by a simple placement classifier — cheap and effective.

Building something that should run AI on-device?

Edgehound designs, compresses, and deploys TinyML models on microcontrollers — from feasibility audit to field-ready firmware.

Related guides