Guides/Techniques

Sensor Fusion on Embedded Devices: Combining IMU, Audio & Environmental Data

How to fuse multiple sensor streams on-device — IMU, acoustic, current, environmental — into decisions a single sensor can't make, from Kalman filters to ML classifiers.

10 min read·

One sensor sees one dimension of reality. Sensor fusion combines streams — motion plus sound plus current plus temperature — into decisions no single channel can make, like telling a machine fault from a forklift driving by.

This guide covers the fusion approaches that actually ship on embedded hardware: classical filters, feature-level fusion, and model-level fusion.

Why fuse at all

Each sensor is ambiguous alone. An accelerometer sees vibration but can't tell a bearing fault from a passing truck; a microphone hears the truck but not the bearing. Together they disambiguate — the false-positive rate of a fused classifier is routinely a fraction of the best single-sensor model.

Fusion also covers gaps: when one channel saturates, clips, or drifts, the others carry the decision. Robustness against sensor degradation is a product feature users notice only when it's missing.

Three fusion levels

Data-level fusion combines raw or lightly processed streams before features — a 6-axis IMU's x/y/z accelerometer and gyroscope channels feeding one network is the everyday example. Feature-level fusion extracts features per stream, concatenates them, and feeds one model: audio MFCCs joined with vibration statistics and current harmonics.

Decision-level fusion runs separate models per sensor and merges outputs — voting, weighted scores, or a small meta-model. It's the most modular (each sensor model ships and validates independently) but leaves cross-sensor patterns unlearned.

  • Data-level — raw streams merged early, model sees everything
  • Feature-level — per-sensor features concatenated, one model
  • Decision-level — separate models, merged outputs, most modular

The classical workhorse: Kalman filtering

Before ML, sensor fusion meant Kalman filters — recursive state estimation that blends predictions with noisy measurements under uncertainty. Orientation estimation (fusing accelerometer + gyroscope + magnetometer into attitude), position tracking, and sensor-noise smoothing are still Kalman's home turf, at kilobyte-scale compute.

The hybrid pattern that wins products: Kalman/state estimation produces cleaned, physically meaningful signals — orientation, velocity, machine state — that then feed an ML classifier. Physics does the estimation; the network does the semantics.

Synchronization and sampling realities

Fusion is only as good as time alignment. Sensors sampled on independent clocks drift — a 100 Hz IMU and a 1 Hz temperature probe need a resampling or timestamping strategy before their windows can be fused. Skew that looks harmless in the lab becomes a systematic error source in the field.

Buffering strategy follows from the fusion level. Feature-level fusion needs aligned windows across sensors — typically handled with ring buffers and a window-complete event that triggers the fused feature computation and inference.

Sizing fused models for MCUs

A fused model sees more input features than any single-sensor model, but it doesn't scale linearly — early layers shared across modalities do the compression. Depthwise-separable convolutions and small attention-free sequence models handle multi-channel inputs in tens of KB.

Budget the full pipeline, not just the model: per-sensor DSP + fused inference + post-processing must fit the cadence and energy envelope. Often the answer is asymmetry — a fast cheap check on one sensor that gates the expensive fused inference.

Frequently Asked Questions

Which fusion level should I start with?

Feature-level fusion is the pragmatic default: it captures cross-sensor patterns without demanding perfectly aligned raw streams, keeps the model small, and makes it easy to add or drop a sensor later by extending the feature vector. Move to data-level when raw cross-channel interactions matter (multi-axis IMU) and decision-level when sensors are developed and validated independently.

Do I still need Kalman filters if I'm using ML?

Often yes — they solve different problems. Kalman filters estimate physical state (orientation, position, smooth signals) under noise with provable behavior; ML classifiers assign semantics to that state. Kalman-cleaned features fed to a classifier is one of the most reliable embedded AI patterns.

How do I handle sensors with different sample rates?

Choose a common decision cadence (say, once per second), then compute each sensor's features over that window at its native rate — the slow sensor contributes fewer samples per window. Align windows by timestamp, and handle missing data explicitly rather than letting interpolation silently invent values.

Does sensor fusion increase power consumption?

It adds per-sensor sampling and DSP cost, but often reduces total system power: a fused classifier is usually gated behind a cheap single-sensor check that runs always-on. The expensive multi-sensor inference fires only when the cheap stage suspects something — keeping the always-on budget near single-sensor levels.

Building something that should run AI on-device?

Edgehound designs, compresses, and deploys TinyML models on microcontrollers — from feasibility audit to field-ready firmware.

Related guides