Guides/Techniques

On-Device Anomaly Detection: Catching Rare Events with Sensor Data

Anomaly detection on microcontrollers — autoencoders, statistical baselines, and one-class models that flag the weird without needing labeled failure data.

10 min read·

Most interesting device events are rare — bearing failures, intrusions, leaks, falls. You can't train a classifier on failures you've never recorded, which is exactly the problem anomaly detection solves: learn what normal looks like, flag everything else.

This guide covers the anomaly-detection approaches that fit on microcontrollers and the evaluation trap that catches almost every first implementation.

Learn normal, flag weird

Anomaly detection inverts the supervised problem. Instead of labeled examples of every failure mode, the model trains on normal-operation data — weeks of healthy vibration, typical current draw, ordinary motion. At inference it scores how unlike normal each new window is; cross a threshold and you have an anomaly.

The practical consequence: you can ship a model before the first failure ever happens, and it catches failure modes nobody anticipated. The cost is semantic — the model says 'abnormal', not 'bearing wear'; interpretation is a downstream problem.

Approaches that fit on an MCU

Statistical baselines are the humble workhorse: running mean/variance per feature, z-score or Mahalanobis distance against a learned normal profile — kilobytes of state, microseconds per window. For many structured-signal tasks (current draw, temperature) they match fancier models at zero ML complexity.

Autoencoders are the ML staple: a small network trained to reconstruct normal inputs, where reconstruction error becomes the anomaly score. Compact AE architectures (a few dense or conv layers, bottleneck latent) run in 30–100 KB and catch subtler patterns than statistical baselines. Clustering approaches (k-means on feature space, distance to nearest normal cluster) sit between the two.

  • Statistical baseline — z-score / Mahalanobis on features, KBs of state
  • Autoencoder — reconstruction error as anomaly score, 30–100 KB
  • Clustering — distance to nearest normal cluster in feature space
  • One-class SVM — classic option, heavier but principled

Thresholds, alarms, and the precision trap

The threshold is the product. Set it from validation data: score your normal test set, pick the operating point that yields the false-alarm rate the product can tolerate — one false alarm a week feels very different from ten a day. Reconstruction error distributions are heavy-tailed, so small threshold moves change alarm rates a lot.

Add persistence before alerting: require N anomalous windows in M to declare an event. A single glitchy sample shouldn't wake the fleet. Temporal smoothing typically removes 90% of false alarms while adding seconds of detection delay — a trade most products happily make.

The evaluation trap

Accuracy is meaningless here — a model that never alarms scores 99.9% on a mostly-normal dataset. Evaluation needs precision and recall on a labeled validation set that includes real or injected anomalies, and the honest metric is alarms-per-time-period in the field.

Plan for the threshold to be tuned per deployment. A model trained on one machine's normal may flag another's as permanently abnormal; per-device baselining or normalization layers handle unit-to-unit variation that fleet training data can't.

From anomaly score to root cause

The score tells you when something changed, not what. Production systems pair the anomaly flag with feature breakdown — which input features contributed most to the score — to hint at cause: high-frequency harmonics rising suggests mechanical wear; baseline drift suggests sensor degradation.

The full loop then becomes: device flags anomaly → sends compact event upstream → fleet analytics correlate across units → humans decide. The edge model's job is catching the event early and cheaply; diagnosis is a fleet problem.

Frequently Asked Questions

How much normal data do I need to train an anomaly detector?

Enough to cover the normal operating envelope — typically days to weeks of representative sensor data across the states the device legitimately experiences (idle, running, start/stop transitions). A model trained only on steady-state running will flag every startup as an anomaly. Coverage of normal variety matters more than raw volume.

What if my normal data already contains anomalies?

Contamination happens — a few percent of anomalous samples in training data usually just makes the model slightly more permissive. Above that, clean the data: score with an initial model or statistical filter, drop the outliers, retrain. Iterative self-cleaning converges quickly and is standard practice.

Anomaly detection or supervised classifier — which is better?

Supervised wins whenever you have enough labeled examples of each class — it detects more accurately and names the class. Anomaly detection wins when failure modes are rare, unlabeled, or unknown: you can't label what hasn't happened. Many products run both — anomaly detection catches everything weird; a classifier (trained on accumulated anomalies) names known ones.

Can an anomaly detector adapt as the machine ages?

Yes, with care — slowly drifting 'normal' (seasonal load, component aging) can be absorbed by periodically refreshing the baseline or using running statistics with long time constants. The danger is adapting fast enough to swallow the very anomaly you're watching for, so adaptation rates must be far slower than fault-evolution rates.

Building something that should run AI on-device?

Edgehound designs, compresses, and deploys TinyML models on microcontrollers — from feasibility audit to field-ready firmware.

Related guides