Skip to content
technical

AI in IoT: what actually runs in production

4 min read
AI in IoT: what actually runs in production

We run a pipeline that touches hundreds of sites, millions of BLE beacon events per day, and a handful of different hub types. “AI-powered IoT” is in our category whether we like it or not, so we’ve had a long time to figure out which parts of that phrase are doing real work.

Here’s the honest breakdown.

Where AI actually helps in an IoT stack

Edge inference for noise filtering

The most underrated AI application in IoT isn’t a dashboard prediction — it’s a small classification model running on the hub itself, deciding what’s worth sending upstream at all.

BLE environments are loud. A hub scanning for asset beacons in a warehouse will pick up phones, headsets, employee wearables, and a half-dozen vendor-specific beacon formats it was never meant to track. Sending all of that upstream and filtering it in the cloud is expensive and slow. Running a small model at the edge that says “this looks like a tracked asset, this looks like noise” cuts transmission volume dramatically and makes everything downstream cheaper.

The models for this aren’t deep neural nets — they’re usually gradient-boosted classifiers or even logistic regression trained on MAC address patterns, RSSI profiles, and advertisement data structure. Small enough to run on a gateway with 256MB of RAM. Accurate enough to be useful. That’s the right shape for edge inference.

Device fingerprinting from raw scan data

When a new beacon shows up in range of a hub, the question “what is this?” matters before you can route the data anywhere useful. Bluetooth advertisement packets contain enough signal — manufacturer-specific data, service UUIDs, advertisement interval, RSSI decay curve — that a trained classifier can make a reasonable guess at device type without any prior configuration.

We use this to handle fleet hardware that isn’t pre-registered in the system. The model doesn’t have to be right 100% of the time. It has to be right enough that a human reviewing unmatched devices sees a short candidate list, not a wall of hex strings.

That’s a real use case. It runs in production. It’s not magic.

Anomaly scoring on telemetry streams

The thing people call “anomaly detection” in IoT is almost always one of two things: a threshold rule, or a Z-score against a rolling window. Both are useful. Neither requires a model.

Where a model starts to add value is when you have enough historical data to learn what “normal” looks like across multiple correlated signals simultaneously — runtime hours, utilization, connectivity, location variance — and you want to score incoming data against that learned baseline without manually defining every threshold.

Even then, the model’s job isn’t to tell you what’s wrong. It’s to surface the three equipment items most worth a human looking at today. The human still looks.

LLMs for querying telemetry data

This is the newest layer and the one we’re most cautious about.

The pitch: a field ops manager types “which generators ran more than 8 hours yesterday at sites with active jobs?” and the system returns a table. No SQL, no dashboard navigation, no waiting for a report.

The reality: this works surprisingly well for straightforward queries on structured data, and it falls apart unpredictably on anything with joins, time zones, or ambiguous business terms. The same LLM that nails “show me runtime by site this week” will confidently hallucinate a number when the question crosses two data models it’s never been shown together.

Our current approach: a small set of known-good query templates the LLM selects and parameterizes, rather than free-form SQL generation. It covers 80% of what ops managers actually ask. The other 20% goes to a real analyst.

What doesn’t run

Unsupervised anomaly detection on raw RSSI sounds like it should work. RSSI is noisy, environment-dependent, and changes with weather, wall materials, and which employees are standing where. Every model we’ve tried needed so much site-specific retraining that it was cheaper to ask the site manager to set a threshold.

Predictive beacon battery failure keeps coming up in customer conversations. The data doesn’t support it. Battery drain is too linear and too slow for early warning to matter much, and the variance across beacon manufacturers swamps the signal.

Natural language configuration — letting customers describe their routing rules in plain English and having an LLM generate the rule — looks great in demos and creates subtle misconfiguration in production. Business rules need to be explicit and auditable. “Route alerts to the night shift supervisor except on weekends” has ambiguities that a language model will silently resolve in a way that pages the wrong person at 2am.

The pattern we’ve landed on

The most productive framing we’ve found: AI at the boundary, deterministic logic in the middle.

  • Edge: small models for classification and noise filtering, where the cost of upstream transmission makes inference worthwhile
  • Stream layer: statistical anomaly scoring to prioritize what humans review
  • Query layer: LLM-assisted template selection for operational questions
  • Routing and delivery: explicit rules, no model involvement

The moment AI touches the part of the system where a wrong answer has a real consequence — missed maintenance, wrong site, paged-the-wrong-person — you want a human in the loop or a rule you can read and audit.

That’s not a limitation. That’s engineering.