Fault detection in building automation: why FDD is harder than it looks
Photo by Aleksandr Lyaptsev on Unsplash.
Fault detection and diagnostics, FDD for short, compares how a piece of building equipment is actually running against how it should be running, using the sensor and BMS data a building already produces. When a gap shows up, FDD flags it before it turns into wasted energy, an uncomfortable floor, or an equipment failure nobody saw coming. Smart building systems generate more sensor readings than any operator can review by hand, which is exactly why FDD exists as a category.
Facilities teams usually meet FDD bundled inside a BMS, a stand-alone analytics platform, or a module inside broader property management software. Wherever it sits, the promise is the same: catch building automation issues while they're still cheap to fix. The technique works in principle. Getting it to hold up in a live building, across years of controls changes and technician turnover, is the hard part. Most FDD deployments stop earning their keep within a year of going live.
The reason usually isn't the concept. It's what happens when FDD logic meets real building data. Sensors drift for years without recalibration. Points get renamed after a retrofit and nobody updates the map. Thresholds get set once at commissioning and stay untouched for a decade. Four problems account for most of why fault detection breaks down in practice, and understanding them is the difference between a system operators trust and one they've muted.
Sensor drift turns FDD into a false-alarm generator
FDD rules are only as good as the readings feeding them. A stuck damper trips the same threshold as a drifting temperature sensor, and rule-based FDD has no way to tell the two apart on its own. A supply-air sensor reading half a degree high doesn't announce itself. It just nudges every rule downstream of it toward the wrong verdict.
Drift is the specific failure that hurts FDD hardest, because it happens slowly enough to look like normal variation and long enough to accumulate into a real error. A rule tuned against last year's calibrated sensor starts firing against this year's drifted one. Or the opposite happens: the drift shifts the baseline just enough that a genuine fault now reads as within the normal band, and the rule stays quiet when it should fire. Either way, the FDD system is grading against a moving target it doesn't know has moved.
We've written separately about the data-quality failures behind this: sensors that lie quietly, gaps in delivery, points nobody can name. FDD inherits every one of them, then adds its own failure mode on top. Instead of just producing a gap in a chart, bad input data turns into a confident, specific, wrong verdict about equipment health.
Missing point context confuses rule-based FDD logic
Most FDD rule sets assume they know what a point measures. A rule that checks whether return air temperature tracks supply air temperature needs to know, with certainty, which tag is the return sensor and which is the supply. In a lot of BAS deployments that mapping was inferred from a point name a controls technician wrote a decade ago, never verified against anything since.
When the mapping is wrong, or was right once and got remapped after a retrofit, the rule doesn't fail loudly. It evaluates against the wrong pair of points and produces a plausible-looking, entirely wrong result. The FDD system reports a fault that isn't there, or misses one that is, and there's no error message. As far as the rule engine is concerned, it did its job correctly.
The root problem sits in the metadata, not in the model. A rule engine can only reason about relationships it's been told exist. FDD deployments that go live on unverified point maps to get alerts running fast tend to lose operator trust just as fast, for reasons that have nothing to do with the detection logic itself.
Alarm fatigue is what actually kills FDD, not the algorithm
Set an FDD threshold tight enough to catch small faults early and it fires on every minor swing a building produces in a normal week. Set it loose enough to stay quiet during normal operation and it misses the early, cheap-to-fix stage of a real fault. Most deployments drift toward loose thresholds within the first few weeks of complaints, which quietly defeats the point of catching anything early.
Operators don't ignore alarms because they're lazy. They ignore them because the ratio of noise to signal makes triage a worse use of their day than just checking the plant in person. Once an operator has muted a fault category twice for a non-issue, the system has effectively lost that category for good, whether or not anyone updates a configuration to reflect it.
A related version of the same problem shows up across correlated points. One physical fault, a stuck valve or a failed damper actuator, often trips several rules at once across sensors that all sit downstream of it. Ten related alarms for one root cause reads as ten problems on a screen. The operator ends up doing the correlation work the system should have done, every single time, until they stop reading the list altogether.
Rule brittleness makes more rules the wrong fix
The standard response to false alarms and missed faults is more rules. An exception for this equipment type. A seasonal adjustment for that climate. A suppression window around known maintenance events. Each addition solves the specific case it was written for and narrows what the rule covers everywhere else.
A rule tuned for a rooftop unit in Malmö in February doesn't generalize to the same unit type in a warmer climate, a different control sequence, or August. FDD rule sets patched for years accumulate exceptions faster than they accumulate coverage. Eventually the technical maintenance burden of keeping the rule set current rivals the burden it was supposed to reduce. At that point another rule isn't the fix. A different way of deciding what counts as a fault is.
What grounded, ranked FDD looks like instead
The alternative isn't a smarter rule engine. It's a system that checks its own inputs before it trusts them, and ranks what it finds instead of listing it. Explore's anomaly detection runs against the physics of the plant rather than fixed thresholds alone, and causal filtering collapses the correlated alarms from one root cause into a single finding before an operator ever sees them.
Where a reading disagrees with the physics around it, that disagreement gets flagged as a drift or data problem, not reported as an equipment fault. That's the grounded-inference discipline in practice: the system only surfaces a finding it can back with the data behind it, and says so plainly when it can't. Nothing invented, nothing padded to fill a dashboard.
Findings that survive that check land in a single ranked queue, priced and evidenced, instead of an alarm panel operators have learned to tune out. The dashboard is the question. The queue is the answer. That's how we run BMS analytics across HVAC, IoT and every monitoring system already installed in a building, without adding hardware or replacing the BAS underneath it, and without burning technical maintenance hours chasing noise.
FAQ
What is fault detection in building automation (FDD)?
FDD compares how a piece of equipment is actually running against how it should be running, using the sensor and BMS data a building already produces, so a fault surfaces early instead of after it wastes energy or fails outright. It runs on standard protocols like BACnet, Modbus and OPC UA, over data most buildings already generate.
Why does FDD produce so many false alarms?
Mostly because rule-based FDD trusts its inputs by default. Sensor drift, ambiguous point mapping and thresholds set once at commissioning all push good rules toward bad conclusions, and the rule has no way to know its input has gone stale.
Can adding more rules fix alarm fatigue?
Rarely for long. Each new rule or exception narrows the coverage it was written for and adds another case to maintain. FDD rule sets built this way tend to accumulate exceptions faster than they gain reliability, which is why the fix belongs at the input and ranking layer, not in another rule.
Does FDD replace a BMS or a CMMS?
No. FDD reads the data a BMS already produces. It doesn't replace the control system, and it isn't a maintenance-scheduling tool. Explore surfaces and ranks what's wrong. Deciding how a work order gets logged and assigned stays with whatever CMMS or process a team already runs.
How does FrostLogic reduce false positives in fault detection?
By checking data against the physics of the building before treating it as a fault. Explore validates readings against the points they should agree with, then runs causal filtering to collapse correlated alarms into one finding, surfacing only what it can back with evidence, ranked by cost rather than dumped into an alarm list.
What's your building not telling you?
If FDD in your buildings has turned into a list nobody opens, tell us what's tripping the most: nuisance alarms, or a rule set nobody trusts anymore. We listen first, then tell you straight whether Explore helps. 30 or 60 minutes, your pick. No commitment either way. Talk it through.
FrostLogic Explore brings sensor intelligence, scenario simulation, and grounded-inference AI to commercial and industrial buildings. Learn more about Sensor Intelligence or talk it through with us.
Curious how this would look on your building?
What's your building not telling you?
Tell us what you're trying to figure out: energy drift, a BMS you don't trust, compliance you're chasing. We listen first, then tell you straight whether Explore helps. 30 or 60 minutes, your pick. No commitment either way.
:quality(80))