Hidden Failure 101: The failures your plant already has and nobody has noticed
By Terrence O'Hanlon · Creator, Uptime® Elements — A Reliability Framework and Asset Management System™
Walk out to the standby generator behind your plant. It has been started monthly for eleven years and never once asked to carry the load. Ask the person who signs that sheet why the interval is thirty days.
You will not get an answer. That is the subject of this article.
What is a hidden failure?
A hidden failure is a failure mode that will not become evident to a person or the operating crew under normal circumstances.
The function is already gone. No alarm sounds, no gauge moves, no work order opens. The plant runs exactly as it did the day before.
Hidden failures live in protective and standby equipment — relief valves, interlocks, trips, backup pumps, emergency lighting, fire systems, standby power. Anything that sits quietly and waits.

Why does this matter so much?
Because the hidden failure alone costs nothing. That is precisely why it survives.
The damage arrives at the second event. The primary function fails, the protective device is called on, and it is not there. That joint event — the multiple failure — is where people get hurt and plants are destroyed.
One published RCM analysis specified forty-three failure-finding tasks for hidden failures. The legacy program it replaced contained four. That gap was live, unmanaged risk inside a plant with a full maintenance department and a clean audit record.
Every plant has a number like that. Almost none have measured theirs.
Why don't the usual strategies work?
Three maintenance strategies dominate most programs. None of them reach a hidden failure.
Time-directed maintenance needs an age-reliability relationship to schedule against. Most failures are random rather than age-related, and protective devices sit squarely in that population. There is no age to overhaul against, and intruding on a healthy device adds risk of its own.
Condition-directed maintenance needs a condition to measure. A hidden failure has already occurred. The degradation curve is behind you, not ahead of you.
Run to failure is legitimate when the consequence is genuinely tolerable. A protective device’s failure appears to carry no consequence at all — which is what makes it look attractive here, and why it is wrong. The consequence is not the device failing. It is the device being absent when demanded. Run to failure on a protective function means quietly accepting the multiple failure.

What is failure finding?
A failure-finding task is a scheduled task that seeks to determine whether a hidden failure has occurred or is about to occur.
Its logic is simple. Is the function there? If not, restore it. No measurement, no renewal, no prediction.

How is that different from predictive maintenance?
This is the distinction most programs get wrong.
Predictive maintenance and condition monitoring look forward at a failure that is coming. They detect degradation and buy time to act. Failure finding looks backward at a failure that has already arrived, and shortens the time you are unknowingly unprotected.
One asks is it failing? The other asks is it there?
Point condition monitoring at standby equipment and it finds nothing — not because the equipment is healthy, but because no condition is left to monitor. Buying analytics to solve a hidden-failure problem is the most expensive way to learn that.
How do you start today?
You need one idea. A hidden failure can occur at any point between two tests, so on average you discover it halfway through the interval. The interval you choose is roughly twice the time you are blind.
That turns an unanswerable question into an answerable one. Stop asking how often should we test this? Ask the person who owns the consequence: how long can we tolerate this protection being gone without knowing? Set the interval at about twice their answer.
Four moves, starting this week:
- Build a hidden function register for one system — every function that can fail unannounced, what it protects, and the interval in force today.
- Add a column for where that interval came from. “Unknown” is the most common honest answer, and counting those entries is your finding.
- Take the five highest-consequence entries to their owner and ask the blind-time question.
- Write one design change. The permanent fix is to make the failure evident, not to hunt it forever.
Where does this live in Uptime® Elements?
Hidden failures are a Reliability Engineering for Maintenance (REM) subject, and they travel a clear line across the framework.
Failure Mode Effects Analysis (Fmea) discovers them. Criticality Analysis (Ca) ranks them by consequence. Reliability Centered Maintenance (Rcm) selects failure finding as the task. Reliability Engineering (Re) y Capital Project Management (Cp) eliminate them by making the function evident.
Then execution crosses domains. Operator Driven Reliability (Odr) is the natural home for these tasks, because operators are already in front of the equipment. Planning and Scheduling (Ps) and the CMMS (Cmms) carry the interval. Executive Sponsorship (Es) decides whether the design change is funded or deferred.


How does it read across the diagnostic lenses?
S-D-I-P-F™ places it exactly. Failure finding sits to the right of F — the failure has happened and you are discovering it. Specify and Design are untouched by any task on the menu, and that is where roughly eighty percent of lifecycle cost is already committed.
That makes the Trim Tab one line in a specification: every protective function shall indicate its own availability. Written once, it retires the task, the interval, and the audit trail that would otherwise run for thirty years.
On the Maturity Matrix, faithfully executed tasks with undocumented intervals sit at Planned. Precision begins when the interval comes from tolerable exposure. High-Reliability Organization is reached when the hidden-failure population is shrinking.
And remember which domain does which job. REM is the only domain that reduces failures. Every other improves how efficiently you manage failures that still occur.
How do you build competency?
Apply 10-20-70. Study the REM Passport and the Rcm and Fmea elements — that is your ten percent. Take the blind-time question to a consequence owner and let them set the number in front of you: twenty percent. Then build the register, derive the intervals, and write the design change. That is the seventy that makes a leader.
The Certified Reliability Leader body of knowledge, domain knowledge badges, and Mastery Belt projects administered by the Association of Asset Management Professionals turn that work into a credential and a portfolio.
A hidden function register is a Mastery Belt project waiting for an owner. Most sites have never built one. It fits on a page.
Your plant has hidden failures right now. The only question is whether anyone has looked.
Know. Be. Do. And create a future that was not going to happen anyway.
Hidden Failure 101 Audiobook
listen to the audiobook on The Reliability Leadership Institute website.
Uptime® Elements, Regenerative Reliability™, and S-D-I-P-F™ are trademarks of Reliability Leadership Institute LLC and are licensed to the Association of Asset Management Professionals for credentialing and professional development.
© 2026 Reliability Leadership Institute LLC. All rights reserved.













