We are living in an era where artificial intelligence is widely promoted as the ultimate cure for operational downtime. Companies are eager to deploy predictive maintenance algorithms, expecting immediate, flawless predictions. However, many quickly encounter a frustratingand often overlooked challenge: they simply do not have the historical failure data required to train these advanced models.
In my experience, this scarcity of failure data happens for one or two reasons.
The first is organisational: failure data is simply not being collected effectively in the field. If an organisation does not systematically and accurately record failures, no amount of algorithmic sophistication can compensate. We must embrace a fundamental operational truth: data work is AI work. If your goal is to generate trusted, actionable intelligence, your underlying failure data must be complete, trustworthy, and usable. Investing in robust failure recording and corrective action systems is not a secondary administrative task; it is thefoundation of any successful AI strategy.
The second reason for data scarcity is a positive engineering reality: some assets simply do not fail very often. Modern critical infrastructure, including the highly reliable cooling systems operating in today's data centres, is designed to exceptionally high standards. As a result, physical failures are rare. While this is excellent from an operational perspective, it creates a challenge for data scientists, who have very little historical failure data available to train traditional machine learning models. After all, how do we predict a failure mechanism that you have never actually observed?
To overcome this challenge, we must adapt our methodology. Instead of waiting for failures to occur, we begin by developing a highly detailed understanding of what normal operating behaviour looks like. Once that baseline has been established, we can generate high-fidelity synthetic data using physics-based simulation models, such as finite element analysis (FEA), to simulate structural degradation and potential failure scenarios. In addition, we can incorporate published industry reliability data and combine it with advanced uncertainty modelling techniques.
For example, when working with cutting-edge data centre cooling systems where historical failure rates are unknown, we can apply structured risk-assessment methodologies such as Failure Modes and Effects Analysis (FMEA). By combining these frameworks with expert engineering judgment and reliability data from adjacent industries, we can estimate failure probabilities and quantify operational risks with a high degree of confidence. We do not need a long history of catastrophic failures to build a reliable predictive system. By combining structured engineering models with synthetic simulations, we can confidently anticipate and prevent downtime before it happens.