With much attention devoted to data center buildout, there is one factor that receives much less attention but could prove just as much of a bottleneck: maintenance.
Outages have become a significant challenge for the burgeoning industry: in the Uptime Institute’s Annual Outage Analysis 2026, half of the respondents say their data center has experienced an impactful outage in the past three years. Among them, 87% believe the incident could have been prevented through better management, processes or configuration.
This number points to an opportunity to improve operational practices, with one area to examine is how teams turn equipment alerts into timely, well-controlled maintenance work.
The challenge is amplified by workforce constraints that make those technical demands harder to absorb through traditional operating models. Uptime reported that 46% of operators struggled to find qualified candidates in 2025 and 37% had difficulty with retention. That means fewer experienced people are available to diagnose faults, execute maintenance, supervise contractors and pass operational knowledge between shifts.
Maintenance practices have yet to catch up with critical-infrastructure status
Despite these advances, it is striking that maintenance practice in parts of the sector remains largely reactive or calendar-based.
Research reported by ETDatacenters in August 2026 found that 21% of 300 U.K. data center managers still relied primarily on reactive maintenance, which means intervening only after something failed. Another 28% used preventive maintenance based largely on fixed schedules, while just 26% reported a fully predictive strategy.
The same research found an average unplanned downtime of six hours over the previous year and only 53% of respondents were fully confident in their maintenance strategy.
That gap becomes harder to sustain as data centers assume a more critical economic role and customers expect near-continuous availability. Reactive maintenance exposes operators to the full consequence of failure, while purely calendar-based maintenance can consume scarce labor on equipment that remains healthy or miss degradation that develops between scheduled interventions. The priority is to match maintenance tasks to asset criticality and failure modes, using condition information where it improves decisions, and retain scheduled work where it remains appropriate. This helps focus limited engineering capacity on the work that most effectively protects availability.
Moving from monitoring to connected execution: the case of a hyperscaler
Addressing this maintenance gap starts with connecting the systems, information and workflows that already exist across the facility.
Verdantix identifies four structural barriers that limit performance: fragmented systems, a continuing divide between IT and facilities, point-function software organized around individual tasks and leaner teams with fewer qualified technicians. The implication is that better monitoring alone will not resolve the problem, because an alarm has limited value if the technician responding to it cannot immediately see the asset’s history, current configuration, relevant procedure and outstanding work.
The first step to escape this predicament is to create a reliable asset hierarchy, a clear maintenance system of record and enough condition information to distinguish routine work from interventions that genuinely protect availability. An Enterprise Asset Management (EAM) platform can provide the asset records and maintenance workflows needed to support this approach, with asset performance management capabilities helping teams use condition information to prioritize work.
A top-five global hyperscaler offers a case in point. As it sought to improve uptime while supporting continued growth, its teams realized that they lacked sufficient insight into the condition of assets such as Uninterruptible Power Supply (UPS) and cooling equipment. They also relied on multiple siloed systems and still used paper-based procedures for some field activities.
Rather than replace every existing platform, the company introduced Octave Attune EAM as a common maintenance environment that could integrate with its wider operational landscape. The move supported the shift to reliability-centered maintenance and root-cause analysis, supported by trusted enterprise asset information, while technicians gained mobile access to checklists, parts requests, follow-on work orders and supporting CAD or BIM information.
Giving the context and procedures needed at the point of work
As in the case of that hyperscaler, the asset registry is the necessary foundation to more effective maintenance. What needs to be built on top of it is consolidated and contextualized data that shortens the path from alarm to correct action.
Connecting asset records with current design and engineering information can help technicians understand equipment dependencies, locate relevant drawings and review maintenance history. A digital twin is one way to bring this information together. The practical value depends on keeping the information accurate, up to date and accessible at the point of work.
The full context here can also mean the appropriate procedures. While it’s tempting to imagine a data center outage as a hardware issue, operators tell a different story. 31% say human error is a major contributor to recent outages they have experienced, with “staff failed to follow procedures” being the primary type (cited by 59%), and “incorrect procedures” the second (36%). Digital workflows, such as those supported by Octave Tempo Operating Procedures, can help teams access current procedures and follow a consistent process at the point of work. Their effectiveness also depends on procedures being accurate, maintained and supported by appropriate training and supervision. That is, of course, particularly needed at a time when operators struggle to staff their teams with experienced employees. Processes and procedures that rely on memory, tribal knowledge and incomplete information instill fragility in operations.
Taken together, these four components build on the same idea: reliable AI infrastructure depends on more than building capacity. It requires four capabilities that work together: enterprise asset management to maintain accurate asset records and maintenance workflows; asset performance management to prioritize work using condition information; information management to put design data, drawings and asset history at the point of work; and digital procedures that teams follow consistently.
None is sufficient on its own, but together they allow operators to go beyond monitoring. They turn equipment alerts into timely, well-controlled maintenance and help protect availability. That becomes increasingly important as the sector scales and is more widely treated as critical infrastructure.