A practical guide to commissioning, monitoring, maintenance, leak response and safe operations for liquid-cooled AI infrastructure.

Professional illustration of a Singapore data-centre operations room monitoring liquid-cooled server racks, with simplified icons for temperature, flow, leak detection, alarms and maintenance.

Singapore’s SS 726:2026 provides guidance for implementing liquid cooling in data centres operating in tropical conditions. Announced by Enterprise Singapore and IMDA on 27 August 2026, the standard arrives as AI workloads increase rack power density and make cooling operations more complex.

For facility managers, the key question is not simply whether liquid cooling is more efficient than conventional cooling. The practical question is whether the building, cooling systems, monitoring architecture, maintenance procedures and engineering teams are ready to operate a liquid-cooled environment safely and consistently.

Liquid cooling introduces new interfaces between IT equipment, cooling distribution equipment, water systems, controls and facility operations. It therefore needs to be managed as an integrated facility system rather than as an isolated technology upgrade.

What SS 726:2026 means for facility management

SS 726:2026 is intended to guide the design, commissioning, operation and maintenance of liquid-cooling systems in tropical data centres. It should not be treated as a replacement for project-specific engineering design, equipment instructions, risk assessments, safe work procedures or contractual requirements.

For facility teams, the standard is especially relevant in five operational areas:

  • Confirming that liquid-cooling systems are properly commissioned before handover.
  • Monitoring temperature, humidity, airflow, flow conditions and other critical parameters.
  • Detecting and responding to leaks or abnormal water conditions.
  • Maintaining cooling equipment and maintaining accurate digital records.
  • Coordinating alarms, redundancy, isolation and controlled shutdowns with IT teams.

1. Commissioning should test the complete operating chain

Commissioning should go beyond confirming that pumps run and that a cooling distribution unit, or equivalent equipment, reaches its nominal operating condition. The commissioning team should test the full chain from the heat-producing IT load through the liquid loop, heat rejection equipment, controls, alarms and standby arrangements.

A practical commissioning plan should establish:

  • Normal operating ranges for supply and return temperatures, flow and pressure.
  • Expected performance at different load conditions, including changes in AI workload.
  • Alarm thresholds, alarm priorities, escalation contacts and acknowledgement procedures.
  • Response to loss of power, loss of flow, pump failure, control failure and cooling equipment fault.
  • Isolation points and the sequence for safely removing equipment from service.
  • Interfaces between building management systems, data-centre infrastructure management platforms and IT monitoring tools.

Functional performance tests should be documented, witnessed by the relevant parties and linked to an operating baseline. This baseline gives the FM team a reference for identifying degradation later.

2. Build a monitoring architecture that reflects the risk

Liquid-cooled facilities require more than a single room temperature reading. Monitoring should cover the conditions that affect both the IT load and the cooling system. Earlier tropical data-centre guidance highlights the importance of rack-level temperature, humidity and airflow monitoring, supported by risk assessment and suitable alert thresholds.

Depending on the design, useful points may include:

  • Rack or row-level inlet temperature, humidity and airflow conditions.
  • Liquid supply and return temperature.
  • Flow, pressure and differential pressure across relevant equipment.
  • Leak detection in areas where liquid could reach electrical or IT equipment.
  • Pump status, valve position, filter condition and equipment run hours.
  • Heat rejection equipment status and available standby capacity.
  • Water quality indicators and treatment-system status where applicable.

The objective is not to collect data for its own sake. Each sensor should support a decision, such as reducing load, dispatching an engineer, switching to standby equipment or initiating a controlled shutdown.

3. Treat leak detection as an operational workflow

A leak alarm should trigger a predefined response, not an improvised investigation. The response plan should identify who receives the alarm, who is authorised to isolate equipment, how electrical and IT risks are assessed, and when the incident is escalated.

Facility teams should define:

  • The locations covered by leak detection and the areas that remain dependent on visual inspection.
  • Alarm verification steps that do not expose staff to unnecessary risk.
  • Isolation valves, access routes and equipment boundaries.
  • Coordination requirements with electrical, mechanical and IT personnel.
  • Temporary containment, drainage and recovery arrangements where designed into the facility.
  • Incident recording, root-cause review and post-event corrective action.

Portable inspection equipment and planned visual checks remain important. Digital detection improves response time, but it does not remove the need for sound housekeeping, accessible pipework and clear labelling.

4. Control water quality and contamination risk

Water quality can affect heat transfer, corrosion, fouling, filters, pumps and valves. The appropriate water treatment approach depends on the selected equipment, materials, loop configuration and supplier requirements. Facility managers should therefore establish a water-quality plan with the design and maintenance teams rather than relying on generic targets.

The plan should document sampling frequency, test parameters, acceptable ranges, treatment responsibilities, filtration requirements and actions when conditions move outside the approved range. Records should be connected to the relevant equipment and loop so that recurring trends can be identified.

Any addition, draining, flushing or chemical treatment of the loop should follow an approved method statement and safe work procedure. Uncontrolled top-ups may introduce contamination or make trend analysis less reliable.

5. Use condition monitoring to support predictive maintenance

Liquid cooling creates opportunities for condition-based and predictive maintenance, provided the data is reliable and connected to engineering decisions. Changes in pump vibration, flow, pressure, temperature difference, filter differential pressure or valve behaviour may indicate developing problems before a critical alarm occurs.

AI and digital automation can help by:

  • Combining BMS, DCIM, sensor and maintenance data in one operational view.
  • Identifying abnormal patterns against a commissioned baseline.
  • Prioritising work orders according to asset condition and operational risk.
  • Summarising alarm history and recurring faults for engineering review.
  • Automating inspection reminders, permit workflows and maintenance records.

Automation should support, not replace, engineering judgement. A model-generated alert needs a clear explanation, a responsible owner and a documented action. Teams should also validate sensor quality, time synchronisation and data continuity before relying on analytics.

6. Plan redundancy, isolation and shutdowns before an incident

High-density AI equipment can respond quickly to thermal changes. A facility should therefore define what happens when a cooling component is unavailable, when a loop must be isolated or when a leak cannot be contained immediately.

Shutdown planning should cover normal maintenance, emergency isolation and loss of critical cooling. It should define decision authority, load-reduction steps, communication channels, expected equipment states and restart checks. Procedures should be tested through tabletop exercises and, where safe and approved, controlled functional tests.

Redundancy should also be understood operationally. Standby equipment is only useful if it is available, correctly configured, tested, powered, connected to controls and supported by trained staff. Maintenance windows should verify that the intended standby path actually works.

7. Make the FM team part of the design conversation

The most effective time to resolve monitoring, access, isolation and maintenance issues is before installation and handover. Facility managers should review sensor locations, valve accessibility, drainage, spare parts, safe access, control integration, alarm ownership and documentation during the design and construction stages.

For Singapore businesses considering AI infrastructure or higher-density computing, the practical lesson is clear: liquid cooling is an operational capability, not only a mechanical installation. SS 726:2026 provides a useful framework, while successful implementation depends on disciplined commissioning, reliable data, clear procedures and coordinated human decision-making.

ISS can support organisations reviewing engineering operations, facility management workflows, monitoring requirements and AI-enabled digital services for complex facilities. Contact ISS to discuss your engineering, facility management or AI automation requirements.

Further reading