Back to Standards
Uptime & ReliabilityGlobal

Tier Standard: Operational Sustainability

TSOS

Certification that validates ongoing operational excellence and management effectiveness.

Issuing Body: Uptime InstituteCode: TIER-OPERATIONAL-SUSTAINABILITYOfficial Website

Purpose

Ensures facilities maintain design tier performance through proper operations, maintenance, and management.

Requirements Overview

Operational procedures; Staff training; Change management; Risk assessments; Maintenance programs; Performance monitoring

Overview

The Uptime Institute's Tier Standard: Operational Sustainability (TSOS) emerged in response to a critical challenge in data center management: the persistent gap between initial infrastructure design and long-term operational performance. Developed as a comprehensive successor to traditional operational assessment methodologies, TSOS addresses the industry-wide problem of data center performance degradation after initial commissioning. Historically, data center certifications focused solely on infrastructure design, providing a snapshot of capabilities at a single point in time. TSOS represents a paradigm shift by introducing a dynamic, ongoing validation process that requires continuous proof of operational excellence. The standard recognizes that infrastructure reliability extends far beyond initial design specifications, demanding rigorous documentation, staff competency, and consistent performance monitoring. At its core, TSOS provides a framework for data center operators to demonstrate sustained operational discipline. Unlike previous standards, it requires comprehensive evidence of procedural execution, staff training, change management, and continuous improvement. The certification has become a critical benchmark for mission-critical facilities, particularly in industries where infrastructure reliability directly impacts business continuity and customer trust. The standard's significance lies in its holistic approach to data center management. It mandates not just what facilities claim to do, but provides a structured method to prove consistent execution of operational protocols. By setting standards comparable to high-reliability industries like aviation and healthcare, TSOS ensures that data centers maintain the highest levels of operational integrity throughout their lifecycle.

Key Requirements

Documented Operational Procedures with Site-Specific Protocols

Data centers must maintain comprehensive, current operational procedure documentation that covers all critical infrastructure systems including cooling management, power distribution, environmental monitoring, security protocols, and emergency response procedures, with each procedure containing clear step-by-step instructions, decision trees for exception handling, and specific parameters for the facility's equipment configurations.

Procedures must be more than generic templates—they must reflect the actual facility design, equipment models, configuration settings, and operational constraints specific to that data center, and must include documented evidence that procedures are reviewed and updated within defined intervals (typically annually or following significant infrastructure changes).

Staff must demonstrate competency by passing procedure-based assessments and technical exams that validate their understanding of site-specific operational requirements.

Formal Change Management and Configuration Control

All changes to critical infrastructure, software configurations, network settings, environmental controls, and security systems must flow through a formal change management process that includes documented change requests with business justification, technical impact assessments, rollback procedures, change windows scheduled during low-risk periods, and post-implementation verification with documented results.

The standard requires maintaining a current configuration baseline document and change log that demonstrates all modifications are tracked, authorized by qualified personnel, tested before implementation, and verified to confirm they did not degrade system performance or redundancy.

This requirement is significantly more stringent than typical IT change management because it encompasses physical infrastructure modifications that could affect cooling airflow, power distribution capacity, or environmental monitoring sensor placement.

Risk Assessment and Mitigation Program Tied to Availability Targets

Data centers must conduct documented risk assessments that identify failure scenarios specific to their infrastructure design, quantify probability and impact of each risk, and implement mitigation controls with measurable effectiveness metrics.

TSOS requires assessment of both hardware risks (component failure rates, supply chain vulnerabilities, single points of failure) and operational risks (procedure gaps, staff skill deficiencies, maintenance scheduling conflicts, change implementation errors), with mitigation strategies documented and implemented with clear accountability assignments.

Facilities must demonstrate that risk mitigation investments directly correlate to achieving their target availability levels, with evidence showing how specific controls reduce identified risks to acceptable levels.

Preventive and Corrective Maintenance Programs with Performance Correlation

Data centers must maintain documented preventive maintenance (PM) schedules for all critical systems (CRAC/CRAH units, chillers, UPS systems, generators, PDU components, monitoring sensors) that specify maintenance intervals based on manufacturer recommendations adjusted for facility-specific operating conditions, with documented records proving execution of all scheduled maintenance activities.

The standard requires tracking maintenance effectiveness through metrics such as mean time between failures (MTBF), maintenance labor hours consumed, spare parts inventory turnover, and correlation of maintenance performance to facility availability metrics.

When corrective maintenance is required, facilities must document root cause analysis, implement corrective actions to prevent recurrence, and verify effectiveness through follow-up monitoring.

Staff Competency Validation and Continuing Education Requirements

All personnel performing critical operational functions must demonstrate documented technical competency through written exams, hands-on practical assessments, and initial qualification periods where performance is monitored by senior staff before independent authorization.

TSOS requires maintaining individual competency records for each staff member with their role-specific qualifications, exam scores, practical assessment results, and continuing education completion dates, ensuring operators maintain current knowledge through periodic recertification (typically every 1-2 years).

The standard specifically requires competency validation in areas such as equipment-specific shutdown procedures, emergency response protocols, environmental monitoring system operation, and site-specific configuration knowledge.

Performance Monitoring and Metrics Tracking with Data Center-Specific KPIs

Facilities must implement continuous automated monitoring systems that capture operational data for all critical infrastructure domains including power distribution efficiency, cooling system performance, environmental conditions (temperature, humidity, air pressure differential), equipment utilization, and system availability metrics, with monitoring data retained for audit review periods (typically minimum 3-5 years).

TSOS requires establishing facility-specific key performance indicators (KPIs) based on design specifications and operational targets, then tracking actual performance against those KPIs with monthly or quarterly reporting that correlates operational activities (maintenance, configuration changes, staffing changes) to performance variations.

The standard mandates that performance data must be analyzed to identify trends, trigger preventive maintenance when metrics approach warning thresholds, and drive continuous improvement decisions.

Infrastructure Inspection and Condition Assessment Programs

Data centers must conduct documented periodic inspections of all critical infrastructure components (cooling system piping for corrosion or leaks, UPS battery terminals for oxidation, PDU connections for loose contacts, sensor calibration verification) with established inspection intervals, documented assessment criteria, photographic or video evidence of conditions, and remediation actions for identified deficiencies.

Inspections must include thermal imaging surveys to identify hot spots in cooling distribution, visual inspection of generator fuel systems and battery connections, verification of cable management compliance with airflow standards, and documentation of any deferred maintenance items with justification and completion target dates.

The standard requires trending of inspection findings to demonstrate that corrective actions prevent recurrence of identified issues.

Emergency Response and Business Continuity Procedures with Documented Testing

Data centers must maintain documented emergency response procedures covering scenarios such as loss of utility power, chiller failure, environmental sensor malfunction, security breach, and natural disasters, with each procedure specifying escalation paths, decision authorities, communication protocols, and specific actions to preserve availability during crisis conditions.

TSOS requires that emergency procedures be tested through documented drills or simulation exercises at least annually, with results recorded showing that staff could execute critical procedures within required timeframes and that procedures successfully achieved their intended outcomes.

Testing must include scenarios such as simulated generator failure, controlled cooling system shutdown to verify failover sequences, and emergency notification system activation to validate that all required personnel receive timely alerts.

Who Uses & Why

TSOS certification becomes essential for data center operators facing significant operational risks or stringent service level agreements (SLAs). Mission-critical environments in finance, healthcare, government, and telecommunications sectors should prioritize this standard to demonstrate operational reliability to customers and stakeholders. Mandatory certification scenarios include: - Facilities with SLA penalties exceeding $10,000 per hour of downtime - Organizations subject to regulatory compliance requirements (e.g., NIS Directive in Europe, HIPAA in healthcare) - Hyperscale cloud and colocation providers supporting enterprise-level customers Optional but beneficial scenarios include: - Facilities experiencing recent operational challenges or staff turnover - Data centers planning infrastructure expansions - Organizations seeking to differentiate their operational capabilities in competitive markets Geographic considerations vary, with European and North American markets showing the most stringent requirements. The standard's global applicability makes it valuable for multinational organizations seeking consistent operational standards across multiple locations. Cost and complexity considerations should be carefully evaluated. While certification requires significant investment in documentation, training, and process improvement, it typically provides substantial long-term benefits in risk mitigation and customer confidence.