Site Supervisor
Supervises day-to-day operations of critical facilities teams ensuring SLA compliance and safety standards
0+ years
Stable
Key Responsibilities
- Oversee daily operations of mission-critical infrastructure across multiple data center zones
- Coordinate and manage cross-functional facilities maintenance teams ensuring 99.99% uptime
- Conduct comprehensive risk assessments and implement proactive mitigation strategies for power, cooling, and mechanical systems
- Manage emergency response protocols and lead incident resolution for infrastructure failures
- Verify compliance with ANSI/NFPA, OSHA, and enterprise safety standards through regular audits and documentation
- Develop and implement preventative maintenance schedules for high-value data center equipment
- Monitor real-time infrastructure performance metrics and escalate potential system anomalies
Skills & Expertise
Required Skills
- Data Center Infrastructure Management (DCIM)
- Building Management Systems (BMS)
- Electrical Power Management Systems (EPMS)
- High-voltage electrical systems knowledge
- Mechanical and HVAC system troubleshooting
- Facility risk assessment and mitigation
- Network infrastructure management
Preferred Skills
- CDCP (Certified Data Center Professional) certification
- OSHA safety certifications
- ServiceNow asset management
- AutoCAD and facility design software
- Predictive maintenance technologies
A Day in the Life
What a typical workday looks like
My day begins at 6:45 AM when I arrive at the hyperscale data center, reviewing the overnight shift's incident log on my SecureView dashboard. By 7:30, I've completed a comprehensive walk-through of critical zones, verifying CRAC Unit 3's performance metrics and checking ambient temperatures in Rows 12-14. I prioritize today's critical tasks: a firmware upgrade for our core switching infrastructure and investigating a potential thermal anomaly near Rack 214's high-density compute cluster. Mid-morning involves rapid response coordination. I field three urgent ServiceNow tickets: troubleshooting a potential power distribution unit fluctuation, coordinating with our cooling systems vendor about a predictive maintenance alert, and scheduling an emergency cable replacement in our east-side meet-me room. Between communications, I draft detailed incident reports and update our operational risk tracking spreadsheet, ensuring every interaction is meticulously documented. During the scheduled maintenance window, I lead a cross-functional team briefing about our upcoming infrastructure optimization project. We review the implementation strategy for next-generation cooling solutions, discussing thermal efficiency improvements and potential energy consumption reductions. My team collaborates on refining deployment timelines and identifying potential implementation risks in our complex, interconnected environment. The afternoon focuses on deep technical investigation. I analyze recent performance data from our DCIM (Data Center Infrastructure Management) platform, identifying subtle performance degradation in our northwest cooling zone. Using advanced thermal mapping tools, I trace potential root causes, coordinating with our engineering team to develop a proactive mitigation strategy before any critical system impact occurs. As the day concludes, I prepare a comprehensive shift handover report, highlighting key operational status updates, ongoing investigations, and potential emerging risks. I brief the incoming night shift supervisor about the thermal anomaly near Rack 214, transfer active incident tracking documentation, and conduct a final facility walkthrough. By 5:15 PM, I've ensured seamless operational continuity and complete situational awareness for the next team.