Critical Infrastructure Manager
Manages critical systems operations teams and ensures 100% uptime for power, cooling, and network infrastructure
0+ years
Stable
Skills & Expertise
Required Skills
- Data Center Infrastructure Management (DCIM)
- Building Management Systems (BMS)
- Electrical Power Management Systems (EPMS)
- High-voltage electrical system design and maintenance
- Mechanical systems engineering (HVAC, cooling infrastructure)
- Redundancy and fault-tolerance architecture
- Risk assessment and critical infrastructure protection
- Compliance with NFPA, ASHRAE, and ISO data center standards
Preferred Skills
- ServiceNow IT service management
- Maximo asset management
- AutoCAD electrical/mechanical design
- CDCP (Certified Data Center Professional) certification
- OSHA safety certification
- Professional Engineer (PE) license
- Advanced thermal management and computational fluid dynamics
- Industrial control systems (ICS) security
A Day in the Life
What a typical workday looks like
My day begins at 7:15 AM, arriving early at the hyperscale data center facility. I immediately pull up the overnight monitoring dashboards in Splunk and review the infrastructure health reports from our DCIM platform. No critical alerts triggered during the night shift, but I notice slight temperature variances in Rack 14's northwest quadrant that require closer investigation. I prioritize a thermal mapping review and schedule a detailed inspection of CRAC Unit 3 before the morning team briefing. By 10:30 AM, I'm deep in vendor coordination mode. A Schneider Electric field technician is on-site to discuss potential UPS system firmware upgrades, and I'm simultaneously triaging an intermittent network latency issue reported by our connectivity team. I document each interaction meticulously in ServiceNow, ensuring our change management logs are comprehensive and traceable. Three separate support tickets require immediate cross-functional communication with our network and facilities teams. The midday maintenance window opens precisely at noon. Our team executes a planned power redistribution across Zones A and B, carefully load-balancing critical infrastructure without service interruption. We leverage our redundant power architecture to migrate workloads seamlessly while performing planned maintenance on backup generator circuits. Real-time collaboration using our secure communication platform keeps everyone aligned and informed during the complex evolution. By 2:30 PM, I'm conducting a deep-dive root cause analysis on yesterday's minor cooling efficiency degradation. Using thermal imaging data and historical performance metrics from our BMS (Building Management System), I develop a predictive model to optimize cooling distribution. The investigation reveals subtle airflow restrictions in specific hot/cold aisle configurations that could impact long-term energy efficiency. As the day winds down, I prepare a comprehensive shift handover report in our incident management system. I brief the night shift technician on all ongoing work streams, potential risk areas, and critical system status. Before departing, I conduct a final walkthrough of the primary data hall, physically verifying critical infrastructure parameters and ensuring all systems are operating within optimal parameters. Another day of 100% uptime successfully managed.