Shift Lead
Leads 24/7 shift operations teams, coordinates work assignments, and escalates critical issues
0+ years
Stable
Key Responsibilities
- Coordinate and assign work tasks for shift team members across data center operational zones
- Monitor and respond to critical infrastructure health alerts within defined SLA parameters
- Perform routine equipment and environmental condition inspections, documenting findings in enterprise management systems
- Manage and escalate technical incidents, maintaining comprehensive communication logs and root cause analysis documentation
- Supervise power distribution, cooling systems, and network infrastructure operational continuity during assigned shift
- Conduct shift handover briefings, ensuring complete knowledge transfer and tracking of ongoing operational issues
- Execute emergency response protocols for facility-level incidents involving power, cooling, or security anomalies
Skills & Expertise
Required Skills
- Data Center Infrastructure Management (DCIM) software proficiency
- Building Management Systems (BMS) monitoring and control
- Energy Power Management Systems (EPMS) configuration
- Electrical systems troubleshooting and maintenance
- Network infrastructure and connectivity understanding
- Server rack and equipment deployment
- OSHA safety compliance and risk management
Preferred Skills
- ServiceNow incident management
- AutoCAD technical documentation
- CDCP (Certified Data Center Professional) certification
- Virtualization platform knowledge
- Asset management systems like Maximo
A Day in the Life
What a typical workday looks like
My day starts promptly at 0700 in the facility's operations center, reviewing the overnight shift's summary report in ServiceNow. I quickly scan critical infrastructure metrics from our DCIM dashboard, noting temperature anomalies in Rack 14's west quadrant and a minor power draw irregularity in Row C. I prioritize a thermal investigation and schedule a detailed review of the rack's environmental controls before the morning management brief. By 1030, I'm deep into vendor coordination, troubleshooting a recurring connectivity issue with our network equipment provider. I'm exchanging detailed diagnostics via Slack, pulling historical performance data from our monitoring platform to substantiate the intermittent bandwidth degradation. Simultaneously, I'm updating our incident tracking system, ensuring every communication and diagnostic step is meticulously documented for future reference. During the maintenance window, I'm coordinating a planned firmware upgrade across three high-density compute clusters. My team is strategically staged, with technicians assigned to specific rack groups to minimize potential service interruptions. We're using our standardized change management protocol, running parallel test environments to validate the upgrade sequence before executing the production implementation. The afternoon brings a deep-dive investigation into a persistent latency issue affecting storage array performance. I'm correlating data from multiple monitoring tools - combining insights from our DCIM, network performance dashboards, and server health metrics. My analysis reveals a potential bottleneck in our interconnect infrastructure, prompting a detailed root cause assessment and preliminary mitigation strategy. As my shift concludes, I prepare a comprehensive handover brief for the evening team. I document all open investigations, highlight potential risk areas, and brief the incoming Shift Lead on critical infrastructure status. My final actions involve updating our shift log, confirming all scheduled tasks are tracked, and ensuring smooth operational continuity for the next team.