Back to All Roles
EngineeringSenior Level

AI Infrastructure Specialist

Manages GPU cluster infrastructure and high-power density deployments for AI and machine learning workloads

Experience

0+ years

Growth

Stable

Skills & Expertise

Required Skills

  • Data Center Infrastructure Management (DCIM)
  • Building Management Systems (BMS)
  • Electrical Power Management Systems (EPMS)
  • High-density cooling infrastructure design
  • Rack and power distribution unit (PDU) configuration
  • Infrastructure redundancy and fault tolerance planning
  • Network and server infrastructure deployment

Preferred Skills

  • ServiceNow asset management
  • AutoCAD infrastructure design
  • Schneider Electric/APC power management tools
  • CDCP (Certified Data Center Professional) certification
  • Advanced thermal modeling and computational fluid dynamics

A Day in the Life

What a typical workday looks like

My alarm sounds at 6:45 AM, and I'm reviewing the overnight telemetry from our AI infrastructure clusters before I've finished my first coffee. The dashboard from our Prometheus monitoring system shows Rack 14's GPU cluster maintained 99.98% uptime, but there's a slight temperature anomaly in CRAC Unit 3 that needs immediate investigation. I prioritize today's tasks: resolving the cooling efficiency issue, preparing for a major firmware update across our NVIDIA DGX systems, and conducting a capacity planning review for next quarter's machine learning workload expansion. By 10:30 AM, I'm deep in vendor coordination, troubleshooting a persistent latency issue with our latest interconnect fabric. Our Mellanox support engineer and I are analyzing packet loss metrics, exchanging diagnostic logs through our secure communication channels. Simultaneously, I'm updating our infrastructure runbooks in Confluence, documenting the root cause analysis process for similar future incidents. A high-priority support ticket from the research team about a stalled training job requires immediate triage. During the scheduled maintenance window, I'm coordinating with the team to perform a rolling update on our Kubernetes cluster. We're carefully migrating workloads between nodes, minimizing any potential service disruption. Our team uses a combination of Ansible playbooks and custom scripts to ensure precise, synchronized updates across the infrastructure. A brief architecture review meeting helps align our infrastructure optimization strategies with the upcoming AI model training requirements. The afternoon is focused on performance forensics. I'm using advanced tracing tools like eBPF to investigate deep system-level performance characteristics of our GPU clusters. Analyzing power consumption patterns and thermal dynamics, I'm developing recommendations for our next-generation rack design. Our custom Python scripts help aggregate and visualize complex infrastructure metrics, revealing potential optimization opportunities in our high-density compute environment. As the shift winds down, I'm preparing a comprehensive handover report for the night team. Logging all system states, pending maintenance tasks, and potential risk areas into our centralized tracking system. I brief the incoming infrastructure specialist about the cooling anomaly in Rack 14, the firmware update preparation, and the ongoing performance investigation. A final systems check confirms all critical AI infrastructure components are stable and ready for the next wave of computational workloads.

Career Progression

Previous Roles

Entry-level position

You Are Here

Engineering

AI Infrastructure Specialist

Senior Level

Next Steps

Senior leadership position

Related Roles