YOUR NEXT CHAPTER
Principal Site Reliability Engineer
About the role
• Shape compute-platform strategy and qualification criteria for x64, ARM, accelerator, and inference hardware across firmware, OS, runtime, and reliability
• Lead provisioning and CI/CD improvements, including reproducible, safe, and scalable bare-metal provisioning, imaging, configuration, and infrastructure delivery pipelines
• Drive reliability and observability by addressing systemic failures, enhancing telemetry and diagnostics, and establishing safe production rollout and recovery safeguards
• Investigate complex hardware/software incidents, identify root causes, guide decision-making, and convert findings into engineering fixes
• Review designs, mentor engineers, establish standards, and resolve multi-domain issues across teams and vendors
• Enable new hardware through provisioning services and provide seamless firmware upgrade models
• Ensure reliable server behavior across Akamai datacenters
• Expertise in Linux systems and server platforms, including boot processes, storage, networking, hardware diagnostics, firmware, BIOS/UEFI, BMCs, and server lifecycles
• Experience designing and operating bare-metal provisioning, configuration management, and scalable infrastructure automation
• Fluency in Python and Bash
• Experience building automation, APIs, deployment pipelines, and debugging multi-layer failures
• Ability to identify whether hardware-test failures originate from firmware, BIOS/UEFI configuration, device firmware, kernel/driver behavior, or the physical test environment
• Expertise in observability, metrics, capacity analysis, incident root-cause investigation, and production readiness
• Practical knowledge of x86 and ARM platforms, accelerators, and inference infrastructure, including drivers, runtimes, and compatibility
• Hands-on experience with Linux virtualization technologies, including KVM, QEMU, and libvirt
• Understanding of nested virtualization, CPU virtualization extensions, VM networking and storage, and Linux-level troubleshooting of virtualized environments
• Health, well-being, finances, and life beyond work benefits
• FlexBase workplace flexibility: work at home, in an office, or a combination of both