← All jobs

OpenStack Infrastructure Engineer

About the role

• Build, manage and scale OpenStack compute and storage infrastructure
• Operate and troubleshoot open-source infrastructure systems
• Design and maintain reliable compute and storage hosting platforms
• Troubleshoot across compute, storage, physical and virtual networking, databases, message queues, containers, hypervisors and guest workloads
• Plan and execute upgrades, migrations and high-impact infrastructure changes
• Apply infrastructure-as-code, configuration management and automation practices
• Document architecture decisions, runbooks, change plans, incident findings and recovery procedures
• Participate in technical reviews and provide detailed feedback
• Work in a distributed, largely asynchronous team
• Participate in Rootly-managed on-call rotation and respond to occasional out-of-hours incidents, maintenance or operational requests
• Deliver customer value through stable, reliable and scalable hosting infrastructure

• Experience operating open-source infrastructure systems; extensive experience expected at senior level
• OpenStack implementation, operations or troubleshooting experience
• Experience with Ceph or similar distributed storage systems, including capacity, replication, failure domains, recovery, latency and performance
• Strong Linux systems administration and troubleshooting skills
• Practical knowledge of data-centre-grade hardware, including enterprise servers, CPUs, memory, storage media, NICs, firmware and out-of-band management
• Understanding of how hardware design, power, cooling, rack layout and component failure affect platform reliability and performance
• Strong networking fundamentals; experience in OVN, OVS, BGP underlays, LACP, Juniper, IPv4, IPv6 and physical or virtual cloud networks advantageous
• Ability to troubleshoot across compute, storage, physical and virtual networking, databases, message queues, containers, hypervisors and guest workloads
• Experience with virtualisation and cloud infrastructure at scale
• Security-minded approach to architecture, automation, access control and operational workflows
• SRE mindset focused on reducing failure probability, recovery time, operational risk and data-loss exposure
• Experience with Infrastructure as Code, configuration management and automation practices
• Preference for repeatable, version-controlled automation over undocumented manual work
• Discipline planning and executing upgrades, migrations and high-impact infrastructure changes
• Strong end-to-end ownership from physical infrastructure and network fabric through OpenStack services to customer workloads
• Ability to reason about capacity and performance across CPU, memory, storage, IOPS, latency, throughput, packet rates and control-plane scale
• Ability to use AI-assisted tools productively while understanding their limitations
• Clear technical documentation skills, including architecture decisions, runbooks, change plans, incident findings and recovery procedures
• Ability to contribute constructively to technical reviews, challenge unsafe assumptions and respond well to detailed feedback
• Ability to work effectively in a distributed, largely asynchronous team
• Willingness to participate in an on-call rotation managed through Rootly and respond to occasional out-of-hours incidents, maintenance or operational requests
• Experience with a reasonable number of Ubuntu Linux, Ansible, OpenStack, Ceph, Terraform, Rundeck, LXC, Nspawn, OVN, OVS, BGP, JunOS, LACP, Docker, Prometheus, Grafana, Git, Jira or similar tools

• Salary negotiable, commensurate with skills and track record
• High level of discretion and autonomy
• Meaningful opportunity focused on autonomy, mastery and purpose