← Todos los empleos

Senior Site Reliability Engineer

Sobre el puesto

• Create solutions that improve automation and efficiency for systems and teams
• Optimize workflows, infrastructure, and applications
• Collaborate on deployment, monitoring, and incident resolution
• Improve system monitoring for faster error detection and remediation
• Enhance the performance and reliability of the virtualization platform
• Develop and maintain automated tools and scripts for system reliability, deployment, and incident response
• Participate in on-call rotations and guide restoration and repair of service-impacting issues
• Write automation and tooling to reduce operational toil and improve deployment safety
• Contribute to capacity planning, autoscaling configuration, and workload scheduling for AI compute infrastructure
• Provide support and mentorship to other engineers
• Promote continuous improvement and operational excellence across systems

• Expert-level experience in a SysAdmin (Linux/Unix Administration), DevOps, or SRE role
• Experience working with large-scale distributed systems
• Expertise in Kubernetes and large-scale containerization systems
• Proficiency in at least one programming language: Python or Golang
• Experience with configuration management using Terraform, SaltStack, or Ansible
• Experience defining SLOs
• Experience with observability tools such as Prometheus and Grafana
• Experience with distributed tracing
• Experience architecting software and infrastructure at scale
• Ability to develop automation and monitoring
• Ability to collaborate effectively with engineering teams unfamiliar with SRE practices

• Support for career development through GROW and Mentoring programs
• Internal development events such as the APEX Expo
• Access to LinkedIn Learning
• Opportunities to build new skills, explore new roles, and try different opportunities
• 15-minute exploratory call with the Recruiter