← Todos los empleos

Senior Site Reliability Engineer

Sobre el puesto

• Design, develop, and manage applications and infrastructure supporting Akamai's Compute products and services
• Create solutions that improve observability and enforce SLAs across internal teams
• Collaborate with operations and application development teams
• Create tooling and software that monitors and improves system reliability
• Solve complex problems through proactive troubleshooting, automation, and systems programming
• Deploy and maintain the observability platform and internal tooling
• Partner across teams to ensure product and service reliability, scalability, and usability
• Guide engineers and developers in assessing service performance
• Collaborate with support, operations, and engineering teams to investigate and troubleshoot complex problems
• Release new applications and modernize existing tooling

• Bachelor's degree in Computer Science or Engineering
• 6 years of experience in Site Reliability Engineering or a related engineering role
• Linux system administration expertise
• Understanding of networking fundamentals, including TCP/IP, DNS, routing/switching, and storage concepts
• Hands-on experience with containerized environments and Kubernetes, including operating and troubleshooting production systems
• Knowledge of CI/CD and DevOps practices
• Hands-on experience using Jenkins, Git, Prometheus, and Grafana
• Experience with Infrastructure as Code and configuration management using Terraform, Ansible, SaltStack, or similar tools
• Exposure to cloud storage systems
• Automation/scripting skills using Python, Bash, Go, Rust, or similar languages

• Health, well-being, financial, and broader life-support benefits
• FlexBase flexible workplace options: work from home, in an office, or a combination of both