TU PRÓXIMO CAPÍTULO
Site Reliability Engineer II
Sobre el puesto
• Deploy and maintain the observability platform and internal tooling
• Partner across teams to ensure product and service reliability, scalability, and usability
• Guide engineers and developers to increase confidence in service performance
• Collaborate with support, operations, and engineering teams to investigate and troubleshoot complex problems
• Improve monitoring and analysis platforms for rapid error detection and remediation, including automated remediation
• Participate in on-call rotations and guide restoration and repair of service-impacting issues
• Maintain system reliability and enhance monitoring, alerting, and log aggregation
• Develop advanced tools using Kubernetes, Kafka, and ClickHouse
• Collaborate with AWS, Azure, and Linode to drive continuous improvements and optimize performance
• 2+ years of relevant experience
• Bachelor's degree in Computer Science or related field
• Ability to design and implement a comprehensive monitoring and observability strategy for complex software products
• Experience utilizing ArgoCD for GitOps-based deployments and Kubernetes application delivery workflows
• Experience creating and managing CI/CD workflows with GitHub Actions, including building, testing, and deploying pipelines
• Experience defining and implementing frameworks and tools
• Expertise with Kubernetes, Docker, and third-party clouds such as AWS, Azure, or Linode
• Experience with Linux-based infrastructure and Bash/Python
• Expertise troubleshooting across network, system, application, and database layers, including network protocols, Linux OS, and SQL/KQL
• Health, well-being, financial, and life benefits
• FlexBase flexible work arrangements: work at home, in an office, or a combination of both