← All jobs

Senior Site Reliability Engineer

About the role

• Implement, optimize, and maintain reliable and scalable CI/CD pipelines
• Develop and maintain Infrastructure as Code scripts for automated infrastructure provisioning and management
• Apply Infrastructure as Code best practices across environments
• Identify and automate routine operational tasks
• Implement automation for deployment, scaling, maintenance, and operational processes
• Make SRE and engineering decisions considering technical debt, system design, stability, reliability, monitoring, observability, and business requirements
• Monitor and maintain CI/CD platform resilience
• Troubleshoot complex platform and pipeline issues with senior and staff engineers
• Write post-mortem documentation for internal and external stakeholders
• Contribute to coding standards, engineering practices, and non-functional requirements
• Participate in code and Pull Request reviews
• Mentor junior and mid-level engineers
• Act as a technical reference for CI/CD, Infrastructure as Code, automation, and platform reliability
• Stay current with emerging technologies and industry practices
• Support and execute Proofs of Concept
• Deliver technical solutions supporting platform resilience, continuous improvement, and strategic squad goals

• Must be based in Brazil
• Proficiency in English at B2 level or above (Upper-Intermediate)
• 5 or more years of relevant work experience with a Bachelor's or Associate’s Degree, or at least 2 years of work experience with an Advanced degree (e.g. Masters, MBA, JD, MD)
• Experience implementing, optimizing, and maintaining CI/CD pipelines in production environments
• Experience with Infrastructure as Code and infrastructure automation practices
• Experience working in DevOps, Site Reliability Engineering, Platform Engineering, Cloud Engineering, or a related technical area
• Ability to develop and maintain automation scripts for infrastructure provisioning, deployment, scaling, and operational activities
• Understanding of SRE principles, including system reliability, availability, resilience, monitoring, observability, and incident management
• Experience supporting reliable and scalable platforms and troubleshooting complex issues in distributed environments
• Knowledge of software engineering standards, coding practices, version control, and Pull Request review processes
• Ability to evaluate technical trade-offs involving system design, technical debt, stability, reliability, maintainability, and business requirements
• Experience identifying opportunities to automate manual operational processes and improve engineering efficiency
• Knowledge of monitoring, logging, metrics, alerting, and observability practices
• Experience creating technical documentation and post-mortem reports for technical and non-technical stakeholders
• Ability to collaborate with engineers across different experience levels and provide constructive technical feedback
• Strong analytical and problem-solving skills, with the ability to work on well-scoped and moderately ambiguous technical challenges
• Ability to contribute to technical discussions and escalate broader or cross-squad decisions to senior and staff engineers when appropriate
• Preferred: Experience maintaining CI/CD platforms or internal developer platforms used by multiple engineering teams
• Preferred: Experience working with critical or mission-critical production systems
• Preferred: Experience with cloud infrastructure, container orchestration, and modern deployment practices
• Preferred: Familiarity with GitOps, Infrastructure as Code, and automated infrastructure management
• Preferred: Experience conducting Proofs of Concept and evaluating new technologies for implementation in production environments
• Preferred: Experience mentoring junior and mid-level engineers
• Preferred: Experience with incident response, root cause analysis, and post-mortem practices
• Preferred: Experience working in global or distributed engineering teams
• Preferred: Relevant cloud, DevOps, Kubernetes, or Infrastructure as Code certifications

• Remote work arrangement
• Opportunity to create impact at scale
• Skills growth and professional development opportunities
• Mentoring and professional development support
• Equal employment opportunity protections