O SEU PRÓXIMO CAPÍTULO
Staff Software Engineer – L4
Sobre a vaga
• Own the reliability posture of production services, including availability, latency, capacity, efficiency, performance, monitoring, and alerting
• Define, instrument, and operate against SLIs and SLOs, using error budgets to drive engineering priorities
• Identify stability risks and mitigate them before they affect customers
• Reduce repair items and prevent recurring incident classes
• Improve detection, response, and recovery times
• Strengthen failure domains, validate recovery paths, and make production changes safer
• Participate in on-call and lead responses when production is degraded
• Write post-mortems, identify root causes, and drive follow-up work to completion
• Identify, diagnose, report, and document production problems
• Write, configure, and deploy maintainable, reviewed, documented, and tested code that improves reliability
• Orchestrate complex changes across systems and services, including design changes, technical decisions, migrations, and upgrades
• Lead debugging, troubleshooting, and analysis of service architecture and design
• Use code review to improve code quality
• Reduce operational overhead for infrastructure and services
• Drive cross-team projects from conception to completion
• Coordinate programs, estimate and communicate delivery timelines, track milestones, and update stakeholders
• Identify and communicate changes that may impact stability
• 8+ years of related engineering experience, with a substantial portion focused on reliability, infrastructure, or platform engineering
• Demonstrated accountability for production systems, including having carried a pager for important services and owned outcomes during failures
• Strong software engineering fundamentals and ability to build and ship production code
• Experience defining and operating against SLIs and SLOs, and using error budgets to inform engineering priorities
• Depth in production operations, including incident command, post-mortem analysis, capacity planning, and observability
• Record of preventing recurrence by reducing incident classes and operational toil
• Experience driving changes across multiple teams and building alignment without formal authority
• Track record of improving engineers through code review, design feedback, and mentorship
• Experience with large-scale distributed systems in a cloud environment
• Familiarity with infrastructure-as-code, container orchestration, and GitOps-style delivery
• Experience with multi-region architecture, failure-domain design, or regional expansion work
• Background in chaos engineering, game days, or other proactive resilience validation
• Competitive pay
• Generous time off
• Ample parental and wellness leave
• Healthcare
• Retirement savings program
• Support for employee volunteering and donation efforts