← Todas as vagas

Senior Site Reliability Engineer

Sobre a vaga

• Contribute to the development and operation of a globally deployed, multi-tenant platform built on Kubernetes and cloud-native technologies
• Build automation and tooling to reduce operational toil, improve deployment safety, and accelerate incident response
• Create and maintain SLOs and KPIs
• Collaborate with Engineering, Product, and Support teams
• Participate in on-call rotations and guide restoration and repair of service-impacting issues
• Provide technical and architectural leadership for designing, building, and scaling the Next Generation Control Plane
• Lead transformation toward a highly available, resilient, distributed microservices architecture supporting massive-scale operations

• 8+ years of relevant experience
• Bachelor's degree in Computer Science or similar field, or equivalent experience
• Experience building and operating highly available, fault-tolerant, scalable production services using Kubernetes and other cloud-native technologies
• Solid understanding of Linux internals, especially containerization and networking
• Code-first approach to operating infrastructure
• Knowledge of Go
• Comfortable using Python or Bash for scripting
• Familiarity with infrastructure-as-code tools such as Crossplane, Pulumi, Terraform, or Ansible
• In-depth experience with observability tooling such as OpenTelemetry, Prometheus, Grafana, Loki, or similar
• Ability to approach complex problems with methodical curiosity

• Health, well-being, financial, and life benefits
• FlexBase flexible work arrangements: work from home, in an office, or a combination of both