O SEU PRÓXIMO CAPÍTULO
SRE Engineer
Sobre a vaga
• Define and track SLIs, SLOs, SLAs, MTTR, and MTTD
• Implement observability, monitoring, alerting, and APM
• Monitor latency, traffic, errors, saturation, availability, and performance
• Prevent, identify, and resolve incidents
• Lead root cause analyses and define actions to prevent recurrence
• Identify risks, bottlenecks, and single points of failure
• Support the design of resilient, scalable, and highly available solutions
• Automate operational activities and reduce manual tasks
• Operate and evolve Kubernetes and Docker environments
• Support capacity planning, business continuity, and disaster recovery strategies
• Participate in deployments and support application stabilization
• Collaborate with teams to improve reliability from the solution design stage
• Create and maintain dashboards, alerts, operational procedures, and documentation
• Foster a culture of reliability, observability, and continuous improvement
• Experience as a Site Reliability Engineer, SRE, or in an equivalent role
• Hands-on experience with cloud environments using GCP, AWS, and/or Azure
• Knowledge of Kubernetes and Docker
• Experience with observability, monitoring, alerting, and APM
• Knowledge of SRE metrics and practices, such as SLI, SLO, SLA, MTTR, and MTTD
• Experience managing, investigating, and resolving incidents
• Knowledge of application and infrastructure troubleshooting
• Experience administering Linux environments
• Knowledge of networking, security, performance, and high availability
• Experience with automation and Infrastructure as Code
• Experience with CI/CD pipelines
• Strong communication skills and the ability to work with multidisciplinary teams
• Analytical, proactive, collaborative, and prevention-oriented mindset
• Preferred: experience with GKE, EKS, or AKS
• Preferred: knowledge of Dynatrace, Datadog, Grafana, Prometheus, or similar tools
• Preferred: experience with the ELK Stack, Elasticsearch, and Kibana
• Preferred: knowledge of Terraform and Ansible
• Preferred: experience with critical environments and distributed systems
• Preferred: experience in financial institutions or regulated environments
• Preferred: experience with capacity management and cloud cost optimization
• Preferred: knowledge of disaster recovery and business continuity
• Preferred: experience defining and managing error budgets
• Preferred: certifications in Cloud, Kubernetes, or SRE
• Meal voucher
• Food allowance
• Home office allowance
• Medical insurance
• Dental insurance
• Life insurance
• Birthday day off
• TotalPass / Wellhub
• Boon Saúde app
• Discount partnerships
• Discounts at partner businesses and educational institutions
• Welcome kit
• Onboarding program
• Verity Learning
• Verity Break
• #VerityComVocê