TU PRÓXIMO CAPÍTULO
SRE Engineer
Sobre el puesto
• Define and track SLIs, SLOs, SLAs, MTTR, and MTTD
• Implement observability, monitoring, alerting, and APM
• Monitor latency, traffic, errors, saturation, availability, and performance
• Prevent, identify, and resolve incidents
• Conduct root cause analyses and define actions to prevent recurrence
• Identify risks, bottlenecks, and single points of failure
• Support the design of resilient, scalable, and highly available solutions
• Automate operational activities and reduce manual tasks
• Operate and evolve Kubernetes and Docker environments
• Support capacity planning, business continuity, and disaster recovery strategies
• Participate in deployments and support application stabilization
• Work with teams to improve reliability from the solution design stage
• Create and maintain dashboards, alerts, procedures, and operational documentation
• Promote a culture of reliability, observability, and continuous improvement
• Experience as a Site Reliability Engineer, SRE, or in an equivalent role
• Hands-on experience with cloud environments using GCP, AWS, and/or Azure
• Knowledge of Kubernetes and Docker
• Experience with observability, monitoring, alerting, and APM
• Knowledge of SRE metrics and practices, such as SLI, SLO, SLA, MTTR, and MTTD
• Experience managing, investigating, and resolving incidents
• Knowledge of application and infrastructure troubleshooting
• Experience administering Linux environments
• Knowledge of networking, security, performance, and high availability
• Experience with automation and Infrastructure as Code
• Experience with CI/CD pipelines
• Strong communication skills and the ability to work with cross-functional teams
• Analytical, proactive, collaborative, and prevention-oriented mindset
• Experience with GKE, EKS, or AKS
• Knowledge of Dynatrace, Datadog, Grafana, Prometheus, or similar tools
• Experience with the ELK Stack, Elasticsearch, and Kibana
• Knowledge of Terraform and Ansible
• Experience with critical environments and distributed systems
• Experience working in financial institutions or regulated environments
• Experience with capacity management and cloud cost optimization
• Knowledge of disaster recovery and business continuity
• Experience defining and managing error budgets
• Cloud, Kubernetes, or SRE certifications
• Meal allowance
• Food allowance
• Home office allowance
• Health insurance
• Dental insurance
• Life insurance
• Birthday day off
• Total Pass / Wellhub
• Boon Saúde app
• Discount partnerships
• Discounts at partner establishments and educational institutions
• Welcome kit
• Onboarding
• Verity Learning
• Verity Break
• #VerityComVocê
• Access to professional development courses
• Great Place to Work-certified workplace