← Toutes les offres

SRE Engineer

À propos du poste

• Define and track SLIs, SLOs, SLAs, MTTR, and MTTD.
• Implement observability, monitoring, alerting, and APM.
• Monitor latency, traffic, errors, saturation, availability, and performance.
• Prevent, identify, and resolve incidents.
• Lead root cause analyses and define actions to prevent recurrence.
• Identify risks, bottlenecks, and single points of failure.
• Support the design of resilient, scalable, and highly available solutions.
• Automate operational activities and reduce manual tasks.
• Operate and evolve Kubernetes and Docker environments.
• Support capacity planning, business continuity, and disaster recovery strategies.
• Participate in deployments and support application stabilization.
• Collaborate with teams to improve reliability from the solution design stage onward.
• Create and maintain dashboards, alerts, procedures, and operational documentation.
• Promote a culture of reliability, observability, and continuous improvement.

• Experience as a Site Reliability Engineer, SRE, or in an equivalent role.
• Hands-on experience with cloud environments using GCP, AWS, and/or Azure.
• Knowledge of Kubernetes and Docker.
• Experience with observability, monitoring, alerting, and APM.
• Knowledge of SRE metrics and practices, such as SLI, SLO, SLA, MTTR, and MTTD.
• Experience managing, investigating, and resolving incidents.
• Knowledge of application and infrastructure troubleshooting.
• Experience administering Linux environments.
• Knowledge of networking, security, performance, and high availability.
• Experience with automation and Infrastructure as Code.
• Experience with CI/CD pipelines.
• Strong communication skills and the ability to work with multidisciplinary teams.
• Analytical, proactive, collaborative, and prevention-oriented mindset.
• Preferred: Experience with GKE, EKS, or AKS.
• Preferred: Knowledge of Dynatrace, Datadog, Grafana, Prometheus, or similar tools.
• Preferred: Experience with the ELK Stack, Elasticsearch, and Kibana.
• Preferred: Knowledge of Terraform and Ansible.
• Preferred: Experience with mission-critical environments and distributed systems.
• Preferred: Experience working in financial institutions or regulated environments.
• Preferred: Experience with capacity management and cloud cost optimization.
• Preferred: Knowledge of disaster recovery and business continuity.
• Preferred: Experience defining and managing error budgets.
• Preferred: Cloud, Kubernetes, or SRE certifications.

• Meal voucher
• Food allowance
• Home office allowance
• Medical insurance
• Dental insurance
• Life insurance
• Birthday day off
• TotalPass / Wellhub app
• Boon Saúde
• Discount partnerships
• Agreements with businesses and educational institutions
• Welcome kit
• Verity onboarding
• Verity Learning Interval
• Great Place to Work certification