TU PRÓXIMO CAPÍTULO
Senior Site Reliability Operations Engineer – Finance
Sobre el puesto
• Lead incident response as Incident Commander, coordinating teams, communications, and service restoration
• Produce executive-level incident reports and run RCAs
• Drive continuous improvement
• Monitor and improve observability using AWS CloudWatch and New Relic
• Reduce alert noise and observability gaps
• Provide hands-on system support across Linux and Windows environments
• Manage and execute deployments via Jenkins, GitLab, or similar CI/CD platforms
• Lead infrastructure initiatives including migrations, upgrades, and process improvements
• Enforce change management and risk assessment for production changes
• Maintain documentation and SOPs
• Liaise between engineering teams and external vendors
• Participate in a 1-week on-call rotation, with possible critical incident call-ins between 6:00 PM and 6:00 AM PT
• 5+ years of experience in Windows and Linux environments with proven troubleshooting capabilities
• Strong knowledge of AWS CloudWatch, New Relic, Nagios, and SumoLogic
• Practical experience with Jenkins, GitLab, CommVault, and AWS Backup
• Strong scripting skills in PowerShell, Python, or equivalent
• Outstanding communication skills, especially under pressure, including executive reporting
• Experience in high-paced environments and with on-call support models
• Ability to manage complex tasks independently
• Must be currently based in Latin America
• 100% Remote Work
• Highly Competitive USD Pay
• Paid Time Off
• Work with Autonomy
• Work with Top American Companies
• Engagement activities
• Work-life balance support
• Collaboration with a diverse, global network
• Work with seasoned senior professionals