TU PRÓXIMO CAPÍTULO
Senior Site Reliability Operations Engineer – Finance
Sobre el puesto
• Lead incident response as Incident Commander, coordinating teams, communications, and service restoration
• Produce executive-level incident reports and conduct root cause analyses
• Drive continuous improvement following incidents
• Monitor and improve observability using AWS CloudWatch, New Relic, Nagios, and SumoLogic
• Reduce alert noise and observability gaps
• Provide hands-on system support across Linux and Windows environments
• Troubleshoot complex infrastructure issues
• Manage and execute deployments through Jenkins, GitLab, or similar CI/CD platforms
• Own infrastructure initiatives including migrations, upgrades, and process improvements
• Enforce change management and risk assessment for production changes
• Maintain documentation and standard operating procedures
• Liaise between engineering teams and external vendors
• Support 24/7 stability of internal IT infrastructure and mission-critical backend systems
• 5+ years of experience in Windows and Linux environments with proven troubleshooting capabilities
• Strong knowledge of AWS CloudWatch, New Relic, Nagios, and SumoLogic
• Practical experience with Jenkins and GitLab CI/CD tools
• Practical experience with CommVault and AWS Backup
• Strong scripting skills in PowerShell, Python, or equivalent
• Outstanding communication skills, especially under pressure, including executive reporting
• Experience in high-paced environments and with on-call support models
• Autonomous and proactive attitude; capable of managing complex tasks independently
• Availability for a 1-week on-call rotation, including potential critical incident call-ins between 6:00 PM and 6:00 AM PT
• 100% Remote Work
• Highly Competitive USD Pay
• Paid Time Off
• Work with Autonomy
• Work with Top American Companies
• Engagement activities
• Work-life balance
• Collaboration with a diverse, multicultural global network
• Work alongside seasoned senior professionals