← All jobs

Senior Software Engineer – Reliability

About the role

• Improve investigation-agent root-cause accuracy so engineers can trust diagnoses without re-checking
• Expand auto-remediation from a handful of playbooks to dozens using staged autonomy
• Build and run fault-injection benchmarks and evaluation harnesses
• Set evaluation pass marks, report results with honest denominators, and validate the harnesses
• Design fail-closed checks, kill switches, approval flows, and blast-radius limits for production actions
• Build new agents that remove categories of operational toil
• Develop clean interfaces, guardrails, and safe defaults for other teams using the reliability platform
• Stop production incidents from worsening and permanently remove recurring failure classes
• Contribute to Atlan's mature, opinionated codebase and improve its documented architecture and eval-gated pull requests
• Collaborate across teams to drive adoption of the reliability platform

• Experience carrying a pager, owning incidents end to end, or working a support or escalation queue
• Demonstrated ability to eliminate recurring operational problems rather than merely optimize runbooks
• Experience measuring outcomes using adoption metrics, real numbers, and honest denominators
• Experience rebuilding a core part of one's work with AI and shipping an AI-native workflow used by others
• Experience architecting agents that take autonomous action, including defining guardrails
• Understanding of false-positive rates, rollback paths, and blast radius
• Experience building platforms for other teams, including clean interfaces, documented failure modes, and safe defaults
• Ability to contribute quickly to a high-rigor existing codebase and improve it without rewriting it
• Ability to write deterministic code for routing, filtering, and safety before generative steps
• Genuine interest in reliability and operational efficiency
• Availability for overlap between roughly 11am and 8pm IST

• Strong base salary
• Performance-based variable pay
• Impact-driven equity for most roles
• Health, dental, vision, and mental health benefits from Day 1
• Flexible health stipends
• Flexible time off
• Modern leave policies
• Accelerated growth and learning opportunities
• Global, remote-first, high-trust work environment
• Work from anywhere with a diverse team across 15+ countries
• Async work environment
• Flexibility and ownership over how you work