← Todos los empleos

Staff Engineer, Core – MLOps

Sobre el puesto

• Architect and advance the control plane, including service registry, schema registry, SLO enforcement, and CLI tooling
• Design and build the context plane to aggregate operational signals and create automated feedback loops for self-healing maintenance
• Own the multi-language service chassis and golden path, including Java and Python client libraries, workload specifications, Helm charts, and deployment pipelines
• Define gRPC and Protocol Buffer inter-service contracts, API gateway transcoding, versioning policies, and schema evolution rules
• Operate the platform substrate across Kubernetes, Terraform, ingress infrastructure, Kafka, billing pipelines, Valkey, and database modernization initiatives
• Lead Requests for Discussion on workflow orchestration, gateway orchestration, multi-cluster routing, and automated failover
• Establish SLOs and error budgets, health-aware traffic isolation, and automated weighted canary deployments
• Participate in shared infrastructure on-call rotation and lead incident post-mortems
• Mentor engineers across squads, review system proposals, and establish engineering best practices
• Serve as technical lead for the Core & MLOps squad and influence architecture and release velocity across five neighboring squads

• 10+ years of experience building scalable distributed backend systems
• Proven track record of authoring internal platforms or core libraries widely adopted by engineering teams
• Expertise in Java, using reactive frameworks such as Vert.x or Netty
• Strong proficiency in Python
• Deep gRPC and Protocol Buffers experience
• Experience managing schema evolution and backward compatibility in mission-critical environments
• Hands-on production Kubernetes expertise at scale
• Experience with Terraform and event streaming with Kafka
• Experience designing automated telemetry pipelines, materialized views, or feature stores that dynamically adapt system behavior based on production data
• Proven success defining SLOs/SLIs, analyzing blast radius, and designing fault-tolerant systems built around rigorous service contracts
• Exceptional technical writing skills
• Strong interpersonal and written communication skills for a globally distributed, remote-first environment
• Experience with Java 21, Helm, CircleCI, OCI, GCP, Hetzner, Servers.com, Confluent Kafka, BigQuery, Valkey, MySQL/PostgreSQL, HBase, Google Pub/Sub, Prometheus, Grafana, Loki, and OpenTelemetry
• Bonus: experience with Temporal, DBOS, model serving, performance monitoring, drift detection, SPIRE, mTLS, Cilium, Istio, Envoy, CLIs, SDKs, project generators, web scraping/crawling, or open-source contributions

• Freedom and flexibility to work from where you do your best work
• Flexible working hours
• Attend conferences and meet with team members from across the globe
• Work with cutting-edge open source technologies and tools
• Remote-first culture
• Continuous innovation and complex technical challenges
• Global community and collaboration with distributed systems engineers and data specialists across the world