O SEU PRÓXIMO CAPÍTULO
ML Infrastructure Operations Engineer
Sobre a vaga
• Configure and maintain ML environments using Kubernetes, Docker, and YAML-based configurations.
• Execute, monitor, and validate distributed ML training workloads across multi-node accelerator clusters.
• Review, modify, and execute Python and Bash scripts to automate workload execution and adjust runtime configurations.
• Troubleshoot workload failures, collect logs, identify infrastructure or configuration issues, and document findings.
• Track workload execution status, communicate test progress, and collaborate with engineering teams to improve reliability and operational efficiency.
• 3+ years of experience in ML Operations, Software Testing, Systems QA, Linux Systems Administration, or a related technical role.
• Strong proficiency with Linux command-line operations and shell environments.
• Experience reading and modifying Python and Bash scripts.
• Familiarity with Kubernetes, Docker, and containerized environments.
• Experience executing and monitoring distributed workloads.
• English: C1
• Sueldo base
• Seguro de Gastos Médicos Mayores (incluye plan dental y visión)
• 15 días de aguinaldo
• 25% de prima vacacional
• 12 días de vacaciones (A partir del primer año)
• Seguro social
• Vales de despensa quincenales