← Todas as vagas

Platform Architect, AI/ML Infrastructure, GCP-focused

Sobre a vaga

• Build and operate model and inference serving infrastructure for real-time and batch inference across multiple tenants
• Manage latency, throughput, autoscaling, and reliability
• Own the ML deployment lifecycle, including model registry, versioning, promotion workflows, rollout strategies, and safe rollback
• Operate agentic and LLM workloads in production, including inference providers, gateways, quotas, throttling, guardrails, prompt/version management, and graceful degradation
• Build reproducible, automated training, evaluation, and deployment pipelines as code
• Extend infrastructure-as-code practices to ML systems using Terraform and multi-project design
• Operate GitOps for ML workloads and own ArgoCD configuration and promotion workflows
• Run ML and AI workloads on multi-tenant Kubernetes/GKE, managing GPU scheduling, workload placement, tenant isolation, and capacity
• Own ML reliability and observability, including inference SLOs, model/data drift detection, regression monitoring, alert quality, on-call ergonomics, and runbooks
• Drive ML cost efficiency through accelerator right-sizing, committed-use and Spot VM capacity management, and cost attribution
• Use agentic coding tools to scaffold environments, generate/review IaC and pipeline code, and accelerate automation
• Identify platform problems proactively and shape platform evolution

• 5+ years in platform engineering, SRE, MLOps, or infrastructure, including meaningful time operating production systems at scale
• Hands-on experience deploying and operating ML or AI workloads in production
• Strong SRE/DevOps foundation, including ownership of reliability, SLOs, post-mortems, and measurable improvements
• Deep Terraform expertise, including complex state, reusable modules, multi-project configurations, and CI-driven plan/apply workflows
• Strong GitOps background with ArgoCD or Flux in production
• Deep Kubernetes knowledge and production cluster operations, including control-plane-level troubleshooting
• Production GKE experience is strongly preferred
• Strong GCP background, including VPC networking, Compute Engine, IAM, Cloud Storage, and multi-project/organization design
• Hands-on production BigQuery experience, including partitioning, clustering, query cost/performance tuning, and dataset-level IAM
• Familiarity with Dataflow, Pub/Sub, or Dataproc
• Hands-on experience building and operating CI/CD pipelines
• Understanding of differences between ML pipelines and standard application CI/CD
• Senior-level automation-first thinking
• Active use of agentic coding tools
• Strong communication skills
• Experience with GPU/accelerator scheduling and node lifecycle management is nice to have
• Experience operating LLM inference at scale is nice to have
• Experience with ML pipeline and orchestration tooling is nice to have
• Experience with model registries, feature stores, and experiment tracking is nice to have
• Familiarity with model and data drift monitoring and ML-specific observability is nice to have
• FinOps background is nice to have
• Familiarity with data infrastructure is nice to have
• Experience with multi-tenant infrastructure is nice to have
• Prior startup scaling experience is nice to have
• Bachelor's Degree or equivalent experience
• Degree in IT or Computer Science or equivalent experience
• English conversational skills

• Payment in USD
• Remote work arrangement
• Working hours aligned with EST time zone