TU PRÓXIMO CAPÍTULO
Linux Infrastructure Engineer – Bare Metal, Storage, AI Factory Infrastructure
Sobre el puesto
• Design, deploy, operate, and troubleshoot large-scale Linux-based infrastructure
• Build and manage infrastructure from the hardware layer, including servers, networking, storage, GPU clusters, and AI-ready platforms
• Deploy and manage GPU-accelerated infrastructure for AI/ML workloads
• Provision, monitor, optimize, and manage the lifecycle of GPU platforms and clusters
• Support AI training environments and high-performance computing workloads
• Design and operate bare metal infrastructure and BMaaS platforms
• Administer enterprise Linux storage and Ceph platforms
• Design and troubleshoot high-performance data center networking and Layer 2/Layer 3 infrastructure
• Support high-availability, clustering, disaster recovery, and mission-critical production environments
• Automate operational tasks using Bash and Python
• Create operational documentation, runbooks, and infrastructure standards
• Expert-level Linux administration; Ubuntu required, Red Hat and SUSE preferred
• Deep expertise in bare metal server deployment, architecture, provisioning, and lifecycle management
• Experience operating Bare Metal as a Service (BMaaS) platforms and large-scale infrastructure environments
• Strong understanding of server hardware, including BIOS/UEFI, RAID controllers, firmware management, iLO/iDRAC/IPMI, NICs and SmartNICs, HBA cards, and hardware diagnostics
• Experience designing, implementing, and supporting enterprise Linux infrastructure at scale
• Experience deploying and managing GPU-accelerated infrastructure for AI/ML workloads
• Understanding of NVIDIA GPU technologies, including A100, H100, H200, B200, or equivalent platforms
• Experience with NVIDIA DGX and OEM GPU servers, GPU provisioning and lifecycle management, monitoring, and performance optimization
• Knowledge of AI Factory architecture and infrastructure requirements
• Experience supporting GPU clusters, AI training environments, and HPC workloads
• Understanding of GPU resource allocation and scheduling, multi-GPU systems, GPU networking, and high-bandwidth, low-latency infrastructure design
• Familiarity with CUDA, NCCL, GPUDirect Storage, NVIDIA Fabric Manager, and NVIDIA Base Command
• Advanced Linux storage administration: LVM, XFS, EXT4, NFS, iSCSI, Fibre Channel SAN, and Multipath I/O
• Strong hands-on experience with Ceph, including MON, OSD, MDS, RBD, CephFS, RGW, capacity planning, performance tuning, and failure recovery
• Experience with high-performance AI storage platforms such as WEKA, VAST Data, Dell PowerScale, Pure Storage FlashBlade, and NetApp
• Understanding of NVMe-over-Fabrics, RDMA, GPUDirect Storage, parallel file systems, and AI data pipelines
• Strong networking knowledge: bonding, VLANs, routing, MTU optimization, DNS, and DHCP
• Experience with 100G/200G/400G Ethernet, RoCE, RDMA, and spine-leaf architectures
• Familiarity with NVIDIA Spectrum-X, Mellanox/NVIDIA ConnectX adapters, or equivalent technologies
• Strong understanding of Layer 2 and Layer 3 infrastructure design and troubleshooting
• Experience with high availability, clustering, and disaster recovery
• Strong troubleshooting skills across Linux operating systems, hardware platforms, GPU infrastructure, networking, and enterprise storage
• Bash and Python scripting for automation and operational efficiency
• Experience creating operational documentation, runbooks, and infrastructure standards
• Nice to have: Kubernetes infrastructure, KVM, VMware, OpenShift Virtualization, Ansible, NVIDIA Base Command Manager, Slurm, Prometheus, Grafana, OpenTelemetry, DCIM tools, IPAM solutions, and AWS, Azure, or hybrid cloud exposure