O SEU PRÓXIMO CAPÍTULO
Linux Infrastructure Engineer – Bare Metal, Storage, AI Factory Infrastructure
Sobre a vaga
• Design, deploy, operate, and troubleshoot large-scale Linux-based infrastructure
• Power traditional enterprise workloads and modern AI/ML environments
• Build and manage infrastructure from the hardware layer up, including servers, networking, storage, GPU clusters, and AI-ready platforms
• Deploy, provision, and manage bare metal servers and BMaaS platforms throughout their lifecycle
• Deploy and manage GPU-accelerated infrastructure, GPU clusters, and AI training/HPC environments
• Monitor and optimize GPU performance and manage GPU resources
• Design and support enterprise storage, Ceph clusters, high-performance AI storage, and parallel file systems
• Design and troubleshoot high-performance data center networking and Layer 2/Layer 3 infrastructure
• Support high-availability, clustering, disaster recovery, and mission-critical production environments
• Troubleshoot operating system, hardware, storage, networking, and AI infrastructure challenges
• Automate operational tasks using Bash and Python
• Create operational documentation, runbooks, and infrastructure standards
• Collaborate with the dedicated DevOps team while focusing on infrastructure engineering rather than DevOps delivery pipelines
• Expert-level Linux administration; Ubuntu required, Red Hat and SUSE preferred
• Deep expertise in bare metal server deployment, architecture, provisioning, and lifecycle management
• Experience operating Bare Metal as a Service (BMaaS) platforms and large-scale infrastructure environments
• Strong understanding of server hardware, including BIOS/UEFI, RAID controllers, firmware management, iLO/iDRAC/IPMI, NICs and SmartNICs, HBA cards, and hardware diagnostics
• Experience designing, implementing, and supporting enterprise Linux infrastructure at scale
• Experience deploying and managing GPU-accelerated infrastructure for AI/ML workloads
• Understanding of NVIDIA GPU technologies, including A100, H100, H200, B200, or equivalent platforms
• Experience with NVIDIA DGX and OEM GPU servers, GPU provisioning, lifecycle management, monitoring, and performance optimization
• Knowledge of AI Factory architecture and infrastructure requirements
• Experience supporting GPU clusters, AI training environments, and HPC workloads
• Understanding of GPU resource allocation and scheduling, multi-GPU systems, GPU networking, and high-bandwidth, low-latency infrastructure design
• Familiarity with CUDA, NCCL, GPUDirect Storage, NVIDIA Fabric Manager, and NVIDIA Base Command
• Advanced Linux storage administration: LVM, XFS, EXT4, NFS, iSCSI, Fibre Channel SAN, and Multipath I/O
• Strong hands-on experience with Ceph, including MON, OSD, MDS, RBD, CephFS, RGW, capacity planning, performance tuning, and failure recovery
• Experience with high-performance AI storage platforms such as WEKA, VAST Data, Dell PowerScale, Pure Storage FlashBlade, and NetApp
• Understanding of NVMe-over-Fabrics, RDMA, GPUDirect Storage, parallel file systems, and AI data pipelines
• Strong networking knowledge: bonding, VLANs, routing, MTU optimization, DNS, and DHCP
• Experience with 100G/200G/400G Ethernet, RoCE, RDMA, and Spine-Leaf architectures
• Familiarity with NVIDIA Spectrum-X, Mellanox/NVIDIA ConnectX adapters, or equivalent technologies
• Strong understanding of Layer 2 and Layer 3 infrastructure design and troubleshooting
• Experience with high availability, clustering, and disaster recovery
• Strong troubleshooting skills across Linux operating systems, hardware platforms, GPU infrastructure, networking, and enterprise storage
• Bash and Python scripting for automation and operational efficiency
• Experience creating operational documentation, runbooks, and infrastructure standards
• Kubernetes infrastructure, virtualization platforms, Ansible, NVIDIA Base Command Manager, Slurm/HPC schedulers, observability platforms, DCIM tools, IPAM solutions, and AWS/Azure or hybrid cloud exposure are nice to have
• Not primarily focused on CI/CD, Terraform, GitOps, application delivery pipelines, cloud-only administration, or software development