← Toutes les offres

Linux Infrastructure Engineer, Bare Metal, Storage, AI Factory Infrastructure

À propos du poste

• Design, deploy, operate, and troubleshoot large-scale Linux-based infrastructure
• Build and manage infrastructure from the hardware layer up, including servers, networking, storage, GPU clusters, and AI-ready platforms
• Deploy and manage bare metal servers and BMaaS platforms
• Design, implement, and support enterprise Linux infrastructure at scale
• Deploy and manage GPU-accelerated infrastructure for AI/ML workloads
• Support GPU clusters, AI training environments, and HPC workloads
• Provision, monitor, optimize, and manage GPU infrastructure lifecycle
• Administer enterprise storage and Ceph platforms, including capacity planning, performance tuning, and failure recovery
• Design and troubleshoot high-performance data center networking and Layer 2/Layer 3 infrastructure
• Support high-availability, clustering, disaster recovery, and mission-critical production environments
• Troubleshoot operating system, hardware, GPU, networking, and storage challenges
• Automate operational work using Bash and Python
• Create operational documentation, runbooks, and infrastructure standards

• Senior-level, highly experienced Linux infrastructure engineering expertise
• Expert-level Linux administration; Ubuntu required, Red Hat and SUSE preferred
• Deep expertise in bare metal server deployment, architecture, provisioning, and lifecycle management
• Experience operating Bare Metal as a Service (BMaaS) platforms and large-scale infrastructure environments
• Strong understanding of server hardware, including BIOS/UEFI, RAID controllers, firmware management, iLO/iDRAC/IPMI, NICs and SmartNICs, HBA cards, and hardware diagnostics
• Experience designing, implementing, and supporting enterprise Linux infrastructure at scale
• Experience deploying and managing GPU-accelerated AI/ML infrastructure
• Understanding of NVIDIA GPU technologies, including A100, H100, H200, B200, or equivalent platforms
• Experience with NVIDIA DGX and OEM GPU servers, GPU provisioning, lifecycle management, monitoring, and performance optimization
• Knowledge of AI Factory architecture and infrastructure requirements
• Experience supporting GPU clusters, AI training environments, and HPC workloads
• Understanding of GPU resource allocation and scheduling, multi-GPU systems, GPU networking, and high-bandwidth, low-latency infrastructure design
• Familiarity with CUDA, NCCL, GPUDirect Storage, NVIDIA Fabric Manager, and NVIDIA Base Command preferred
• Advanced Linux storage administration: LVM, XFS, EXT4, NFS, iSCSI, Fibre Channel SAN, and Multipath I/O
• Strong hands-on Ceph experience, including cluster architecture, MON, OSD, MDS, RBD, CephFS, RGW, capacity planning, performance tuning, and failure recovery
• Experience with high-performance AI storage platforms such as WEKA, VAST Data, Dell PowerScale, Pure Storage FlashBlade, or NetApp
• Understanding of NVMe-over-Fabrics, RDMA, GPUDirect Storage, parallel file systems, and AI data pipelines
• Strong networking knowledge: bonding, VLANs, routing, MTU optimization, DNS, and DHCP
• Experience with 100G/200G/400G Ethernet, RoCE, RDMA, and Spine-Leaf architectures
• Familiarity with NVIDIA Spectrum-X, Mellanox/NVIDIA ConnectX adapters, or equivalent technologies
• Strong understanding of Layer 2 and Layer 3 infrastructure design and troubleshooting
• Experience with high availability, clustering, and disaster recovery
• Strong troubleshooting skills across Linux operating systems, hardware platforms, GPU infrastructure, networking, and enterprise storage
• Bash and Python scripting for automation and operational efficiency
• Experience creating operational documentation, runbooks, and infrastructure standards
• Kubernetes infrastructure, virtualization platforms, Ansible, NVIDIA Base Command Manager, Slurm/HPC workload schedulers, observability platforms, DCIM tools, IPAM solutions, and AWS/Azure/hybrid cloud exposure listed as nice to have
• Not primarily a CI/CD, Terraform, GitOps, application delivery, cloud-only, or software development candidate