VOTRE PROCHAIN CHAPITRE
Linux Infrastructure Engineer – Bare Metal, Storage, AI Factory Infrastructure
À propos du poste
• Design, deploy, operate, and troubleshoot large-scale Linux-based infrastructure
• Power traditional enterprise workloads and modern AI/ML environments
• Build and manage infrastructure from the hardware layer up, including servers, networking, storage, GPU clusters, and AI-ready platforms
• Deploy and manage GPU-accelerated infrastructure for AI/ML workloads
• Support GPU clusters, AI training environments, and high-performance computing (HPC) workloads
• Operate Bare Metal as a Service (BMaaS) platforms and large-scale infrastructure environments
• Design, implement, and support enterprise Linux infrastructure at scale
• Administer and optimize enterprise storage and high-performance AI storage platforms
• Manage Ceph cluster architecture, capacity planning, performance tuning, and failure recovery
• Design and troubleshoot high-performance data center networking and Layer 2/Layer 3 infrastructure
• Support high availability, clustering, disaster recovery, and mission-critical production environments
• Perform troubleshooting across Linux operating systems, hardware, GPU infrastructure, networking, and storage
• Use Bash and Python for automation and operational efficiency
• Create operational documentation, runbooks, and infrastructure standards
• Expert-level Linux administration; Ubuntu required, Red Hat and SUSE preferred
• Deep expertise in bare metal server deployment, architecture, provisioning, and lifecycle management
• Experience operating Bare Metal as a Service (BMaaS) platforms and large-scale infrastructure environments
• Strong understanding of server hardware, including BIOS/UEFI, RAID controllers, firmware management, iLO/iDRAC/IPMI, NICs and SmartNICs, HBA cards, and hardware diagnostics
• Experience designing, implementing, and supporting enterprise Linux infrastructure at scale
• Experience deploying and managing GPU-accelerated infrastructure for AI/ML workloads
• Understanding of NVIDIA GPU technologies, including A100, H100, H200, B200, or equivalent platforms
• Experience with NVIDIA DGX and OEM GPU servers, GPU provisioning, lifecycle management, monitoring, and performance optimization
• Knowledge of AI Factory architecture and infrastructure requirements
• Experience supporting GPU clusters, AI training environments, and HPC workloads
• Understanding of GPU resource allocation and scheduling, multi-GPU systems, GPU networking, and high-bandwidth, low-latency infrastructure design
• Familiarity with CUDA, NCCL, GPUDirect Storage, NVIDIA Fabric Manager, and NVIDIA Base Command
• Advanced Linux storage administration: LVM, XFS, EXT4, NFS, iSCSI, Fibre Channel, SAN, and Multipath I/O
• Strong hands-on experience with Ceph, including MON, OSD, MDS, RBD, CephFS, RGW, capacity planning, performance tuning, and failure recovery
• Experience with high-performance AI storage platforms such as WEKA, VAST Data, Dell PowerScale, Pure Storage FlashBlade, and NetApp
• Understanding of NVMe-over-Fabrics (NVMe-oF), RDMA, GPUDirect Storage, parallel file systems, and AI data pipelines
• Strong networking knowledge: bonding, VLANs, routing, MTU optimization, DNS, and DHCP
• Experience with 100G/200G/400G Ethernet, RoCE, RDMA, and Spine-Leaf architectures
• Familiarity with NVIDIA Spectrum-X, Mellanox/NVIDIA ConnectX adapters, or equivalent technologies
• Strong understanding of Layer 2 and Layer 3 infrastructure design and troubleshooting
• Experience with high availability, clustering, and disaster recovery
• Strong troubleshooting skills across Linux operating systems, hardware platforms, GPU infrastructure, networking, and enterprise storage
• Bash and Python scripting for automation and operational efficiency
• Experience creating operational documentation, runbooks, and infrastructure standards
• Kubernetes infrastructure, virtualization platforms, Ansible, Slurm/HPC schedulers, observability platforms, DCIM tools, IPAM solutions, and AWS/Azure or hybrid cloud exposure are nice to have
• Not a DevOps-focused role; primarily CI/CD, Terraform, GitOps, application delivery, cloud-only administration, or software development experience is not sought