VOTRE PROCHAIN CHAPITRE
AI Solution Architect
À propos du poste
• Own end-to-end architecture for AI Factory and enterprise AI solutions from requirements through production readiness.
• Assess AI/ML workload requirements for training, fine-tuning, inference, batch processing and high-performance computing.
• Design GPU compute architectures including NVIDIA HGX/DGX/OEM platforms, multi-GPU systems, NVLink/NVSwitch and GPU resource allocation.
• Design high-performance AI networking using 100/200/400/800G Ethernet, EVPN/VXLAN and leaf-spine architectures.
• Design AI storage and data architectures using object storage and parallel file systems such as Ceph and WEKA.
• Define AI platform architecture across Kubernetes, HPC, container runtimes, model-serving platforms and enterprise AI frameworks.
• Establish architecture standards for security, identity, tenant isolation, data protection, observability, disaster recovery and operational resilience.
• Develop reference architectures, high-level/low-level designs, capacity models, bills of materials, technology evaluations and implementation roadmaps.
• Lead technical evaluations, proof-of-concepts, vendor assessments and architecture review boards.
• Collaborate with infrastructure, network, security, storage, cloud, data, application and operations teams.
• Define performance, availability, scalability, security and cost objectives and validate architecture against measurable acceptance criteria.
• Provide technical leadership during deployment, migration, integration, troubleshooting and production transition.
• Produce AI Factory reference architectures, solution blueprints, HLDs, LLDs, architecture diagrams, capacity/performance/scalability models, technology evaluations, security architecture inputs, BOMs, sizing, migration strategies, implementation roadmaps, operational readiness checklists, runbooks and acceptance criteria.
• 10+ years of infrastructure, cloud, enterprise architecture or solution architecture experience, with significant AI/GPU infrastructure exposure.
• Proven experience designing large-scale enterprise platforms and translating business requirements into technical architectures.
• Hands-on understanding of physical infrastructure, GPU systems, networking, storage and Linux platforms.
• Bachelor's degree in Computer Science, Engineering, Information Technology or related field preferred.
• NVIDIA certifications or equivalent GPU/AI infrastructure credentials preferred.
• AWS Solutions Architect / Azure Solutions Architect certification preferred.
• TOGAF or equivalent enterprise architecture certification preferred.
• CCNP/CCIE or equivalent networking certification preferred.
• CISSP or equivalent security certification preferred.
• Kubernetes certifications such as CKA/CKAD preferred.
• Red Hat / Linux certifications preferred.
• AI/ML architecture expertise, including NVIDIA AI Enterprise, NGC, CUDA, NCCL, DCGM, GPU Operator and AI platform ecosystems.
• PyTorch, TensorFlow and JAX, with operational understanding of training and inference workloads.
• GPU scheduling, multi-tenancy, MIG/vGPU, GPU utilization and workload placement.
• LLM, generative AI, RAG, fine-tuning, model serving and inference architecture.
• NVIDIA A100/H100/H200/B200 or equivalent GPU platforms; NVLink, NVSwitch, PCIe topology and multi-GPU performance architecture.
• DGX/HGX/OEM GPU server architecture and lifecycle management.
• AI Factory capacity planning, rack density, power, cooling, commissioning and lifecycle strategy.
• 100/200/400/800G Ethernet, InfiniBand, RoCEv2, RDMA, Netris, BGP, EVPN/VXLAN, VRF, ECMP, VLAN, MTU, PFC, ECN and QoS.
• NVIDIA ConnectX/SuperNIC, Spectrum/Spectrum-X, Quantum and BlueField DPU technologies.
• GPU east-west traffic, GPUDirect RDMA and network performance troubleshooting.
• Parallel file systems, object storage, NFS, NVMe/NVMe-oF and high-throughput data pipelines.
• Ceph, WEKA, VAST, Dell PowerScale, Pure FlashBlade, NetApp or equivalent technologies.
• Data lake/lakehouse concepts, metadata, lineage, data movement and data lifecycle.
• GPUDirect Storage and storage/network performance optimization.
• Kubernetes, GPU Operator, container runtimes, Kubernetes GPU scheduling and HPC or equivalent workload schedulers.
• Model serving/inference platforms, MLOps platform architecture, API gateways, service discovery, secrets management and platform integration.
• AWS and/or Azure AI infrastructure and security services.
• Hybrid cloud connectivity, IAM, private networking, cloud storage, workload placement, cloud cost optimization, capacity planning and FinOps.
• Zero Trust, network segmentation, IAM/RBAC, PAM, workload identity, GPU/DPU/container/Kubernetes/firmware/supply-chain security.
• Encryption at rest/in transit, secrets management, audit logging, compliance controls, data/model protection, tenant isolation and secure model access.
• Prometheus, Grafana, OpenTelemetry, NVIDIA DCGM and infrastructure telemetry.
• Monitoring across GPU, CPU, memory, network, storage, power and thermal domains.
• High availability, backup/restore, disaster recovery, business continuity, failure-domain design, performance engineering, bottleneck analysis, SLO/SLA design and capacity forecasting.