TU PRÓXIMO CAPÍTULO
GPU Kernel Engineer – CUDA, Triton, Accelerator Performance
Sobre el puesto
• Review GPU and accelerator kernel implementations for correctness
• Compare outputs against reference implementations
• Evaluate numerical tolerance thresholds
• Review kernel benchmarks and determine whether comparisons are fair
• Identify performance bottlenecks and optimization opportunities
• Assess whether performance targets are realistic given hardware limits
• Review kernel translations and hardware migrations
• Identify compilation, driver, memory, shape, and runtime issues
• Determine whether technical tasks are genuinely difficult or incorrectly configured
• Provide clear, actionable technical feedback
• Implement and debug kernels
• Optimize CUDA and Triton kernels
• Translate between kernel frameworks
• Perform hardware migration and operator fusion
• Profile and benchmark performance
• Verify numerical correctness
• Debug compilation and runtime issues
• Optimize memory hierarchy and kernel-level AI workload performance
• 3+ years of hands-on experience developing, optimizing, or debugging GPU or accelerator kernels
• Strong experience with at least two of: CUDA; Triton; NKI / AWS Neuron; Pallas / JAX
• Strong understanding of GPU performance optimization
• Experience with kernel profiling tools such as Nsight, NCU, roofline analysis, or framework-native profilers
• Understanding of memory bandwidth, compute throughput, GPU occupancy, shared memory, register pressure, memory coalescing, and bank conflicts
• Strong understanding of floating-point numerical correctness and tolerance thresholds
• Experience debugging kernel compilation and runtime issues
• Ability to distinguish software defects, environment problems, and genuine optimization challenges
• Experience writing kernels from technical specifications, translating kernels between frameworks, migrating kernels across hardware platforms, debugging incorrect implementations, optimizing kernel performance, and fusing multiple operations into optimized kernels
• Experience across both NVIDIA GPU and custom accelerator ecosystems (nice to have)
• Experience with AWS Trainium, TPU, JAX, or other accelerators (nice to have)
• Compiler engineering experience (nice to have)
• Familiarity with MLIR, XLA, or intermediate representation lowering (nice to have)
• Contributions to GPU or ML kernel libraries (nice to have)
• Experience with cuBLAS, cuDNN, Triton community kernels, or JAX/XLA custom calls (nice to have)
• Experience with AI model evaluation, RLHF, or technical benchmark development (nice to have)
• $65 per hour compensation
• Part-time, project-based consulting engagement
• Remote work