← Todos los empleos

Senior Software Engineer – Open Source, SWE-Bench Evaluation

Sobre el puesto

• Review ML challenges involving experiment design, model selection, datasets, metrics, preprocessing, distribution shift, contamination, label noise, feature leakage, hyperparameter tuning, train/validation/test methodology, reproducibility, and statistical significance
• Determine whether challenges are technically sound, reproducible, appropriately difficult, and require strong machine learning reasoning
• Evaluate whether datasets contain meaningful and learnable signals
• Identify unintended shortcuts or artifacts in synthetic datasets
• Determine whether tasks require genuine diagnosis of underlying ML problems rather than brute-force model selection or large hyperparameter searches
• Review evaluation metrics and improvement thresholds
• Detect metric gaming, data leakage, and evaluation flaws
• Verify reproducibility across the complete data-to-model-to-evaluation pipeline
• Assess whether challenge difficulty is appropriately calibrated
• Provide recommendations for improving, recalibrating, or excluding problematic tasks
• Analyze ML experiments, datasets, metrics, and pipelines for applied machine learning model-training and evaluation challenges

• 3+ years of hands-on applied machine learning experience
• Strong experience with ML experiment design, model selection, hyperparameter tuning, model evaluation, data preprocessing, and validation
• Strong understanding of train, validation, and test splits
• Ability to identify data leakage, label noise, distribution shift, spurious correlations, feature leakage, and data contamination
• Experience evaluating whether performance improvements are statistically meaningful rather than random fluctuations
• Strong understanding of ML evaluation metrics and when different metrics are appropriate
• Experience debugging ML workloads across CPU and GPU environments
• Ability to analyze technical problems and provide clear written feedback
• Experience creating or participating in Kaggle, DrivenData, or similar ML competitions is nice to have
• Experience designing benchmark datasets or ML challenges is nice to have
• Background in data-centric AI or dataset quality is nice to have
• Experience with synthetic data generation and validation is nice to have
• Familiarity with statistical testing, confidence intervals, and effect sizes is nice to have
• Experience with ML evaluation pipelines, RLHF, or AI model evaluation is nice to have
• Experience developing ML curricula or technical assessments is nice to have
• Understanding of shortcut learning, spurious correlations, Goodhart’s Law, Simpson’s paradox, and metric gaming is nice to have
• Applicants must select a programming language or library for the interview and provide a location

• Remote work
• Part-time, project-based consulting engagement