← All jobs

Senior Software Engineer – Open Source, SWE-Bench Evaluation

About the role

• Review software engineering tasks derived from real GitHub issues and pull requests
• Assess real-world bug fixes and feature implementations
• Evaluate open-source repositories, pull requests, unit tests, and test coverage
• Review repository setup and dependency management
• Assess reproducibility and environment configuration
• Evaluate task difficulty, complexity, and multi-file or cross-module code changes
• Determine whether tasks are clearly specified, technically solvable, appropriately tested, and representative of professional software engineering problems
• Identify flaky tests, missing dependencies, ambiguous requirements, version conflicts, and environment-specific behavior
• Provide clear recommendations on whether tasks should be accepted, improved, or excluded
• Provide technical feedback and quality assessments

• 3+ years of professional software engineering experience
• Strong experience working with large, multi-file codebases
• Experience reviewing pull requests, debugging issues, and maintaining production code
• Strong understanding of unit testing and test coverage
• Ability to evaluate whether tests correctly validate a solution without unnecessarily restricting implementation approaches
• Experience with dependency management, environment setup, and reproducibility
• Strong understanding of Git and GitHub-based development workflows
• Ability to analyze complex technical problems and provide clear written feedback
• Contributions to or maintenance of open-source projects (nice to have)
• Experience with SWE-Bench, SWE-Bench Verified, or similar coding benchmarks (nice to have)
• Experience with major Python open-source projects such as Django, Flask, scikit-learn, SymPy, matplotlib, requests, or pytest (nice to have)
• Experience with Docker, CI/CD, pip, conda, or dependency pinning (nice to have)
• Knowledge of test fixtures, test isolation, or property-based testing (nice to have)
• Experience designing technical assessments or reviewing coding challenges (nice to have)
• Experience with AI/ML evaluation, data curation, RLHF, or benchmark development (nice to have)

• $65 per hour compensation
• Remote work
• Part-time, project-based consulting engagement