← All jobs

Data Scientist

About the role

• Build and evaluate ML approaches for company/entity matching
• Develop embedding and LLM-based matching approaches
• Develop scoring and ranking methodologies to identify true matches and distinguish them from duplicates, lookalikes, and unrelated entities
• Work with messy data, including names, aliases, domains, websites, firmographic attributes, multilingual records, and data hierarchies
• Define benchmark datasets, metrics, baselines, and error-analysis processes
• Design and execute experiments to validate hypotheses
• Compare LLM-assisted approaches against lower-cost alternatives
• Analyze model behavior, edge cases, and trade-offs
• Consider inference economics and scalability from the beginning
• Communicate experimental findings and recommendations to engineering and business stakeholders
• Independently establish experimental pipelines and research approaches
• Clearly document both successful and unsuccessful experiments

• 5+ years of professional Data Science / Machine Learning experience
• Strong applied Machine Learning fundamentals
• Excellent Python and SQL skills
• Hands-on experience with embeddings and semantic similarity
• Practical experience applying LLMs to real-world problems
• Experience with supervised and unsupervised learning
• Strong experience with classification and NLP
• Working knowledge of neural networks and transformer architectures
• Hands-on experience with TensorFlow, PyTorch, PyCaret, or equivalent ML frameworks
• Experience retraining or maintaining classification models in production
• Strong experimental design and model evaluation skills
• Experience defining baselines, metrics, test sets, and error-analysis processes
• Ability to evaluate model quality and demonstrate measurable improvements
• Strong understanding of scalability and ML inference costs
• Strong English communication skills
• Nice-to-have: entity resolution, record linkage, or deduplication experience
• Nice-to-have: ranking and similarity scoring
• Nice-to-have: retrieval, clustering, or candidate-generation techniques
• Nice-to-have: LLM/embedding solutions designed for cost and scale constraints
• Nice-to-have: Spark, Snowflake, Databricks, or BigQuery
• Nice-to-have: experience with company, domain, website, or firmographic data
• Nice-to-have: experience working with multilingual datasets
• Strong analytical and experimental mindset
• Intellectual honesty and willingness to communicate negative results
• Strong autonomy and self-direction
• Excellent written and verbal communication
• Ability to defend technical recommendations with stakeholders
• Strong problem-solving skills
• Comfort working with ambiguity and large-scale datasets
• Ability to balance model quality, cost, and scalability

• Remote work option
• Full-time employment