← Todas as vagas

Data Scientist

Sobre a vaga

• Build and evaluate ML approaches for company/entity matching
• Develop embedding and LLM-based matching approaches
• Develop scoring and ranking methodologies to identify true matches and distinguish them from duplicates, lookalikes, and unrelated entities
• Work with messy data, including names, aliases, domains, websites, firmographic attributes, multilingual records, and data hierarchies
• Define benchmark datasets, metrics, baselines, and error-analysis processes
• Design and execute experiments to validate hypotheses
• Compare LLM-assisted approaches against lower-cost alternatives
• Analyze model behavior, edge cases, and trade-offs
• Consider inference economics and scalability from the beginning
• Communicate experimental findings and recommendations to engineering and business stakeholders
• Independently establish experimental pipelines and research approaches
• Clearly document both successful and unsuccessful experiments

• 5+ years of professional Data Science / Machine Learning experience
• Strong applied Machine Learning fundamentals
• Excellent Python and SQL skills
• Hands-on experience with embeddings and semantic similarity
• Practical experience applying LLMs to real-world problems
• Experience with supervised and unsupervised learning
• Strong experience with classification and NLP
• Working knowledge of neural networks and transformer architectures
• Hands-on experience with TensorFlow, PyTorch, PyCaret, or equivalent ML frameworks
• Experience retraining or maintaining classification models in production
• Strong experimental design and model evaluation skills
• Experience defining baselines, metrics, test sets, and error-analysis processes
• Ability to evaluate model quality and demonstrate measurable improvements
• Strong understanding of scalability and ML inference costs
• Strong English communication skills
• Entity resolution, record linkage, or deduplication experience (nice-to-have)
• Ranking and similarity scoring (nice-to-have)
• Retrieval, clustering, or candidate-generation techniques (nice-to-have)
• LLM/embedding solutions designed for cost and scale constraints (nice-to-have)
• Spark, Snowflake, Databricks, or BigQuery (nice-to-have)
• Experience with company, domain, website, or firmographic data (nice-to-have)
• Experience working with multilingual datasets (nice-to-have)
• Strong analytical and experimental mindset
• Intellectual honesty and willingness to communicate negative results
• Strong autonomy and self-direction
• Excellent written and verbal communication
• Ability to defend technical recommendations with stakeholders
• Strong problem-solving skills
• Comfort working with ambiguity and large-scale datasets
• Ability to balance model quality, cost, and scalability

• Remote work option
• Opportunity to learn fast and take ownership
• Collaboration with strong teams
• Investment in modern ways of working