← Toutes les offres

Principal Data Engineer – RWE

À propos du poste

• Develop data processes for automated ongoing generation of patient-level data products for dashboards, reports, and studies
• Manipulate datasets onboarded from data vendors and partners into usable structures for RWE studies, dashboards, and other outputs
• Transform heterogeneous healthcare datasets into reusable data models for observational research and epidemiology studies
• Convert bespoke datasets to OMOP format where appropriate and account for residual data that cannot be converted
• Build FAIR data pipelines and semantic data engineering frameworks
• Create AI-ready datasets supporting generative AI use cases
• Engage with epidemiologists, statisticians, market access and health economists to scope requirements and translate needs into actionable data structures
• Collaborate with the RWE programming team to build data structures and support study outputs
• Liaise with IT to ensure inbound datasets are fit for agreed purposes
• Liaise with technical staff from analysis software vendors such as Databricks
• Maintain documentation of data flows, schemas, pipelines, and processes
• Design and perform data validation and monitoring to ensure accuracy and reliability
• Troubleshoot data loading, extraction, and transformation issues
• Collaborate with three other Data Engineering team members and provide workload support as needed

• Strong understanding of Real World Data (RWD) and Real World Evidence (RWE) concepts
• Ability to assess business requirements and recommend appropriate real-world healthcare datasets for analytical use cases
• Deep understanding of healthcare data models and healthcare data ecosystems
• Strong expertise in OMOP CDM v5.4, v6, including extensions
• Knowledge of SNOMED CT, RxNorm, ICD-10, LOINC, and HCPCS/CPT
• Strong experience building scalable ETL/ELT pipelines
• Expertise in Databricks, PySpark, Spark SQL, SQL, and Delta Lake
• Experience working with large-scale healthcare and patient-level datasets
• Strong understanding of Semantic Data Engineering principles
• Experience building FAIR-compliant data pipelines
• Experience with cloud-based data platforms and distributed processing frameworks
• Strong Power BI development and data modelling skills
• Ability to create reusable analytical datasets for dashboards and studies
• Experience designing AI-ready datasets and analytics data products
• Experience implementing automated data quality frameworks
• Strong data profiling, validation, and monitoring skills
• Understanding of healthcare data quality assessment methodologies
• Excellent stakeholder management and communication skills
• Ability to translate complex business requirements into technical solutions
• Experience working with cross-functional global teams
• Exposure to one or more of: Oncology, Respiratory, Immunology & Inflamation, Infectious Diseases

• Equal opportunities employer
• Inclusive and diverse working environment
• B Corp accredited company focused on social and environmental performance, transparency, and accountability
• AI-assisted recruitment with human-led assessments, selection decisions, and hiring outcomes