← All jobs

Data Architect, AI

About the role

• Design end-to-end data architectures on the Databricks Lakehouse Platform
• Build reliable, production-ready Bronze, Silver, and Gold data pipelines
• Troubleshoot and tune large-scale Apache Spark jobs
• Design governance models for data and AI assets
• Implement role-based, row-level, and column-level access controls and data lineage
• Work with AWS, Azure, or GCP cloud infrastructure
• Optimize Databricks clusters, DBU usage, performance, and cloud costs
• Potentially build, deploy, and monitor GenAI applications and machine learning models
• Automate Databricks workflows and support streaming data ingestion and downstream BI integration

• Expertise in designing end-to-end data architectures using Bronze, Silver, and Gold layers
• Deep understanding of distributed computing and Databricks/Spark internals
• Ability to troubleshoot and tune large-scale Spark jobs using caching, partitioning, and broadcast joins
• Understanding of Delta Lake ACID transactions, schema enforcement, time travel, and Z-ordering
• Ability to design unified governance models for data and AI assets, including RBAC, row-level security, column-level security, and data lineage
• Strong grasp of AWS, Azure, or GCP native cloud ecosystems, including VNet setups and IAM roles
• Advanced SQL skills
• Fluency in Python or Scala
• Ability to monitor DBUs, right-size serverless and multi-node clusters, and optimize cloud costs
• Advanced English
• Databricks certifications, GenAI/MLflow, CI/CD and DevOps, streaming data, and BI experience are nice to have, not required