Sr Data Engineer -
Design, build, and optimize scalable data and machine learning platforms using Databricks and Spark. Manage end-to-end data pipelines, ML workflows, and production AI systems while ensuring data quality and security.
- On-site
- Toronto, ON
- Posted Jun 9, 2026
- 1 position
More jobs you can apply to directly
Similar opportunities posted by employers hiring on Jobs.ca, with no external application form.
Forgeahead Solutions Corporation
Technical Lead and Senior Software Engineer
- On-site
Alcohol and Gaming Commission of Ontario (AGCO)
Information Management Lead / Responsable de la gestion de l’information
- On-site
Dalfen Ltée
Accounting Technician
- On-site
Job summary
Job Summary We are seeking a highly skilled Databricks Engineer with AI/ML experience to design, build, and optimize scalable data and machine learning platforms on Databricks. The role involves end-to-end ownership of data pipelines, ML workflows, and production AI systems. Key Responsibilities Design and implement scalable ETL pipelines using Databricks & Spark Build Lakehouse architecture using Delta Lake Develop and deploy ML models using MLflow Implement MLOps pipelines for training, testing, and serving models Optimize cluster performance and reduce compute cost Build RAG and LLM-based solutions using Mosaic AI Integrate analytics with BI tools (Power BI, Tableau) Implement data governance using Unity Catalog Collaborate with Data Scientists and Business teams Ensure data quality, security, and compliance Required Skills Mandatory 5+ years of Databricks & Apache Spark Strong Python & PySpark Experience with Delta Lake & Lakehouse MLflow & MLOps experience Cloud platform (AWS/Azure/GCP) Git & CI/CD Preferred Experience with LLMs & Generative AI RAG pipelines & Vector Databases Deep Learning frameworks Databricks certifications Power BI integration Requirements Job Summary We are seeking a highly skilled Databricks Engineer with AI/ML experience to design, build, and optimize scalable data and machine learning platforms on Databricks. The role involves end-to-end ownership of data pipelines, ML workflows, and production AI systems. Key Responsibilities Design and implement scalable ETL pipelines using Databricks & Spark Build Lakehouse architecture using Delta Lake Develop and deploy ML models using MLflow Implement MLOps pipelines for training, testing, and serving models Optimize cluster performance and reduce compute cost Build RAG and LLM-based solutions using Mosaic AI Integrate analytics with BI tools (Power BI, Tableau) Implement data governance using Unity Catalog Collaborate with Data Scientists and Business teams Ensure data quality, security, and compliance Required Skills Mandatory 5+ years of Databricks & Apache Spark Strong Python & PySpark Experience with Delta Lake & Lakehouse MLflow & MLOps experience Cloud platform (AWS/Azure/GCP) Git & CI/CD Preferred Experience with LLMs & Generative AI RAG pipelines & Vector Databases Deep Learning frameworks Databricks certifications Power BI integration
What you’ll do
Design, build, and optimize scalable data and machine learning platforms using Databricks and Spark. Manage end-to-end data pipelines, ML workflows, and production AI systems while ensuring data quality and security.
Requirements
Requires 5+ years of experience with Databricks and Apache Spark along with strong Python and PySpark skills. Proficiency in Delta Lake, MLOps, and cloud platforms is mandatory, with preference for experience in Generative AI and RAG pipelines.
Listed skills
- Microsoft AzurePreferred
- CI/CDPreferred
- Amazon Web ServicesPreferred
- Google CloudPreferred
- GitPreferred
- PythonPreferred
Other relevant skills
Identified from the job description. Confirm important requirements above.
- Databricks
- Apache Spark
- Python
- PySpark
- Delta Lake
- Lakehouse
- MLflow
- MLOps
- AWS
- Azure
- GCP
- Git
- CI/CD
- Generative AI
- RAG
- Vector Databases
- Pipelines
- Vector Database
- MLOps (Machine Learning Operations)
- Generative Artificial Intelligence
- Workflow Management
- Git (Version Control System)
- Unity Engine
- Artificial Intelligence
- Amazon Web Services
- Microsoft Azure
- Business Intelligence
- Data Governance
- Extract Transform Load (ETL)
- Data Quality
- Scalability
- Python (Programming Language)
- Machine Learning
- Power BI
- Tableau (Business Intelligence Software)
- Deep Learning
- Data Pipelines
- Artificial Intelligence Infrastructure
Job areas
- Data & Analytics
- Software
- Technology
- Engineering
- Data Engineer
- Generative Artificial Intelligence Engineer
- Software Developers
- Computer and Information Research Scientists
Additional details
- Minimum experience
- 5+ years
- Posting language
- English
- Working hours
- 40 hours per week