Machine Learning Data Engineer
- Montréal, QC
- On-site
- Posted Sep 2, 2026
- 1 position
Opens an external site
- Employment type
- Full-time
- Experience level
- Senior · 5+ years
- Apply by
- Oct 2, 2026
- Posting language
- English
- Working hours
- 40 hours per week
- Seniority
- Mid-Senior level
- Application method
- Direct apply is available
Job summary
Design and scale infrastructure to transform raw, web-scale unstructured text into high-quality datasets for training advanced AI models. Develop automated data-quality systems, filtering mechanisms, and internal tools to support AI researchers.
Job details
Senior Machine Learning Data Engineer Location: Canada The Opportunity This is an opportunity to work alongside a highly accomplished AI research team tackling fundamental problems in advanced machine learning. Rather than maintaining established pipelines, you’ll be helping develop new approaches to training-data engineering and curation where established playbooks often do not yet exist. The organization is building a substantial technical team in Montreal, and we are hiring multiple people across this area. If you’ve worked on LLM training data, large-scale NLP pipelines, foundation-model infrastructure or web-scale data processing, I’d be interested in speaking with you. We’re working with an ambitious AI research organization building next-generation machine learning systems and are looking for multiple Senior Machine Learning Data Engineers to join its growing technical team. This role sits at the intersection of data engineering, machine learning, NLP, and large-scale training data curation. You’ll be responsible for building and scaling the infrastructure that transforms raw, web-scale data into high-quality datasets used to train advanced AI models. The core challenge is engineering data quality at enormous scale: developing filtering systems, model-based quality scoring, dataset transformations, contamination detection, and tooling that allows researchers to understand and work with training corpora effectively. What You’ll Work On Design, build and scale pipelines that transform raw web-scale data into high-quality datasets for training large machine learning models. Process extremely large unstructured text corpora, including datasets at trillion-token scale. Build sophisticated data-processing systems covering: Deduplication, Model-based quality scoring, Heuristic filtering, Toxicity and content-safety filtering, PII detection and removal, Metadata extraction, Dataset transformations, Versioning and provenance tracking Develop automated data-quality systems using a combination of heuristics, ML classifiers, LLM-based evaluators and human-in-the-loop workflows. Build monitoring and guardrails to identify data-quality regressions before they impact downstream model training. Work closely with AI researchers to understand evolving training-data requirements and identify gaps within existing datasets. Design robust evaluation contamination and data-leakage detection systems. Build internal tools that allow researchers to easily explore, query and understand large datasets. Optimize large-scale distributed processing systems for throughput, reliability and infrastructure cost. What We’re Looking For 5+ years of experience across machine learning engineering, data processing, NLP, data infrastructure or related areas. Experience working with very large unstructured text datasets, ideally at web or foundation-model scale. Strong production-level Python. Hands-on experience with distributed data-processing technologies such as: Spark, Ray, Flink Experience designing and optimizing high-throughput distributed pipelines. Experience with areas such as: PII scrubbing, Content-safety filtering, Dataset quality, Evaluation contamination, Training-data governance Experience with workflow orchestration frameworks such as Airflow, Prefect or Dagster. Ability to work closely with both research and engineering teams and translate research requirements into scalable infrastructure. Particularly Relevant Experience We’d be especially interested in people who have worked on: Pretraining or post-training data pipelines for LLMs Foundation-model training infrastructure Web-scale text processing Common Crawl or comparable large-scale datasets Dataset curation for large language models LLM-based data-quality evaluation ML classifiers for filtering or scoring training data Large-scale data acquisition Dataset deduplication and contamination detection Distributed ML/data infrastructure Experience with vLLM, SGLang, Docker, Kubernetes, infrastructure-as-code, experiment tracking or open-source NLP/data tooling would also be valuable.
What you’ll do
Design and scale infrastructure to transform raw, web-scale unstructured text into high-quality datasets for training advanced AI models. Develop automated data-quality systems, filtering mechanisms, and internal tools to support AI researchers.
Requirements
Requires 5+ years of experience in machine learning engineering or data infrastructure with strong production-level Python skills. Must have hands-on experience with distributed processing technologies like Spark or Ray and working with large-scale unstructured text datasets.
Listed skills
- Kubernetes · Preferred
- Docker · Preferred
- Python · Preferred
Other relevant skills
Identified from the job description. Confirm important requirements above.
- Python
- Spark
- Ray
- Flink
- NLP
- LLM Training Data
- Distributed Data Processing
- Airflow
- Prefect
- Dagster
- Data Curation
- Kubernetes
- Docker
- vLLM
- SGLang
- Data Pipeline Design
Job areas
- Data & Analytics
- Technology
- Software
- Engineering
- Science & Research
More jobs you can apply to directly
Similar opportunities posted by employers hiring on Jobs.ca, with no external application form.
Desjardins
Analyste d'affaires système(BSA) Guidewire
Direct employerEasy Apply- Hybrid
- Lévis, QC
- Posted Sep 21, 2026
Desjardins
Administrateur ou administratrice de plateforme TI - Senior
Direct employerEasy Apply- Hybrid
- Montréal, QC
- Posted Sep 17, 2026