Back to job search
kadence logo
kadenceVerified Job Source

Machine Learning Data Engineer

  • Montréal, QC
  • On-site
  • Posted Sep 2, 2026
  • 1 position

Opens an external site

Employment type
Full-time
Experience level
Senior · 5+ years
Apply by
Oct 2, 2026
Posting language
English
Working hours
40 hours per week
Seniority
Mid-Senior level
Application method
Direct apply is available

Job summary

Design and scale infrastructure to transform raw, web-scale unstructured text into high-quality datasets for training advanced AI models. Develop automated data-quality systems, filtering mechanisms, and internal tools to support AI researchers.

Job details

Senior Machine Learning Data Engineer Location: Canada The Opportunity This is an opportunity to work alongside a highly accomplished AI research team tackling fundamental problems in advanced machine learning. Rather than maintaining established pipelines, you’ll be helping develop new approaches to training-data engineering and curation where established playbooks often do not yet exist. The organization is building a substantial technical team in Montreal, and we are hiring multiple people across this area. If you’ve worked on LLM training data, large-scale NLP pipelines, foundation-model infrastructure or web-scale data processing, I’d be interested in speaking with you. We’re working with an ambitious AI research organization building next-generation machine learning systems and are looking for multiple Senior Machine Learning Data Engineers to join its growing technical team. This role sits at the intersection of data engineering, machine learning, NLP, and large-scale training data curation. You’ll be responsible for building and scaling the infrastructure that transforms raw, web-scale data into high-quality datasets used to train advanced AI models. The core challenge is engineering data quality at enormous scale: developing filtering systems, model-based quality scoring, dataset transformations, contamination detection, and tooling that allows researchers to understand and work with training corpora effectively. What You’ll Work On Design, build and scale pipelines that transform raw web-scale data into high-quality datasets for training large machine learning models. Process extremely large unstructured text corpora, including datasets at trillion-token scale. Build sophisticated data-processing systems covering: Deduplication, Model-based quality scoring, Heuristic filtering, Toxicity and content-safety filtering, PII detection and removal, Metadata extraction, Dataset transformations, Versioning and provenance tracking Develop automated data-quality systems using a combination of heuristics, ML classifiers, LLM-based evaluators and human-in-the-loop workflows. Build monitoring and guardrails to identify data-quality regressions before they impact downstream model training. Work closely with AI researchers to understand evolving training-data requirements and identify gaps within existing datasets. Design robust evaluation contamination and data-leakage detection systems. Build internal tools that allow researchers to easily explore, query and understand large datasets. Optimize large-scale distributed processing systems for throughput, reliability and infrastructure cost. What We’re Looking For 5+ years of experience across machine learning engineering, data processing, NLP, data infrastructure or related areas. Experience working with very large unstructured text datasets, ideally at web or foundation-model scale. Strong production-level Python. Hands-on experience with distributed data-processing technologies such as: Spark, Ray, Flink Experience designing and optimizing high-throughput distributed pipelines. Experience with areas such as: PII scrubbing, Content-safety filtering, Dataset quality, Evaluation contamination, Training-data governance Experience with workflow orchestration frameworks such as Airflow, Prefect or Dagster. Ability to work closely with both research and engineering teams and translate research requirements into scalable infrastructure. Particularly Relevant Experience We’d be especially interested in people who have worked on: Pretraining or post-training data pipelines for LLMs Foundation-model training infrastructure Web-scale text processing Common Crawl or comparable large-scale datasets Dataset curation for large language models LLM-based data-quality evaluation ML classifiers for filtering or scoring training data Large-scale data acquisition Dataset deduplication and contamination detection Distributed ML/data infrastructure Experience with vLLM, SGLang, Docker, Kubernetes, infrastructure-as-code, experiment tracking or open-source NLP/data tooling would also be valuable.

What you’ll do

Design and scale infrastructure to transform raw, web-scale unstructured text into high-quality datasets for training advanced AI models. Develop automated data-quality systems, filtering mechanisms, and internal tools to support AI researchers.

Requirements

Requires 5+ years of experience in machine learning engineering or data infrastructure with strong production-level Python skills. Must have hands-on experience with distributed processing technologies like Spark or Ray and working with large-scale unstructured text datasets.

Listed skills

  • Kubernetes · Preferred
  • Docker · Preferred
  • Python · Preferred

Other relevant skills

Identified from the job description. Confirm important requirements above.

  • Python
  • Spark
  • Ray
  • Flink
  • NLP
  • LLM Training Data
  • Distributed Data Processing
  • Airflow
  • Prefect
  • Dagster
  • Data Curation
  • Kubernetes
  • Docker
  • vLLM
  • SGLang
  • Data Pipeline Design

Job areas

  • Data & Analytics
  • Technology
  • Software
  • Engineering
  • Science & Research

More jobs you can apply to directly

Similar opportunities posted by employers hiring on Jobs.ca, with no external application form.

Browse all Easy Apply jobs