Back to job search
AceStack logo
AceStackVerified Job Source

Data Engineer – Kafka / PySpark / Hadoop : Toronto, ON (FTP)

Expired
  • Toronto, ON
  • On-site
  • Posted Sep 11, 2026
  • 1 position
Employment type
Full-time
Experience level
Lead · 10+ years
Apply by
Oct 10, 2026
Posting language
English
Working hours
40 hours per week
Seniority
Mid-Senior level
Application method
Direct apply is available

This job has expired

This position at AceStack is no longer accepting applications. The original posting remains below for reference.

Expired Sep 16, 2026

Original job posting

Design, develop, and maintain scalable batch and real-time data pipelines using Python, PySpark, and Kafka. Optimize Spark jobs and SQL queries while collaborating with architects and analysts in an Agile environment.

Job details

Job Title: Data Engineer – Kafka / PySpark / Hadoop Location: Toronto, ON Work Model: Onsite Employment Type: Full-Time (FTE) Experience: 10+ Years Job Summary "We are looking for an experienced Data Engineer with strong hands-on expertise in Kafka, PySpark, Python, and Hadoop. The ideal candidate will have experience designing, developing, and supporting scalable batch and real-time data pipelines and working with large-scale distributed data processing environments. Must-Have Skills: • Strong hands-on experience with Python for data engineering and automation • Strong experience with PySpark / Apache Spark • Hands-on experience with Apache Kafka for real-time data ingestion and streaming • Strong experience with Hadoop ecosystem and distributed data processing • Strong SQL skills and experience working with large datasets • Experience developing and maintaining ETL/ELT data pipelines • Experience with data ingestion, transformation, cleansing, and integration • Strong understanding of distributed computing and data processing concepts Key Responsibilities: • Design, develop, and maintain scalable batch and real-time data pipelines • Develop data processing applications using Python and PySpark • Build and support Kafka-based data ingestion and streaming pipelines • Work with Hadoop and related technologies to process large volumes of data • Perform data transformation, validation, and integration • Troubleshoot data pipeline failures, data discrepancies, and performance issues • Analyze and optimize Spark/PySpark jobs and SQL queries • Monitor data pipelines and resolve production issues • Collaborate with data architects, developers, analysts, and business teams • Participate in Agile development, testing, deployment, and production support Good to Have: • Hive • Databricks • AWS / Azure / GCP • Git / CI-CD • Unix/Linux • Airflow / Autosys • Relational and NoSQL databases"

What you’ll do

Design, develop, and maintain scalable batch and real-time data pipelines using Python, PySpark, and Kafka. Optimize Spark jobs and SQL queries while collaborating with architects and analysts in an Agile environment.

Requirements

Requires over 10 years of experience with strong hands-on expertise in the Hadoop ecosystem, PySpark, and Kafka. Candidates must be proficient in Python and SQL for processing large-scale distributed datasets.

Listed skills

  • SQL · Preferred
  • Python · Preferred

Other relevant skills

Identified from the job description. Confirm important requirements above.

  • Python
  • PySpark
  • Apache Spark
  • Apache Kafka
  • Hadoop
  • SQL
  • ETL/ELT
  • Data Ingestion
  • Distributed Computing
  • Data Transformation
  • Data Cleansing
  • Data Integration

Job areas

  • Data & Analytics
  • Technology
  • Software
  • Engineering

Current jobs at AceStack

These verified opportunities are still accepting applications.