Data Engineer – Kafka / PySpark / Hadoop : Toronto, ON (FTP)
- Toronto, ON
- On-site
- Posted Sep 11, 2026
- 1 position
- Employment type
- Full-time
- Experience level
- Lead · 10+ years
- Apply by
- Oct 10, 2026
- Posting language
- English
- Working hours
- 40 hours per week
- Seniority
- Mid-Senior level
- Application method
- Direct apply is available
This job has expired
This position at AceStack is no longer accepting applications. The original posting remains below for reference.
Expired Sep 16, 2026
Original job posting
Design, develop, and maintain scalable batch and real-time data pipelines using Python, PySpark, and Kafka. Optimize Spark jobs and SQL queries while collaborating with architects and analysts in an Agile environment.
Job details
Job Title: Data Engineer – Kafka / PySpark / Hadoop Location: Toronto, ON Work Model: Onsite Employment Type: Full-Time (FTE) Experience: 10+ Years Job Summary "We are looking for an experienced Data Engineer with strong hands-on expertise in Kafka, PySpark, Python, and Hadoop. The ideal candidate will have experience designing, developing, and supporting scalable batch and real-time data pipelines and working with large-scale distributed data processing environments. Must-Have Skills: • Strong hands-on experience with Python for data engineering and automation • Strong experience with PySpark / Apache Spark • Hands-on experience with Apache Kafka for real-time data ingestion and streaming • Strong experience with Hadoop ecosystem and distributed data processing • Strong SQL skills and experience working with large datasets • Experience developing and maintaining ETL/ELT data pipelines • Experience with data ingestion, transformation, cleansing, and integration • Strong understanding of distributed computing and data processing concepts Key Responsibilities: • Design, develop, and maintain scalable batch and real-time data pipelines • Develop data processing applications using Python and PySpark • Build and support Kafka-based data ingestion and streaming pipelines • Work with Hadoop and related technologies to process large volumes of data • Perform data transformation, validation, and integration • Troubleshoot data pipeline failures, data discrepancies, and performance issues • Analyze and optimize Spark/PySpark jobs and SQL queries • Monitor data pipelines and resolve production issues • Collaborate with data architects, developers, analysts, and business teams • Participate in Agile development, testing, deployment, and production support Good to Have: • Hive • Databricks • AWS / Azure / GCP • Git / CI-CD • Unix/Linux • Airflow / Autosys • Relational and NoSQL databases"
What you’ll do
Design, develop, and maintain scalable batch and real-time data pipelines using Python, PySpark, and Kafka. Optimize Spark jobs and SQL queries while collaborating with architects and analysts in an Agile environment.
Requirements
Requires over 10 years of experience with strong hands-on expertise in the Hadoop ecosystem, PySpark, and Kafka. Candidates must be proficient in Python and SQL for processing large-scale distributed datasets.
Listed skills
- SQL · Preferred
- Python · Preferred
Other relevant skills
Identified from the job description. Confirm important requirements above.
- Python
- PySpark
- Apache Spark
- Apache Kafka
- Hadoop
- SQL
- ETL/ELT
- Data Ingestion
- Distributed Computing
- Data Transformation
- Data Cleansing
- Data Integration
Job areas
- Data & Analytics
- Technology
- Software
- Engineering
Current jobs at AceStack
These verified opportunities are still accepting applications.