Opens an external site
- Employment type
- Full-time
- Experience level
- Senior · 5+ years
- Apply by
- Oct 28, 2026
- Posting language
- English
- Working hours
- 40 hours per week
- Seniority
- Mid-Senior level
- Application method
- Direct apply is available
Job summary
Build and maintain large-scale ETL/ELT pipelines and high-performance data processing jobs using platforms and tools such as Hadoop, Databricks, Spark, PySpark, and Python. Optimize pipeline performance and cost, contribute to CI/CD practices, and collaborate with Product Managers to scope and develop features.
Job details
Job Summary Work with cutting-edge big data platforms (e.g., Databricks, Apache Spark) at large scale, pushing the boundaries of data processing and model enablement. • Build and maintain robust ETL/ELT pipelines for ingestion, transformation, and aggregation of large-scale datasets on Hadoop and enterprise data platforms. • Develop high-performance data processing jobs using PySpark/Spark, Python on data platforms such as cloudera and databricks. • Optimize pipeline performance and cost through partitioning, file formats, compute tuning, and efficient query patterns • Contribute to CI/CD for data workflows (testing, code reviews, deployment automation), promoting engineering best practices and maintainable codebases. • Partner with Product Managers to develop a deep understanding of users and use cases and apply that knowledge to scoping and building new modules and features Ideal Candidate Qualifications: • Strong hands-on experience in data engineering building production-grade pipelines on big data platforms (Hadoop ecosystem and cloud data platforms - databricks). • High proficiency in using Python, Spark, Hadoop platforms & tools (Hive, Impala, Airflow, NiFi), SQL to build Big Data products. • Hands-on experience with cloud data platforms such as databricks, snowflake (databricks preferred) • Experience with orchestration/integration tools such as Apache Airflow, Apache NiFi, or Talend. • Working knowledge of DevOps/CI-CD practices: version control (Git), automated testing, release pipelines, and observability. • Strong problem-solving skills with the ability to debug complex data issues and communicate clearly with technical and non-technical stakeholders. • Experience developing Java based applications is an added advantage. LONGDESCRIPTION section. 2 of 6. Section Title: Key Responsibilities Key Responsibilities Work with cutting-edge big data platforms (e.g., Databricks, Apache Spark) at large scale, pushing the boundaries of data processing and model enablement. • Build and maintain robust ETL/ELT pipelines for ingestion, transformation, and aggregation of large-scale datasets on Hadoop and enterprise data platforms. • Develop high-performance data processing jobs using PySpark/Spark, Python on data platforms such as cloudera and databricks. • Optimize pipeline performance and cost through partitioning, file formats, compute tuning, and efficient query patterns • Contribute to CI/CD for data workflows (testing, code reviews, deployment automation), promoting engineering best practices and maintainable codebases. • Partner with Product Managers to develop a deep understanding of users and use cases and apply that knowledge to scoping and building new modules and features Ideal Candidate Qualifications: • Strong hands-on experience in data engineering building production-grade pipelines on big data platforms (Hadoop ecosystem and cloud data platforms - databricks). • High proficiency in using Python, Spark, Hadoop platforms & tools (Hive, Impala, Airflow, NiFi), SQL to build Big Data products. • Hands-on experience with cloud data platforms such as databricks, snowflake (databricks preferred) • Experience with orchestration/integration tools such as Apache Airflow, Apache NiFi, or Talend. • Working knowledge of DevOps/CI-CD practices: version control (Git), automated testing, release pipelines, and observability. • Strong problem-solving skills with the ability to debug complex data issues and communicate clearly with technical and non-technical stakeholders. • Experience developing Java based applications is an added advantage. LONGDESCRIPTION section. 3 of 6. Section Title: Skill Requirements Skill Requirements Work with cutting-edge big data platforms (e.g., Databricks, Apache Spark) at large scale, pushing the boundaries of data processing and model enablement. • Build and maintain robust ETL/ELT pipelines for ingestion, transformation, and aggregation of large-scale datasets on Hadoop and enterprise data platforms. • Develop high-performance data processing jobs using PySpark/Spark, Python on data platforms such as cloudera and databricks. • Optimize pipeline performance and cost through partitioning, file formats, compute tuning, and efficient query patterns • Contribute to CI/CD for data workflows (testing, code reviews, deployment automation), promoting engineering best practices and maintainable codebases. • Partner with Product Managers to develop a deep understanding of users and use cases and apply that knowledge to scoping and building new modules and features Ideal Candidate Qualifications: • Strong hands-on experience in data engineering building production-grade pipelines on big data platforms (Hadoop ecosystem and cloud data platforms - databricks). • High proficiency in using Python, Spark, Hadoop platforms & tools (Hive, Impala, Airflow, NiFi), SQL to build Big Data products. • Hands-on experience with cloud data platforms such as databricks, snowflake (databricks preferred) • Experience with orchestration/integration tools such as Apache Airflow, Apache NiFi, or Talend. • Working knowledge of DevOps/CI-CD practices: version control (Git), automated testing, release pipelines, and observability. • Strong problem-solving skills with the ability to debug complex data issues and communicate clearly with technical and non-technical stakeholders. • Experience developing Java based applications is an added advantage. SKILL section. 4 of 6. Section Title: Must Have Skills Must Have Skills Click Enter to show the proficiency description of Data EngineeringData Engineering Click Enter to show the proficiency description of DatabricksDatabricks Click Enter to show the proficiency description of Apache SparkApache Spark Click Enter to show the proficiency description of PySparkPySpark Click Enter to show the proficiency description of PythonPython Click Enter to show the proficiency description of CI/CDCI/CD SKILL section. 5 of 6. Section Title: Good to have Skills Good to have Skills LONGDESCRIPTION section. 6 of 6. Section Title: Other Requirements Other Requirements Work with cutting-edge big data platforms (e.g., Databricks, Apache Spark) at large scale, pushing the boundaries of data processing and model enablement. • Build and maintain robust ETL/ELT pipelines for ingestion, transformation, and aggregation of large-scale datasets on Hadoop and enterprise data platforms. • Develop high-performance data processing jobs using PySpark/Spark, Python on data platforms such as cloudera and databricks. • Optimize pipeline performance and cost through partitioning, file formats, compute tuning, and efficient query patterns • Contribute to CI/CD for data workflows (testing, code reviews, deployment automation), promoting engineering best practices and maintainable codebases. • Partner with Product Managers to develop a deep understanding of users and use cases and apply that knowledge to scoping and building new modules and features Ideal Candidate Qualifications: • Strong hands-on experience in data engineering building production-grade pipelines on big data platforms (Hadoop ecosystem and cloud data platforms - databricks). • High proficiency in using Python, Spark, Hadoop platforms & tools (Hive, Impala, Airflow, NiFi), SQL to build Big Data products. • Hands-on experience with cloud data platforms such as databricks, snowflake (databricks preferred) • Experience with orchestration/integration tools such as Apache Airflow, Apache NiFi, or Talend. • Working knowledge of DevOps/CI-CD practices: version control (Git), automated testing, release pipelines, and observability. • Strong problem-solving skills with the ability to debug complex data issues and communicate clearly with technical and non-technical stakeholders. • Experience developing Java based applications is an added advantage
What you’ll do
Build and maintain large-scale ETL/ELT pipelines and high-performance data processing jobs using platforms and tools such as Hadoop, Databricks, Spark, PySpark, and Python. Optimize pipeline performance and cost, contribute to CI/CD practices, and collaborate with Product Managers to scope and develop features.
Requirements
Requires hands-on experience building production-grade data pipelines on big data and cloud data platforms, with strong proficiency in Python, Spark, Hadoop tools, and SQL. Experience with orchestration tools and CI/CD practices is expected; strong problem-solving and communication skills are required, while Java application development is an advantage.
Listed skills
- SQL · Preferred
- CI/CD · Preferred
- Git · Preferred
- Python · Preferred
Other relevant skills
Identified from the job description. Confirm important requirements above.
- Data Engineering
- Databricks
- Apache Spark
- PySpark
- Python
- CI/CD
- Hadoop
- SQL
- ETL/ELT
- Hive
- Impala
- Apache Airflow
- Apache NiFi
- Talend
- Snowflake
- Git
Job areas
- Data & Analytics
- Technology
- Software
- Engineering
More jobs from HCLTech
Senior Full Stack Engineer (Azure Cloud Platform)
- Hybrid
- ON
- Posted Sep 26, 2026
Azure DevOps Lead Technical Consultant
- Hybrid
- ON
- Posted Sep 25, 2026
Data Engineer
- On-site
- Toronto, ON
- Posted Sep 24, 2026
Technicien Support Informatique
- Hybrid
- Baie-Comeau, QC
- Posted Sep 19, 2026