Back to job search
IS
Iris Software Inc.Verified Job Source

Data Engineer

Lead the development of PySpark data engineering capabilities and feature engineering pipelines from the ground up. Translate and migrate legacy SAS-based workloads into a modern, scalable PySpark environment.

  • Hybrid
  • Toronto, ON
  • Posted Aug 18, 2026
  • Apply by Sep 17, 2026
  • 1 position

More jobs you can apply to directly

Similar opportunities posted by employers hiring on Jobs.ca, with no external application form.

Job summary

Iris Software is looking to hire a Data Engineer for a full-time opportunity in Toronto, ON (hybrid position). Please respond back with your most recent resume if you would be interested...! Job Title: Data Engineer Location: Toronto, ON (2 days onsite) Full-time with Iris Software for one of the banks in downtown Toronto The client is looking for: - feature engineering - data wrangling - model automation - Tool stack probably in this order – Python, PySpark, Glue, Sagemaker, SAS We are seeking a highly skilled Data Engineer to help build a net-new PySpark data engineering capability from the ground up. The initial focus of this role will be designing and developing net-new feature engineering pipelines to support downstream data science initiatives. As the platform capability matures, you will play a critical role in a large-scale modernization effort, tasked with translating and migrating legacy SAS-based workloads into a modern, scalable PySpark environment managed via Anaconda. The ideal candidate possesses strong core data engineering expertise, experience building data preparation pipelines, and the analytical ability to reverse-engineer legacy business logic into optimized open-source code. Key Responsibilities Foundation & Feature Engineering: Lead the foundational setup of Python/PySpark development environments, actively managing complex package dependencies and virtual environments using Conda/Anaconda. Pipeline Development: Design, build, and deploy net-new data pipelines focused heavily on feature engineering and data preparation to feed downstream machine learning and analytics use cases. Code Translation: Analyze and reverse-engineer existing SAS programs (developed by business users and data scientists) to accurately extract business rules, data transformations, and calculations. Platform Modernization: Translate extracted SAS logic into efficient, scalable, and functionally equivalent PySpark code. Performance Tuning: Optimize PySpark code for performance, scalability, and maintainability within distributed data processing environments. Stakeholder Collaboration: Collaborate closely with business users and data scientists to clarify requirements, validate feature outputs, and resolve discrepancies during the migration process. Quality Assurance: Perform robust unit testing, data reconciliation, and automated validation to ensure absolute data parity between legacy SAS outputs and the new PySpark pipelines. Documentation: Document technical designs, code lineage, testing results, and migration methodologies for future team scaling. Required Qualifications 5+ years of hands-on experience in data engineering, data pipeline development, or building data platforms. Strong programming expertise in Python and PySpark for distributed data processing. Proven experience building data pipelines specifically for feature engineering, data curation, and advanced data preparation. Ability to read, interpret, and reverse-engineer legacy SAS code (such as SAS data steps, procedures, and macros) to extract complex business logic. (Note: Deep SAS development experience is a plus, but the ability to translate it is the core requirement). Experience establishing and managing Python environments, ensuring reproducibility, and handling library dependencies using Anaconda/Conda. Advanced SQL skills and experience working with large-scale structured datasets. Experience with code migration, platform modernization, or translating legacy codebases into modern open-source stacks. Strong analytical, problem-solving, and communication skills, with a track record of successfully interfacing directly with business stakeholders. Preferred Qualifications Experience working with cloud-based data platforms (Azure, AWS, or GCP). Familiarity with version control tools (Git) and collaborative development workflows. Experience with CI/CD pipelines, automated testing, and MLOps principles. Exposure to financial services, banking, or other highly regulated industries. Knowledge of data governance, data cataloging, and data quality best practices. Key Skills Python & PySpark Feature Engineering & Data Pipelines Anaconda / Conda Environment Management SAS Interpretation & Code Translation Distributed Computing & Spark Optimization SQL & Data Reconciliation Legacy Platform Modernization Agile Delivery About Iris Software Inc. With 4,000+ associates and offices in India, U.S.A. and Canada, Iris Software delivers technology services and solutions that help clients complete fast, far-reaching digital transformations and achieve their business goals. A strategic partner to Fortune 500 and other top companies in financial services and many other industries, Iris provides a value-driven approach - a unique blend of highly skilled specialists, software engineering expertise, cutting-edge technology, and flexible engagement models. High customer satisfaction has translated into long-standing relationships and preferred-partner status with many of our clients, who rely on our 30+ years of technical and domain expertise to future-proof their enterprises. Associates of Iris work on mission-critical applications supported by a workplace culture that has won numerous awards in the last few years, including Certified Great Place to Work in India; Top 25 GPW in IT & IT-BPM; Ambition Box Best Place to Work, #3 in IT/ITES; and Top Workplace NJ-USA. Kind regards, Amarpreet Singh 200 Bay Str. Toronto, ON, M5H4E9 200 Metroplex Drive, Suite #300 Edison, NJ 08817 [email protected] | www.irissoftware.com Iris Software ranked 38th among 1200 companies spanning 20+ industries in the country to become India’s Best Companies to Work For 2022!

What you’ll do

Lead the development of PySpark data engineering capabilities and feature engineering pipelines from the ground up. Translate and migrate legacy SAS-based workloads into a modern, scalable PySpark environment.

Requirements

Requires over 5 years of experience in data engineering with strong expertise in Python, PySpark, and SQL. Must be able to reverse-engineer legacy SAS code and manage Python environments using Anaconda/Conda.

Listed skills

  • SQLPreferred
  • PythonPreferred

Other relevant skills

Identified from the job description. Confirm important requirements above.

  • Python
  • PySpark
  • Feature Engineering
  • Data Pipelines
  • Anaconda
  • Conda
  • SAS Translation
  • Distributed Computing
  • Spark Optimization
  • SQL
  • Data Reconciliation
  • Platform Modernization
  • Glue
  • Sagemaker
  • Data Wrangling
  • Model Automation

Job areas

  • Data & Analytics
  • Technology
  • Software
  • Engineering
  • Consulting

Additional details

Minimum experience
5+ years
Apply by
Sep 17, 2026
Posting language
English
Working hours
40 hours per week
Office presence
2 days per week
Seniority
Mid-Senior level
Application method
Direct apply is available