View all jobs

Junior Data Engineer

  • Vaughan, Ontario

About Us

​We are dedicated to building a cleaner, more sustainable future. By applying innovative technology and data-driven insights to green environmental initiatives, we help measure, analyze, and reduce environmental impact at scale. We are looking for a passionate, forward-thinking Junior Data Engineer to join our data team and help build the data pipelines powering our eco-focused solutions.

Position Overview

As a fresh graduate joining our team, you will work closely with senior data engineers and analysts to design, build, and maintain high-volume data pipelines. You will transform raw environmental datasets—such as energy metrics, carbon emissions data, and resource usage—into actionable insights. This role is ideal for a recent Computer Science graduate eager to apply modern data tooling (AWS, PySpark, Python) toward solving meaningful sustainability challenges.

Key Responsibilities

  • Pipeline Development: Design, build, and maintain automated batch and real-time ETL/ELT pipelines to ingest, clean, and transform large-scale environmental data.

  • Data Processing: Utilize Python and Apache Spark (PySpark) to process structured and unstructured datasets efficiently.

  • Cloud Infrastructure: Help manage and expand our cloud data infrastructure using core AWS services (e.g., S3, Glue, EMR, Redshift, Lambda).

  • Data Quality & Governance: Implement automated testing, validation, and monitoring to ensure data accuracy, reliability, and security.

  • Cross-Functional Collaboration: Partner with Data Scientists, Business Analysts, and Sustainability Specialists to deliver clean, structured data for reporting and machine learning applications.

Required Qualifications

  • Education: Bachelor’s degree in Computer Science (or a closely related core computing field, such as Computer Engineering or Software Engineering) completed within the last 0–12 months.

  • Core Programming: Strong foundation in Python and fundamental software engineering principles (OOP, data structures, algorithms, version control with Git).

  • Distributed Computing: Academic or hands-on project experience using Apache Spark / PySpark to process large datasets.

  • Cloud Fundamentals: Working knowledge or project experience with AWS core services (S3, EC2, IAM, Lambda, or managed data services).

  • Databases & SQL: Solid grasp of relational databases, SQL query writing, data modeling concepts, and basic schema design.

Nice-to-Have / Preferred Qualifications

  • Coursework, internship, or personal project focus on environmental data, sustainability, clean energy, or IoT telemetry data.

  • Exposure to workflow orchestration tools (e.g., Apache Airflow, Dagster).

  • Familiarity with containerization technologies (Docker, Kubernetes).

  • Knowledge of CI/CD practices for data infrastructure.

What We Offer

  • Mission-Driven Impact: Direct involvement in projects that combat climate change and advance sustainable practices.

  • Mentorship & Growth: A collaborative environment with dedicated mentorship from experienced senior data engineers.