Junior Data Engineer
About Us
We are dedicated to building a cleaner, more sustainable future. By applying innovative technology and data-driven insights to green environmental initiatives, we help measure, analyze, and reduce environmental impact at scale. We are looking for a passionate, forward-thinking Junior Data Engineer to join our data team and help build the data pipelines powering our eco-focused solutions.
Position Overview
As a fresh graduate joining our team, you will work closely with senior data engineers and analysts to design, build, and maintain high-volume data pipelines. You will transform raw environmental datasets—such as energy metrics, carbon emissions data, and resource usage—into actionable insights. This role is ideal for a recent Computer Science graduate eager to apply modern data tooling (AWS, PySpark, Python) toward solving meaningful sustainability challenges.
Key Responsibilities
-
Pipeline Development: Design, build, and maintain automated batch and real-time ETL/ELT pipelines to ingest, clean, and transform large-scale environmental data.
-
Data Processing: Utilize Python and Apache Spark (PySpark) to process structured and unstructured datasets efficiently.
-
Cloud Infrastructure: Help manage and expand our cloud data infrastructure using core AWS services (e.g., S3, Glue, EMR, Redshift, Lambda).
-
Data Quality & Governance: Implement automated testing, validation, and monitoring to ensure data accuracy, reliability, and security.
-
Cross-Functional Collaboration: Partner with Data Scientists, Business Analysts, and Sustainability Specialists to deliver clean, structured data for reporting and machine learning applications.
Required Qualifications
-
Education: Bachelor’s degree in Computer Science (or a closely related core computing field, such as Computer Engineering or Software Engineering) completed within the last 0–12 months.
-
Core Programming: Strong foundation in Python and fundamental software engineering principles (OOP, data structures, algorithms, version control with Git).
-
Distributed Computing: Academic or hands-on project experience using Apache Spark / PySpark to process large datasets.
-
Cloud Fundamentals: Working knowledge or project experience with AWS core services (S3, EC2, IAM, Lambda, or managed data services).
-
Databases & SQL: Solid grasp of relational databases, SQL query writing, data modeling concepts, and basic schema design.
Nice-to-Have / Preferred Qualifications
-
Coursework, internship, or personal project focus on environmental data, sustainability, clean energy, or IoT telemetry data.
-
Exposure to workflow orchestration tools (e.g., Apache Airflow, Dagster).
-
Familiarity with containerization technologies (Docker, Kubernetes).
-
Knowledge of CI/CD practices for data infrastructure.
What We Offer
-
Mission-Driven Impact: Direct involvement in projects that combat climate change and advance sustainable practices.
-
Mentorship & Growth: A collaborative environment with dedicated mentorship from experienced senior data engineers.