About the Role
We are looking for an experienced Data Engineer with 4+ years of hands-on experience in designing, developing, and maintaining scalable data pipelines and data platforms.
The ideal candidate will have strong expertise in Python, SQL, PySpark/Spark, ETL/ELT, workflow orchestration, cloud platforms, and data modeling.
You will work closely with Data Scientists, Analysts, Software Engineers, ML Engineers, and business teams to build reliable, scalable, and production-ready data solutions.
Key Responsibilities
Design, develop, and maintain scalable batch and real-time data pipelines.Build and optimize ETL/ELT workflows using Python, SQL, and PySpark/Spark.Develop data ingestion and transformation pipelines for structured and semi-structured data.Work with large-scale datasets and optimize distributed data processing for performance and cost.Design and implement data models, schemas, data lakes, and data warehouses.Develop and manage workflow orchestration using Apache Airflow or similar tools.Implement data quality checks, validation, monitoring, and error-handling mechanisms.Work with cloud data services and storage platforms such as AWS, Azure, or GCP.Collaborate with Data Scientists, Analysts, ML Engineers, and Product teams to deliver production-ready datasets.Troubleshoot pipeline failures and improve reliability, scalability, and performance.Follow engineering best practices including Git, CI/CD, testing, documentation, and code reviews.Required Qualifications
4+ years of professional experience in Data Engineering or a closely related field.Strong programming skills in Python.Strong proficiency in SQL, including complex joins, aggregations, CTEs, and window functions.Hands-on experience with Apache Spark / PySpark.Strong understanding of ETL/ELT pipelines and data engineering concepts.Experience with Apache Airflow or another workflow orchestration tool.Professional experience with at least one major cloud platform: AWS, Azure, or GCP.Strong understanding of data warehousing and data modeling.Experience with relational databases such as PostgreSQL, MySQL, SQL Server, or similar.Familiarity with Git and CI/CD practices.Strong problem-solving, analytical, and communication skillsPreferred Qualifications
Experience with Kafka or other streaming technologies.Experience with Databricks, Snowflake, BigQuery, Redshift, or similar platforms.Knowledge of Delta Lake, Apache Iceberg, or modern Lakehouse architectures.Experience with Docker and cloud-based DevOps practices.Knowledge of data governance, security, and data quality frameworks.Experience building real-time/streaming data pipelines.Familiarity with dbt or similar transformation frameworks.Technology Stack
Programming: Python, SQLData Processing: Apache Spark, PySparkOrchestration: Apache AirflowCloud: AWS / Azure / GCPDatabases: PostgreSQL, MySQL, SQL Server or similarData Platforms: Snowflake, BigQuery, Databricks, RedshiftDevelopment: Git, CI/CD, DockerBenefits
Opportunity to work on large-scale data engineering projects.Exposure to modern cloud, data processing, and data platform technologies.Work with experienced engineering and data teams.Opportunity to build and optimize production-grade data pipelines.Competitive compensation based on experience and skills.Hybrid working environment in Pune.What We're Looking For
We are looking for a Data Engineer who can independently own data engineering projects, make sound technical decisions, and build production-grade data pipelines.
The ideal candidate should be comfortable working with large datasets, distributed processing frameworks, cloud platforms, data warehouses, and cross-functional teams.
Candidates should have strong ownership, problem-solving abilities, attention to data quality, and the ability to work in a fast-paced engineering environment.