The posting
Key Responsibilities
- Design, develop and maintain scalable data pipelines and data ingestion frameworks for large-volume datasets.
- Develop data transformation and processing applications using Apache Spark, PySpark, Scala and Python.
- Build and optimize data pipelines using Azure Databricks, Azure Data Factory, AWS EMR and related cloud services.
- Work with Hadoop, HDFS, Hive, Snowflake, Teradata and Data Lake environments.
- Develop batch and real-time data processing solutions using Spark Structured Streaming and Kafka.
- Perform data extraction, transformation and loading across heterogeneous source and target systems.
- Develop and optimize Spark SQL, HiveQL and SQL queries for performance and cost efficiency.
- Design data models, partitioning strategies and scalable data storage architectures.
- Build and manage workflow orchestration using Apache Airflow.
- Implement CI/CD pipelines and automated testing using tools such as Jenkins, Docker, GitHub Actions and pytest.
- Troubleshoot data pipeline, performance and production issues and implement sustainable solutions.
- Collaborate with business stakeholders, architects and technology teams to understand requirements and deliver data engineering solutions.
- Ensure data quality, reliability, security and operational stability across enterprise data platforms.
Required Skills
- 6+ years of experience in Data Engineering / Big Data Engineering.
- Strong hands-on experience with Apache Spark / PySpark.
- Strong programming skills in Python and/or Scala.
- Good experience with Hadoop, HDFS and Hive.
- Experience developing ETL/ELT and data ingestion pipelines.
- Strong SQL and data processing skills.
- Experience with Azure Databricks, Azure Data Factory, AWS EMR or equivalent cloud data platforms.
- Experience with Kafka / real-time streaming is an advantage.
- Hands-on experience with Airflow and data pipeline orchestration.
- Experience with Snowflake, Teradata, SQL Server or other enterprise databases.
- Good understanding of Data Lake, Delta Lake, Data Warehousing and Data Modelling.
- Experience with Git, CI/CD, Docker and automated testing.



