The posting
Job Summary
We are seeking a Senior Data Engineer to design, develop, and optimize scalable data pipelines and processing solutions using Big Data, cloud, and AI technologies. Collaborate with architects, data scientists, and engineering teams to deliver high-performance, production-ready data systems.
Responsibilities
- Design, develop, and maintain scalable batch and real-time data pipelines using Apache Spark/PySpark, SQL, and Python to support large-scale data processing
- Develop and optimize ETL/ELT pipelines for efficient data ingestion and transformation
- Build data ingestion solutions integrating databases, APIs, files, and streaming platforms to ensure reliable data flow
- Utilize AWS services including S3, Glue, EMR, Redshift, Kinesis, Lambda, and DynamoDB to deploy and manage cloud-based data solutions
- Develop and optimize data processing workflows using Hadoop, Hive, Spark, Kafka, Cloudera, and Databricks for high throughput and low latency
- Perform SQL and Spark performance tuning to enhance large-scale data workload efficiency
- Design and implement data models, data warehouses, and data lake architectures to support analytics and reporting needs
- Develop CI/CD pipelines and support containerized deployments using Docker, Kubernetes, and OpenShift to automate delivery and scaling
- Integrate Generative AI and NLP capabilities into enterprise data applications to enhance data insights and automation
- Develop solutions leveraging LLM frameworks, RAG, vector databases, and AI APIs to incorporate advanced AI functionalities
- Collaborate with solution architects, data scientists, software engineers, and business stakeholders to deliver robust, production-ready data solutions
- Participate in system design, development, testing, deployment, and provide production support to ensure system reliability
Required competencies and certifications
- Bachelor’s or Master’s degree in Computer Science, Information Technology, Engineering, or related field
- Minimum 8 years of experience in Data Engineering or Big Data with expertise in large-scale data processing
- Proficient programming skills in Python, Java, and/or Scala for data engineering tasks
- Hands-on experience with Apache Spark/PySpark, SQL, Hadoop, and Hive for data processing
- Experience working with cloud platforms, particularly AWS services
- Knowledge of data warehouses, relational databases, and data lake technologies
- Experience with Kafka, Airflow, Jenkins, Git, and CI/CD practices for data pipeline orchestration and automation
Preferred competencies and qualifications
- Experience with Docker, Kubernetes, or OpenShift for containerization and orchestration (marked advantageous in original JD)
- Familiarity with Databricks, Snowflake, and Cloudera platforms (marked advantageous in original JD)
- Exposure to Generative AI, LLMs, RAG, LangChain/LangGraph, and vector databases (marked advantageous in original JD)
- Strong analytical, problem-solving, and communication skills



