The posting
Key Responsibilities
- Design, develop and maintain scalable ETL/ELT data pipelines for large-volume structured and unstructured datasets.
- Develop high-performance data processing solutions using Python, PySpark, Apache Spark, Spark SQL and Scala.
- Build batch and real-time data pipelines using Kafka, Kinesis, Spark Streaming, AWS Glue and Airflow.
- Develop data ingestion frameworks integrating RDBMS, APIs, files, cloud storage and streaming platforms.
- Design and implement data lakes, data warehouses and cloud-based data processing platforms.
- Work with Databricks, Hadoop, Hive, Cloudera, Presto and Snowflake for large-scale data processing and analytics.
- Perform data modelling, data transformation, data quality, query optimisation and performance tuning.
- Develop and optimise SQL solutions across Oracle, SQL Server, PostgreSQL, Teradata, MongoDB and cloud databases.
- Design and implement data migration solutions involving large-scale enterprise datasets.
- Develop and support real-time and batch processing architectures for enterprise applications.
- Integrate data platforms with REST APIs, GraphQL and enterprise applications.
- Implement CI/CD and DevOps practices using Jenkins, Git, Docker, Kubernetes and OpenShift.
- Develop cloud-native data solutions using AWS and Azure, including S3, Glue, EMR, Redshift, Kinesis, Lambda, RDS and DynamoDB.
- Develop AI/GenAI-enabled data solutions involving LLMs, NLP, RAG, Agentic AI and vector databases.
- Integrate LLM services and AI platforms such as Azure OpenAI, OpenAI APIs, Hugging Face and Google Gemini/ADK.
- Develop NLP pipelines for text processing, embeddings, summarisation, sentiment analysis, voice-to-text and speaker diarisation.
- Design and implement vector search and retrieval solutions using Redis, ChromaDB and FAISS.
- Develop AI-powered APIs and applications using FastAPI, Gradio and Python.
- Collaborate with architects, data scientists, software engineers, business analysts and product teams to deliver enterprise data solutions.
- Participate in Agile SDLC activities including requirements analysis, architecture, development, testing, deployment and production support.
- Troubleshoot complex data, application and platform issues and provide scalable technical solutions.
Required Technical Skills
Data Engineering: Python, PySpark, Apache Spark, Spark SQL, Scala, Hadoop, Hive, Kafka, Presto, Databricks, Cloudera, Snowflake
Cloud Technologies: AWS, Azure, S3, Glue, EMR, Redshift, Kinesis, Lambda, RDS, DynamoDB, OpenSearch
Databases: SQL Server, Oracle, PostgreSQL, Teradata, MongoDB, Redis
Programming: Python, Java, Scala, SQL, Shell Scripting, Node.js
AI / GenAI / NLP: Generative AI, LLM, NLP, RAG, Agentic RAG, LangChain, LangGraph, LlamaIndex, Hugging Face Transformers, Azure OpenAI, OpenAI API, Google Gemini/ADK, PyTorch
Vector & AI Search: Redis Vector Database, ChromaDB, FAISS, Embeddings, Hybrid Search, Semantic Search
DevOps & Deployment: Docker, Kubernetes, OpenShift, Jenkins, Git, Terraform, CI/CD
Data Integration & APIs: REST APIs, GraphQL, FastAPI, API Gateway, AWS Lambda, CDC, Debezium
Data Visualisation: Power BI, Data Modelling, Reporting and Analytics
Qualifications
- Bachelor's or Master's degree in Computer Science, Information Technology, Programming & Systems Analysis, Computer Studies or a related discipline.
- Strong professional experience in Big Data, Cloud Computing or related technology domains.
- Experience working with large-scale enterprise data platforms and production data pipelines.
- Strong programming and SQL skills.
- Experience with cloud-based data engineering and modern data processing frameworks.
- Experience with AI/ML, NLP or Generative AI will be highly advantageous.
Preferred Experience
- Enterprise Banking / Financial Services experience.
- Experience working with large-scale customer, transaction or financial datasets.
- Experience with data migration and legacy ETL modernisation.
- Experience implementing AI/GenAI solutions within enterprise data platforms.
- Experience with production deployments and CI/CD environments.
- Strong understanding of data governance, security, data quality and performance optimisation.



