The posting
DevOps Engineer –Kubernetes/OpenShift, OpenTelemetry & Enterprise Observability
Successful applicants must be comfortable with weekends and shifts as and when required.
Job Description
We are seeking a hands-on DevOps Engineer to support the deployment, automation, monitoring and operational reliability of enterprise microservices within a large-scale banking technology environment.
The role focuses on Kubernetes/OpenShift container platforms, CI/CD engineering and distributed observability, including application metrics, logs, traces and telemetry pipelines.
The successful candidate will collaborate with development, platform engineering, infrastructure and production support teams to implement observability solutions, troubleshoot complex deployment and telemetry issues, and improve application reliability across development, testing and production environments.
Key Responsibilities
Kubernetes, OpenShift & Microservices
- Deploy, configure, maintain and troubleshoot containerised microservices running on Kubernetes and OpenShift platforms.
- Support Kubernetes workloads, Helm deployments, namespaces, resource management, service connectivity and container lifecycle operations.
- Investigate pod failures, deployment issues, application health problems, resource constraints and service communication errors.
- Support application onboarding, environment configuration and deployment automation across SIT, UAT and production environments.
CI/CD & Automation Engineering
- Develop, maintain and optimise CI/CD pipelines using Jenkins, GitLab CI/CD, GitHub Actions or equivalent platforms.
- Integrate monitoring configurations, application instrumentation and deployment validation into automated delivery workflows.
- Automate environment provisioning, configuration management, application deployment and operational tasks.
- Troubleshoot pipeline failures, container deployment issues and environment-specific release problems.
- Support production deployment preparation, release execution, cutover planning and post-deployment verification.
Enterprise Observability & Telemetry
- Implement, configure and support observability platforms using Grafana, Prometheus and related monitoring technologies.
- Support OpenTelemetry Collectors and application instrumentation for metrics, logs and distributed traces.
- Work with telemetry storage and analytics platforms such as Grafana Mimir, Grafana Tempo and ClickHouse.
- Support Kafka-based telemetry ingestion, event streaming and integration with distributed applications.
- Investigate missing telemetry, incomplete traces, ingestion failures, log-processing issues and monitoring inconsistencies.
- Develop operational dashboards, monitoring alerts and service-health visibility across containerised applications.
- Collaborate with development teams to integrate application observability into CI/CD and production delivery processes.
- Support performance validation, monitoring optimisation and enterprise observability adoption.
Production Reliability & Troubleshooting
- Investigate and resolve application, Kubernetes, infrastructure and observability-related production incidents.
- Perform root-cause analysis, log investigation, service-health validation and performance troubleshooting.
- Support production stabilisation, deployment verification and post-implementation activities.
- Maintain technical documentation, operational procedures, incident records and troubleshooting runbooks.
- Coordinate with development, infrastructure and support teams to resolve complex technical issues.
Essential Technical Requirements
- 3–5 years of relevant experience in DevOps Engineering, Platform Engineering, Site Reliability Engineering or related roles.
- Hands-on experience deploying and supporting containerised applications using Kubernetes or OpenShift.
- Experience developing and maintaining CI/CD pipelines.
- Practical experience with Grafana, Prometheus or comparable observability technologies.
- Working understanding of application instrumentation, metrics, log aggregation and distributed tracing.
- Experience with Docker, Helm and Linux-based application environments.
- Experience troubleshooting microservices deployments and production application issues.
- Working knowledge of Python, Bash or Node.js for scripting and automation.
- Familiarity with infrastructure automation using Terraform or Ansible.
- Strong troubleshooting, analytical skills.
Specialised Technical Experience –Advantageous
- OpenTelemetry SDKs, OpenTelemetry Collectors and OTLP-based telemetry ingestion.
- Grafana Mimir for metrics storage and Grafana Tempo for distributed tracing.
- ClickHouse for high-volume telemetry analytics.
- Apache Kafka for telemetry ingestion and event-streaming pipelines.
- Prometheus exporters, distributed trace correlation and metrics optimisation.
- Application instrumentation involving Python or Node.js microservices.
- GitOps deployment workflows using ArgoCD.
- Enterprise-scale production monitoring, alerting and reliability engineering.
Qualifications & Additional Requirements
- Bachelor's degree in Computer Science, Information Technology, Software Engineering or a related discipline, or equivalent relevant experience.
- Experience in banking, financial services or other regulated enterprise environments is advantageous.
- Relevant Kubernetes, cloud, DevOps or observability certifications are advantageous.
- Ability to work independently and collaborate effectively with technical stakeholders.
- Strong understanding of production support, incident management and change-management processes.
Interested applicants, please email your resume to Karin Chan Wei Kien
Email: [email protected]
CEI Reg No: R1104584
Recruit Express Pte Ltd
UEN: 199601303W
EA Licence No:99C4599



