The posting
Job Description
We are seeking a hands-on Production Support to join a growing technology team supporting critical front-office business applications. This role is ideal for someone who enjoys troubleshooting complex issues, investigating root causes, driving resolutions, and improving platform reliability through automation and operational excellence.
This is a business-facing role that combines production support responsibilities with Site Reliability Engineering (SRE) practices in a modern cloud-based environment. The successful candidate will work closely with business users, application owners, and technology teams to ensure the smooth operation of key investment applications.
Key Responsibilities
- Provide production support for critical business applications, ensuring high availability and operational stability.
- Investigate and resolve production incidents, performing detailed root cause analysis before escalating where necessary.
- Monitor applications, infrastructure, and scheduled jobs, proactively identifying and addressing potential issues.
- Troubleshoot issues across cloud, Kubernetes, Linux, databases, and application layers.
- Support and manage batch scheduling platforms such as Control-M or Autosys.
- Communicate effectively with business users during incidents, providing timely updates and managing expectations.
- Drive continuous improvement initiatives by automating repetitive operational tasks and enhancing support processes.
- Partner with technology teams and system owners to improve platform reliability, resiliency, and operational efficiency.
- Participate in an on-call and weekend support rotation as required.
- Contribute to SRE initiatives, including monitoring, observability, automation, and operational best practices.
Requirements
- 5+ years of experience in Production Support, Application Support, or SRE environment.
- Strong incident management, problem management, and troubleshooting capabilities.
- Proven ability to independently investigate issues using logs, monitoring tools, database queries, and system diagnostics.
- Hands-on experience with Linux/Unix environments and Bash/Shell scripting.
- Experience supporting Kubernetes-based applications, including troubleshooting pods and containerized workloads.
- Knowledge of cloud platforms such as AWS and/or Azure.
- Experience with Control-M, Autosys, or similar enterprise scheduling tools.
- Strong SQL and database troubleshooting skills.
- Excellent communication and stakeholder management skills, with experience supporting business users.
- Ability to perform effectively in a fast-paced environment and take ownership of issues through to resolution.
Preferred Skills
- Experience with monitoring and observability tools such as Datadog.
- Infrastructure-as-Code experience using Terraform.
- Automation experience using Python or similar scripting languages.
- Familiarity with AI-assisted operational tooling or agentic AI concepts.
- Exposure to financial services or investment technology environments.
Ideal Candidate
- Currently working in a hands-on Production Support or SRE role.
- Demonstrates curiosity and persistence when troubleshooting complex issues.
- Takes ownership and accountability rather than acting as a ticket dispatcher.
- Comfortable interfacing directly with business stakeholders.
- Proactive, adaptable, and eager to learn new technologies.
- Balances operational support responsibilities with a continuous improvement mindset.
We regret to inform you that only shortlisted candidates will be notified.
EA Reference: CHONG LI NING PHOEBE, R25157679
Allegis Group Singapore Pte Ltd, Company Reg No. 200909448N, EA License No. 10C4544



