Skip to content

Open nowPosted 22 days ago

Site Reliability Engineer (SRE) / Production Support Engineer

MyCareersFuture94,028 open roles

Pay
SGD 8,000 – SGD 10,000 a month
Where
Central, Singapore
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowSite Reliability Engineer (SRE) / Production Support EngineerMyCareersFuture · Central, Singapore
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on MyCareersFuture's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

7.7% of postings close within 7 days. Measured by our own scanner across the market.

Share of postings closed within
  1. 1.6%1 day
  2. 3.3%3 days
  3. 7.7%7 days
  4. 14.0%14 days
  5. 33.7%30 days
This job: posted 22 days ago

The posting

Job Summary

We are seeking an experienced Site Reliability Engineer (SRE) / Production Support Engineer with 10+ years of experience in enterprise infrastructure, production support, system administration, and SRE operations. The ideal candidate will have strong hands-on experience in AIX, Linux/RHEL, application production support, monitoring, automation, incident management, and cloud/DevOps environments.

The candidate will be responsible for ensuring high availability, reliability, performance, and security of critical enterprise applications and infrastructure in a 24x7 production support environment. The role requires close collaboration with development, infrastructure, middleware, security, and business teams to identify and resolve production issues and continuously improve service reliability.

Key Responsibilities

Site Reliability Engineering

  • Define and establish SLIs, SLOs, SLAs, error budgets, MTTD, and MTTR for enterprise applications and services.
  • Monitor application availability, latency, performance, capacity, and overall reliability.
  • Identify and implement Golden Signals and observability best practices across applications and infrastructure.
  • Analyze production incidents and recurring issues to identify root causes and implement permanent remediation.
  • Develop automation and engineering solutions to reduce manual operational activities and improve reliability.
  • Participate in 24x7x365 production support and on-call operations.
  • Perform capacity planning and proactively identify infrastructure and application resource requirements.
  • Support highly available and resilient application architectures.
  • Participate in disaster recovery planning, testing, and implementation.

Production & Infrastructure Support

  • Provide L2/L3 production and infrastructure support for critical enterprise applications.
  • Perform administration, troubleshooting, configuration, maintenance, and performance tuning of IBM AIX and RHEL/Linux servers.
  • Support server patching, upgrades, maintenance, and DR activities.
  • Troubleshoot application, middleware, operating system, connectivity, and infrastructure-related issues.
  • Perform middleware administration and restart activities for WebSphere, JBoss, IBM HTTP Server, and Apache Tomcat.
  • Support firewall changes, SSL certificate renewals, SCP configuration, and system-to-system key exchanges.
  • Support application deployment activities across Blue/Green environments.
  • Coordinate application-related changes and deployments across development, infrastructure, middleware, and business teams.

Monitoring & Observability

  • Configure and maintain monitoring and alerting for infrastructure and application services.
  • Work with monitoring and observability platforms such as:
  • Grafana
  • Centreon
  • AppDynamics
  • ELK / Kibana
  • Application and infrastructure logging platforms
  • Analyze system and application metrics, logs, and performance trends.
  • Develop appropriate alerts to proactively identify service degradation and failures.
  • Promote observability practices and help development teams implement effective monitoring.

Incident & Change Management

  • Manage and resolve incidents within defined SLA/OLA timelines.
  • Participate in major incident and emergency response activities.
  • Perform incident investigation, troubleshooting, root-cause analysis, and problem management.
  • Raise and manage Change Requests (CRs) for application deployments, BAU fixes, infrastructure changes, middleware changes, and maintenance activities.
  • Create and manage service requests and incident tickets.
  • Prepare RCA and corrective/preventive action plans for recurring and high-priority incidents.
  • Follow ITIL-based Incident, Change, Problem, and Service Request Management processes.

Automation & DevOps

  • Develop automation scripts using Python and Shell scripting to improve operational efficiency.
  • Automate repetitive infrastructure and production support activities.
  • Work with DevOps tools and technologies including:
  • Git / GitHub / Bitbucket
  • Jenkins
  • Ansible
  • Chef
  • Docker
  • Maven
  • JFrog / Nexus Repository
  • SonarQube
  • Fortify / Nexus IQ
  • Support CI/CD pipelines and application deployment processes.
  • Collaborate with development teams to integrate monitoring, logging, and reliability controls into deployment pipelines.

Cloud & Application Support

  • Provide support for applications hosted on AWS and Pivotal Cloud Foundry (PCF) environments.
  • Work with enterprise application technologies including:
  • IBM WebSphere
  • IBM HTTP Server
  • JBoss
  • Apache Tomcat
  • Java applications
  • Support microservices-based applications and their associated infrastructure.
  • Assist with application migration, optimization, and adoption of new technologies where required.

Security & Compliance

  • Apply system security best practices across AIX and Linux environments.
  • Support security hardening, patching, access management, and vulnerability remediation.
  • Work with security and access management tools such as IBM Tivoli Access Manager.
  • Support SSL certificate management and renewal activities.
  • Prepare quarterly operational reports and documentation required for audit, risk, and compliance activities.
  • Ensure production changes and operational activities comply with organizational security and governance standards.

Required Technical Skills

Operating Systems

  • IBM AIX 5.x / 6.x / 7.x
  • Red Hat Enterprise Linux (RHEL)
  • Linux
  • Windows Server

SRE / Monitoring

  • SLI / SLO / SLA
  • Error Budgets
  • MTTD / MTTR
  • Observability
  • Golden Signals
  • Grafana
  • Centreon
  • AppDynamics
  • ELK / Kibana

Cloud & Containers

  • AWS
  • Pivotal Cloud Foundry (PCF)
  • Docker

Middleware & Application Technologies

  • IBM WebSphere Application Server
  • IBM HTTP Server
  • JBoss
  • Apache Tomcat
  • Java

Databases

  • DB2 UDB
  • SQL
  • MySQL / MariaDB

Automation & Scripting

  • Python
  • Shell Scripting
  • Groovy

DevOps / CI-CD

  • Git / GitHub
  • Bitbucket
  • Jenkins
  • Ansible
  • Chef
  • Maven
  • JFrog / Nexus Repository
  • SonarQube
  • Fortify
  • Nexus IQ

ITSM / Ticketing

  • ServiceNow
  • BMC Remedy
  • IBM ISM
  • JIRA

Soft Skills

  • Strong analytical and troubleshooting skills.
  • Excellent incident management and problem-solving capabilities.
  • Ability to work effectively in a high-pressure 24x7 production environment.
  • Strong communication and stakeholder management skills.
  • Ability to collaborate effectively with Development, Infrastructure, Middleware, Security, and Business teams.
  • Strong ownership and accountability for production services.
  • Ability to prioritize critical issues and coordinate resolution during major incidents.

Education

  • Bachelor's degree in engineering / technology / computer science or equivalent.
  • B.Tech / B.E. preferred.

Experience

  • 10+ years of overall IT experience in SRE, Production Support, Infrastructure Support, or System Administration.
  • Strong hands-on experience with AIX and Linux/RHEL administration.
  • Experience supporting enterprise applications in banking, financial services, or other mission-critical environments is preferred.
  • Experience working in 24x7x365 production support and on-call environments.

Preferred Profile

The ideal candidate will be a hands-on SRE/Production Support professional with strong expertise in AIX/Linux administration, enterprise application support, observability, incident management, automation, DevOps, and cloud technologies, with a proven ability to improve application reliability, reduce operational effort, and maintain high service availability.

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against MyCareersFuture's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on MyCareersFuture's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    MyCareersFuture's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.