Skip to content

Open nowPosted 13 days ago

System Engineer

MyCareersFuture94,028 open roles

Pay
SGD 5,000 – SGD 7,500 a month
Where
Central, Singapore
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowSystem EngineerMyCareersFuture · Central, Singapore
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on MyCareersFuture's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

7.7% of postings close within 7 days. Measured by our own scanner across the market.

Share of postings closed within
  1. 1.6%1 day
  2. 3.3%3 days
  3. 7.7%7 days
  4. 14.0%14 days
  5. 33.7%30 days
This job: posted 13 days ago

The posting

We are seeking an experienced AI System Engineer to design, deploy, operate, and optimize AI training clusters, GPU computing platforms, and supporting infrastructure. The ideal candidate should possess strong expertise in Linux systems, GPU computing environments, container platforms, and AI/HPC cluster architectures.

Key Responsibilities

AI Cluster deployment and operation

· Deploy and operate AI training and HPC clusters;

· Install, configure, and optimize operating systems on GPU servers;

· Manage cluster resources and capacity;

· Perform system upgrades, patch management, and change implementation;

· Develop and maintain standardized operational procedures.

Linux System Management

· Manage large-scale Linux environments;

· Perform system performance tuning;

· Analyze system logs and kernel issues;

· Troubleshoot system stability problems;

· Manage user access and security policies.

GPU Platform Support

· Manage NVIDIA GPU computing platforms;

· Deploy and maintain CUDA, NVIDIA Drivers, and Fabric Manager;

· Troubleshoot GPU, NV Link, and NV Switch-related issues;

· Optimize GPU cluster performance;

· Support customer in resolving training environment issues.

Container Platform & Orchestration System

· Build and maintain Kubernetes clusters;

· Support AI workload scheduling;

· Deploy and manage container runtime environments;

· Optimize GPU utilization within containers;

· Manage Kubernetes high-availability architectures.

AI infrastructure management

· Manage distributed storage platforms;

· Operate high-speed networking environments;

· Collaborate with datacenter teams for troubleshooting;

· Monitor infrastructure health and performance;

· Improve system reliability and availability.

Automated operation and maintenance and platform development

· Develop infrastructure automation tools;

· Create deployment and health-check scripts;

· Build monitoring and observability platforms;

· Implement alerting and self-healing mechanisms;

· Improve operational efficiency through automation.

  • Good communication, teamwork, and ownership mindset.
  • Willing to participate in on-call rotation, maintenance windows, and emergency incident response, willing to accept short-term business trips.

Required Qualifications

  • Bachelor's degree or above in Computer Engineering, Electrical Engineering, Telecommunications, or related fields.
  • Linux System · 3+years of Linux administration experience; · Strong knowledge of Ubuntu, Rocky Linux, and RHEL; · Familiarity with system boot process, kernel, filesystems, and performance tuning; · Ability to troubleshoot complex system issues independently.
  • GPU & AI Platform Strong understanding of NVIDIA GPU architecture; Experience with NVIDIA GPU products H100, H200, B200 B300 ,GB200 NVL72 and GB300 NVL72
  • Familiar with CUDA, NCCL, NV Link, NV Switch, GPU Direct RDMA
  • Understanding of distributed AI training architectures.

Container and Cloud Native

· Hands-on experience with Kubernetes;

· Familiarity with Docker and Containerd;

· Experience with Helm;

· Knowledge of GPU Operator;

· Understanding of Kubernetes GPU scheduling.

Networking and Storage

· Strong understanding of TCP/IP networking;

· Experience with InfiniBand and RoCE;

· Familiarity with RDMA architectures;

· Experience with one or more storage systems Lustre, BeeGFS, Ceph

and NFS

Automation capabilities

· Strong scripting skills in Shell and Python;

· Know about Ansible;

· Ability to build infrastructure automation scripts.

Preferred Qualities

· Experience supporting Large Language Model(LLM) training platforms;

· Knowledge of Slurm workload manager;

· Experience with Ray and Kubeflow;

· Familiarity with NVIDIA Base Command Manager(BCM);

· Experience with NVIDIA NIM;

· Knowledge of PXE deployment solutions;

· Experience on using DDN product;

· Experience operating large-scale GPUclusters;

· Experience supporting global datacenteroperations.

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against MyCareersFuture's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on MyCareersFuture's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    MyCareersFuture's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.