Skip to content

Open nowPosted 23 days ago

TECHNICAL MANAGER-- GPU CLOUD & AI INFRASTRUCTURE

MyCareersFuture94,028 open roles

Pay
SGD 15,000 – SGD 20,000 a Monthly
Where
Singapore
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowTECHNICAL MANAGER-- GPU CLOUD & AI INFRASTRUCTUREMyCareersFuture · Singapore
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on MyCareersFuture's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

7.8% of postings close within 7 days. Measured by our own scanner across the market.

Share of postings closed within
  1. 1.6%1 day
  2. 3.4%3 days
  3. 7.8%7 days
  4. 14.3%14 days
  5. 33.7%30 days
This job: posted 23 days ago

The posting

TECHNICAL MANAGER - GPU CLOUD & AI INFRASTRUCTURE

Location: Singapore Employment Type: Full-Time, Permanent Monthly Salary: S$15,000-S$20,000 Reporting To: Head of GPU Cloud and AI Infrastructure Travel Requirement: Regional travel within Southeast Asia

ROLE OVERVIEW

We are seeking an experienced Technical Manager to lead the architecture, development, deployment and technical operation of our GPU Cloud and AI infrastructure platform.

Based in Singapore, the successful candidate will serve as the technical lead for GPU infrastructure architecture, AI computing clusters, cloud-platform development, customer solution design, project delivery and platform operations.

The role will connect GPU servers, data center infrastructure, high-speed networking, parallel storage, cloud-native platforms and AI model workloads to create a stable, scalable, schedulable, measurable and commercially viable GPU Cloud service.

This is a technical leadership position. Candidates must possess hands-on experience in GPU clusters, AI infrastructure, high-performance computing or large-scale parallel-computing platforms.

KEY RESPONSIBILITIES

1. GPU Infrastructure Architecture

• Lead the overall technical architecture and solution design of GPU servers and AI computing clusters.

• Evaluate high-performance GPU platforms, server configurations, CPUs, GPUs, memory, high-speed local storage, network interfaces and rack configurations.

• Design or review InfiniBand, RoCE and high-speed Ethernet network architectures.

• Evaluate NV Link, NV Switch, RDMA and multi-node, multi-GPU communication solutions.

• Work with data center teams to confirm rack power density, liquid or air-cooling requirements, PDU allocation, network connectivity and cabling specifications.

• Establish technical standards for GPU cluster deployment, performance testing, production acceptance and capacity expansion.

2. GPU Cloud Platform Development

• Lead the technical planning, implementation and operation of GPU resource pools, computing clusters and GPU Cloud platforms.

• Build and maintain GPU resource scheduling and management platforms using Kubernetes and/or Slurm.

• Manage GPU drivers, acceleration libraries, container runtimes, NVIDIA GPU Operator and related infrastructure software stacks.

• Design multiple GPU service models, including bare metal, virtual machines, containers and GPU instances.

• Implement full-GPU allocation, MIG partitioning, GPU sharing, quota management and multi-tenant resource isolation.

• Support user management, access control, resource provisioning, order activation, resource recovery and API integration.

• Work with product and commercial teams to develop standardized and scalable GPU Cloud services.

3. AI Workload Support

• Support customers with large language model training, fine-tuning, inference deployment and AI workload optimization.

• Analyze customer requirements relating to model size, dataset size, concurrency, throughput and latency.

• Recommend appropriate GPU models, node quantities, network topologies and storage configurations.

• Support distributed training environments using PyTorch, TensorFlow an dother mainstream AI frameworks.

• Deploy and optimise inference solutions using vLLM, NVIDIA Triton, TensorRT-LLM or similar AI inference frameworks.

• Lead customer proof-of-concept projects, performance benchmarks, technical testing and acceptance processes.

4. Platform Operations and SLA Management

• Establish monitoring and alerting systems covering GPUs, computing nodes, networks and storage using technologies such as DCGM, Prometheus and Grafana.

• Monitor GPU utilization, GPU memory usage, power consumption, temperature errors and idle capacity.

• Develop accurate usage-metering logic based on GPU hours, instance hours or Token consumption.

• Establish service-level agreements, incident-classification standards, response procedures, escalation mechanisms and disaster-recovery plans.

• Continuously improve platform availability, GPU utilisation and commercial output per GPU.

5. Customer Solutions and Pre-Sales Support

• Work with sales and commercial teams to understand customers’ AI and infrastructure requirements.

• Translate customer requirements into solution architectures, technical proposals, RFP or RFQ responses, bills of materials and implementation plans.

• Lead technical presentations, architecture discussions, proof-of-concept validation, customer onboarding and capacity-expansion activities.

• Define service boundaries, delivery standards, SLA commitments and technical acceptance criteria.

6. Project and Vendor Management

• Manage server OEMs, data center providers, network and storage vendors, and system integrators.

• Oversee equipment delivery, rack installation, cabling, environment initialization, cluster commissioning and final acceptance.

• Manage project schedules, technical risks, issue lists and remediation activities.

• Establish technical documentation, operational procedures, deployment standards and maintenance manuals.

• Recruit and manage infrastructure and operations engineers as the business expands.

REQUIREMENTS

• Bachelor’s degree in Computer Science, Electronic Engineering, Telecommunications, Software Engineering or a related discipline.

• At least 7 years of relevant experience in cloud computing, high-performance computing, AI infrastructure or data center environments.

• At least 3 years of hands-on experience building or operating GPU clusters, AI Cloud platforms or large-scale parallel-computing environments.

• Strong practical knowledge of Linux, Docker, Kubernetes, Helm and cloud-native technologies.

• Strong understanding of GPU hardware architecture, GPU drivers, acceleration libraries and AI server software stacks.

• Hands-on experience with Kubernetes GPU scheduling and/or Slur cluster-resource management.

• Good understanding of InfiniBand, RoCE, RDMA, NV Link, NV Switch and high-speed networking technologies.

• Practical experience supporting large language model training, fine-tuning or inference workloads.

• Experience with platform monitoring, incident management, capacity planning and SLA management.

• Strong customer communication, cross-functional collaboration, technical documentation and vendor-management capabilities.

• Excellent written and spoken English, as English will be used for regional operations, technical documentation, customer communication and vendor management.

• Professional proficiency in Mandarin is required for regular technical coordination with China-based engineering teams, technology suppliers and Mandarin-speaking customers.

• Willing and able to travel within Southeast Asia when required.

PREFERRED QUALIFICATIONS

• Experience building or managing NVIDIA HGX-based or other enterprise-scale AI computing clusters.

• Experience with Slurm, HPC or large-scale distributed training platforms.

• Experience implementing GPU virtualization, MIG partitioning, GPU sharing or usage-metering systems.

• Hands-on deployment experience with vLLM, NVIDIA Triton, TensorRT-LLM or similar AI inference frameworks.

• Experience building GPU-hour, instance-hour or Token-based metering and billing systems.

• Previous experience with a hyperscale, AI Cloud provider, GPU computing service provider, data center operator or major server OEM.

• Relevant professional certifications such as CKA, CKAD, Linux, cloud computing or networking certifications.

KEY PERFORMANCE INDICATORS

• On-time and successful launch of GPU clusters and GPU Cloud platforms.

• Platform availability and customer SLA achievement.

• Average GPU utilization and idle-capacity management.

• Customer PoC conversion and technical acceptance success rate.

• Mean Time to Recovery for critical incidents.

• Improvement in billable GPU usage and commercial output per GPU.

• Quality and completeness of technical documentation, operating procedures and delivery standards.

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against MyCareersFuture's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on MyCareersFuture's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    MyCareersFuture's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.