Skip to content

Open nowPosted 3 hours agoWe saw it 5 min after it went up

Principal Engineer, AI Cloud Software

Firmus Technologies69 open roles

Where
Singapore
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowPrincipal Engineer, AI Cloud SoftwareFirmus Technologies · Singapore
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Firmus Technologies's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

8.2% of postings close within 7 days. Measured by our own scanner across the market. Firmus Technologies postings stay open a median of 12 days.

Share of postings closed within
  1. 1.8%1 day
  2. 3.6%3 days
  3. 8.2%7 days
  4. 15.2%14 days
  5. 34.0%30 days
This job: posted 3 hours ago

Firmus Technologies median: 12 days open

The posting

Firmus Technologies

Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific.

Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability.

At Firmus, we are unique in our approach. We design, build, and operate a new class of digital infrastructure – the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally.

Firmus AI Cloud

Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers.

It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale.

ROLE SUMMARY

Firmus Technologies is seeking a Senior AI Infrastructure Engineer, Observability, to join our Engineering and Technology team. You will establish how we measure, validate and communicate the health of GPU infrastructure used for customer and internal workloads. You will define trusted health signals and service-readiness criteria, and turn them into reusable dashboards, alerts, queries, diagnostic checks and operational guidance. Your work will help commissioning, infrastructure and operations teams bring capacity online safely, identify degradation early and recover from failures quickly. You will also make knowledge self-service by publishing clear reference implementations, runbooks and AI-ready operational knowledge that other teams can use and extend.

KEY RESPONSIBILITIES

  • The Health Standard Define GPU and host health criteria for customer and internal workloads, and the service-readiness gates that repair and capacity workflows depend on. Cover what stops or corrupts AI jobs: GPUs falling off the bus, XID events, ECC and memory faults, NVLink/NVSwitch degradation, thermal and power capping, NCCL and collective failures, silent data corruption, and stragglers running below fleet baseline.
  • Reference Implementations Publish golden dashboards, alerts, PromQL/LogQL queries and health checks that other teams adopt and extend. Your output is the standard and the examples. The alerts must be specific, low-noise, and with a clear next action.
  • Diagnostics with Real Pass/Fail Criteria DCGM checks, NCCL and bandwidth tests, stress and burn-in, and validation jobs that confirm a server matches expected performance. Others must be able to run them without you.
  • Fault Isolation Separate a bad GPU from a cooling, host, network or power-limit problem using host and BMC/management-interface telemetry together, including when the same signature appears across many servers. A clean management view means nothing if the host is throwing faults.
  • Detection of the Failures that Don't Crash

Rising correctable ECC counts, NVLink retries, thermal slowdown, XID patterns, wrong results with no error. Keep the knowledge current: what each signal means, what to do next on the machine, and where the operation team must make the final decision.

  • Partnership and Escalation Work with commissioning on bring-up and acceptance, operations on break-fix, infrastructure on what healthy hardware looks like, and the telemetry owners on making your signals production-grade. Run technical sessions with customer teams on the telemetry they need. Join GPU and host incidents, including debugging on the server and convert every finding into a reusable check, alert or runbook. Give engineering leadership a straight read on fleet health and risk.

SKILLS AND EXPERIENCE

  • Bachelor's degree in computer science or a related technical field, or equivalent practical experience.
  • 7+ years in GPU, HPC, AI infrastructure or closely related large-scale systems engineering environments, including ownership of health monitoring or diagnostics used by customers or internal teams.
  • Experience with production GPU fault diagnosis. You've found genuine faults through XID events, ECC/memory errors, NVLink issues, power or thermal limits and caught at least one before the job died.
  • Strong Linux and server fundamentals. You can debug on the machine and reason across GPU, CPU, memory, PCIe, power and cooling. You automate in Python or similar.
  • Hands-on with GPU diagnostics and validation such as DCGM, NCCL/collective tests, stress testing and you've turned them into checks other people run.
  • Experience analysing infrastructure telemetry, building trusted dashboards and alerts. Savvy with PromQL, LogQL, Grafana or comparable tooling. Able to follow an unexpected signal until there is a cause and can turn that into an alert or check others will trust.
  • Uses AI tools as a normal part of analysis and build work. Can structure health knowledge and build AI skills so operation team with AI assistants can use it effectively to recover from incidents.
  • Willing to take part in the incident-response on-call rotation.
  • Willing to travel overseas occasionally when the role requires it.
  • Clear and effective written and verbal communication in English.

Highly Desirable Experiences

  • Production experience on large GPU systems with tight GPU-to-GPU interconnect.
  • Multiple sites or GPU generations, especially preparing health monitoring before a new GPU class entered production.
  • Background at an AI cloud, GPU cloud or HPC centre running customer workloads.
  • Experience working with data or ML engineers on GPU/host failure prediction from health signals.
  • Built a structured operational knowledge so AI assistants can use it effectively during incident recovery.

Location & Reporting

  • Location: Singapore
  • Report to: Senior Manager, Platform Engineering

Employment Basis

Full-time

Diversity

At Firmus, we are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions.

Join us in our mission to revolutionize the AI industry through sustainable practices and cutting-edge engineering. Apply now to be part of shaping the future of sustainable AI infrastructure.

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Firmus Technologies's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Firmus Technologies's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Firmus Technologies's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.