Skip to content

Open nowPosted 2 days ago

Network Engineer, AI Cluster Commissioning

Firmus Technologies63 open roles

Where
Melbourne, Victoria, Australia
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowNetwork Engineer, AI Cluster CommissioningFirmus Technologies · Melbourne, Victoria, Australia
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Firmus Technologies's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

7.9% of postings close within 7 days. Measured by our own scanner across the market. Firmus Technologies postings stay open a median of 11 days.

Share of postings closed within
  1. 1.6%1 day
  2. 3.4%3 days
  3. 7.9%7 days
  4. 14.2%14 days
  5. 34.1%30 days
This job: posted 2 days ago

Firmus Technologies median: 11 days open

The posting

Infrastructure Automation Engineer, AI Cluster Commissioning

Firmus Technologies

Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific.

Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability.

At Firmus, we are unique in our approach. We design, build, and operate a new class of digital infrastructure – the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally.

Firmus AI Cloud

Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers.

It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale.

Why Firmus?

As an NVIDIA Cloud and Engineering partner in Asia Pacific, you will gain skills, experience, and exposure across the AI industry and be part of shaping what this industry looks like for decades to come.

We are founder-led, not a big corporate. Decisions happen fast, our leaders are accessible, and there's minimum bureaucracy between you and the work. Ownership comes early. Whatever your role, you will have a direct line to outcomes, helping shape how the business grows as we scale nationally across a long-term, large-scale roadmap.

Work alongside founders and experts in AI infrastructure, energy systems and next-generation compute.

What we build here has impact beyond the business. Our AI Factories are designed to operate as assets to the energy grid to actively strengthen the communities and regions they operate in rather than drawing from them.

Considering applying? You don't need a perfect background to join our team. If you're driven and curious, there's a path for you. We back our people to grow into new domains and take on challenges beyond their previous experience.

Role Summary

Firmus Technologies is seeking a skilled Network Engineer to join our Commissioning team. This position will play a crucial role in the deployment, commissioning, and configuration of our network designs for AI infrastructure projects. This role offers an exciting opportunity to work at the forefront of AI networking technology and contribute to the growth of AI infrastructure.

Key Responsibilities

Deploy and Commission High-Performance Networks

  • Deploy, test, and commission low-latency, high-throughput interconnects (e.g.: Ethernet 100/200/400/800 GbE) for AI workloads.
  • Configure, test, diagnose, remediate, and benchmark performance across multi-node clusters, high-speed storage fabrics and parallel computing AI environments.
  • Work with partners and vendors to respond to and resolve network issues identified during the bring-up and commissioning of large-scale AI platforms.
  • Analyse logs, run diagnostics and coordinate with internal teams, partners, and vendors as required.
  • Develop and maintain monitoring tools to proactively identify bottlenecks, errors and abnormal behaviours.
  • Test configurations, performance, redundant paths and other aspects as set out in the commissioning test plan and acceptance tests.

Ethernet and RDMA Networking

  • Design, monitor and troubleshoot Ethernet fabrics and RDMA-enabled transport layers.
  • Maintain network configurations and ensure fabric health and topology visibility using vendor and open source tools.
  • Perform verification and acceptance tests for the new network fabric.
  • Understand various RoCE optimisation protocols and mechanisms (e.g.: SHARP, CollNet) and use performance tools (e.g.: ibstat, perfquery, ib_write_bw, nccl, etc) to monitor fabric health, congestion and link errors.
  • Enable and test optimised networking for AI frameworks and work closely with partners and vendors to ensureefficient multi-node communication.
  • Perform deep dive diagnostics to resolve layer 1-4 issues across HPC and AI workloads.

Network as Code and Automation

  • Develop and maintain automated network configurations using Infrastructure as Code (IaC) tools (e.g.: Ansible, Netbox, bash and Python scripts).
  • Implement CI/CD pipelines for network changes to improve speed, consistency, and auditability.
  • Automate routine tasks such as provisioning, backups and compliance checks.

Project Management and Stakeholder Management

  • Support the deployment team in their project management and resource allocation for the network portion of AI cluster installations.
  • Collaborate and work closely with the Global Operations Centre, Software Defined Infrastructure team, Data Centre Infrastructure team and Solution Architects to integrate new deployments.
  • Work closely with both the Firmus Engineering and Operations teams to align network infrastructure with customers’ requirements.
  • Facilitate knowledge sharing and communication between teams and create and maintain comprehensive technical documentation.
  • Maintain and build strong relationships with key technology partners and vendors and proactively manage and coordinate partner engagement on site.

Network Security

  • Implement authentication and access control for fabric and out-of-band management networks.
  • Implement secure configurations for RDMA/RoCEv2 fabrics, including multi-tenant workload isolation.
  • Support vulnerability management: coordinate scanning, patching, and remediation tracking for network infrastructure.
  • Collaborate with Security and Risk team to enforce policies and respond to security incidents.
  • Design and implement zero-trust network architecture across corporate and AI compute fabric, including segmentation and least-privilege access.
  • Participate in security incident response — detection, containment, root cause analysis, and post-incident reporting — with the Security and Risk team.
  • Contribute to compliance and audit activity (e.g. ISO 27001, SOC 2) relating to network controls.
  • Integrate network telemetry and logs with SIEM/observability tooling for security monitoring.

Technology Expertise

  • Physical network hardware and advanced networking technologies, including NVIDIA Spectrum Ethernet Platform, RDMA over Converged Ethernet (RoCE), DPU/SmartNICs.
  • Familiarity with open-source network operating systems such as Cumulus Linux and Sonic as well as network simulation environments like NVIDIA Air.
  • Testing and benchmark tools and processes to validate network topologies, network performance, and acceptance tests.
  • Provide technical support and troubleshooting for advanced networking technologies, escalating to vendors as needed.

Skills & Experience

  • Bachelor’s degree in network engineering, computer science, or a related technical field.
  • 5+ years of experience in network engineering.
  • Experience in Linux systems especially host network configuration.
  • Experience with high performance Ethernet networks, IPv4, IPv6, BGP, RoCE.
  • Strong project management skills and experienced in complex technical projects.
  • Excellent problem-solving and analytical skills.
  • Ability to work independently and as part of a team.
  • Strong communication skills, both written and verbal.
  • Willingness to undertake international and/or domestic travel for on-site deployments and commissioning as required.
  • Solid understanding of advanced networking technologies, particularly those related to AI would be highly advantageous.
  • Hands-on experience with NVIDIA Spectrum Ethernet Platform and RDMA over Converged Ethernet (RoCE) preferred.
  • Willingness to undertake international and/or domestic travel for on-site deployments and commissioning as required.

Highly Desirable Experiences

  • NVIDIA Infiniband networking technology, subnet manager configuration, UFM, multi-tenancy configurations.
  • Network security, firewall configuration and management, network segmentation, secure network design, zerotrust architecture
  • Authentication and acess controls

Location & Reporting

  • This role is based in Australia or Singapore with regular visits to current and future project sites in Australia and SE Asia.
  • Report to: Head of AI Cluster Commissioning
  • Employment Basis: Full-time
From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Firmus Technologies's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Firmus Technologies's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Firmus Technologies's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.