Skip to content

Open nowPosted 43 days ago

Staff DevOps Engineer

Nexxa.AI18 open roles

Where
Toronto - Canada
Work mode
Remote
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowStaff DevOps EngineerNexxa.AI · Toronto - Canada
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Nexxa.AI's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

8.2% of postings close within 7 days. Measured by our own scanner across the market.

Share of postings closed within
  1. 1.8%1 day
  2. 3.8%3 days
  3. 8.2%7 days
  4. 15.2%14 days
  5. 34.2%30 days
This job: posted 43 days ago

The posting

Nexxa http://Nexxa.ai is building the best AI systems for heavy industries — enabling machines, systems, and operations to think, decide, and act autonomously across manufacturing, large-scale infrastructure, logistics, and legacy environments.

Our mission is to translate deep technical breakthroughs into operational reality, solving some of the hardest systems-level problems in industry.

ABOUT THE ROLE

We're looking for a Senior/Staff DevOps Engineer who has spent the last several years building and operating the infrastructure that lets AI and industrial systems run reliably at scale. You understand what it takes to keep production ML and data workloads fast, observable, and resilient — from GPU-backed training and inference clusters to the pipelines that connect them to real-world industrial environments.

This role is ideal for candidates who want deep infrastructure ownership at a company where uptime, latency, and reliability directly affect physical operations — not just software. You'll partner closely with AI, data, and product engineering teams to make sure the systems they build can actually run in production, safely and at scale.

WHAT YOU'LL DO

- Own and evolve Nexxa's core infrastructure — compute, networking, storage, and deployment systems — end-to-end

- Design and operate CI/CD pipelines that support fast, safe iteration across AI, data, and product engineering teams

- Build and maintain infrastructure-as-code (e.g., Terraform, Pulumi) for reproducible, auditable environments across cloud and on-prem/edge deployments

- Architect and manage Kubernetes-based platforms for training, inference, and application workloads, including GPU scheduling and autoscaling

  • Partner with data and AI teams to support the infrastructure behind:
  • Data warehouses and lakehouse architectures (e.g., Snowflake, BigQuery, Redshift, Databricks)
  • Feature stores, embedding indices, and retrieval pipelines
  • Model training, evaluation, and serving infrastructure

- Define and drive observability practices — metrics, logging, tracing, and alerting — across distributed systems

- Establish and enforce reliability practices: SLOs/SLIs, incident response, postmortems, and on-call rotations

- Design for security and compliance across cloud infrastructure, secrets management, and access control, particularly relevant to industrial and legacy-environment integrations

- Make pragmatic tradeoffs across cost, latency, reliability, and developer velocity

- Collaborate with engineering leadership to define infrastructure roadmap and platform strategy

- Mentor engineers on infrastructure best practices and raise the bar for operational excellence across the org

REQUIRED QUALIFICATIONS

- 6+ years of experience in DevOps, Site Reliability Engineering, Platform Engineering, or infrastructure-focused software engineering roles

  • Deep hands-on experience with:
  • Cloud platforms (AWS, GCP, or Azure) at production scale
  • Kubernetes in production, including GPU workload scheduling
  • Infrastructure-as-code tooling (Terraform, Pulumi, or equivalent)
  • CI/CD systems (e.g., GitHub Actions, GitLab CI, CircleCI, Jenkins, ArgoCD)

- Strong track record designing and operating observability stacks (e.g., Prometheus, Grafana, Datadog, OpenTelemetry)

- Experience supporting ML/AI infrastructure — training clusters, model serving, data pipelines — a strong plus

- Excellent scripting/programming skills (Python, Go, or Bash) for automation and tooling

- Proven ability to independently scope and lead infrastructure projects from design through production rollout

- Strong incident management instincts — you can lead through an outage calmly and drive toward root cause

PREFERRED QUALIFICATIONS

- Experience operating infrastructure that bridges cloud and edge/on-prem environments, especially in industrial or manufacturing contexts

- Familiarity with data warehouse/lakehouse platforms (Snowflake, BigQuery, Redshift, Databricks)

- Experience with service mesh, zero-trust networking, or compliance frameworks relevant to industrial/critical infrastructure (e.g., SOC 2, IEC 62443)

- History of building internal developer platforms or self-service infrastructure tooling

- Experience scaling infrastructure teams or setting technical direction at a Staff level

WHAT SUCCESS LOOKS LIKE

- You can own ambiguous, high-stakes infrastructure problems end-to-end

- Systems you build stay reliable as usage and scale grow — you design for the next order of magnitude, not just today

- You bring strong technical judgment on tradeoffs between reliability, cost, and speed

- You raise the bar for operational rigor and engineering discipline across the team

- You help define what's next for the platform, not just execute what's known

WHY JOIN NEXXA.AI http://Nexxa.ai?

- Innovative Environment: Play a critical role in transforming heavy industries through groundbreaking AI and automation technologies

- Collaborative Culture: Be part of a team that values innovation, discipline, and continuous improvement

- Professional Growth: Benefit from significant opportunities for career development and advancement

- Competitive Compensation: Enjoy a comprehensive salary and equity package reflective of your expertise and contributions

If you're passionate about building the infrastructure that powers advanced AI solutions in the real world, we'd love to connect.

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Nexxa.AI's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Nexxa.AI's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Nexxa.AI's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.