Skip to content

Open nowPosted 43 days ago

QA Engineer (AI Systems)

Nexxa.AI18 open roles

Where
Toronto - Canada
Work mode
Hybrid
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowQA Engineer (AI Systems)Nexxa.AI · Toronto - Canada
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Nexxa.AI's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

8.2% of postings close within 7 days. Measured by our own scanner across the market.

Share of postings closed within
  1. 1.8%1 day
  2. 3.8%3 days
  3. 8.2%7 days
  4. 15.2%14 days
  5. 34.2%30 days
This job: posted 43 days ago

The posting

Nexxa is building the best AI systems for heavy industries — enabling machines, systems and operations to think, decide and act autonomously across manufacturing, large-scale infrastructure, logistics and legacy environments.

Our mission is to translate deep technical breakthroughs into operational reality, solving some of the hardest systems-level problems in industry.

ROLE OVERVIEW

We're looking for a Lead / Senior / Staff QA Engineer to own quality for Nexxa's AI agent systems — products that plan, call tools, and take multi-step actions autonomously in industrial environments. This isn't traditional UI testing: you'll be designing evaluation frameworks for non-deterministic, tool-using systems, building golden datasets, catching regressions in reasoning quality, and stress-testing agent behavior under adversarial and real-world edge-case conditions.

You'll work closely with ML engineers, backend engineers, and Forward Deployed Engineers to define what "good" looks like for an agent operating in high-stakes industrial settings, then build the infrastructure and processes to measure it continuously.

KEY RESPONSIBILITIES

- Design and build evaluation harnesses and regression suites for LLM-based agents, covering reasoning quality, tool-call correctness, task completion, and multi-turn coherence.

- Develop golden datasets and labeled test sets, including edge cases, ambiguous inputs, and adversarial prompts specific to industrial and operational contexts.

- Define and track quality metrics beyond simple accuracy — groundedness, hallucination rate, task success rate, latency/cost tradeoffs, and safety violations.

- Build automated pipelines that run evals on every model, prompt, or tool-integration change, and integrate them into CI/CD.

- Conduct structured red-teaming and adversarial testing (prompt injection, jailbreaks, tool misuse, unsafe actions) in partnership with security teams.

- Test agent behavior across the full action loop — planning, tool selection, tool execution, error recovery, and final output — not just the final response.

- Investigate and triage failures where the root cause could be the model, the prompt, the tool/API, or the orchestration logic.

- Partner with ML and backend engineers to translate eval failures into actionable, reproducible bug reports.

- Establish quality bars and sign-off criteria for new agent capabilities before they reach customer environments.

- Mentor other engineers on testing strategies specific to probabilistic, LLM-driven systems.

- Advocate for testability and observability in agent architecture from day one.

QUALIFICATIONS

- 5+ years in QA/SDET roles, with demonstrated ownership of test strategy for complex systems.

- Hands-on experience testing LLM-based products, chatbots, or AI agents — you understand why traditional deterministic test assertions break down for generative systems.

- Practical experience with eval frameworks or tooling (e.g., promptfoo, DeepEval, RAGAS, LangSmith) or a track record of building your own.

- Strong scripting/programming ability (Python preferred) to build test automation, data pipelines, and eval tooling.

- Understanding of how LLM agents work: prompting, tool/function calling, context management, RAG, memory, and orchestration frameworks.

- Experience designing test data and labeled datasets, including sourcing, sampling, and managing dataset drift over time.

- Familiarity with LLM-specific failure modes: hallucination, prompt injection, context poisoning, tool misuse, goal drift, and non-determinism.

- Comfortable operating in ambiguity — defining what "correct" means for a task when there's no single right answer.

- Strong written communication skills for turning fuzzy quality signals into clear, actionable findings for engineering and product stakeholders.

PREFERRED

- Experience with human-in-the-loop evaluation workflows (labeling pipelines, inter-rater reliability, rubric design).

- Background in ML/data science sufficient to read model evals and statistical significance.

- Experience red-teaming or doing adversarial/security testing on ML systems.

- Familiarity with observability/tracing tools for LLM applications (e.g., LangSmith, Arize, Langfuse, Weights & Biases).

- Experience testing AI systems in industrial, IoT, or operational technology (OT) environments.

- Prior experience setting up eval infrastructure from scratch at a startup or fast-moving team.

- What We're Looking For A QA engineer who wants to define what quality means for autonomous, real-world AI systems.

- Someone who can build rigorous evaluation infrastructure for problems that don't have a single right answer.

- A systems thinker who enjoys turning ambiguous agent behavior into measurable, trustworthy signals.

- A strong collaborator who partners well with ML engineers, backend engineers, and Forward Deployed teams.

WHY JOIN NEXXA.AI http://Nexxa.AI?

Innovative Environment: Play a critical role in transforming heavy industries through groundbreaking AI and automation technologies.

Collaborative Culture: Be part of a team that values innovation, discipline, and continuous improvement.

Professional Growth: Benefit from significant opportunities for career development and advancement.

Competitive Compensation: Enjoy a comprehensive salary and equity package reflective of your expertise and contributions.

If you're passionate about AI quality and eager to help define what trustworthy autonomous systems look like in heavy industry, we'd love to connect.

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Nexxa.AI's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Nexxa.AI's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Nexxa.AI's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.