Skip to content

Open nowPosted 11 hours agoWe saw it 39 min after it went up

QA Engineer-AI Native Quality

Newton Research7 open roles

Where
Greater Boston Area
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowQA Engineer-AI Native QualityNewton Research · Greater Boston Area
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Newton Research's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

7.8% of postings close within 7 days. Measured by our own scanner across the market. Newton Research postings stay open a median of 38 days.

Share of postings closed within
  1. 1.7%1 day
  2. 3.5%3 days
  3. 7.8%7 days
  4. 14.6%14 days
  5. 34.1%30 days
This job: posted 11 hours ago

Newton Research median: 38 days open

The posting

QA Engineer, AI-Native Quality Newton Research · Research & Development · Boston / Needham, MA

Company Description Newton Research is a fast-growing software start-up founded by repeat entrepreneurs and well-funded by blue chip venture capital firms. We are building the next generation of the closed loop media lifecycle, developing AI agents that leverage the latest in LLMs and generative AI with specialized knowledge. Our products generate actionable business insights for our customers and partners, assisting in each step of the media planning, buying and measurement lifecycle.

About the Role Newton ships on a sprint cadence through a develop, stage and customer-environment pipeline, and the product surface is wide: conversations, blueprints, connectors, scheduled tasks, permissions and sharing, SSO, and AI agents whose behavior is not fully deterministic. A missed regression lands in front of a media planner or a customer's security review. We run everything through an AI-first lens, because it is the only way quality scales. If a quality task is repeatable, an agent does it and you supervise; if it takes judgment, that is where you spend your time. The gap this hire fills: evals and skill-change testing. Code changes already have CI and review, including PRs written by Claude. What has no safety net is behavior change: an edit to a skill, a prompt, a tool definition or a model version can silently change what our agents do, and nothing tests that today. You own that layer: curated eval sets, scoring and regression tracking, so any change to how an agent behaves is measured before it ships.

  • Measure the AI: agent output varies run to run, so “correct” is a range you define with evals, rubrics and scoring, not an exact-match assertion.
  • Let agents run the checks: agents run suites, triage failures and draft repro-ready defects; you design their roles and guardrails and review what they produce.
  • Make Newton verifiable by agents: you keep the product and pipeline observable and fixture-friendly so agents can verify it without a human in the loop.

You are also a release-readiness partner (are we good?), alongside our existing QA lead, and that judgment stays human. But you are not hired to test Claude-driven PRs line by line.

What You Will Do

  • Own the eval suite for Newton's agents: curated datasets of inputs and expected behaviors, rubric and LLM-as-judge scoring, regression tracking across model, prompt, skill and agent releases; validate judges against human-labeled examples
  • Test skill and prompt changes before they ship: every edit to a skill, system prompt or tool definition runs against the relevant evals in CI, with a before/after comparison a reviewer can read in one glance; gate releases on the results
  • Cover the agent flows end to end: tool-call correctness, task completion, multi-turn coherence, blueprint creation from conversations, code generation, scheduled tasks; flag output that is wrong, empty or silently degraded
  • Handle non-determinism with rigor: repeated runs, pass-rate thresholds and simple statistics, so a flaky agent is a measured finding, not an anecdote
  • Turn production and customer signal into evals: mine logs, error tracking and customer reports so every escaped bad-behavior case becomes a permanent eval, ideally drafted by an agent and reviewed by you
  • Probe AI-specific risk: prompt injection, data leakage across users, projects and permissions, hallucinated or ungrounded numbers in analytics output, cost and latency regressions
  • Automate before you repeat: any check you do by hand twice becomes a Playwright test, an eval or an agent workflow; drive manual regression time down every sprint
  • Design agentic QA workflows in CI: agents that run suites, triage failures, draft defects and re-verify fixes, with guardrails, cost limits and escalation rules you define
  • Keep human-judgment work sharp: exploratory testing across roles, feature flags and environments, SSO and connector authorization, and the cross-feature bugs only a person thinks to look for
  • Write bug reports that are machine- and human-readable, and partner with engineering on testability (observability, seedable data, stable interfaces)

What Makes You a Great Fit

  • 4+ years in QA or test engineering on a complex web product, ideally B2B SaaS shipping frequently
  • Hands-on experience building evals for LLM or agent products (datasets, rubrics, LLM-as-judge, regression tracking), or a clear track record of getting there fast
  • Working knowledge of how agents work: prompting, skills, tool calling, context, RAG and orchestration, well enough to tell where a failure originates
  • An automation-first instinct: you can show what you removed from a manual process
  • Python (or similar) for eval tooling and data checks; Playwright or similar for end-to-end tests
  • Statistical literacy: pass rates, variance and sample size when outputs are not deterministic
  • Daily user of AI coding and testing assistants, with the judgment to verify their output rather than trust it
  • Strong exploratory instincts and excellent written communication
  • Nice to have: adtech, martech or marketing analytics domain; SSO / SAML / OAuth flows; data-connector testing; eval or observability tooling

How Success Is Measured

  • Eval suite covering Newton's core agent flows, run on every model, prompt, skill or agent change, with judge accuracy checked against human labels
  • Share of skill and prompt changes that ship with a before/after eval result (target: all of them)
  • Escaped bad-behavior cases converted into permanent evals
  • Share of regression coverage run automatically or by agents, rising every sprint
  • Release calls that hold up: few surprises after release

Salary range: $115,000-130,000 + Equity

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Newton Research's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Newton Research's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Newton Research's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.