Skip to content

Open nowPosted today

Senior Product Software Engineer - AI Quality Engineering

Brightflag378 open roles

Where
USA - Cape Girardeau, MO
Work mode
Hybrid
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowSenior Product Software Engineer - AI Quality EngineeringBrightflag · USA - Cape Girardeau, MO
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Brightflag's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

8.1% of postings close within 7 days. Measured by our own scanner across the market. Brightflag postings stay open a median of 31 days.

Share of postings closed within
  1. 1.7%1 day
  2. 3.6%3 days
  3. 8.1%7 days
  4. 15.0%14 days
  5. 33.9%30 days
This job: posted today

Brightflag median: 31 days open

The posting

We are seeking an experienced Senior AI Quality Engineer to own and implement quality strategy, evaluation systems, and reliability practices for AI-powered product experiences in Wolters Kluwer Tax and Accounting. In this role, you will architect, build, and maintain evaluation systems that measure retrieval accuracy, citation correctness, response quality, safety, performance, and overall alignment with customer intent. You will work hands-on across test automation, AI evaluation, APIs, data pipelines, observability, and release readiness while partnering with software engineers, product managers, UX, and stakeholders to ensure our generative AI systems are trustworthy, observable, secure, and ready for production use in professional tax and accounting workflows.

This role goes beyond validating that features function as designed. You will own the practical quality engineering implementation for non-deterministic AI systems: defining how quality is measured, building automated evaluation and regression capabilities, implementing quality gates in delivery pipelines, driving release-readiness decisions, mentoring team members in AI quality practices, and establishing the standards customers can trust.

Key Responsibilities:

  • Own and implement the quality engineering approach for enterprise-scale generative AI applications, including RAG pipelines, LLM orchestration, agentic workflows, APIs, ingestion pipelines, and AI-driven user experiences.
  • Architect and implement evaluation harnesses that measure retrieval accuracy, citation correctness, answer relevance, groundedness, hallucination rate, safety, latency, cost, and end-to-end system behavior.
  • Design automated regression suites that detect quality drift across prompts, embeddings, chunking strategies, model versions, retrieval configuration, orchestration logic, and application workflows.
  • Partner with developers to validate AI system architecture, service integrations, API behavior, data ingestion, observability, resilience, security controls, and production readiness.
  • Lead root-cause analysis for AI quality failures using logs, traces, evaluation outputs, feedback signals, customer scenarios, and domain-specific acceptance criteria.
  • Define, implement, and maintain measurable quality gates for AI releases, and advise product and engineering leaders on risk, readiness, and tradeoffs across accuracy, reliability, latency, cost, and user trust.
  • Develop test automation in Python and related frameworks for APIs, services, chat workflows, data pipelines, and AI evaluation workflows within CI/CD pipelines.
  • Build and maintain observability practices and dashboards that track model behavior, retrieval performance, operational health, quality trends, and customer-impacting failure modes.
  • Contribute to responsible AI practices including bias and fairness checks, content safety validation, prompt injection testing, data privacy validation, and compliance-oriented evidence collection.
  • Collaborate with product managers, UX, subject matter experts, and stakeholders to translate customer workflows and business requirements into testable AI quality criteria.
  • Mentor engineers and quality team members in AI evaluation methods, non-deterministic testing strategies, automation design, and production-quality engineering practices.
  • Stay current with emerging AI quality, LLM evaluation, RAG assessment, and LLMOps practices through research, experimentation, and community engagement.

Key Requirements:

  • 5+ years of experience in quality engineering, SDET, software engineering, test automation, reliability engineering, or equivalent engineering roles.
  • 2+ years of hands-on experience testing, evaluating, or validating AI/ML, LLM, RAG, or other non-deterministic systems in production or production-like environments.
  • Strong proficiency with Python test automation frameworks such as Pytest, unittest, or equivalent, including experience building reusable test harnesses and regression suites.
  • Experience validating REST APIs, microservices, service integrations, data ingestion pipelines, and end-to-end web or chat-based workflows.
  • Practical understanding of RAG systems, vector search, embeddings, prompt behavior, model evaluation, groundedness, citation quality, and AI failure modes.
  • Experience defining quality metrics and success criteria where outputs are probabilistic rather than strictly deterministic.
  • Knowledge of CI/CD pipelines, Git-based workflows, automated regression testing, release gates, and modern engineering practices.
  • Experience with observability, logging, tracing, dashboards, and metrics used to monitor production behavior and diagnose system failures.
  • Ability to navigate ambiguity, decompose complex AI behaviors into testable risks, and communicate findings clearly to technical and non-technical stakeholders.
  • Strong collaboration skills in Agile environments, with demonstrated ability to influence engineering quality practices across product, design, development, and domain expert teams.
  • Commitment to responsible AI quality practices, including trust, transparency, security, privacy, fairness, and customer impact.

Desired Qualifications:

  • Experience with LLM observability, tracing, and evaluation platforms such as LangSmith, Langfuse, TruLens, Arize/Phoenix, or custom evaluation and monitoring tooling.
  • Experience testing RAG pipelines using vector databases or search platforms such as Azure AI Search, Pinecone, Weaviate, Elasticsearch, or similar technologies.
  • Experience with LLM application frameworks such as LangChain, LlamaIndex, Semantic Kernel, or custom orchestration approaches.
  • Experience with Azure OpenAI Service, Azure AI Studio, Azure cloud services, or enterprise AI deployments.
  • Experience with containerized environments and cloud-native delivery practices using Docker, Kubernetes, and related tooling.
  • Experience with performance, load, or resiliency testing tools such as Locust, k6, JMeter, or equivalent.
  • Experience with production system observability tooling such as Azure Application Insights, Grafana, Prometheus, Datadog, OpenTelemetry, or equivalent.
  • Knowledge of LLMOps, model monitoring, A/B testing, human-in-the-loop evaluation, or production feedback-loop design.
  • Experience validating AI applications that handle sensitive, regulated, or professional-domain data.
  • Background in tax, accounting, legal, financial, or other domains where answer accuracy and explainability are critical.

Our Interview Practices

To maintain a fair and genuine hiring process, we kindly ask that all candidates participate in interviews without the assistance of AI tools or external prompts. Our interview process is designed to assess your individual skills, experiences, and communication style. We value authenticity and want to ensure we’re getting to know you—not a digital assistant. To help maintain this integrity, we ask to remove virtual backgrounds and include in-person interviews in our hiring process. Please note that use of AI-generated responses or third-party support during interviews will be grounds for disqualification from the recruitment process.

Applicants may be required to appear onsite at a Wolters Kluwer office as part of the recruitment process.

Compensation:

$83,400.00 - $145,650.00 USD

This role is eligible for Bonus.

Compensation range listed is based on primary location of the position. Actual base salary offer is influenced by a wide array of factors including but not limited to skills, experience and actual hiring location. Your recruiter can share more information about the specific offer for the job location during the hiring process.

Additional Information:

Wolters Kluwer offers a wide variety of competitive benefits and programs to help meet your needs and balance your work and personal life, including but not limited to: Medical, Dental, & Vision Plans, 401(k), FSA/HSA, Commuter Benefits, Tuition Assistance Plan, Vacation and Sick Time, and Paid Parental Leave. Full details of our benefits are available upon request.

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Brightflag's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Brightflag's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Brightflag's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.