Skip to content

Open nowPosted 11 hours ago

Senior Data Scientist AI Evaluation

Jobgether4,188 open roles

Where
Canada
Work mode
Remote
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowSenior Data Scientist AI EvaluationJobgether · Canada
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Jobgether's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

8.2% of postings close within 7 days. Measured by our own scanner across the market. Jobgether postings stay open a median of 6 days.

Share of postings closed within
  1. 1.8%1 day
  2. 3.8%3 days
  3. 8.2%7 days
  4. 15.2%14 days
  5. 34.2%30 days
This job: posted 11 hours ago

Jobgether median: 6 days open

The posting

This position is listed on behalf of a partner company, which manages all applications and next steps. Our partner is looking for a Senior Data Scientist AI Evaluation based in Canada.

We are seeking a senior-level data scientist to establish and advance the evaluation practices that measure the quality, reliability, and correctness of AI models and intelligent agents. In this role, you will transform complex quality questions into measurable standards, trusted ground truth, and rigorous evaluation frameworks. You will design repeatable testing processes that identify performance regressions and support safer, more confident AI releases. Working closely with Product, Engineering, Analytics Engineering, and business stakeholders, you will turn evaluation results into practical improvements. This is a high-impact individual contributor position offering ownership of a developing AI evaluation practice within a fast-growing financial technology environment. You will help shape how AI quality is measured, validated, documented, and continuously improved across the organization.

Accountabilities

  • AI Evaluation Design: Define ground truth, quality metrics, scoring methodologies, and evaluation frameworks to assess the accuracy, consistency, and reliability of AI models and agents.
  • Evaluation Pipeline Development: Build repeatable evaluation loops that monitor model quality over time, identify performance regressions, and support reliable release decisions.
  • Statistical Measurement and Validation: Apply rigorous statistical methods to evaluation design, including sample sizing, confidence intervals, significance testing, and techniques for handling non-deterministic model outputs.
  • Automated Grader Validation: Compare automated evaluation methods with human assessments to establish reliability, identify limitations, and improve scoring accuracy.
  • Evaluation Infrastructure Collaboration: Partner with Engineering and Analytics Engineering to operationalize evaluation frameworks, integrate testing into development workflows, and support scalable evaluation infrastructure.
  • Performance Analysis and Improvement: Interpret evaluation results, identify weaknesses and quality gaps, and provide actionable recommendations to improve model and agent performance.
  • Quality Standards and Documentation: Establish evaluation guidelines, documentation practices, review processes, and consistent quality standards across AI development initiatives.
  • Cross-Functional Partnership: Collaborate with Product, Engineering, Analytics Engineering, and business stakeholders to define success criteria and align evaluation priorities with business needs.
  • Continuous Improvement and Mentorship: Promote evaluation best practices, share technical insights, and help foster a culture of measurable AI quality, accountability, and evidence-based decision-making.
  • Safe AI Deployment: Support teams in making informed release decisions through independent, reliable assessments of model performance and readiness.
  • Approximately 6–10 years of experience in quantitative data science, machine learning, or a related technical discipline, with focused experience in measurement, evaluation, experimentation, or model validation.
  • Strong quantitative and statistical foundations, including experience designing rigorous experiments and interpreting results with appropriate uncertainty measures.
  • Demonstrated experience defining meaningful evaluation metrics, establishing ground truth, and assessing the quality of complex or ambiguous model outputs.
  • Proficiency in Python and SQL, with experience applying these skills to data analysis, model assessment, and evaluation workflows.
  • Experience evaluating machine learning models in production environments and translating findings into practical improvements.
  • Ability to validate automated graders against human judgments and account for variability, non-determinism, and potential measurement bias.
  • Strong problem-solving and analytical judgment, particularly when developing evaluation approaches for new or evolving AI systems.
  • Excellent communication skills, with the ability to explain evaluation results, trade-offs, and recommendations to technical teams and business stakeholders.
  • Proven ability to collaborate across functions while independently owning complex analytical projects in a fast-paced environment.
  • A quantitative degree in data science, statistics, mathematics, computer science, or a related field is an advantage; equivalent practical experience is also welcome.
  • Hands-on experience evaluating large language models (LLMs) or AI agents in production, including evaluation harnesses, LLM-as-judge calibration, and continuous integration regression gates, is a plus.
  • Experience evaluating text-to-SQL systems, analytics agents, or other AI applications where outputs can be verified against underlying data is an advantage.
  • Background in fintech, brokerage, financial services, or other domains where incorrect outputs can create significant business or risk consequences is desirable.
  • Familiarity with AI tools used in research, analysis, and engineering workflows is beneficial.
  • Competitive compensation: Salary package designed to reflect experience and expertise.
  • Stock options: Opportunity to participate in the company's long-term growth.
  • Health benefits: Benefits designed to support employee health and well-being.
  • Home-office setup allowance: One-time allowance of USD $500 to support your remote workspace.
  • Monthly stipend: USD $150 per month provided through a Brex card.
  • Remote work environment: Opportunity to work remotely within the Americas, with Canada as the target location for this listing.
  • Technical ownership: Lead the development of evaluation practices and establish standards that influence AI quality across the organization.
  • Cross-functional impact: Work closely with product, engineering, analytics, and business teams to turn rigorous measurement into better AI systems.
  • Professional development: Build expertise in AI evaluation, model validation, and the responsible deployment of intelligent systems in a growing technology environment.

How Jobgether works:

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Why Apply Through Jobgether?

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

#LI-CL1

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Jobgether's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Jobgether's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Jobgether's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.