Skip to content

Open nowPosted 10 days ago

AI Evaluation Engineer

Joblogic37 open roles

Where
Pakistan
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowAI Evaluation EngineerJoblogic · Pakistan
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Joblogic's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

7.4% of postings close within 7 days. Measured by our own scanner across the market.

Share of postings closed within
  1. 1.5%1 day
  2. 3.5%3 days
  3. 7.4%7 days
  4. 13.1%14 days
  5. 34.4%30 days
This job: posted 10 days ago

The posting

The Joblogic Story

Established in 1998, Joblogic is the UK’s #1 Field Service Management (FSM) software platform. We are a global business with offices in the UK, Pakistan, and Vietnam. Since our management buy-out in 2013, we have grown from ~£500K ARR to ~£35M+ ARR and expanded our team from 11 to 500+ people.

Recently, we secured a strategic growth investment from Vista Equity Partners — a global technology investor specialising in enterprise software. This investment includes over £100 million in new primary capital and will fuel our next phase of growth by accelerating our AI-first roadmap, expanding our platform into CAFM (Computer-Aided Facilities Management) capabilities, and supporting our expansion across Europe and beyond.

With Vista’s backing, we’re transforming from a successful UK business into a global scaling SaaS rocket ship, and we’d love for you to join us on our journey to £100M ARR across international markets.

Joblogic provides software to service contractors who install and maintain the built environment. Our platform helps businesses streamline operations, improve profitability, ensure compliance, and achieve rapid growth. With over 100,000 users across industries including HVAC, plumbing, electrical maintenance, facilities management, and building fabric maintenance, we are entering a new era of intelligent automation, predictive maintenance, and data-driven decision-making for service firms.

About the Role

We are building Joblogic’s AI Agent Platform — a multi-tenant system for designing, versioning, evaluating, and running AI agents that work across email, voice, SMS, WhatsApp, and CRM channels on behalf of our customers. The platform is built on a LangGraph runtime with retrieval over Azure AI Search, a real-time voice stack, human-in-the-loop review queues, and an evaluation harness backed by LangSmith and PromptFoo.

We are looking for an AI Evaluation Engineer to own quality for everything we ship that has a model in it. You will define what “good” means, measure it rigorously, find out why it is not met, and drive the fixes. The role is deliberately hybrid: you design the rubrics, datasets, and judges, and you build the harnesses and release gates that run them in CI and against live production traffic. Evaluation is not a reporting function here — it is the mechanism by which agent quality improves release over release, and you own it.

You will work closely with the engineers building agents, the data team, and product, and your work will directly determine what tens of thousands of field-service businesses experience when an agent answers on their behalf.

What You’ll Do

  • Own release sign-off — gate every agent version on a green regression suite, judge-scored evaluations within agreed error bars, and a red-team pass — no promotion without them.
  • Run error analysis — review sampled production traces across chat, email, WhatsApp, and voice every week, maintain the failure taxonomy, and turn new failure modes into dataset examples within the sprint.
  • Build datasets and rubrics — design and maintain golden datasets and grading rubrics for multi-turn, tool-using agents, promoting interesting production runs into datasets in LangSmith.
  • Build and calibrate LLM judges — align judges to human labels, re-label a held-out set each cycle, publish agreement per rubric, and retire or retrain judges that drift.
  • Run offline and online evaluation — regression suites in CI alongside online scoring of sampled production traffic, with clear pass/fail thresholds and drift tracking.
  • Triage regressions to root cause — attribute failures to prompt, tool, retrieval, model upgrade, or speech provider using paired statistics rather than aggregate deltas, and hand engineers concrete fixes.
  • Own voice quality metrics — define and gate call-outcome metrics — task success, containment, barge-in recovery, word error rate under noise — alongside latency budgets.
  • Evaluate classical ML models — set acceptance criteria, slice-level thresholds, calibration checks, and drift alerts for the vision, speech, and predictive models the platform depends on.
  • Run adversarial testing — probe for prompt injection through inbound email and messaging content, jailbreaks, PII leakage, and tool misuse, and add regression coverage for every finding.
  • Run the human evaluation programme — own annotation queues, reviewer guidelines, and inter-annotator agreement as an ongoing operation rather than a one-off study, and keep human and automated scores connected.
  • Collaborate & ship — work in a cross-functional team using tools such as Jira and Slack, write clear documentation, and ship iteratively with a strong quality bar.

Essential Experience and Skills

  • 3+ years in roles where evaluating models was the core of the job — ML evaluation, ML quality, applied ML, or data science with an evaluation focus.
  • Strong Python engineering skills, with experience building test harnesses and clean, well-tested code.
  • Practical experience evaluating LLM-powered applications or AI agents: building datasets, defining heuristic and LLM-as-judge rubrics, running evaluations, and interpreting results to improve a system. This is a core requirement.
  • Experience grading tool-using agents on both trajectory and outcome — tool-call correctness, expected-trajectory match, end-state checks — and reporting reliability over repeated trials.
  • Experience building LLM judges and aligning them to human labels, including agreement statistics and mitigation of position, verbosity, and self-preference bias.
  • Solid classical ML evaluation foundations: classification, regression, and ranking metrics, cross-validation, calibration, slice-based evaluation, and drift monitoring.
  • Statistics for small evaluation sets: paired comparisons, confidence intervals, power analysis, and minimum detectable effect — you know roughly how many examples a claim needs before you make it.
  • Experience with RAG evaluation: faithfulness, groundedness, context precision and recall.
  • Hands-on experience with an LLM observability and evaluation platform (LangSmith, MLflow, or equivalent): datasets, experiments, custom evaluators, feedback, and wiring evaluations into release gates.
  • Strong data analysis skills using Pandas, NumPy, and SQL to quantify behaviour and communicate findings.
  • Working knowledge of how agents are built — prompts, tools, retrieval, and memory — and how they interact, sufficient to root-cause a failure rather than only report it.
  • Awareness of AI safety and adversarial risk: prompt injection, jailbreaks, data leakage, tool misuse, and responsible-AI practice.
  • Awareness of the compliance side of evaluation: handling conversation data under UK GDPR, keeping evaluation evidence auditable, and emerging record-keeping expectations for AI systems.
  • Committed to continuous learning, proactive problem-solving, and timely issue identification, with a keen interest in staying current with a fast-moving field.
  • Strong communicator, experienced in collaborating with cross-functional teams using tools such as Jira and Slack.
  • Creative and innovative thinker, consistently contributing fresh ideas and solutions in alignment with current technological trends.

Nice to Have

  • Experience with LangGraph and LangSmith specifically (tracing, datasets, online evaluators, judge alignment).
  • Experience with evaluation tooling such as PromptFoo, DeepEval, RAGAS, or Inspect.
  • Experience with red-team tooling such as PromptFoo red team, Microsoft PyRIT, or NVIDIA Garak, and familiarity with the OWASP Top 10 for LLM and Agentic Applications.
  • Experience evaluating voice agents: call simulation, turn-taking and latency budgets, speech recognition accuracy under noise.
  • Experience evaluating computer-vision or speech models (mAP/IoU, WER/CER) with sample-level failure mining.
  • Experience with Databricks or AWS SageMaker and experiment tracking with MLflow.
  • Experience designing human-in-the-loop annotation programmes and measuring inter-annotator agreement.
  • Experience with Python web frameworks (FastAPI / Flask), pytest-native harnesses, and CI/CD.
  • Publications, open-source contributions, or public evaluation work.

What We Offer

  • Professional Working environment
  • Market Competitive Salary
  • Life Insurance & Medical Insurance (Including Family)
  • OPD
  • Provident Fund
  • Gym Facility
  • Maximum 45 Weekly Hours (Monday–Friday)
  • Remote Working (During Pandemic Situation)
  • Company trip
  • 29 Annual Leaves
  • 8 Sick & uncapped Compassionate Leaves (As per Company Policy)
  • Have a chance to work onsite with the UK team
From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Joblogic's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Joblogic's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Joblogic's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.