Skip to content

Open nowPosted 18 hours ago

Senior ML Engineer – Speech & Voice (NXJ-212)

Newxel10 open roles

Where
Poland (Hybrid/Entity)
Work mode
Hybrid
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowSenior ML Engineer – Speech & Voice (NXJ-212)Newxel · Poland (Hybrid/Entity)
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Newxel's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

8.2% of postings close within 7 days. Measured by our own scanner across the market. Newxel postings stay open a median of 13 days.

Share of postings closed within
  1. 1.8%1 day
  2. 3.6%3 days
  3. 8.2%7 days
  4. 15.2%14 days
  5. 34.0%30 days
This job: posted 18 hours ago

Newxel median: 13 days open

The posting

THE ROLE

GPU-based microservices handle the core speech pipeline—STT consumption, word-level alignment, diarization, and speaker identification. The primary challenge isn't just serving models, but establishing rigorous data-driven evaluation to separate real accuracy gains from benchmark noise on difficult sports audio. The Senior ML Engineer will own the speech services, lead fair model bake-offs, and hold full authority over which models reach production.

ABOUT THE PRODUCT

The platform delivers real-time AI video processing and automated content generation for professional sports leagues globally. The underlying pipeline operates in high-noise live broadcast environments, demanding tight latency budgets, high throughput, and robust handling of overlapping speech and crowd noise.

Technology Stack: The platform delivers real-time AI video processing and automated content generation for professional sports leagues globally. The underlying pipeline operates in high-noise live broadcast environments, demanding tight latency budgets, high throughput, and robust handling of overlapping speech and crowd noise.

WHAT YOU’LL BE DOING

- Own and scale the core speech pipeline covering alignment, diarization, speaker identification, and enrollment signature-matching workflows.

- Optimize GPU inference performance for throughput, memory footprint, and cost through batching, precision tuning, and compilation.

- Build and maintain labeled sports test benchmarks stratified by speaker count, audio quality, language, and background noise.

- Define and track production evaluation metrics including DER, WER, word-level speaker attribution, speaker-count error, latency, and GPU cost.

- Conduct rigorous model bake-offs and write clear decision memos detailing trade-offs, licensing constraints, and statistical uncertainty.

- Maintain a shadow or A/B deployment framework to gate every production release with empirical evidence.

WHAT WE EXPECT

MUST-HAVE

- 3–5+ years of experience shipping ML systems into production, with 2+ years dedicated to speech or audio pipelines.

- Deep Python and PyTorch expertise, including hands-on GPU profiling, memory optimization, and inference acceleration.

- Hands-on experience with at least two key speech domains: diarization (pyannote, NeMo/Sortformer, VBx), speaker embeddings (ECAPA-TDNN), or ASR and forced alignment (Whisper, Parakeet, wav2vec-family).

- Strict evaluation rigor: demonstrated experience building test sets, computing DER/WER/EER correctly (handling collars and reference pitfalls), and running statistical hypothesis tests.

- Demonstrated ability to critically analyze academic papers or model cards and reproduce published claims on internal datasets.

- Hands-on experience deploying containerized GPU services in production environments using Docker and queue/API architectures.

NICE-TO-HAVE

- Hands-on experience with Azure ecosystem tools (Service Bus, Blob Storage) and event-driven microservices.

- Experience with domain adaptation or fine-tuning speech models on noisy, broadcast, or far-field sports audio.

- Experience handling multilingual STT pipelines, particularly with Hebrew, Arabic, or Spanish.

- Experience designing shadow deployment paths, experiment tracking workflows, and annotation team processes.

WHY THIS ROLE IS WORTH YOUR TIME

- Direct technical ownership of production GPU microservices where benchmark evaluations directly determine what reaches live users.

- Applied research opportunity on challenging sports audio conditions (crowd noise, overlapping commentators, PA systems) rather than clean laboratory datasets.

- Clear, evidence-based engineering culture where architectural and model choices are driven by data-backed decision memos rather than hype.

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Newxel's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Newxel's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Newxel's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.