Skip to content

Open nowPosted 332 days ago

Backend / ML-Ops Engineer — Speech Model Deployment & Inference Optimization

outcomesai3 open roles

Where
Bengaluru
Work mode
On site
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowBackend / ML-Ops Engineer — Speech Model Deployment & Inference Optimizationoutcomesai · Bengaluru
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on outcomesai's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

7.9% of postings close within 7 days. Measured by our own scanner across the market.

Share of postings closed within
  1. 1.6%1 day
  2. 3.6%3 days
  3. 7.9%7 days
  4. 14.9%14 days
  5. 34.2%30 days
This job: posted 332 days ago

The posting

OutcomesAI is a healthcare technology company building an AI-enabled nursing platform designed to augment clinical teams, automate routine workflows, and safely scale nursing capacity.

Our solution combines AI voice agents and licensed nurses to handle patient communication, symptom triage, remote monitoring, and post-acute care — reducing administrative burden and enabling clinicians to focus on direct patient care.

Our core product suite includes:

● Glia Voice Agents – multimodal conversational agents capable of answering patient calls, triaging symptoms using evidence-based protocols (e.g., Schmitt-Thompson), scheduling visits, and delivering education and follow-ups.

● Glia Productivity Agents – AI copilots for nurses that automate charting, scribing, and clinical decision support by integrating directly into EHR systems such as Epic and Athena.

● AI-Enabled Nursing Services – a hybrid care delivery model where AI and licensed nurses work together to deliver virtual triage, remote patient monitoring, and specialty patient support programs (e.g., oncology, dementia, dialysis).

Our AI infrastructure leverages multimodal foundation models — incorporating speech recognition (ASR), natural language understanding, and text-to-speech (TTS) — fine-tuned for healthcare environments to ensure safety, empathy, and clinical accuracy. All models operate within a HIPAA-compliant and SOC 2–certified framework. OutcomesAI partners with leading health systems and virtual care organizations to deploy and validate these capabilities at scale. Our goal is to create the world’s first AI + nurse hybrid workforce, improving access, safety, and efficiency across the continuum of care.

Own the infrastructure and pipelines for integrating trained ASR/TTS/Speech-LLM models into production. Focus on scalable serving, GPU optimization, monitoring, and continuous improvement of inference latency and reliability.

What You’ll Do

  • Containerize and deploy speech models using Triton Inference Server with TensorRT/FP16 optimizations.
  • Develop and manage CI/CD pipelines for model promotion (staging → production).
  • Configure autoscaling on Kubernetes (GPU pools) based on active calls or streaming sessions.
  • Build health and observability dashboards: latency, token delay, WER drift, SNR/packet loss monitors.
  • Integrate LM bias APIs, failover logic, and model switchers for fallback to larger/cloud models.
  • Implement on-device or edge inference paths for low-latency scenarios.
  • Collaborate with AI team to expose APIs for context biasing, rescoring, and diagnostics.
  • Optimize GPU/CPU utilization, cost optimization, and memory footprint for concurrent ASR/TTS/Speech LLM workloads.
  • Maintain data and model versioning pipelines with MLflow, DVC, or internal registries.

Desired Skills

  • Experience with Triton, TensorRT, Docker, Kubernetes, and GPU scheduling.
  • Familiarity with speech inference (streaming ASR, TTS pipelines).
  • Proficient in Python, Bash, and cloud services (AWS/GCP/Azure).
  • Understanding of observability stacks (Prometheus, Grafana, ELK).
  • Knowledge of DevSecOps, access policies, and PHI-safe environments.
  • Interest in inference optimization, mixed precision, and quantization.

Qualifications

  • B.Tech / M.Tech in Computer Science or related field.
  • 4–6 years of experience in backend or ML-ops; at least 1–2 years with GPU inference pipelines.
  • Proven experience deploying models to production environments with measurable latency gains.
From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against outcomesai's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on outcomesai's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    outcomesai's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.