Skip to content

Open nowPosted 30 days ago

Agentic Engineer II - Voice AI

Workable (global search)108,016 open roles

Where
Shenzhen, Guangdong Province, China
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowAgentic Engineer II - Voice AIWorkable (global search) · Shenzhen, Guangdong Province, China
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Workable (global search)'s own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

7.9% of postings close within 7 days. Measured by our own scanner across the market. Workable (global search) postings stay open a median of 7 days.

Share of postings closed within
  1. 1.6%1 day
  2. 3.6%3 days
  3. 7.9%7 days
  4. 14.9%14 days
  5. 34.0%30 days
This job: posted 30 days ago

Workable (global search) median: 7 days open

The posting

About Wati

Started as a WhatsApp team inbox in 2020, Wati has evolved into an AI-powered customer engagement platform that goes beyond a single channel. Designed for businesses that sell, support, and grow through conversations, Wati observes customer intent in real time, decides the next best revenue action, and executes it across marketing, sales, and support — on WhatsApp, Instagram, Facebook, TikTok, SMS, and more.

Trusted by over 16,000 customers across 190+ countries, Wati simplifies complex operations and business conversations with a unified inbox, no-code automation, and our intelligent AI layer, Astra.

Proudly backed by Tiger Global, Sequoia Capital, DST Global, and Shopify, and recognised as a Premium Partner of Meta and Google.

About the Role

We are looking for an Agentic Engineer – Voice AI to build and scale Wati's real-time voice AI capabilities on WhatsApp.

In this role, you will develop the systems that allow AI agents to listen, think, and speak in real time over WhatsApp voice calls. This includes building and optimizing the real-time media pipeline (WebRTC, LiveKit), integrating frontier AI models (OpenAI Realtime API, Google Gemini Live), and engineering the cascade architecture that connects speech recognition, language models, and speech synthesis into a seamless conversational experience.

You will work on latency-critical infrastructure where every millisecond matters — from audio transport and voice activity detection to model inference and text-to-speech delivery. You will also contribute to the broader AI agent stack, including tool calling, context management, and multi-turn conversation orchestration.

This role sits at the intersection of real-time communication systems, AI model integration, and conversational voice experiences.

What You Will Do

• Design, build, and optimize real-time voice AI pipelines — from WebRTC media transport to LLM inference and speech synthesis

• Integrate and orchestrate frontier AI models including OpenAI Realtime API, Google Gemini multimodal live, and cascade architectures (ASR → LLM → TTS)

• Build and maintain the media infrastructure: LiveKit-based audio routing, Opus codec handling, RTP/RTCP transport, and voice activity detection

• Develop agent capabilities for voice interactions — tool calling, function execution, context engineering, and multi-turn conversation management

• Optimize end-to-end latency across the voice pipeline, from audio capture to AI response playback

• Collaborate with product and platform teams to deliver production-grade voice AI experiences on WhatsApp

• Ensure reliability, performance, and scalability of voice AI infrastructure serving customers across 190+ countries

Requirements

• 3+ years of software engineering experience, with strong backend development skills (Go or Python preferred)

• Experience with real-time communication technologies: WebRTC, RTP/RTCP, audio codecs, or media server infrastructure

• Familiarity with AI/LLM integration — model APIs, tool calling, prompt engineering, or agent orchestration

• Experience with or strong interest in speech technologies: ASR, TTS, voice activity detection, or audio processing pipelines

• Understanding of distributed systems, microservices, and cloud-native architectures (GCP preferred)

• Comfortable working with PostgreSQL, Redis, and pub/sub messaging systems

• Strong problem-solving ability and ability to work in fast-paced, ambiguity-rich environments

• Ability to debug complex, latency-sensitive systems by reading code, traces, and real-time metrics

Nice to Have

• Hands-on experience with OpenAI Realtime API, Gemini multimodal live, or similar real-time AI model APIs

• Experience with AI agent frameworks (Dify, LangChain, CrewAI, etc.)

• Familiarity with MCP (Model Context Protocol) or other agent integration standards

• Contributions to open-source projects in the AI or real-time communication space

Behavioural Expectations

• Strong ownership and bias for action in a fast-moving environment

• Proactive, self-driven, and comfortable working across teams to drive outcomes

• AI-native mindset, actively using AI tools in daily engineering workflow and experimenting with agent frameworks and emerging technologies

• Curiosity-driven — eager to push the boundaries of what voice AI can do in real business contexts

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Workable (global search)'s own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Workable (global search)'s form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Workable (global search)'s answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.