Skip to content

Open nowFirst seen 3 hours ago

Software Engineer, Multimedia & Multimodal AI

Meta1,033 open roles

Where
Menlo Park, CA; Remote, US
Work mode
Remote
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowSoftware Engineer, Multimedia & Multimodal AIMeta · Menlo Park, CA; Remote, US
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Meta's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

8.0% of postings close within 7 days. Measured by our own scanner across the market. Meta postings stay open a median of 35 days.

Share of postings closed within
  1. 1.6%1 day
  2. 3.6%3 days
  3. 8.0%7 days
  4. 15.0%14 days
  5. 34.2%30 days
This job: first seen 3 hours ago

Meta median: 35 days open

The posting

Applied AI (AAI) is Meta’s organization focused on making our AI models best-in-class, starting with coding. Within AAI, the Multimedia & MultiModality team covers the multimedia domain across every modality, on both the input and the output side of a model: image, video, audio, speech and music. We work directly with research, model-training and engineering partners across MSL, TBD and FAIR. Current problems include evaluating video experiences, diagnosing multimedia model behavior, producing domain-expert agent tasks, and building the data and measurement pipelines multimodal capabilities are trained and judged against. About the role You will take a modality or a capability area, decide what data is worth producing and how it should be measured, and carry it from an open question through to a pipeline that runs and a measurement the org relies on.

This is a multimodal role, not a text-only role. You will work across image, video, audio and speech, as model inputs and as model outputs, and the data and evaluations you own will cover media, not text alone.You will choose where the pod invests, own outcomes beyond your individual contribution, set standards other engineers build against, and raise quality without becoming the review bottleneck.

Responsibilities

  • Design and build agentic workflows and pipelines, including human-in-the-loop and expert-in-the-loop designs, to automate data production and scale output past what manual authoring supports.
  • Design and own data pipelines at scale: ingestion, filtering, pseudo-labeling and captioning with attribute classifiers, and provenance tracking for audio corpora.
  • Build evaluation infrastructure: objective metrics (speaker/style similarity, codec and generator quality), human listening-test pipelines, and the correlation analysis that ties the two together.
  • Improve training efficiency and reliability — distributed training, GPU utilization, codec and tokenizer retraining, experiment management.
  • Reproduce and extend state-of-the-art research: implement new methods from papers into our codebases and run rigorous ablations.
  • Mentor engineers on the team, contribute to hiring and onboarding, and raise the bar on evaluation and quality practice.
  • Build and train generative and representation models for speech, sound, and music — including text-, audio-, and video-conditioned generation, infilling, editing, and style transfer.

Minimum Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
  • 6+ years of programming experience in a relevant language or 3+ years of experience + PhD
  • 3+ years building ML systems in production or research settings
  • strong Python and PyTorch
  • Demonstrated experience with speech, audio, or music ML - ASR, TTS, audio codecs, music information retrieval, self-supervised audio representation learning, or audio generative modeling
  • Experience with large-scale data pipelines and distributed training
  • Track record of translating research ideas into working, measurable systems

Preferred Qualifications

  • Publications at top venues (ICASSP, Interspeech, ISMIR, NeurIPS, ICML, ICLR) in speech, audio, or music
  • Generative modeling of continuous data (diffusion / flow matching, audio or vision), and demonstrated ability to switch domains and ramp quickly
  • Audio DSP depth — pitch detection, FFT, real-time signal processing
  • Experience with disentangled or controllable generation (voice, emotion, style, instrumentation)
  • Experience building evaluation harnesses and human-eval pipelines for generative audio
  • Music domain expertise: stem separation, mixing, lyrics/vocal conditioning
  • Experience designing benchmarks or evaluations for model capability, with attention to grading reliability, reproducibility and label quality
  • Experience building data pipelines for image, video, audio, speech or complex media formats, including versioning, lineage and provenance
  • Experience designing AI agents, orchestration, or human-in-the-loop systems
  • Hands-on experience evaluating or red-teaming multimodal models, or creating the data used to improve them
  • Understanding of Responsible AI practices and building quality controls into AI output
  • Experience with zero-to-one work: forming a charter and standing up process while priorities are still moving
  • Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
  • Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
  • Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies

US: $154,003/year to $217,000/year + bonus + equity

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Meta's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Meta's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Meta's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.