Skip to content

Open nowPosted 121 days ago

Member of Technical Staff (MTS) - Multimodal Foundation Models

Workable (global search)107,990 open roles

Where
Fremont, CA, United States
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowMember of Technical Staff (MTS) - Multimodal Foundation ModelsWorkable (global search) · Fremont, CA, United States
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Workable (global search)'s own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

7.9% of postings close within 7 days. Measured by our own scanner across the market. Workable (global search) postings stay open a median of 6 days.

Share of postings closed within
  1. 1.6%1 day
  2. 3.6%3 days
  3. 7.9%7 days
  4. 14.9%14 days
  5. 34.0%30 days
This job: posted 121 days ago

Workable (global search) median: 6 days open

The posting

Focus

Multimodal Foundation Models · Representation Learning · Method Innovation

We are looking for strong technical builders and researchers who deeply understand foundation models and representation learning beyond simply applying existing frameworks.

Ideal candidates should have:

  • Strong experimental rigor
  • Solid systems and modeling intuition
  • Hands-on engineering ability
  • Interest in scalable multimodal AI systems for real-world autonomy

We value people who can bridge research and production, and who care about robustness, scalability, efficiency, and practical deployment in large-scale autonomous driving systems.

Responsibilities

1. Large-Scale Foundation Model Pretraining

  • Develop scalable pretraining pipelines for large-scale multimodal driving data
  • Design and optimize training strategies for:
  • Vision-language-action models
  • Video foundation models
  • Long-context temporal modeling
  • Multimodal representation alignment
  • Improve:
  • Training stability
  • Data efficiency
  • Scaling efficiency
  • Representation robustness
  • Work on distributed training systems and large-scale model optimization using frameworks such as:
  • PyTorch Distributed
  • DeepSpeed
  • Megatron-LM

2. Representation Learning & Method Innovation

  • Design and improve self-supervised and multimodal learning methods for real-world autonomous driving systems
  • Conduct architecture-level research on:
  • Vision Transformers (ViT)
  • Video / temporal architectures
  • Multimodal fusion and alignment
  • Embedding and retrieval systems
  • Long-context and memory-efficient architectures
  • Explore and improve:
  • Pretraining objectives
  • Loss functions
  • Training paradigms
  • Generalization and robustness
  • Analyze model behavior through:
  • Rigorous ablation studies
  • Failure case analysis
  • Representation probing and evaluation

3. Efficient Foundation Models & Scalable Deployment

  • Improve the efficiency, scalability, and deployability of large multimodal foundation models for real-world autonomous driving systems
  • Work on areas such as:
  • Model quantization
  • Knowledge distillation
  • Efficient attention mechanisms
  • Sparse architectures and Mixture-of-Experts (MoE)
  • Long-context and memory-efficient modeling
  • Inference acceleration and serving optimization
  • Training and inference system efficiency
  • Optimize model throughput, latency, memory usage, and deployment performance for large-scale production environments

Requirements

  1. MS or PhD in:
  2. Computer Vision
  3. Machine Learning
  4. Robotics
  5. Computer Science
  6. Related fields
  7. Strong understanding of:
  8. Foundation models
  9. Self-supervised learning
  10. Representation learning
  11. Multimodal learning
  12. Large-scale pretraining
  13. Hands-on experience with methods such as:
  14. CLIP
  15. DINO / DINOv2
  16. MAE
  17. Contrastive learning
  18. Masked modeling
  19. MoE or scalable transformer architectures
  20. Experience with one or more of the following is highly valued:
  21. Video foundation models
  22. Long-context modeling
  23. Retrieval systems
  24. Efficient inference
  25. Distributed training
  26. Model compression and deployment optimization
  27. Strong publication record in top-tier venues is preferred:
  28. CVPR
  29. ICCV
  30. ECCV
  31. NeurIPS
  32. ICLR
  33. ICML
From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Workable (global search)'s own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Workable (global search)'s form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Workable (global search)'s answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.