Skip to content

Open nowPosted 46 days ago

AI Research Scientist, Real-Time Video Understanding & Physical AI

Workable (global search)108,016 open roles

Where
Mountain View, CA, United States
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowAI Research Scientist, Real-Time Video Understanding & Physical AIWorkable (global search) · Mountain View, CA, United States
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Workable (global search)'s own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

7.9% of postings close within 7 days. Measured by our own scanner across the market. Workable (global search) postings stay open a median of 7 days.

Share of postings closed within
  1. 1.6%1 day
  2. 3.6%3 days
  3. 7.9%7 days
  4. 14.9%14 days
  5. 34.0%30 days
This job: posted 46 days ago

Workable (global search) median: 7 days open

The posting

The Problem

As robots powered by learned policies enter real manufacturing environments, making them trustworthy becomes as important as making them capable. Learned policies can be confidently wrong — executing smoothly while doing something other than what was intended — and understanding what a robot is actually doing on a factory floor, in real time and from observation alone, is an open research frontier.

Our lab researches the AI systems that make robot fleets trustworthy in production: real-time perception and reasoning over robot behavior — and, at its core, detecting when a robot is doing something wrong — built on multi-camera video understanding and multimodal signals from the operating environment and the robot itself.

Samsung SDS builds and operates the systems behind Samsung's global manufacturing — thousands of production lines worldwide. This research is being developed with that environment as its destination.

This is not a monitoring dashboard project. It is a frontier problem in video understanding: not just recognizing what a robot is doing, but reliably detecting when it is doing it wrong — on continuous, real-world behavior, in real time, under production latency and cost constraints.

The Team

You would join a small, hands-on lab of PhD-level researchers — no layers between you and the research. Your research happens alongside real robots that our lab operates end-to-end: we collect our own data through teleoperation, train open-source robot foundation models on our GPUs, and deploy them to humanoid robots to test their behavior. Ongoing work extends to dexterous manipulation. The work that proves out in the lab has a path to pilot deployment in real manufacturing settings.

We run on a simple contract: the mission is fixed; the method is yours. The lab's direction is clear and executive-sponsored, and everyone's work compounds toward it — but how you get there (which architectures, which formulations, which experiments) is your call to make and defend.

At this size, a new researcher is not headcount. Your technical judgment shapes how we get there from your first week.

What You Will Work On

  • Detecting anomalous robot behavior from real-time streaming video of humanoid robots at work — fusing external multi-camera views and, potentially, the robot's own egocentric video
  • Efficient VLM research: adapting vision-language models to achieve low-latency, low-cost on-prem deployment without sacrificing reasoning quality
  • Video-language grounding: connecting continuous visual observations of robot behavior with language and structured task knowledge
  • Multimodal fusion beyond vision: combining camera streams with robot state signals and manufacturing context data into a unified representation for judgment
  • World models and video prediction: moving beyond detection — reasoning about the causes and dynamics of robot behavior, and anticipating what happens next

Why This Role

  • Own the technical agenda. This is an executive-sponsored research effort at an early, formative stage. You will shape the architecture, the research questions, and the evaluation standards — not inherit them.
  • Ship into the physical world at scale. Samsung SDS operates the systems behind Samsung's global manufacturing. When this research succeeds, it does not end as a paper or a demo — the deployment path runs onto real production lines, at a scale almost no research organization anywhere can offer. Papers are a milestone here, not the finish line.
  • A data setting few others have. Synchronized multi-camera video of robots at work, paired with rich operational context from real manufacturing environments — a combination that few academic labs or frontier AI labs can match.
  • Real robots, every day. The lab trains and runs learned policies on its own physical platforms — from teleoperation-based learning to dexterous manipulation — so your research has a living testbed: real robots executing real learned behaviors, generating the kind of behavioral data most video researchers never get to touch. Publishing and patenting are part of how we work.

Requirements

Minimum Qualifications

  • PhD in Computer Vision, Machine Learning, Robotics, or a related field, or equivalent industry research experience
  • First-author publications at top venues (CVPR, ICCV, ECCV, NeurIPS, ICML, ICLR, CoRL, RSS, or comparable)
  • Hands-on research experience in video understanding: temporal action detection/segmentation, video anomaly detection, streaming/online perception, or long-form video reasoning

Preferred Qualifications

  • Video-language grounding, instructional/procedural video understanding, or video question answering
  • Multimodal learning combining vision with non-visual signals (sensor time-series, structured/tabular context)
  • World models, video prediction, or causal reasoning over temporal data
  • Exposure to robot learning (VLA models, imitation learning) — useful for understanding how learned policies behave and fail
  • Hands-on experience building perception or ML systems that run on real-world data streams — cameras, sensors, robots, or production data — not only curated benchmarks

Benefits

Compensation

This role offers competitive compensation package including base salary, bonus, and benefits.

Expected salary range for this role in Mountain View, CA:

$230,000 – $270,000 base salary, depending on experience, interview assessment results, skills and qualification. On top of base salary this role may participate in performance bonus plan.

Samsung SDSA offers a comprehensive suite of programs to support our employees:

  • Top-notch medical, dental, vision and prescription coverage
  • Wellness program
  • Parental leave
  • 401K match and savings plan
  • Flexible spending accounts
  • Life insurance
  • Paid Holidays
  • Paid Time off
  • Additional benefits

Samsung SDS America, Inc. is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, race, color, religion, sex, sexual orientation, gender identity or expression, national origin, disability, status as a protected veteran, marital status, genetic information, medical condition, or any other characteristic protected by law.

We are committed to providing reasonable accommodations to participate in the job application or interview process for candidates with disabilities. Please let your recruiter know if you need an accommodation at any point during the interview process.

Certain roles are eligible for additional rewards, including annual bonus. U.S.-based employees have access to medical, dental, and vision insurance, a 401(k) plan and company match, short-term and long-term disability coverage, basic life insurance, and wellbeing benefits, among others. U.S.-based employees also receive, per calendar year, up to 10 scheduled paid holidays, and Paid Time Off.

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Workable (global search)'s own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Workable (global search)'s form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Workable (global search)'s answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.