Skip to content

Open nowPosted 32 days ago

AI Intern – Vision-Language-Action (VLA) & Data

rivr35 open roles

Where
Zürich
Work mode
On site
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowAI Intern – Vision-Language-Action (VLA) & Datarivr · Zürich
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on rivr's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

7.8% of postings close within 7 days. Measured by our own scanner across the market.

Share of postings closed within
  1. 1.6%1 day
  2. 3.4%3 days
  3. 7.8%7 days
  4. 14.3%14 days
  5. 33.7%30 days
This job: posted 32 days ago

The posting

RIVR, part of Amazon is a robotics company pioneering Physical AI through real-world doorstep delivery. Founded in 2024 as an ETH Zurich spin-off, RIVR, part of Amazon developed wheeled-legged robots designed to operate in complex, unstructured environments such as stairs, gates, doors, and uneven urban terrain. We believe that achieving general physical intelligence requires solving real customer problems in the real world, where robots can learn from rich operational data at scale.

Following our acquisition by Amazon in March 2026, we are continuing this mission with greater reach and speed. By combining custom robot hardware, onboard autonomy, and cloud-based coordination, RIVR, part of Amazon is building the next generation of safe, reliable autonomous robots for last-mile delivery.

Important Notice: For this position, we can unfortunately only accept applications from citizens of Schengen Area countries. This restriction does not apply to ETHZ and EPFL students who are required to complete compulsory internships as part of their studies.

Job Description

As an AI Intern - VLA & Data, you will assist the AI engineering team in developing data pipelines and training Vision-Language-Action (VLA) models for robotic systems. In this role, you will sit at the intersection of data engineering and model training, assisting the team in solving the "data bottleneck" in embodied AI. Additionally, your responsibilities will include building tools to analyze datasets, visualize model predictions, and debug performance. You will work closely with senior engineers to understand, curate, and visualize the massive amounts of multi-modal data our fleet generates. This role offers hands-on experience within a dynamic and collaborative environment. We are committed to finding and nurturing exceptional talent; our internships are a key pathway to recruiting outstanding graduates who can make a significant impact in our team.

What you’ll be doing

  • Support the team in processing, curating, and analyzing multi-modal sensor data for training VLA models.
  • Get hands-on experience with state-of-the-art VLA models.
  • Assist in developing software tools to visualize data and debug model performance.
  • Work closely with senior engineers to implement data strategies that improve the robustness of our robotic systems.
  • Help integrate software components to evaluate algorithms in simulation and on hardware.
  • Engage in continuous learning and gain exposure to recent developments in VLA, self-supervised learning, and generative AI.

What you must have

  • At least BSc in Computer Science, Robotics, Machine Learning, or a related field.
  • Proficiency in Python and experience with deep learning frameworks (preferably PyTorch).
  • Hands-on experience (either through coursework, previous internships, or other projects) with deep learning for computer vision.
  • Familiarity with data manipulation libraries (e.g., NumPy, Pandas, OpenCV).
  • Strong problem-solving skills and an eagerness to work with complex, real-world data.
  • Eagerness to learn and contribute in a collaborative team environment.

Get some bonus points

  • An MSc or ongoing PhD in a related field.
  • Previous experience working with large-scale image or video datasets.
  • Knowledge or experience with transformers, Vision-Language Models and/or VLA’s.
  • Experience with 3D geometry, camera projections, or sensor fusion.
  • Experience working with robotic systems
  • Experience in motion/action prediction in robotics context.

RIVR, part of Amazon is committed to building a diverse and inclusive team that values every perspective. If you’re passionate about driving innovation in robotics and creating meaningful impact, we encourage you to apply and bring your unique self to our team.

We believe the best work is done when collaborating and therefore require in-person presence in our office locations.

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against rivr's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on rivr's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    rivr's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.