Skip to content

Open nowPosted 35 days ago

Master Thesis Project - 2027

modulai5 open roles

Where
Stockholm, Sweden; Göteborg, Sweden
Work mode
Hybrid
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowMaster Thesis Project - 2027modulai · Stockholm, Sweden; Göteborg, Sweden
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on modulai's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

7.9% of postings close within 7 days. Measured by our own scanner across the market.

Share of postings closed within
  1. 1.6%1 day
  2. 3.6%3 days
  3. 7.9%7 days
  4. 14.9%14 days
  5. 34.2%30 days
This job: posted 35 days ago

The posting

1. Bridging the sim-to-real Gap: Domain randomization and synthetic data for robotic arm control (STHLM)

Background & Description

We offer a master's thesis project on closing the gap between simulation and physical hardware in robot learning. Policies trained purely on simulated data, including Vision-Language-Action (VLA) models, often fail on real hardware due to mismatches in visuals, physics, and sensing. Domain randomization addresses this by varying simulation parameters during data generation so the model learns features invariant to the sim-real gap rather than simulator artifacts (Tobin et al., 2017).

This thesis treats synthetic data production as an optimization problem: which parameter distributions, quantities, and sim/real mixtures maximize real-world performance per unit of data and compute? Recent work shows that even simple sim/real co-training recipes substantially improve manipulation success rates (Maddukuri et al., 2025), but principled, mathematically grounded strategies remain an open question.

Core idea: synthetic data is generated via a mathematically justified and data-efficient randomization strategy, a VLA or comparable model is trained on it, and the resulting policy is deployed on physical robotic arm hardware. The theoretical contribution is the mathematical analysis behind the strategy, for example coverage guarantees, sample complexity, or framing parameter selection as an optimization problem. The applied contribution is validating the policy on real hardware.

Students will work with a physical robotic arm, GPU compute, and guidance from Modulai's ML engineers

Example directions

  • Formal analysis of how domain randomization ranges should be chosen relative to the true (unknown) distribution of real-world conditions, and what guarantees this gives on real-world generalization
  • Optimizing the mixture and scheduling of simulated versus real demonstration data during VLA fine-tuning
  • Automatic or learned domain randomization, where randomization parameters are adapted based on validation performance rather than fixed by hand
  • End-to-end evaluation: train in simulation only, train with a randomization strategy, and train with sim+real co-training, then compare real-arm task success rates

ML Techniques and Tools

  • Python, PyTorch, Git, Hugging Face
  • Robotics simulators (e.g. MuJoCo, Isaac Sim, or similar) for synthetic data generation
  • Domain randomization and sim-to-real transfer methods
  • Vision-Language-Action models and other end-to-end control architectures
  • Statistical and optimization methods for data generation strategy design
  • Real-time control and deployment on physical robotic arm hardware

References

Tobin et al., Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World, 2017. arXiv:1703.06907 - https://arxiv.org/abs/1703.06907

Maddukuri et al., Sim-and-Real Co-Training: A Simple Recipe for Vision-Based Robotic Manipulation, 2025. arXiv:2503.24361 - https://arxiv.org/abs/2503.24361

2. Image representations in property valuation (Computer vision/Tabular) (Sthlm/Gbg)

Background & Description

We offer a master's thesis project on using image data to improve the accuracy of automated property valuation. The project is run together with a growing startup that is building a state-of-the-art valuation engine, with guidance from Modulai's ML engineers.

Automated valuation models traditionally rely on tabular data: living area, number of rooms, location, construction year and historical transactions. Two apartments with near-identical records can still differ substantially in market value because of condition, renovation standard, light, layout and view. Much of that residual signal is present in listing photographs, floor plans and aerial imagery, but is rarely exploited beyond coarse heuristics. Early work showed that a learned "luxury level" derived from interior and exterior photos, combined with metadata, can outperform established metadata-only estimates (Poursaeed et al., 2017), and later studies confirm that visual features add predictive power on top of strong tabular baselines (Kostic & Jevremović, 2021).

This thesis treats the image side as a representation and integration problem: which visual representations carry the signal that tabular features miss, and how should they be fused into a production valuation model without hurting robustness, calibration or explainability? The candidate representations span the full toolbox - image classification (room type, condition, renovation standard), semantic segmentation (materials, surfaces, greenery, floor-plan geometry), object detection (fireplaces, appliances, balconies) and general-purpose embeddings from pretrained vision or vision-language backbones.

Core idea: Explore image model approaches to represent the image information as efficiently as possible, while keeping explainability of the model. Investigate how such methods may contribute to improved valuation accuracy for apartments and house listings, The methodological contribution is the comparison and fusion strategy; the applied contribution is a validated improvement in a system that is actually shipped.

Students will work with large-scale real listing data, GPU compute, and close guidance from both the company's ML team and Modulai's ML engineers. It is also a domain that is unusually easy to relate to - everyone lives somewhere.

Example directions

  • Systematic comparison of representation families - classification heads, segmentation masks, detected objects and raw embeddings - on equal footing, measured as marginal accuracy over a strong tabular baseline
  • Off-the-shelf vision-language models used as attribute extractors versus purpose-trained models and learned embeddings: accuracy, cost and latency per valuation
  • Fusion architecture: late fusion of pooled embeddings into a gradient-boosted model versus end-to-end multimodal training, including how to aggregate a variable number of images per property
  • Weak supervision from price residuals: learning image representations directly against the part of the price the tabular model cannot explain, instead of relying on generic pretrained features
  • Floor plans as structured input: extracting layout, room adjacency and geometry, and testing whether structure beats appearance
  • Robustness to presentation: staged photos, wide-angle lenses, HDR and photographer differences - separating property quality from marketing quality
  • Uncertainty and explainability: does image data mainly shift the point estimate or tighten prediction intervals, and can per-image contributions be attributed in a way a human valuer would accept?

ML Techniques and Tools

  • Python, PyTorch, Git, Hugging Face
  • CNN and vision-transformer backbones; pretrained embeddings (e.g. CLIP, DINOv2) and vision-language models
  • Semantic segmentation and object detection (e.g. SAM, Mask R-CNN, YOLO-family models)
  • Multimodal fusion and gradient boosting (LightGBM/XGBoost) over combined tabular and image features
  • Explainability (SHAP, attention and saliency maps) and uncertainty quantification (quantile regression, conformal prediction)
  • Cloud GPU compute, experiment tracking and evaluation on real data

References

Poursaeed et al., Vision-based Real Estate Price Estimation, 2017. arXiv:1707.05489 - https://arxiv.org/abs/1707.05489

Kostic & Jevremović, What Image Features Boost Housing Market Predictions?, 2021. arXiv:2107.07148 - https://arxiv.org/abs/2107.07148

Zillow neural network estimate https://www.zillow.com/news/building-the-neural-zestimate/

3. Open Application within Applied Machine Learning

Applied Machine Learning projects encompass a wide range of domains, including healthcare, finance, natural language processing, computer vision, and more. This open application invites students to choose projects aligned with their interests and career goals. Do you have an idea - let us know what it's about by describing it.

Required Skills

Finishing a master's in machine learning or a master's in another field but with courses in machine learning and programming added

Please include the following in your application:

  • Link to relevant GitHub account if available.
  • Grades for bachelor's and master's.
  • Updated CV or an updated LinkedIn profile.

*Suitable candidates will be called to one interview before making a final decision. The last date for application will be the 31th of October, but if suitable candidates apply, the process will end beforehand.

About Modulai

Modulai’s clients range from startups to multinational companies. They all share that machine learning is central to how they operate, compete, and create value.

Our services range from advisory projects and feasibility studies to end-to-end development and refinement of machine learning systems and products.

We use state-of-the-art techniques, always focusing on maximizing business impact, delivering solutions in areas such as credit risk, fraud detection, dynamic pricing, recommendation systems, computer vision, natural language processing, opportunity spotting, logistics optimization, up-sell, cross-sales, smart building optimization, predictive maintenance, and route planning.

Other

When doing a master thesis project at Modulai, you are invited to all team activities such as daily stand-ups, weekly learning breakfasts, monthly AWs, and other team activities. We look forward to having you as part of our team!

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against modulai's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on modulai's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    modulai's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.