Skip to content

Open nowPosted 108 days ago

Research Engineer/Scientist - Machine Learning RL & Optimisation (Contractor)

huaweiuk56 open roles

Where
London, United Kingdom
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowResearch Engineer/Scientist - Machine Learning RL & Optimisation (Contractor)huaweiuk · London, United Kingdom
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on huaweiuk's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

7.8% of postings close within 7 days. Measured by our own scanner across the market.

Share of postings closed within
  1. 1.7%1 day
  2. 3.5%3 days
  3. 7.8%7 days
  4. 14.6%14 days
  5. 34.1%30 days
This job: posted 108 days ago

The posting

About Huawei Research and Development UK Limited

Founded in 1987, Huawei is a leading global provider of information and communications technology (ICT) infrastructure and smart devices. We have 207,000 employees and operate in over 170 countries and regions, serving more than three billion people around the world.

Our vision and mission is to bring digital to every person, home and organization for a fully connected, intelligent world. To this end, we will drive ubiquitous connectivity and promote equal access to networks; bring cloud and artificial intelligence to all four corners of the earth to provide superior computing power where you need it, when you need it; build digital platforms to help all industries and organizations become more agile, efficient, and dynamic; redefine user experience with AI, making it more personalized for people in all aspects of their life, whether they’re at home, in the office, or on the go.

This spirit of innovation has led Huawei to work in close partnership with leading academic institutions in the UK to develop and refine the latest technologies. With a shared commitment to innovation and progress, both parties have worked together to achieve common goals and establish a strong partnership. The partnership between UK and Huawei help to develop the technologies of the future that will transform the way we all communicate, work and live.

For the past 30 years we have maintained an unwavering focus, rejecting shortcuts and easy opportunities that don't align with our core business. With a practical approach to everything we do, we concentrate our efforts and invest patiently to drive technological breakthroughs.

This strategic focus is a reflection of our core values:

  • Staying customer-centric,
  • Inspiring dedication,
  • Persevering,
  • Growing by reflection.

Huawei Research and Development UK Limited Overview

Huawei’s vision is a fully connected, intelligent world. To achieve this, we work to inspire passion for basic research around the world. Our combined passion drives development across the global innovation value chain. Huawei has the largest Research and Development organization in the world with 96,000+ employees in research centers around the globe. In the UK, we already have design centers in Cambridge, London, Edinburgh and Ipswich. We continue to explore and define new research directions and new services. We have expanded our collaborations with academic researchers; researched new network architectures, integration of communications and key enabling technologies; and developed the fundamental theories of these technologies. We invite you to join us on this exciting journey and drive your career forward.

Job Summary

Research and develop large-scale machine learning systems, alignment workflows, and optimization infrastructure to advance LLM reasoning and post-training capabilities. Design and execute scaled reinforcement learning pipelines (e.g., PPO, GRPO) utilizing distributed training frameworks (verl, trl, DeepSpeed, FSDP) integrated with high-performance inference engines (vLLM). Optimize low-level training throughput, kernel performance, and memory utilization across heterogeneous hardware clusters using expressive hardware DSLs (e.g., TileLang, Triton). Advance the LLM orchestration loop and leverage Bayesian optimization to automate the search, generation, and continuous improvement of high-performance NPU kernels.

Key Responsibilities:

  • Design and execute scaled RL finetuning workflows (e.g., PPO, GRPO) to enhance LLM reasoning, instruction-following, and alignment.
  • Architect and manage large-scale distributed training experiments across multi-node GPU, optimizing for maximum throughput and hardware utilization.
  • Develop and maintain training infrastructure using advanced parallelization frameworks (verl, trl, DeepSpeed, FSDP) to support rapidly evolving research needs.
  • Integrate high-performance inference engines like vLLM directly into RL generation loops to reduce rollout latency and accelerate training cycles.
  • Implement robust profiling and debugging pipelines to diagnose bottlenecks in GPU memory, compute, and inter-node communication.
  • Collaborate with data and evaluation teams to design dense reward functions and synthetic data generation pipelines.
  • Design, benchmark, and deploy highly optimized custom tensor operators (e.g., FlashAttention, GEMM) across heterogeneous hardware architectures using modern Domain-Specific Languages (DSLs) and AI compilers.

This job description is only an outline of the tasks, responsibilities and outcomes required of the role. The jobholder will carry out any other duties as may be reasonably required by his/her line manager. The job description and personal specification may be reviewed on an ongoing basis in accordance with the changing needs of Huawei Research and Development UK Limited.

Person Specification:

Required:

  • Master's or PhD (or equivalent industry research experience) in Machine Learning, Computer Science, Data Science, or a highly quantitative field with a heavy focus on Machine Learning.
  • Deep proficiency in PyTorch and experience writing custom training loops and data pipelines.
  • Hands-on experience with RLHF/RLVF methods (e.g., PPO, GRPO) and an understanding of policy optimization dynamics.
  • Technical familiarity with at least two of the following: DeepSpeed, FSDP, verl or trl.
  • Production or heavy research experience utilizing vLLM or similar high-throughput inference serving engines for generation.
  • Ability to thrive in a fast-paced, iterative environment where research and production infrastructure deeply intersect.

Desired:

  • Proven track record of running scaled GPU experiments across multi-node clusters.
  • Experience implementing next-gen alignment and reasoning paradigms, such as GRPO or Monte Carlo Tree Search (MCTS).
  • Deep understanding of GPU architectures, kernels, FlashAttention, and profiling tools.
  • Familiarity with cluster environments and schedulers like Slurm or Kubernetes.
  • Hands-on experience developing within NPU stack or ecosystem integrations.
  • Practical knowledge of hardware-expressive Domain-Specific Languages (e.g., TileLang, Triton) to optimize low-level memory placement, layout propagation, and thread-block scheduling.
  • Research publications at top-tier AI/ML conferences (NeurIPS, ICLR, ICML) or a strong open-source GitHub footprint in LLM training/infrastructure.

What we offer

  • 33 days annual leave entitlement per year (including UK public holidays)
  • Group Personal Pension
  • Life insurance
  • Private medical insurance
  • Medical expense claim scheme
  • Employee Assistance Program
  • Cycle to work scheme
  • Company sports club and social events
  • Additional time off for learning and development
From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against huaweiuk's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on huaweiuk's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    huaweiuk's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.