Skip to content

Open nowPosted 3 days ago

Data Engineer – Applied ML

Similarweb67 open roles

Where
Tel Aviv, Israel
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowData Engineer – Applied MLSimilarweb · Tel Aviv, Israel
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Similarweb's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

7.8% of postings close within 7 days. Measured by our own scanner across the market. Similarweb postings stay open a median of 36 days.

Share of postings closed within
  1. 1.7%1 day
  2. 3.5%3 days
  3. 7.8%7 days
  4. 14.6%14 days
  5. 34.1%30 days
This job: posted 3 days ago

Similarweb median: 36 days open

The posting

Similarweb is the leading digital intelligence platform used by over 3500 global customers. Our wide range of solutions power the digital strategies of companies like Google, eBay, and Adidas.

We help our customers succeed in today’s digital world by giving them access to data-driven insights, competitive benchmarks, strategic analysis, and more.

In 2021, we went public on the New York Stock Exchange, and we haven’t stopped growing since!

We’re looking for a Data Engineer with a strong applied ML focus to join our R&D department!

Why is this role so important at Similarweb?

Our Retail Intelligence products help leading brands and retailers understand how their products, brands and categories perform online. Behind them is data collected from retailers and marketplaces around the world: product pages, brands and categories, each described differently by every site.

Your mission is to turn that data into a single, trusted view: classifying products into a unified taxonomy, normalizing brands and attributes, and matching the same entities across sources. And doing it at scale, across a catalog of more than a billion records that keeps growing and changing every day.

This is an applied ML role within data engineering. You’ll build with LLMs, agentic frameworks such as LangGraph, embeddings and classical ML, and ship them as production pipelines. It’s hands-on work, not research for its own sake, but it takes a real understanding of classification and NLP methods to choose the right tool for each problem and prove that it works.

So, what will you be doing all day?

Your daily responsibilities may include:

  • Designing and building LLM-powered and ML-based pipelines that classify, normalize, structure and match product, brand and category data
  • Building agentic workflows (LangGraph or similar) that automate complex data tasks end to end
  • Choosing the right approach for each problem (LLMs, embeddings, fine-tuned models, classical classifiers or rules), balancing accuracy, cost and latency
  • Scaling solutions to run efficiently over billions of records, using Spark, Databricks and our cloud infrastructure
  • Building evaluation frameworks: ground-truth datasets, labeling processes, quality metrics and ongoing monitoring
  • Taking solutions from POC to production, and owning them after launch
  • Working closely with Product to define requirements and shape the roadmap
  • Collaborating with data engineers, data scientists and other R&D teams on infrastructure and best practices

This is the perfect job for someone who:

  1. Holds a B.Sc. or M.Sc. in Computer Science, Data Science, Mathematics or another relevant field
  2. Has 4+ years of hands-on experience as a data engineer, ML engineer or data scientist, with solutions running in production
  3. Has strong Python skills and writes production-quality code
  4. Has hands-on experience building LLM-based applications in production (prompt engineering, structured outputs, RAG, embeddings, evaluation)
  5. Has worked with the modern LLM stack: LLM provider APIs (OpenAI, Anthropic, etc.), LangGraph or LangChain, Hugging Face and vector stores
  6. Has a solid grasp of text classification and NLP methods, both classical and modern, and knows when to use each
  7. Has experience processing large-scale data with Spark/PySpark, Databricks or similar, on AWS or another cloud
  8. Understands evaluation and data quality well: precision/recall trade-offs, building ground truth, error analysis
  9. Is pragmatic and delivery-focused, comfortable with ambiguity, and communicates clearly with Product and business stakeholders
  10. Has experience with taxonomies, entity resolution or product/e-commerce data (advantage)
  11. Has experience with fine-tuning or deploying open-source models (advantage)

At Similarweb, collaborating with our colleagues in-office creates a more connected, unified culture. Our best work is a product of our face-to-face collaboration, with the ability to work partially from home.

Why you’ll love being a Similarwebber:

You’ll actually love the product you work with: Our customers aren’t our only raving fans. When we asked our employees why they chose to come work at Similarweb, 99% of them said “the product.” Imagine how exciting your job is when you get to work with the most powerful digital intelligence platform in the world.

You’ll find a home for your big ideas: We encourage an open dialogue and empower employees to bring their ideas to the table. You’ll find the resources you need to take initiative and create meaningful change within the organization.

We offer competitive perks & benefits: We take your well-being seriously, and offer competitive compensation packages to all employees. We also put a strong emphasis on community, with regular team outings and happy hours.

You can grow your career in any direction you choose: Interested in becoming a VP or want to transition into a different department? Whether it’s Career Week, personalized coaching, or our ongoing learning solutions, you’ll find all the tools and opportunities you need to develop your career right here.

Diversity isn’t just a buzzword: People want to work in a place where they can be themselves. We strive to create a workplace that is reflective of the communities we serve, where everyone is empowered to bring their full, authentic selves to work. We are committed to inclusivity across race, gender, ethnicity, culture, sexual orientation, age, religion, spirituality, identity and experience. We believe our culture of equality and mutual respect also helps us better understand and serve our customers in a world that is becoming more global, more diverse, and more digital every day.

We will handle your application and information related to your application in accordance with the Applicant Privacy Policy available here: https://www.similarweb.com/corp/legal/applicant-privacy-policies/

We will handle your application and information related to your application in accordance with the Applicant Privacy Policy available here.

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Similarweb's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Similarweb's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Similarweb's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.