Skip to content

Open nowPosted 64 days ago

Senior Data Engineer

Simulmedia6 open roles

Where
Lviv, Kyiv
Work mode
Remote
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowSenior Data EngineerSimulmedia · Lviv, Kyiv
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Simulmedia's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

7.8% of postings close within 7 days. Measured by our own scanner across the market.

Share of postings closed within
  1. 1.6%1 day
  2. 3.4%3 days
  3. 7.8%7 days
  4. 14.3%14 days
  5. 33.7%30 days
This job: posted 64 days ago

The posting

Simulmedia is looking for an experienced and dynamic Data Engineer with a curious and creative mindset to join our Data Services team. The ideal candidate will have a strong background in Python, SQL and large-scale data pipelines. This is an opportunity to join a team of amazing engineers, data scientists, product managers and designers who are obsessed with building the most advanced TV and streaming advertising platform in the market. As a Data Engineer you will design, build and operate the data platform that powers the company: pipelines that ingest and transform very large datasets from external data partners, data models that the whole company queries, and services that make that data available to internal products. You will work on a team that empowers the other teams to use our huge amount of data efficiently. Using a large variety of technologies and tools, you will solve complicated technical problems and build solutions to make our pipelines robust and fault tolerant and our data easily accessible throughout the company.

Location: Ukraine is mandatory. Our offices are located in Kyiv and Lviv. Teams are located in Kyiv and Lviv and primarily work remotely with occasional offline meetings.

Responsibilities:

  • Design and build batch data pipelines that ingest, validate and transform multi-billion-row datasets from external data providers and internal systems
  • Model complex real-world data: dimensional models, reference data, and temporal data whose attributes change over time (e.g. slowly changing dimensions), and evolve those models safely as upstream sources change their schemas and semantics
  • Develop and operate workloads on our lakehouse platform (Databricks / Spark / Delta) and our data warehouse (Redshift), including migrating existing pipelines from the warehouse to the lakehouse
  • Orchestrate pipelines with Airflow: scheduling, dependencies, retries, backfills and alerting
  • Prove correctness, not just completion: design parity checks and reconciliation queries when replacing an existing pipeline, run large historical backfills, and investigate data discrepancies down to the row level
  • Build and maintain Python services and REST APIs that serve data to internal products
  • Optimize for performance and cost: query tuning, table design, workload management and right-sizing compute
  • Own what you ship: monitor production pipelines, participate in incident triage and root-cause analysis, and harden systems so the same failure does not happen twice
  • Collaborate cross-functionally with product managers, data scientists and stakeholders across the company to deliver on product roadmap
  • Work within an Agile team that releases cutting-edge new features regularly
  • Take a high degree of ownership and freedom to experiment with new technologies to improve our software

Qualifications:

  • Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
  • 7+ years of work experience as a data engineer
  • Proficiency in Python and using it as the primary development language in recent years
  • Expert-level SQL: comfortable writing, reading and tuning complex analytical queries against very large tables, and debugging why two result sets disagree
  • Hands-on experience with a distributed data processing platform (Spark/Databricks strongly preferred; EMR, Snowflake or BigQuery also relevant) and with a columnar data warehouse (Redshift, Snowflake, BigQuery, ClickHouse, etc)
  • Ability to design complex data models: normalized, dimensional and temporal (slowly changing dimensions, effective-dated records, point-in-time correctness)
  • Experience with workflow orchestration tools (Airflow or similar): building DAGs, managing dependencies and running backfills
  • Experience integrating third-party data feeds: handling schema drift, late or missing deliveries, vendor data-quality defects and versioned reference data
  • Experience building REST services in Python (FastAPI, Flask, etc)
  • Experience developing, maintaining, and debugging problems in large server-side code bases
  • Working knowledge of AWS (S3, IAM, ECS or similar compute) and Docker
  • Good knowledge of engineering best practices and testing (unit test, integration test, code review, CI/CD)
  • The desire to take a high level of ownership of the things you work on
  • Ability to learn new things quickly, maintain a high bar for quality, and be pragmatic
  • Must be able to communicate with U.S based teams
  • Experience with Delta Lake / medallion lakehouse architectures is a plus
  • Experience migrating legacy pipelines between platforms with strict parity requirements is a plus
  • Experience with advertising, media or measurement industry data is a plus
  • Ability to communicate effectively with the U.S.-based teams and work 11:00 AM — 8:00 PM EEST (11:00 - 20:00).

Our Tech Stack:

  • Almost everything we run is on AWS (S3, ECS, EMR, RDS and more)
  • Python is our primary language; SQL is everywhere
  • Databricks (Spark, Delta Lake) is our lakehouse platform; Redshift and Postgres are our warehouses and operational databases
  • Airflow orchestrates our pipelines
  • Docker for packaging; GitHub Actions and Jenkins for CI/CD
  • Grafana, Sentry and OpenSearch for observability
  • Datasets measured in billions of rows

Interview Process:

  1. Pre-screening (30 mins)
  2. Technical Interview (1h)
  3. System Design (1.5h)
  4. Product Interview (30 mins)
  5. Offer
From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Simulmedia's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Simulmedia's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Simulmedia's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.