Skip to content

Open nowPosted 13 hours ago

Senior Data Engineer - Data and Analytics

Pattern98 open roles

Where
Pune, India
Work mode
Hybrid
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowSenior Data Engineer - Data and AnalyticsPattern · Pune, India
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Pattern's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

8.0% of postings close within 7 days. Measured by our own scanner across the market. Pattern postings stay open a median of 25 days.

Share of postings closed within
  1. 1.6%1 day
  2. 3.5%3 days
  3. 8.0%7 days
  4. 15.0%14 days
  5. 34.1%30 days
This job: posted 13 hours ago

Pattern median: 25 days open

The posting

What makes this role different

Most data engineering ends at a table. Pattern's ad-tech output leaves the warehouse and spends a client's advertising budget within the hour. A silently wrong join or an unguarded backfill is a customer-facing incident, not a dashboard discrepancy - so correctness, idempotency and data-quality gating are the job, not paperwork after the job.

The system you'll work on

Roles and Responsibilities

  • Develop, deploy, and support automated, scalable batch data pipelines from a variety of sources into the lakehouse.
  • Own and extend Airflow orchestration for a multi-DAG, cross-triggered daily pipeline and a 15-minute action pipeline - including branching, parallel task groups, cross-DAG triggers, backfill and full-refresh paths, and safe reruns.
  • Write and tune large analytical SQL: multi-hundred-column joins, window functions, incremental merges, and the warehouse-sizing and query-profile work needed to keep a daily run inside its window and its budget.
  • Extend the feature store - add new features and labels, wire them through the join layer, and preserve the leakage and data-completeness conventions that make the models trainable.
  • Orchestrate model training and batch inference on SageMaker from Airflow: build training and scoring datasets, manage S3 and Parquet round-trips, containerized training images, instance sizing, and loading predictions and metrics back into the warehouse.
  • Develop and implement data auditing strategies and processes to ensure data quality - including blocking data-quality checks in front of outward writes - and set thresholds that catch bad data without needlessly halting live bidding.
  • Identify and resolve problems in large-scale data processing workflows; maintain pipeline processes and troubleshoot failures, including on-call triage when a run breaks before market open.
  • Guard the safety properties of an outward-writing system: idempotency, new-data detection, action validation and invalidation, and audit trails for every change pushed to marketplace.
  • Collaborate with data scientists, advertising strategists, and platform teams to specify data requirements and provide access to data.
  • Translate business and analytics requirements - ROAS targets, budget pacing, playbook rules, branded versus non-branded strategy - into a comprehensive data model and pipelines.
  • Foster data expertise and own data quality for assigned areas of ownership; work with data infrastructure to triage issues and drive to resolution.
  • Mentor and provide technical direction to other data engineers, and review their SQL and DAG changes.

What "basics of machine learning" means here

  • Build training and evaluation datasets correctly - train/test splits over time, holdout windows, and a working instinct for target leakage in rolling-window features.
  • Reason about class imbalance and resampling (many keyword-hours have no clicks), and about clamping or bounding predictions before they drive a bid.
  • Read regression metrics - MAE, RMSE, MAPE, WMAPE - plus feature importances, and tell “the model got worse” apart from “the upstream data got worse”.
  • Operate the model lifecycle: retraining cadence, hyperparameters as configuration, prediction and metric persistence, validation tables, and drift monitoring.
  • Understand how model outputs compose into a decision - here, predicted clicks, conversion rate, cost per click and basket revenue combining into an expected ROAS per bid, net of cannibalization.

Required qualifications

  • Bachelor's degree in Data Science, Data Analytics, Information Management, Computer Science, Information Technology, a related field, or equivalent professional experience.
  • 4+ years of overall professional experience.
  • 4+ years of hands-on experience with SQL and Python, including advanced SQL - window and analytic functions, complex joins, incremental merges, and query tuning.
  • 3+ years building production data pipelines on modern data architectures, with real ownership of scheduling, dependencies, retries and backfills, at scale and across many source systems.
  • 2+ years working with cloud data warehouses such as Snowflake, Redshift or BigQuery.
  • Production experience with a workflow orchestrator - Airflow strongly preferred - including debugging failed runs in a live system.
  • Experience orchestrating ML training and batch inference from a scheduler, on SageMaker or an equivalent platform.
  • Working knowledge of applied machine learning fundamentals as described above: dataset construction, leakage, evaluation metrics, and model lifecycle operations.
  • Comfort with AWS - at minimum S3 and IAM - and with columnar file formats.
  • Demonstrated ownership of data quality: testing, monitoring, alerting, and root-cause analysis on pipelines other people depend on.
  • Excellent software engineering and scripting practice - version control, code review, modular and reviewable changes.
  • Strong communication skills, in both presentation and comprehension, with the aptitude for cross-collaboration across data management, data science and analytics domains.
  • Ability to lead and mentor a team of data engineers.

Preferred Qualification

  • Experience with digital advertising, bidding or auction systems - Amazon Ads, Google Ads, or a demand-side platform.
  • Advanced Snowflake - streams and tasks, stored procedures, UDFs, clustering, cost and performance tuning.
  • Experience with time-series data and forecasting, and with hourly or day-parted grains.
  • Background in big data, non-relational databases, machine learning or data mining.
  • Experience with data-quality frameworks such as Soda, Great Expectations or dbt tests.
  • Experience with open-source and distributed data platforms: Spark, Hive, Trino/Presto, Cassandra, DynamoDB or Elasticsearch.
  • Broader cloud experience: SNS, SQS, SES, Lambda, Glue, ECR and containerized workloads.
  • Expertise in data governance.
  • Experience working productively with AI coding agents on a large existing codebase.

Your First 90 days

  • Days 1-30 - Read the pipeline end to end and shadow a daily run. Ship small SQL and DAG fixes, take your first on-call triage with support, and be able to explain how a bid becomes an edit on marketplace.
  • Days 31-60 - Own a stage. Add features to the feature store and wire them through, tune a slow task that threatens the run window, and add or re-threshold a data-quality check that catches something real.
  • Days 61-90 - Lead a change that spans the pipeline and the action layer - a new signal, a new playbook rule, or a reliability improvement - with the tests, monitoring and rollback story that make it safe to leave running.
  • The company is a rocket ship experiencing phenomenal growth
  • We have tailwinds and a long runway; we're barely scratching the surface
  • We have big opportunities that will get you energized and excited
  • Great benefits including time off, insurance, competitive pay
From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Pattern's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Pattern's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Pattern's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.