The posting
What makes this role different
Most data engineering ends at a table. Pattern's ad-tech output leaves the warehouse and spends a client's advertising budget within the hour. A silently wrong join or an unguarded backfill is a customer-facing incident, not a dashboard discrepancy - so correctness, idempotency and data-quality gating are the job, not paperwork after the job.
The system you'll work on
Roles and Responsibilities
- Develop, deploy, and support automated, scalable batch data pipelines from a variety of sources into the lakehouse.
- Own and extend Airflow orchestration for a multi-DAG, cross-triggered daily pipeline and a 15-minute action pipeline - including branching, parallel task groups, cross-DAG triggers, backfill and full-refresh paths, and safe reruns.
- Write and tune large analytical SQL: multi-hundred-column joins, window functions, incremental merges, and the warehouse-sizing and query-profile work needed to keep a daily run inside its window and its budget.
- Extend the feature store - add new features and labels, wire them through the join layer, and preserve the leakage and data-completeness conventions that make the models trainable.
- Orchestrate model training and batch inference on SageMaker from Airflow: build training and scoring datasets, manage S3 and Parquet round-trips, containerized training images, instance sizing, and loading predictions and metrics back into the warehouse.
- Develop and implement data auditing strategies and processes to ensure data quality - including blocking data-quality checks in front of outward writes - and set thresholds that catch bad data without needlessly halting live bidding.
- Identify and resolve problems in large-scale data processing workflows; maintain pipeline processes and troubleshoot failures, including on-call triage when a run breaks before market open.
- Guard the safety properties of an outward-writing system: idempotency, new-data detection, action validation and invalidation, and audit trails for every change pushed to marketplace.
- Collaborate with data scientists, advertising strategists, and platform teams to specify data requirements and provide access to data.
- Translate business and analytics requirements - ROAS targets, budget pacing, playbook rules, branded versus non-branded strategy - into a comprehensive data model and pipelines.
- Foster data expertise and own data quality for assigned areas of ownership; work with data infrastructure to triage issues and drive to resolution.
- Mentor and provide technical direction to other data engineers, and review their SQL and DAG changes.
What "basics of machine learning" means here
- Build training and evaluation datasets correctly - train/test splits over time, holdout windows, and a working instinct for target leakage in rolling-window features.
- Reason about class imbalance and resampling (many keyword-hours have no clicks), and about clamping or bounding predictions before they drive a bid.
- Read regression metrics - MAE, RMSE, MAPE, WMAPE - plus feature importances, and tell “the model got worse” apart from “the upstream data got worse”.
- Operate the model lifecycle: retraining cadence, hyperparameters as configuration, prediction and metric persistence, validation tables, and drift monitoring.
- Understand how model outputs compose into a decision - here, predicted clicks, conversion rate, cost per click and basket revenue combining into an expected ROAS per bid, net of cannibalization.
Required qualifications
- Bachelor's degree in Data Science, Data Analytics, Information Management, Computer Science, Information Technology, a related field, or equivalent professional experience.
- 4+ years of overall professional experience.
- 4+ years of hands-on experience with SQL and Python, including advanced SQL - window and analytic functions, complex joins, incremental merges, and query tuning.
- 3+ years building production data pipelines on modern data architectures, with real ownership of scheduling, dependencies, retries and backfills, at scale and across many source systems.
- 2+ years working with cloud data warehouses such as Snowflake, Redshift or BigQuery.
- Production experience with a workflow orchestrator - Airflow strongly preferred - including debugging failed runs in a live system.
- Experience orchestrating ML training and batch inference from a scheduler, on SageMaker or an equivalent platform.
- Working knowledge of applied machine learning fundamentals as described above: dataset construction, leakage, evaluation metrics, and model lifecycle operations.
- Comfort with AWS - at minimum S3 and IAM - and with columnar file formats.
- Demonstrated ownership of data quality: testing, monitoring, alerting, and root-cause analysis on pipelines other people depend on.
- Excellent software engineering and scripting practice - version control, code review, modular and reviewable changes.
- Strong communication skills, in both presentation and comprehension, with the aptitude for cross-collaboration across data management, data science and analytics domains.
- Ability to lead and mentor a team of data engineers.
Preferred Qualification
- Experience with digital advertising, bidding or auction systems - Amazon Ads, Google Ads, or a demand-side platform.
- Advanced Snowflake - streams and tasks, stored procedures, UDFs, clustering, cost and performance tuning.
- Experience with time-series data and forecasting, and with hourly or day-parted grains.
- Background in big data, non-relational databases, machine learning or data mining.
- Experience with data-quality frameworks such as Soda, Great Expectations or dbt tests.
- Experience with open-source and distributed data platforms: Spark, Hive, Trino/Presto, Cassandra, DynamoDB or Elasticsearch.
- Broader cloud experience: SNS, SQS, SES, Lambda, Glue, ECR and containerized workloads.
- Expertise in data governance.
- Experience working productively with AI coding agents on a large existing codebase.
Your First 90 days
- Days 1-30 - Read the pipeline end to end and shadow a daily run. Ship small SQL and DAG fixes, take your first on-call triage with support, and be able to explain how a bid becomes an edit on marketplace.
- Days 31-60 - Own a stage. Add features to the feature store and wire them through, tune a slow task that threatens the run window, and add or re-threshold a data-quality check that catches something real.
- Days 61-90 - Lead a change that spans the pipeline and the action layer - a new signal, a new playbook rule, or a reliability improvement - with the tests, monitoring and rollback story that make it safe to leave running.
- The company is a rocket ship experiencing phenomenal growth
- We have tailwinds and a long runway; we're barely scratching the surface
- We have big opportunities that will get you energized and excited
- Great benefits including time off, insurance, competitive pay



