Skip to content

Open nowPosted 91 days ago

Senior Data Engineer - Gurugram

CLANX25 open roles

Where
Gurgaon (Gurugram), Haryāna, India
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowSenior Data Engineer - GurugramCLANX · Gurgaon (Gurugram), Haryāna, India
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on CLANX's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

7.8% of postings close within 7 days. Measured by our own scanner across the market. CLANX postings stay open a median of 1 days.

Share of postings closed within
  1. 1.6%1 day
  2. 3.4%3 days
  3. 7.8%7 days
  4. 14.3%14 days
  5. 33.7%30 days
This job: posted 91 days ago

CLANX median: 1 days open

The posting

Mechademy is hiring a Senior Data Engineer to build and scale reliable data platforms, pipelines, and models that power enterprise AI, machine learning, and analytics for industrial asset monitoring and predictive maintenance.

Company Details

Mechademy is an enterprise AI company building real-time monitoring, diagnostics, and predictive maintenance solutions for industrial equipment. The company serves clients across oil & gas, power generation, and LNG sectors through production-grade AI and physics-informed machine learning systems.

Website: https://mechademy.com/

Responsibilities

What You’ll Own

  1. Lakehouse Pipelines & Ingestion (35%)

Design and own batch ETL/ELT and CDC pipelines that bring sensor and operational data into the lakehouse, orchestrated in Dagster

Build for reliability: idempotent, incremental, backfill-safe pipelines with sane retry and failure handling, that still produce correct output when a worker is killed mid-run or a message is delivered twice

Onboard new client data sources: schema and tag mapping, time-series normalization, resampling, gap handling at scale

2. Modeling & Serving (25%)

Model raw data into well-structured, documented tables that downstream ML and analytics can trust

Build and maintain the datasets behind ML feature pipelines and the lakehouse layer powering self-serve analytics

Write performant Spark/PySpark and SQL; optimize partitioning, storage formats, and query cost

3. Data Quality & Reliability (10%)

Own data quality: validation, freshness/SLA monitoring, and observability so bad data is caught before it reaches consumers

Make the data layer debuggable: lineage, tests, and alerting that tell you what broke and where

Reason about failure modes across the whole path (queue, worker, orchestrator, database, object store) and design so that a partial failure leaves the system in a state you can recover from

4. Relational & Operational Data (30%)

Contribute to the schema, indexing, and query performance of the relational database the product runs on

Design tables and constraints so that correctness is enforced at the database layer, and diagnose slow queries from their plans

Own retention and the boundary between the operational database and the lakehouse: what stays, what moves, and how it gets there

What Success Looks Like

First 30 days: Productive in the codebase and orchestration layer. First pipeline change merged.

First 90 days: Independently shipping and owning pipelines. Onboarded at least one new data source end-to-end.

First 6 months: Owning a lakehouse data domain, its ingestion, models, and quality, that ML and analytics teams rely on you to drive.

Requirement

Must-Have

  1. 4+ years building production data pipelines: real systems with real consumers, not just one-off scripts
  2. Strong data engineering fundamentals: data modeling, batch vs. streaming, idempotency, incremental processing, partitioning.
  3. Expert SQL and strong Python: query optimization, window functions, clean production-quality code
  4. Relational database depth: you’ve designed schemas for a production PostgreSQL (or equivalent) system and understand normalization and when to break it, indexing strategies, transactions and isolation levels, locking, and how to read a query plan and fix the query
  5. Distributed systems fundamentals: at-least-once delivery and idempotent consumers, partitioning and its effect on ordering, consistency and durability trade-offs, retries, timeouts, and backpressure. You can explain what happens to in-flight work when a worker or a database node dies
  6. Hands-on with a distributed processing engine (Spark/PySpark or equivalent) on non-trivial data volumes
  7. Experience with an orchestrator (Dagster, Airflow, Prefect, or equivalent) and a cloud platform (AWS/Azure)
  8. Data-quality mindset: you build validation and monitoring into pipelines, not after something breaks

Strong Signals (Nice-to-Have)

  1. Time-series or high-frequency sensor data at scale
  2. TimescaleDB or another time-series database (hypertables, continuous aggregates, compression, retention policies)
  3. Warehouse/lakehouse modeling (Delta/Iceberg/Snowflake/Redshift or equivalent) and file-format/partition tuning (Parquet)
  4. CDC / database-replication pipelines
  5. Message brokers or task queues in production (Kafka, RabbitMQ, or equivalent)
  6. Building data for ML: feature pipelines, training datasets, serving consistency
  7. dbt or similar transformation/modeling frameworks
  8. Docker, Terraform/IaC, CI/CD for data
  9. IoT, energy, or industrial sector experience. Not required, but it compresses your ramp

Job Details

Gurugram - Hybrid (2–3 days on-site)

Interview Process

  • Technical Round (Python & SQL)
  • System Design Round
  • Culture Fit Round

Important Note

ClanX is a recruitment partner, helping Mechademy hire a Senior Data Engineer.

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against CLANX's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on CLANX's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    CLANX's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.