Skip to content

Open nowPosted 6 days ago

Senior AI Data Engineer

Strategic Systems International10 open roles

Where
Lahore, Pakistan
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowSenior AI Data EngineerStrategic Systems International · Lahore, Pakistan
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Strategic Systems International's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

8.2% of postings close within 7 days. Measured by our own scanner across the market.

Share of postings closed within
  1. 1.8%1 day
  2. 3.6%3 days
  3. 8.2%7 days
  4. 15.2%14 days
  5. 34.0%30 days
This job: posted 6 days ago

The posting

Senior AI Data Engineer

Job Summary

We are seeking a Senior AI Data Engineer to design and operate the data foundation that our AI systems depend on. This role owns the movement, modeling, and quality of data from source systems through the warehouse and into the retrieval and feature layers that power LLM pipelines, agentic workflows, and analytical products.

The ideal candidate is a rigorous software engineer first and a data specialist second: someone who models a warehouse deliberately, writes production Python that other engineers can extend, and treats pipelines as versioned, tested, observable software rather than scripts.

This role partners closely with the AI/ML Data Scientist, who owns model behavior and retrieval strategy.

The boundary: you own the pipeline, the schema, and the guarantees; they own the algorithm, the prompt, and the evaluation.

Key Responsibilities

Data Warehousing & Dimensional Modeling

  • Design and evolve dimensional models using Kimball methodology - star schemas, conformed dimensions, an enterprise bus matrix, and explicit fact-table grain declarations.
  • Implement transaction, periodic snapshot, and accumulate snapshot fact tables as the business process warrants and defend the choice of grain.
  • Manage slowly changing dimensions (Type 1 / 2 / 3, and hybrid variants) with correct effective-dating, surrogate key strategy, and late-arriving dimension handling.
  • Model semi-structured and unstructured sources - documents, transcripts, event streams-into queryable structures without discarding provenance.
  • Maintain a governed semantic layer so that both human analysts and agentic consumers resolve the same metric to the same number.

Pipeline & ETL/ELT Engineering

  • Build and operate batch and streaming pipelines with orchestration frameworks such as Airflow, Prefect, or Dagster, including backfill, replay, and idempotent-retry semantics.
  • Implement ELT transformation layers with tested, documented, version-controlled SQL.
  • Own data contracts between producing and consuming systems: schema evolution, compatibility rules, and breaking-change procedure.
  • Instrument pipelines for observability - freshness, volume, distribution, and schema-drift checks - with alerting that distinguishes a real incident from ordinary variance.
  • Maintain end-to-end lineage from source record to warehouse fact to retrieved chunk.

Data Engineering for AI

  • Build embedding generation and refresh pipelines: chunk materialization, embedding jobs, incremental re-embedding on source change, and index lifecycle management.
  • Operate vector stores (Pinecone, Weaviate, Chroma, Milvus, pgvector) as production data systems - capacity, index build strategy, upsert/delete correctness, and staleness SLAs.
  • Build extraction pipelines over unstructured sources using OCR, document parsers, and vision-language model outputs, treating extraction confidence as a first-class column.
  • Implement the preprocessing and feature pipelines that back model training and inference, with train/serve consistency as a design requirement.
  • Expose data to AI services and applications through well-specified FastAPI or gRPC interfaces.

Engineering Craft

  • Apply SOLID principles and domain-driven design: bounded contexts that mirror the business domains, ubiquitous language shared with stakeholders, aggregates and repositories that keep domain logic out of transport and persistence layers.
  • Maintain meaningful test coverage (unit, contract, and data-quality assertions) and treat an untested pipeline as an unfinished one.
  • Own CI/CD, containerization, and environment promotion for data services.
  • Contribute to code review, architectural decision records, and internal standards.

Governance & Cost

  • Implement access control, PII handling, retention, and audit requirements at the data layer.
  • Manage warehouse and pipeline cost: partitioning, clustering, materialization strategy, and storage tiering.

Collaboration

  • Translate business problems into data models with product and client stakeholders.
  • Mentor junior and mid-level data engineers.

Required Qualifications

5-10+ years in software engineering or data engineering, with substantial time in production data platform work.

Data warehousing: demonstrable command of Kimball dimensional modeling - not just familiarity with the vocabulary, but the judgment to choose a grain, resolve a many-to-many relationship, and know when to denormalize. Working knowledge of alternative approaches (Data Vault, One Big Table, Inman) and the tradeoffs against Kimball.

SQL: expert-level - window functions, CTEs, query plan reading, and performance tuning on a columnar warehouse.

Python: expert-level, production-grade - typing, packaging, dependency management, testing.

Design: SOLID and domain-driven design applied in real systems; with examples you can walk through.

Orchestration: Airflow, Prefect, Dagster, or equivalent, in production.

Cloud: expert-level on AWS, Azure, or GCP - storage, compute, IAM, networking, and cost management.

Platform: containerization, Kubernetes (EKS/AKS/GKE), and CI/CD.

Experience with Lakehouse table formats (Iceberg, Delta Lake, Hudi) and their maintenance characteristics: compaction, snapshot expiry, schema and partition evolution.

Preferred Qualifications

Experience building the data layer beneath production RAG systems, including hybrid search infrastructure and index freshness guarantees.

Streaming systems: Kafka, Kinesis, Flink, or Spark Structured Streaming.

dbt or an equivalent transformation and testing framework.

Data quality tooling (Great Expectations, Soda, or similar) and catalog/lineage platforms.

Familiarity with the model-facing side of the stack - MLflow, Weights & Biases, feature stores - sufficient to collaborate credibly with data scientists.

Working knowledge of a second language: TypeScript, Java, Go, Scala, or Rust.

Experience with AI security, governance, and compliance frameworks.

Open-source contributions to data or AI infrastructure projects.

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Strategic Systems International's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Strategic Systems International's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Strategic Systems International's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.