Skip to content

Open nowPosted 14 days ago

Distributed Systems Engineer III

Workable (global search)108,016 open roles

Where
Egypt
Work mode
Remote
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowDistributed Systems Engineer IIIWorkable (global search) · Egypt
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Workable (global search)'s own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

7.9% of postings close within 7 days. Measured by our own scanner across the market. Workable (global search) postings stay open a median of 7 days.

Share of postings closed within
  1. 1.6%1 day
  2. 3.6%3 days
  3. 7.9%7 days
  4. 14.9%14 days
  5. 34.0%30 days
This job: posted 14 days ago

Workable (global search) median: 7 days open

The posting

About Mozn

MOZN is a leading Enterprise AI company enabling organizations to make informed decisions in two critical domains: Financial Crime Prevention and Enterprise Knowledge Intelligence.

We’re a diverse, collaborative team of innovators united by a shared purpose: to build AI that delivers tangible business value, builds trust, and empowers people and organizations with augmented intelligence. Our culture is built on the relentless pursuit of excellence and meaningful impact.

If you’re passionate about working alongside exceptional talent on world-class AI, and you want the autonomy and runway to do the best work of your career, join us in shaping the future of intelligent enterprises.

About the role

We are looking for a highly motivated Distributed Systems Engineer III to join our Cloud Platform Engineering team.

This role focuses on building and operating reliable, scalable, and resilient cloud platforms for distributed and data-intensive workloads. You will work across Kubernetes, cloud infrastructure, messaging systems, databases, automation, and platform reliability.

The role requires strong hands-on experience with Kafka, Kubernetes, and at least one relational database such as MySQL or PostgreSQL, along with a solid understanding of distributed-systems fundamentals.

As our platform evolves, you will also contribute to AI and data infrastructure, helping build the underlying platform capabilities required to run data-intensive and AI

What you'll do

Cloud Platform & Distributed Systems

  • Build, operate, and continuously improve production cloud-native platforms running distributed workloads.
  • Work hands-on with Kubernetes, including upgrades, node pools, workload lifecycle, troubleshooting, and platform operations.
  • Deploy and manage workloads using ArgoCD, Helm, GitOps, Terraform, and automation.
  • Design and operate systems with a focus on scalability, availability, resilience, performance, and operational simplicity.
  • Troubleshoot complex issues across Kubernetes, cloud infrastructure, networking, storage, applications, and distributed services.

Kafka & Data Infrastructure

  • Operate and troubleshoot Apache Kafka in production across high-throughput and distributed workloads.
  • Work with topics, partitions, replication, consumer groups, retention, throughput, latency, and failure recovery.
  • Integrate Kafka with databases and applications using technologies such as Kafka Connect, Debezium, or similar CDC/event-streaming platforms.
  • Operate and troubleshoot MySQL and/or PostgreSQL, including replication, high availability, backup, recovery, performance, and migrations.
  • Support data-intensive workloads and analytical platforms such as StarRocks, ClickHouse, Apache Doris, or similar technologies.

Reliability, DR & Multi-Tenant Platforms

  • Design and operate platforms that remain resilient across node, service, zone, and infrastructure failures.
  • Implement and validate backup, recovery, disaster recovery, failover, and business-continuity capabilities.
  • Understand the fundamentals of multi-zone, multi-region, and active-active architectures and apply them where appropriate.
  • Build platforms that support multiple tenants and workloads, with appropriate isolation, scalability, resource management, and reliability.
  • Participate in DR exercises, failure simulations, migrations, and other resilience initiatives.
  • Understand distributed-system trade-offs involving replication, consistency, availability, partitioning, fault tolerance, latency, and throughput.

Automation, Observability & Operations

  • Automate infrastructure and platform lifecycle operations using Terraform, Python, Bash, Go, or similar technologies.
  • Build reliable deployment and GitOps workflows and reduce manual operational effort.
  • Implement effective monitoring, logging, alerting, and observability for distributed workloads.
  • Participate in production incident response, root-cause analysis, and long-term reliability improvements.

AI & Emerging Platform Infrastructure

  • Contribute to the infrastructure needed to support AI, machine-learning, and data-intensive workloads.
  • Help evolve cloud and Kubernetes platforms to support AI workloads, data pipelines, model-serving infrastructure, and associated platform services.
  • Understand the infrastructure requirements around compute, GPUs, networking, storage, data movement, observability, and workload isolation for AI platforms.
  • Work with engineering teams to build reusable platform capabilities that enable AI and data workloads to run reliably at scale.
  • Stay current with emerging infrastructure patterns across AI platforms, distributed data systems, and cloud-native technologies.

Requirements

  • 4–7 years of experience in Platform Engineering, Infrastructure Engineering, Distributed Systems, SRE, Backend Engineering, Data Infrastructure, or a related field.
  • Strong production experience with Apache Kafka — mandatory.
  • Hands-on production experience with Kubernetes — mandatory.
  • Strong experience with at least one of MySQL or PostgreSQL — mandatory.
  • Solid understanding of distributed-systems fundamentals including replication, partitioning, consistency, availability, fault tolerance, scalability, and failure recovery.
  • Experience with ArgoCD/GitOps and infrastructure-as-code such as Terraform.
  • Experience operating workloads on a public cloud such as GCP, OCI, AWS, or Azure.
  • Strong production troubleshooting and incident-resolution skills.
  • Experience with automation or scripting using Python, Bash, Go, Java, or similar languages.
  • Understanding of high-availability, disaster-recovery, and multi-tenant architecture fundamentals.
  • Experience with observability and operational tooling such as Prometheus, Grafana, OpenSearch/ELK, LGTM, or equivalent.
  • Strong understanding of infrastructure and networking fundamentals in cloud-native environments.

Good to Have

  • Experience with Kafka Connect, Debezium, Kafka Streams, or CDC platforms.
  • Experience with distributed analytical databases such as StarRocks, ClickHouse, Apache Doris, or similar.
  • Experience with Flink, Spark, or other distributed data-processing systems.
  • Experience executing large-scale data, database, application, or infrastructure migrations.
  • Experience with active-active, multi-zone, or multi-region systems.
  • Experience operating stateful workloads on Kubernetes.
  • Experience supporting AI/ML infrastructure or GPU-based workloads.
  • Experience with cloud networking, service mesh, ingress, load balancing, or storage platforms.
  • Contributions to Kubernetes, Kafka, distributed-systems, or other open-source infrastructure projects.

Benefits

  • You will be at the forefront of an exciting time for the Middle East, joining a high-growth rocket-ship in an exciting space.
  • You will be given a lot of responsibility and trust. We believe that the best results come when the people responsible for a function are given the freedom to do what they think is best.
  • The fundamentals will be taken care of: competitive compensation, top-tier health insurance, and an enabling culture so that you can focus on what you do best
  • You will enjoy a fun and dynamic workplace working alongside some of the greatest minds in AI.
  • We believe strength lies in difference, embracing all for who they are and empowered to be the best version of themselves.
From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Workable (global search)'s own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Workable (global search)'s form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Workable (global search)'s answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.