Skip to content

Senior/Staff AI Engineer

DDN

Remote - California

WHAT YOU’LL DO

- Build and optimize LLM serving and inference systems for production environments

- Improve performance across GPU and CPU pathways

- Work on KV cache, memory, storage, and throughput bottlenecks

- Design and scale systems that support RAG and retrieval-heavy AI workloads

- Contribute to infrastructure where storage architecture and systems efficiency materially affect AI performance

- Solve engineering problems at the intersection of AI, high-performance systems, and distributed infrastructure

WHAT WE’RE LOOKING FOR

- An engineer who has spent meaningful time building or optimizing production AI systems, not just experimenting with models

- Someone who understands how inference performance is shaped by the interaction between compute, memory, storage, and serving architecture

- Deep hands-on experience working close to the systems layer — for example, improving how workloads run across GPU and CPU resources, reducing bottlenecks, or tuning infrastructure for better throughput and latency

- Evidence of real ownership in areas like model serving, retrieval, caching, storage, or distributed performance, rather than purely application-layer AI work

- The ability to move comfortably between architecture decisions and hands-on implementation, especially in environments where efficiency and scale matter

- A background that suggests you can operate in technically demanding environments, whether that comes from AI infrastructure, high-performance systems, storage platforms, or adjacent distributed systems work

- PhD preferred, but far less important than having built serious systems in the real world

WHY THIS ROLE IS COMPELLING

- This is not a “prompt engineering” job.

- This is not an “AI wrapper” job.

- This is not a generic backend role with AI sprinkled on top.

- This is a chance to work on the infrastructure that determines whether modern AI systems are fast, scalable, efficient, and commercially viable.

- If you want to work on the real mechanics of AI performance — serving, retrieval, compute efficiency, memory behavior, storage architecture, and inference at scale — this is where that work happens.

WHO WILL LOVE THIS ROLE

- Engineers who enjoy deep systems problems

- Builders who care about performance, scale, and architecture

- People who want to work where AI meets infrastructure

- Candidates who would rather solve hard technical bottlenecks than ship surface-level AI features

WHO SHOULD NOT APPLY

This role is not for:

- Purely academic researchers without meaningful production ownership

- Generic software engineers without clear AI systems or inference depth

- Candidates focused mainly on prompt engineering or lightweight application integrations

- MLOps generalists who have not worked deeply on serving, storage, or performance-critical AI systems

-

Seen 19 days ago · DDN postings close after a median of 13 days.

Original posting on DDN's site ↗

Posting text belongs to the employer. Removal requests: contact us.

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

One job at a time

One posting. One CV. $25.

Pick the job you actually want and we write for it.

Get my CV for this job