Skip to content

Open nowPosted 7 hours ago

Forward Deployed Engineer - AI Inference

FriendliAI18 open roles

Where
San Francisco
Work mode
Hybrid
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowForward Deployed Engineer - AI InferenceFriendliAI · San Francisco
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on FriendliAI's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

8.2% of postings close within 7 days. Measured by our own scanner across the market. FriendliAI postings stay open a median of 49 days.

Share of postings closed within
  1. 1.8%1 day
  2. 3.6%3 days
  3. 8.2%7 days
  4. 15.2%14 days
  5. 34.0%30 days
This job: posted 7 hours ago

FriendliAI median: 49 days open

The posting

About the job

FriendliAI is seeking a Forward Deployed Engineer to assist enterprises in deploying, scaling, and operating generative and agentic AI workloads on FriendliAI infrastructure. You will work directly with customers to solve and implement production-grade applications using our products, such as Serverless Endpoints, Dedicated Endpoints, or Container.

Friendli Container is our service that allows customers to download our inference engine as Docker images and deploy it in their chosen environment, such as private clouds or on-premises. Our Friendli Container can be adopted directly to AWS EKS clusters using our EKS add-on product.

You will work directly on our customers’ projects, collaborating with their engineering teams to solve AI inference challenges like scaling, orchestration, and monitoring. This is a hands-on, customer-embedded role. If you have worked in DevOps, platform engineering, or SRE for AI applications, this is your ideal position.

Key Responsibilities

- Design and implement large-scale deployment architectures for LLM and multimodal inference

- Deploy and manage containerized workloads across Kubernetes clusters

- Diagnose production issues, such as performance bottlenecks, and implement temporary fixes as needed

- Collaborate with customers’ DevOps teams to integrate FriendliAI’s infrastructure into their CI/CD workflows

- Develop scripts, Helm charts, and Terraform modules that simplify repeated deployments

- Contribute field insights to shape our platform reliability, observability, and scaling strategies

- Lead workshops, technical sessions, or webinars to help customers master infrastructure best practices.

Qualifications

- 3+ years of experience in cloud infrastructure, DevOps, or reliability engineering

- Bachelor’s or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent

- Proficiency with Kubernetes, Docker, Terraform, and Helm

- Strong foundation in distributed systems, networking, and performance tuning

- Experience with GPU-based computing and generative AI model serving workloads

- Strong technical background in backend systems or AI tooling

- Experience operating workloads on AWS, GCP, or OCI

- Excellent problem-solving and debugging skills in real-world environments

Preferred Experience

- Experience deploying large models (LLMs, diffusion models) on GPUs or clusters

- Familiarity with inference frameworks (Triton, vLLM, TensorRT, DeepSpeed-Inference)

- Familiarity with observability stacks (Prometheus, Grafana, Loki, ELK, OTEL)

- Understanding of networking security and compliance frameworks (e.g., SOC 2)

- Experience supporting on-prem or hybrid-cloud deployments

Benefits

- A front-row seat to the generative AI infrastructure revolution

- Competitive compensation and benefits package

- Daily lunch and dinner provided; unlimited snacks and beverages

- Health check-up and top-tier hardware support

- Flexible working hours and a highly collaborative environment

About us

FriendliAI is the fastest inference cloud for agents, built to run frontier open-weight models in production at scale. It delivers up to 7x faster output token speed, up to 90% lower inference costs, and 99.99% uptime across the most demanding agent workloads — long-context inference, real-time streaming, and accurate tool calling.

We are a small, fast-moving team doing work that matters at one of the most exciting moments in the history of technology. With our world-class inference stack, we are building the platform teams can actually rely on.

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against FriendliAI's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on FriendliAI's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    FriendliAI's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.