Skip to content

Open nowPosted today

Senior Machine Learning Platform Engineer — Model Hosting & MLOps

HP794 open roles

Where
Spring Texas United States of America
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowSenior Machine Learning Platform Engineer — Model Hosting & MLOpsHP · Spring Texas United States of America
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on HP's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

8.1% of postings close within 7 days. Measured by our own scanner across the market. HP postings stay open a median of 23 days.

Share of postings closed within
  1. 1.7%1 day
  2. 3.6%3 days
  3. 8.1%7 days
  4. 15.1%14 days
  5. 34.0%30 days
This job: posted today

HP median: 23 days open

The posting

Senior Machine Learning Platform Engineer — Model Hosting & MLOps

Description -

Role summary

We are hiring a Senior Machine Learning Platform Engineer to build the infrastructure and engineering workflows that take custom models from development into reliable production use. This is a hands-on role for someone who has hosted LLMs on GPU infrastructure, developed the cloud platform around model serving, and built MLOps pipelines that make releases repeatable and observable.

You will work with model developers, application engineers, security, and cloud platform teams to turn a model artifact into a secure, scalable inference service. The work spans serving architecture, infrastructure as code, deployment automation, model lifecycle management, and production operations across AWS and Azure, with potential integration into on-premises environments. You will also help shape agent workflows that use hosted models, so serving choices support the needs of multi-step applications. Success means teams can deploy, update, monitor, and troubleshoot models through clear, reusable platform patterns.

What you will do

  • Design and build hosting for custom ML and AI models across AWS and Azure, with particular focus on GPU-backed LLM inference and real-time endpoints; support batch inference where appropriate.
  • Package models and their dependencies into reproducible serving workloads; choose and implement suitable managed services, containers, or Kubernetes-based patterns based on throughput, latency, security, reliability, and cost.
  • Provision and configure Kubernetes clusters or other suitable hosting infrastructure, including compute, storage, API access, identity and access controls, secrets, networking integration, observability, and environment configuration.
  • Help design secure connections and deployment patterns between cloud and on-premises environments as hosting needs evolve.
  • Develop infrastructure as code and deployment automation so model-hosting environments can be provisioned, reviewed, promoted, and maintained consistently.
  • Build MLOps workflows for model registration, versioning, validation, release, rollback, and retirement. Connect training or model preparation to deployment through automated pipelines and appropriate quality gates.
  • Partner with application teams to design and prototype agent workflows, including model and tool orchestration, state handling, failure recovery, and evaluation.
  • Translate agent workload patterns into hosting decisions about model selection, context length, concurrency, latency, cost, tool access, and end-to-end tracing.
  • Establish production monitoring for service health, latency, throughput, errors, GPU and other resource use, and model behavior. Share responsibility for diagnosing incidents and improving capacity, reliability, and cost.
  • Create reusable deployment templates, reference architectures, documentation, and onboarding paths that help other teams ship models safely.
  • Partner with model and application teams on practical tradeoffs such as online versus batch inference, managed versus self-hosted serving, scaling, evaluation, data handling, and operational ownership.

Required experience

  • Hands-on experience hosting LLM inference on GPU infrastructure in a production environment. You can explain what you personally built, how models reached production, and how you managed throughput, latency, utilization, reliability, and cost.
  • Experience building the surrounding model-serving platform for custom models, such as inference runtimes, deployment patterns, endpoint access, scaling, and operational tooling.
  • Strong software engineering skills, especially Python, with experience building services, automation, and maintainable production code.
  • Experience developing infrastructure for ML workloads across AWS and Azure, with deep hands-on delivery in at least one and practical ability to work in the other. You have used infrastructure as code such as Terraform or an equivalent tool.
  • Experience provisioning, configuring, and maintaining Kubernetes or another production hosting platform for containerized inference workloads.
  • Experience building or operating ML deployment pipelines with versioned artifacts, automated validation, CI/CD, environment promotion, and rollback.
  • Familiarity with LLM agent patterns, including model invocation, tool calls, and multi-step workflows, and how they affect serving capacity, reliability, and access controls.
  • Working knowledge of production concerns for inference services: scaling, latency, availability, logging and metrics, shared incident response, access control, and cost.
  • Ability to work across model development, application, platform, and security teams; turn ambiguous requirements into a working design; and document the resulting operational approach.

Helpful experience

  • Advanced GPU inference optimization, including capacity planning, batching, autoscaling, memory use, model loading, and performance tuning.
  • Serving generative or other compute-intensive custom models beyond LLMs; selecting and tuning inference servers and runtime configurations.
  • AWS SageMaker or EKS; Azure Machine Learning or AKS; or equivalent managed and self-hosted model platforms.
  • Hybrid or on-premises model hosting, including connectivity, security boundaries, hardware constraints, and operational handoff.
  • Model registries, experiment tracking, data or feature pipelines, scheduled retraining, model evaluation, and drift or quality monitoring.
  • Secure enterprise deployment patterns such as private networking, IAM/RBAC, secrets management, auditability, and handling sensitive data.
  • Building shared ML platform capabilities or self-service workflows used by multiple engineering teams.
  • Hands-on development of agent workflows, including tool integration, evaluation, tracing, or guardrails.

Who will thrive here

You are an engineer who has taken responsibility for what happens after a model is trained: how it is packaged, deployed, secured, scaled, observed, updated, and supported. You are comfortable writing code and infrastructure, investigating production failures, and making clear tradeoffs with partner teams.

Pay & Benefits

The pay range for this role is $147,050 to $230,850 USD annually with additional

opportunities for pay in the form of bonus and/or equity (applies to United

States of America candidates only). Pay varies by work location, job-related

knowledge, skills, and experience.

Benefits:

HP offers a comprehensive benefits package for this position, including:

  • Health insurance
  • Dental insurance
  • Vision insurance
  • Long term/short term disability insurance
  • Employee assistance program
  • Flexible spending account
  • Life insurance
  • Generous time off policies, including;
  • 4-12 weeks fully paid parental leave based on tenure
  • 11 paid holidays
  • Additional flexible paid vacation and sick leave
  • US benefits overview https://hpbenefits.ce.alight.com/

The compensation and benefits information is accurate as of the date of this

posting. The Company reserves the right to modify this information at any time,

with or without notice, subject to applicable law.

Job -

Software

Schedule -

Full time

Shift -

No shift premium (United States of America)

Travel -

No

Relocation -

No

Equal Opportunity Employer (EEO) -

HP, Inc. provides equal employment opportunity to all employees and prospective employees, without regard to race, color, religion, sex, national origin, ancestry, citizenship, sexual orientation, age, disability, or status as a protected veteran, marital status, familial status, physical or mental disability, medical condition, pregnancy, genetic predisposition or carrier status, uniformed service status, political affiliation or any other characteristic protected by applicable national, federal, state, and local law(s).

Please be assured that you will not be subject to any adverse treatment if you choose to disclose the information requested. This information is provided voluntarily. The information obtained will be kept in strict confidence.

For more information, review HP’s EEO Policy or read about your rights as an applicant under the law here: “Know Your Rights: Workplace Discrimination is Illegal"

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against HP's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on HP's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    HP's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.