Skip to content

Open nowPosted 45 days ago

AI Platform Operations Manager

stackinfra46 open roles

Pay
$128,260 – $146,018 a year
Where
Denver, CO
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowAI Platform Operations Managerstackinfra · Denver, CO
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on stackinfra's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

8.2% of postings close within 7 days. Measured by our own scanner across the market. stackinfra postings stay open a median of 7 days.

Share of postings closed within
  1. 1.9%1 day
  2. 3.8%3 days
  3. 8.2%7 days
  4. 15.2%14 days
  5. 34.1%30 days
This job: posted 45 days ago

stackinfra median: 7 days open

The posting

THE COMPANY:

STACK INFRASTRUCTURE (STACK) provides digital infrastructure to scale the world’s most innovative companies. We are an award-winning industry leader in building, owning, and operating highly efficient, cost-effective wholesale, colocation, and cloud data centers. Each of our national facilities meets or exceeds the highest industry standards in all operational categories of availability, security, connectivity, and physical resilience.

STACK offers the scale and geographic reach that rapidly growing hyperscale and enterprise companies need. The world runs on data. Data runs on STACK.

THE POSITION:

The DevOps Engineer, AI Platform is responsible for automating, deploying, and operating the infrastructure and delivery pipelines that support STACK’s enterprise AI platform on Azure. This is a hands-on engineering role focused on build and run — not oversight.

Reporting to Head of AI, Enterprise AI & Data Strategy org this individual owns the infrastructure-as-code, CI/CD, containerization, observability, and release automation that allow AI engineers and enterprise application teams to ship agentic AI solutions, RAG pipelines, and integration services reliably and repeatably. The role sits at the intersection of cloud infrastructure, platform engineering, and MLOps — turning platform architecture into automated, governed, observable, and cost-efficient environments that teams across the organization build on.

KEY RESPONSIBILITIES

Infrastructure Automation & Infrastructure as Code

  • Build, maintain, and version infrastructure-as-code modules for Azure environments using Terraform, Bicep, or ARM, including compute, networking, storage, identity, and AI platform resources.
  • Automate provisioning of AI platform components — Azure AI Foundry, Azure OpenAI Service, Azure AI Search, Cosmos DB, ADLS Gen2, and Databricks — as reusable, parameterized deployment patterns.
  • Maintain environment parity across development, test, and production, including configuration management, drift detection, and remediation.
  • Implement and enforce tagging, naming, and resource organization standards that support governance, chargeback, and lifecycle management.
  • Automate routine platform operations — patching, certificate rotation, key and secret rotation, backup validation, and disaster recovery testing.

CI/CD & Release Engineering

  • Design, build, and operate CI/CD pipelines in Azure DevOps or GitHub Actions for application code, infrastructure code, container images, and AI/agent deployments.
  • Implement automated build, test, security scanning, artifact management, and promotion gates across environments.
  • Establish branching strategies, code review standards, and release management practices in partnership with AI engineering and enterprise application teams.
  • Build deployment automation for agentic AI services, MCP (Model Context Protocol) servers, and integration workloads running on Azure Container Apps and Azure Kubernetes Service (AKS).
  • Support model and prompt release workflows — versioning, staged rollout, evaluation gates, and rollback procedures for LLM-based applications.

Container Platform & AI Workload Operations

  • Operate and tune AKS and Azure Container Apps, including cluster upgrades, node pool sizing, autoscaling, ingress, networking, and workload isolation.
  • Build and maintain container images, base image standards, and registry governance in Azure Container Registry.
  • Manage compute scheduling and scaling for AI workloads, including GPU-backed and inference-heavy workloads where required.
  • Implement resiliency patterns — health probes, retries, throttling, quota management, and failover — for AI endpoints and integration services.

Observability, Reliability & Incident Response

  • Instrument platform and AI services with logging, metrics, tracing, and alerting using Azure Monitor, Log Analytics, Application Insights, and equivalent open-source tooling.
  • Build dashboards and service-level indicators covering platform availability, latency, throughput, error rates, token consumption, and model endpoint performance.
  • Participate in on-call rotation, lead incident triage and resolution for platform issues, and drive root cause analysis and corrective actions.
  • Develop and maintain runbooks, operational documentation, and automated remediation for recurring issues.

Security, Governance & Cost Optimization

  • Implement DevSecOps practices — secrets management in Azure Key Vault, managed identity usage, least-privilege access, dependency and container vulnerability scanning, and policy-as-code.
  • Partner with Information Security to ensure pipelines and environments meet enterprise security, data residency, and compliance requirements.
  • Support Azure FinOps practices through cost visibility, rightsizing, reserved capacity, and automated controls on non-production and idle resources.
  • Maintain audit trails and change records for infrastructure and release activity.

Delivery & Cross-Functional Collaboration

  • Work directly with AI engineers, data engineers, and enterprise application teams to remove deployment friction and improve time-to-production for AI solutions.
  • Translate platform architecture and standards into automated, self-service capabilities that teams can consume without deep infrastructure knowledge.
  • Contribute to platform engineering standards, reference implementations, and internal documentation.
  • Provide technical escalation support for build, deployment, and environment issues.

THE DETAILS:

  • Location: Denver, CO
  • Travel: <10%
  • Benefits: Healthcare, Dental Care, Vision Insurance, Life Insurance, Paid Time Off, and Paid Leave Programs
  • Must be eligible to work in the United States
  • Must pass comprehensive background and drug screening

MUST-HAVE QUALIFICATIONS:

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related field, or equivalent practical experience.
  • 5+ years of hands-on experience in DevOps, site reliability engineering, or platform engineering roles with a strong delivery track record.
  • Strong proficiency in Infrastructure as Code — Terraform, Bicep, or ARM — including module design, state management, and reusable patterns.
  • Proven experience building and operating CI/CD pipelines in Azure DevOps or GitHub Actions.
  • Hands-on experience with containerization and orchestration — Docker, Azure Kubernetes Service (AKS), and Azure Container Apps or equivalent.
  • Solid working knowledge of Azure core services — compute, networking (VNet, NSG, Private Endpoints), storage, identity (Entra ID), and Key Vault.
  • Strong scripting and automation skills in Python, PowerShell, or Bash.
  • Experience with monitoring and observability tooling — Azure Monitor, Log Analytics, Application Insights, Prometheus, or Grafana.
  • Working knowledge of Git-based workflows, code review practices, and artifact/registry management.
  • Demonstrated ability to troubleshoot production issues across infrastructure, network, and application layers.

PREFERRED QUALIFICATIONS:

  • Microsoft Certified: DevOps Engineer Expert (AZ-400), Azure Administrator (AZ-104), or Certified Kubernetes Administrator (CKA).
  • Experience deploying and operating AI/ML workloads — model endpoints, RAG pipelines, vector databases, or agentic services in production.
  • Familiarity with MLOps tooling and practices — Azure Machine Learning, MLflow, Databricks, or equivalent model lifecycle platforms.
  • Experience deploying MCP (Model Context Protocol) servers or similar integration services connecting AI agents to enterprise systems.
  • Exposure to agentic AI frameworks such as Semantic Kernel, LangGraph, or AutoGen from a deployment and operations perspective.
  • Experience with GPU compute provisioning, quota management, and inference cost optimization.
  • Knowledge of FinOps frameworks and Azure cost optimization practices.
  • Experience integrating with enterprise systems such as Microsoft 365, Freshworks ITSM, Workday, NetSuite, or Procore.
  • Experience in data center, hyperscale, or infrastructure-intensive industry environments.

Compensation Range:

$128,260.00 - $146,017.59

THIS MIGHT BE RIGHT FOR YOU IF:

  • You are a strong communicator, you are persuasive and clear, blending analytics with experience in decision-making.
  • You do not get flustered easily. You can juggle multiple priorities while balancing urgent requests with shifting timelines and deliverables.
  • You are a team builder. You take the time to understand and develop the strengths of your resources while formulating long-term plans for the growth and success of the team.
  • You are naturally curious and driven toward continual improvement. While you celebrate your successes, you take time to review and analyze campaigns for future learning.

WHY STACK?

  • We offer a competitive compensation package with strong benefits, including medical, dental, and vision insurance, a 401K program, flexible spending accounts – even a cell phone subsidy.
  • We foster a culture of appreciation, including peer-to-peer recognition and rewards programs.
  • Fun is part of our DNA, with events, game nights, happy hours, and barbecues.
  • We’re growing – this is a great time to join and make an impact!

STACK is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity and expression, age, national origin, mental or physical disability, genetic information, veteran status, or any other status protected by federal, state, or local law

Note to external agencies: We are not accepting any blind submissions or resumes/cvs from recruitment agencies. Any candidates sent to STACK Infrastructure, Inc. will not be accepted or considered as a submission without a signed agreement in place. Fees will not be paid in the event a candidate submitted by a recruiter without an agreement in place is hired; such resumes will be deemed the sole property of STACK Infrastructure, Inc.

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against stackinfra's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on stackinfra's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    stackinfra's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.