Skip to content

Open nowPosted 147 days ago

Senior DevOps / Platform Reliability Engineer

zingtree3 open roles

Where
East Coast - United States
Work mode
Remote
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowSenior DevOps / Platform Reliability Engineerzingtree · East Coast - United States
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on zingtree's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

7.8% of postings close within 7 days. Measured by our own scanner across the market.

Share of postings closed within
  1. 1.6%1 day
  2. 3.4%3 days
  3. 7.8%7 days
  4. 14.3%14 days
  5. 33.7%30 days
This job: posted 147 days ago

The posting

About Zingtree

Zingtree is the next-generation intelligent process automation platform reimagining customer experience operations for the world’s top support leaders. With 500+ customers, including Optum, Corpay, Sony, SharkNinja, and Allianz, we transform self-service, surface automation opportunities, and turn every agent into an expert.

The Role

We’re hiring a Senior DevOps / Platform Reliability Engineer to own the platform that powers our agentic CX product. You’ll build the CI/CD, infrastructure, and observability backbone that enables us to ship multi-agent systems safely to enterprise customers.

If you want to operate a production AI platform and use AI to help operate it, this role is for you.

In this role, you will collaborate with development, operations, and infrastructure teams to automate and streamline processes, build and maintain tools for deployment, monitoring, and operations, and troubleshoot issues across development and production environments.

What You'll Do

  • Own and evolve CI/CD pipelines using GitHub Actions and OIDC-based authentication for microservices and agentic workloads, with safe, fast, and reversible deployments.
  • Automate infrastructure provisioning using Infrastructure as Code (IaC) tools such as Terraform and CloudFormation.
  • Operate and scale our Kubernetes platform (EKS + Argo CD), including autoscaling, ingress, external-dns, cert-manager, External Secrets Operator, backups, runtime guardrails, and multi-tenant isolation for enterprise customers.
  • Manage the edge and network perimeter, including Cloudflare (CDN, WAF, Bot Management, DDoS protection, Zero Trust / Access), CloudFront, API Gateway, ALB/NLB, Route 53, and network security controls.
  • Operate the data and event tier, including Aurora MySQL, ElastiCache/Redis, S3, and MSK (Kafka), with responsibility for backups, point-in-time recovery (PITR), and multi-AZ disaster recovery aligned to defined RTO/RPO objectives.
  • Build and maintain Lambda workloads where event-driven or serverless architectures are the right fit.
  • Build observability as a product using Prometheus, Grafana, and OpenTelemetry, including telemetry for LLM and agentic systems such as token cost, tool-call latency, evaluation signals, and prompt/version tracking.
  • Strengthen our security and compliance posture for SOC 2 and HIPAA, including least-privilege IAM, SCPs, secrets management, SAST/DAST, dependency and container scanning, image signing, AWS Config, Security Hub, GuardDuty, Inspector, and evidence automation.
  • Drive FinOps initiatives, including tagging standards, Savings Plans and Reserved Instances, per-tenant and per-workload cost attribution, and LLM cost controls.
  • Build and evolve our AI-native DevOps capabilities (see section below).
  • Partner with engineering teams to define platform standards, service templates, deployment best practices, and operational SLOs.
  • Monitor system performance and ensure reliability, scalability, and security across infrastructure and services.
  • Collaborate with software engineering teams to support continuous integration and continuous delivery best practices.
  • Document infrastructure, deployment processes, and operational standards to support knowledge sharing across the team.

Agentic AI in DevOps

  • Design and operate auto-remediation agents for common production toil such as certificate rotation, noisy pods, infrastructure drift, and flaky CI pipelines, with human-in-the-loop (HITL) controls for any destructive or customer-impacting actions.
  • Use LLMs for incident triage and root cause analysis, including log and trace summarization, signal correlation, and first-draft postmortems that are always reviewed by humans.
  • Connect AI agents to internal systems through the Model Context Protocol (MCP), including GitHub, Jira, PagerDuty, AWS, Kubernetes, Terraform, and related platforms, using scoped credentials, audit logging, and allow-listed access.
  • Apply AI-driven observability techniques, including anomaly detection on metrics, LLM-based log clustering, and alert deduplication and summarization on top of Prometheus and OpenTelemetry.
  • Establish operational guardrails such as prompt/version pinning, evaluation frameworks for agent behavior, cost and rate-limit controls, policy-as-code (OPA/Conftest) for AI-generated infrastructure changes, and clearly defined blast-radius controls.
  • Define best practices for AI coding assistants such as GitHub Copilot, Claude, and Amazon Q in infrastructure repositories, including review workflows, prompt design, and restrictions on auto-merged changes.
  • Treat AI components as production systems with SLOs, observability, on-call readiness, runbooks, and rollback strategies for agents and prompts.

About You

  • 5+ years of experience in DevOps, SRE, or Platform Engineering operating production systems on AWS.
  • Strong experience with CI/CD pipelines and tools such as GitHub Actions, GitLab CI, Jenkins, or CircleCI.
  • Hands-on experience operating production EKS environments, including autoscaling, ingress, secrets management, and cluster upgrades.
  • Strong AWS networking experience, including multi-account VPC design, subnets, routing, security groups, NACLs, Route 53, ACM, and load balancers.
  • Deep experience with Terraform and GitHub Actions, ideally using OIDC-based cloud authentication.
  • Experience with Aurora/RDS MySQL, Redis (ElastiCache), and S3, including backups, PITR, migrations, and lifecycle management.
  • Strong observability experience using Prometheus, Grafana, and OpenTelemetry.
  • Experience operating Argo CD at scale.
  • Experience with Infrastructure as Code tools such as Terraform, CloudFormation, or Ansible.
  • Experience managing Cloudflare services including WAF, Bot Management, Rate Limiting, and Zero Trust / Access, along with CloudFront.
  • Experience operating Kafka/MSK at scale, including topics, consumer groups, and schema registries.
  • Experience with Lambda and event-driven architectures.
  • Comfortable working with Python, Bash, and Linux systems.
  • Strong understanding of security best practices across IAM, KMS, secrets management, networking, and software supply chain security.
  • Familiarity with vulnerability scanning and compliance tooling.
  • Experience operating LLM or ML workloads in production, including LiteLLM, Bedrock, pgvector, prompt caching, or evaluation systems.
  • Experience building or integrating MCP servers or deploying agent frameworks such as LangGraph or CrewAI in production environments.

How We Work

  • We bias toward automation over toil. If you do it twice, script it. If it pages twice, fix it.
  • We’re a small team with high ownership. You’ll help define standards, not just follow them.
  • Humans stay in the loop for anything risky. AI accelerates decision-making but does not replace judgment.
  • We value blameless incident reviews, documented decisions, and fast feedback loops.

What We Offer

  • Competitive compensation packages
  • Comprehensive health benefits: 100% of employee premiums covered 75%–80% of dependent premiums covered for most health, dental, and vision plans 401(k) plans to support retirement planning (no employer matching currently) Paid parental leave Unlimited PTO Flexible remote work from anywhere Up to $200/month co-working reimbursement Home office stipend: Up to $500 for home office setup $100/month for internet, phone, and related expenses

Zingtree Values

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against zingtree's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on zingtree's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    zingtree's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.