The posting
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability Engineering Lead based in the United States.
This leadership role is responsible for building reliable, scalable, and secure cloud platforms while developing a high-performing SRE team. You will lead engineers, establish priorities, and drive initiatives that improve service reliability, resilience, automation, and operational efficiency. The role combines people leadership with hands-on technical direction across modern cloud and infrastructure environments. You will work closely with Development, Security, Product, and other engineering teams to strengthen platform performance and incident response. A key focus will be reducing operational toil through automation, observability, infrastructure as code, and self-healing capabilities. You will also guide post-incident reviews, root-cause analysis, and continuous improvement efforts across services and infrastructure. This is an opportunity to shape SRE practices at scale while supporting engineers in their technical and professional growth.
Accountabilities
- Lead, mentor, and develop a small to medium-sized team of Site Reliability Engineers through regular 1:1s, performance reviews, career planning, and ongoing coaching.
- Own hiring, onboarding, team capacity, resourcing, and workforce planning decisions to ensure the team can effectively support business and platform priorities.
- Establish team objectives, prioritize the engineering backlog, coordinate planning, and ensure projects and operational tasks remain aligned with reliability goals.
- Lead reliability initiatives across infrastructure and services, improving availability, scalability, resilience, security, and operational performance.
- Drive incident response activities and facilitate blameless post-incident reviews, ensuring timely root-cause analyses and actionable follow-up.
- Partner with Development, Security, Product, and other engineering teams to resolve cross-functional issues and strengthen collaboration.
- Champion automation and operational excellence by reducing manual work, eliminating recurring toil, and introducing self-healing systems and infrastructure automation.
- Support the design and evolution of scalable, secure, cloud-native environments and continuously identify opportunities to improve performance, reliability, and cost efficiency.
- Demonstrated experience in SRE, DevOps, infrastructure engineering, or a related discipline, including experience leading engineering teams.
- Expert-level knowledge of Kubernetes, including cluster architecture, upgrades, autoscaling, security hardening, and large-scale troubleshooting.
- Advanced experience with Terraform, including modular infrastructure-as-code design, state management, multi-environment provisioning, and policy-as-code.
- Deep knowledge of Azure Cloud services, including compute, networking, identity and access management, storage, and cost optimization.
- Experience designing and scaling CI/CD pipelines using GitHub Actions, release strategies, and automated rollback approaches.
- Strong knowledge of observability platforms such as Prometheus, Grafana, and OpenTelemetry, along with experience managing SLOs, SLAs, and error budgets.
- Strong automation capabilities and advanced proficiency in Python, Bash, and/or PowerShell for infrastructure tooling and operational automation.
- Deep understanding of networking fundamentals, including TCP/IP, DNS, load balancing, VPNs, and cloud-native networking.
- Proven experience leading incident response, conducting root-cause analysis, and implementing measurable reliability improvements.
- Strong people leadership, communication, prioritization, and problem-solving skills, with the ability to support engineers while coordinating effectively across technical teams.
- U.S. national base salary range of $118,300–$219,800, with geographic differentials potentially applying depending on location.
- Eligibility for an annual incentive bonus.
- Country- and location-specific employee benefits designed to support overall health and well-being.
- Support for an accessible and inclusive hiring process, including reasonable accommodations where required.
- Remote/home-based opportunities available in multiple U.S. locations, including Florida, Connecticut, New Jersey, New York, and Pennsylvania.
- Opportunity to lead a technically sophisticated SRE function focused on cloud platforms, automation, resilience, and operational excellence.
- Professional development opportunities through team leadership, cross-functional collaboration, and exposure to large-scale cloud environments.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1



