The posting
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site-Reliability Engineer based in United States.
This is a senior infrastructure engineering role focused on the reliability, availability, and performance of Windows-based production environments. You will bridge development and operations to build highly available services and maintain strong operational standards across hybrid infrastructure. The role combines infrastructure automation, configuration management, observability, incident response, and continuous improvement. You will work closely with development teams to strengthen deployment and release processes while establishing measurable reliability objectives. The position offers an opportunity to work with modern cloud, automation, and monitoring technologies in a mission-critical environment. The role is remote and requires the ability to obtain the appropriate security clearance.
Accountabilities:
- Design, implement, and maintain scalable infrastructure using Infrastructure as Code (IaC) practices across production environments.
- Develop and maintain automation scripts using PowerShell, Python, Ruby, and other scripting languages for operating system provisioning, configuration management, and recurring operational tasks.
- Implement and manage configuration management solutions such as Terraform, Puppet, and/or Chef across hybrid infrastructure environments.
- Monitor system health, performance, reliability, and availability using established observability tools and operational best practices.
- Establish, maintain, and enforce Service Level Agreements (SLAs), Service Level Objectives (SLOs), and error budgets for production services.
- Participate in an on-call rotation, respond to production incidents, restore services rapidly, and conduct thorough root cause analysis.
- Partner with development teams to improve deployment pipelines, release processes, and overall service reliability.
- Create and maintain operational documentation, including procedures, runbooks, and architectural decisions.
- Lead or contribute to post-mortem reviews and implement corrective actions designed to prevent recurring incidents.
- Troubleshoot complex infrastructure and application issues across multiple technology layers while maintaining a strong focus on operational excellence.
- Bring at least 5 years of experience in Systems Administration, DevOps, Site-Reliability Engineering, or a closely related infrastructure role.
- Have strong hands-on expertise with Windows Server environments, including Windows Server 2016 or later, Active Directory, IIS, and Microsoft SQL.
- Demonstrate strong cloud infrastructure skills, with AWS experience preferred.
- Possess advanced scripting capabilities, including the development of reusable modules and integrations with REST APIs.
- Have hands-on experience with Terraform for infrastructure provisioning and Puppet or Chef for configuration management.
- Be experienced with monitoring and observability platforms such as Prometheus, Grafana, Datadog, or New Relic.
- Have a solid understanding of networking fundamentals, including DNS, TCP/IP, load balancing, and VPN technologies.
- Demonstrate strong analytical and problem-solving skills, with the ability to troubleshoot complex issues spanning multiple technology layers.
- A Bachelor's degree in Computer Science, Information Technology, or a related discipline is preferred, although equivalent professional experience may be considered.
- Relevant certifications such as AWS Solutions Architect, Microsoft certifications, or HashiCorp Certified: Terraform Associate are desirable.
- Experience with containerization technologies such as Docker and Kubernetes, as well as CI/CD tools including GitLab Pipelines, Jenkins, or GitHub Actions, is a plus.
- Knowledge of security best practices, compliance frameworks, and log aggregation and analysis tools such as the ELK Stack or Splunk is desirable.
- Be able to obtain the required security clearance for the position.
- Annual salary range of $105,000–$140,000, with actual compensation determined by experience, qualifications, skills, geographic location, contract requirements, and business needs.
- Medical, dental, and vision insurance for eligible employees.
- Life, AD&D, and disability insurance.
- Paid time off and 11 company holidays.
- 401(k) retirement plan with company matching.
- Additional employee benefits and wellness resources, subject to applicable eligibility requirements and plan terms.
- Remote work arrangement based in the United States.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1



