Skip to content

Open nowPosted 6 days ago

Cloud Systems Engineer - Site Reliability

Workable (global search)108,016 open roles

Where
Philadelphia, PA, United States
Work mode
Remote
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowCloud Systems Engineer - Site ReliabilityWorkable (global search) · Philadelphia, PA, United States
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Workable (global search)'s own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

7.9% of postings close within 7 days. Measured by our own scanner across the market. Workable (global search) postings stay open a median of 7 days.

Share of postings closed within
  1. 1.6%1 day
  2. 3.6%3 days
  3. 7.9%7 days
  4. 14.9%14 days
  5. 34.0%30 days
This job: posted 6 days ago

Workable (global search) median: 7 days open

The posting

About Us

TherapyNotes is the go-to superhero for behavioral health Practice Management and EHR software! Our top-notch SaaS solution handles scheduling, billing, documenting, telehealth, and more so clinicians can focus on awesome patient care.

We're a dynamic team of pros who love to innovate and push the envelope, keeping our software cutting-edge. Join us, and let's revolutionize behavioral health software together while making a real difference!

About The Position

We are seeking a Site Reliability Engineer to improve the reliability and operability of the production services and shared platforms supporting our growing 24×7 SaaS environment. In this role, you will apply software and systems engineering practices to improve availability, performance, scalability, resilience, observability, incident response, and operational automation. You will partner with software development, infrastructure, database, security, and other technology teams to establish measurable reliability goals, reduce operational toil, and ensure services are supportable throughout their lifecycle. If you are passionate about building reliable systems, solving complex production problems, and driving continuous improvement, we want to hear from you.

Requirements

  • BS degree in Information Systems, Engineering, or equivalent experience.
  • 5+ years of engineering experience in Systems Engineering, Cloud or Platform Engineering, DevOps, Software Engineering, and/or SRE.
  • Experience designing and operating production systems using cloud-based compute, storage, networking, and containerization technologies; Azure and Kubernetes preferred.
  • Strong Linux systems and networking fundamentals, with experience troubleshooting complex distributed systems in production.
  • Expertise with an observability platform; Datadog experience strongly preferred. Experience with Prometheus, Grafana, New Relic, or equivalent platforms is also valuable.
  • Experience with scripting and operational automation using tools such as Bash, PowerShell, or Python, along with infrastructure-as-code and configuration-management practices.
  • Experience participating in production on-call rotations, incident response, root cause analysis, and post-incident improvement.
  • Experience working in Agile/DevOps environments and operating production services using ITSM practices where applicable.
  • Prior software development experience—or experience investigating application behavior through code, logs, and distributed traces—is a plus

Responsibilities

  • Own and continuously improve how we use Datadog to make reliability visible and actionable across metrics, logs, traces, dashboards, monitors, alerts, and service-level views.
  • Design, implement, and maintain high-availability, high-throughput, data- and compute-intensive critical systems supporting a growing 24×7 SaaS platform.
  • Partner with service owners to define and improve reliability through meaningful SLIs, SLOs, error budgets, actionable alerting, and operational-readiness practices.
  • Participate in and help drive incident management for production events, serving as an incident commander or technical responder as needed. Coordinate triage, service restoration, escalation, communication, incident documentation, root cause analysis, and completion of corrective actions.
  • Partner with development teams to investigate issues across the infrastructure and application layers using metrics, logs, distributed traces, and code-level context.
  • Improve deployment safety and service resilience through automated validation, recovery and rollback capabilities, reliability testing, and analysis of system failure modes.
  • Partner with other technical leaders to ensure all newly introduced systems are supportable and maintainable by both development and operations.
  • Provide escalated technical guidance and support to other technology teams throughout the organization.
  • Provide on-call coverage for production support and other duties as required.
  • Ensure supported systems and operational activities comply with organizational security, HIPAA, and operating policies.
  • Identify and eliminate repetitive operational toil using Bash, PowerShell, Python, or Ansible. Manage infrastructure as code using Terraform/OpenTofu and configuration automation using Ansible.

Benefits

  • Competitive salary - $110,000-$150,000
  • Employer sponsored health, dental, vision, life, and disability insurance
  • Retirement plan with company contribution
  • Annual company profit sharing
  • Personal development/training budget
  • Open, collaborative work environment
  • Extensive 2-week onboarding plan
  • Comprehensive mentorship program

Equal Opportunity Employer Statement & Applicant Rights TherapyNotes LLC is an Equal Opportunity Employer and does not discriminate based on race, color, religion, sex, national origin, age, disability, genetic information, or any other protected status under federal, state, or local law. We are committed to providing a workplace free of discrimination and harassment.For more information about your rights under federal employment laws, please review the following:

  • Know Your Rights: Workplace Discrimination is Illegal
  • Family and Medical Leave Act (FMLA): Employee Rights Under FMLA

If you require a reasonable accommodation during the application process, please contact [email protected].

9/17/2026

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Workable (global search)'s own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Workable (global search)'s form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Workable (global search)'s answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.