Skip to content

Open nowPosted 9 days ago

Senior Site Reliability Engineer (AUS)

Climavision3 open roles

Where
Australia, Remote
Work mode
Remote
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowSenior Site Reliability Engineer (AUS)Climavision · Australia, Remote
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Climavision's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

8.2% of postings close within 7 days. Measured by our own scanner across the market.

Share of postings closed within
  1. 1.8%1 day
  2. 3.6%3 days
  3. 8.2%7 days
  4. 15.2%14 days
  5. 34.0%30 days
This job: posted 9 days ago

The posting

Senior Site Reliability Engineer

Remote | Australia

About Climavision

At Climavision, we’re rebuilding climate technology from the ground up and changing the way we see weather. We merge the power of our proprietary, high-resolution weather radar and satellite network with advanced weather prediction modelling and decades of industry expertise to reduce existing coverage gaps and drastically improve forecasting ability. Our revolutionary new approach to climate technology weather solutions is poised to help reduce the economic risks of climate change on companies, governments, and societies alike. We are backed by The Rise Fund, the world’s largest global impact platform committed to achieving measurable, positive social and environmental outcomes alongside competitive financial returns. Climavision is headquartered in Louisville, KY, with research and development operations in Raleigh, NC.

The Work

Are you an experienced Site Reliability Engineer who thrives at the intersection of software engineering and production operations? Do you take pride in keeping mission-critical customer systems reliable under real-world operational pressure? Are you looking for an opportunity to own production reliability for a modern hybrid infrastructure platform spanning cloud, colocation, and edge environments?

If so, we have an exceptional opportunity for you.

Climavision is seeking a Senior Site Reliability Engineer to contribute towards reliability, operational excellence, and production resilience across the company's platform and data services. This role sits on a shared SRE team that supports the full business rather than a single product line, covering both the radar network and the weather intelligence sides of the company as priorities shift. A central focus of this role is building the observability layer that puts the health of the full fleet in one place, and then automating recovery so that systems heal themselves. Multi-cluster and multi-replica high availability across our distributed edge fleet remains a core part of the work.

This is a hands-on engineering role for someone who is equally comfortable troubleshooting Kubernetes clusters, leading incident response, and improving operational maturity across the organization. The successful candidate will combine deep production operations expertise with a disciplined approach to reliability engineering and strong automation skills.

Climavision operates a hybrid infrastructure footprint spanning Microsoft Azure, colocation data centers, and edge Kubernetes clusters, deployed alongside weather radar systems. This role will drive production reliability across Azure, colocation, and edge environments. Right-sizing cluster resources and migrating workloads off Azure to reduce spend are active priorities for the team.

35% Kubernetes Platform Reliability and Operations

30% Production Reliability Engineering and Incident Response

20% Observability, Monitoring, and Alerting

15% Automation, Recovery, and Cost Optimization

Primary Responsibilities:

• Own production reliability for Climavision's customer-facing platform and data services across Azure, colocation, and edge Kubernetes environments.

• Work as part of a shared SRE function supporting the whole company rather than a single product line, taking on work across both the radar network and the weather intelligence sides of the business as priorities shift.

• Contribute to the definition and improvement of SLIs, SLOs, alerting standards, and operational metrics used to measure platform reliability.

• Build and own the observability layer for the fleet. Today the underlying metrics exist but are only reachable from the command line inside each cluster. This role is responsible for surfacing that data in shared dashboards and building the alerting that tells the team something is going wrong before a customer does.

• Design and build automated recovery and self-healing for production systems, so that common failure modes are detected and remediated without human intervention.

• Optimize cluster resourcing and cost, including right-sizing workloads and nodes and supporting the migration of workloads off Azure to reduce spend.

• Support and coordinate production incident response efforts, including troubleshooting, mitigation, communication, and postmortem analysis.

• Diagnose and resolve complex production issues across application services, Kubernetes infrastructure, storage, and distributed systems.

• Drive multi-replica and multi-cluster high availability across Climavision's services, including workload placement, scheduling, and deployment patterns that allow services to run safely as multiple replicas across multiple clusters.

• Contribute to the multi-cluster high-availability strategy across Climavision's hybrid fleet, including active-active and active-passive failover behavior, traffic routing, data replication considerations, and graceful degradation when a cluster becomes unavailable.

• Operate and improve Climavision's self-managed Kubernetes platform spanning cloud-hosted, colocation, and edge clusters, with a focus on availability, resiliency, recovery, and operational performance.

• Ensure Kubernetes platform lifecycle activities including upgrades, patching, cluster health, node management, and production change management are executed in a manner that preserves service availability and minimizes customer-facing risk.

• Improve reliability and operational maturity of production platform services, including observability, autoscaling, ingress, and distributed storage. Partner with the teams responsible for the underlying networking and security primitives rather than owning those areas directly.

• Design and validate Kubernetes workloads for resiliency, scalability, and operational efficiency, including autoscaling behavior, workload placement, resource management, and graceful degradation strategies.

• Partner with software engineering teams across the company to improve production readiness, resiliency patterns, deployment safety, and operational visibility before services reach production.

• Maintain and improve deployment pipelines, Helm charts, Kubernetes manifests, and infrastructure automation supporting safe and repeatable production releases.

• Support and evolve Climavision's observability platform, including metrics, logging, distributed tracing, dashboarding, and alerting.

• Conduct performance engineering and capacity-planning efforts for customer-facing services during peak weather-event demand.

• Help facilitate blameless postmortem reviews and drive operational follow-up items through completion.

• Improve disaster recovery, failover, and business continuity capabilities across cloud, colocation, and edge environments.

• Drive operational excellence initiatives, including automation, reduction of operational toil, game days, production readiness reviews, and reliability best practices.

• Contribute as a senior technical resource and mentor on reliability engineering and production operations practices.

On-Call Expectation:

Climavision operates customer-facing production systems under contractual SLAs that do not pause outside business hours. The Senior Site Reliability Engineer will participate in a rotating on-call schedule made up of two separate rotations:

• A weekday rotation. The engineer on a weekday shift is the first point of contact for production incidents and for engineering teams needing support during the business week.

• A separate weekend rotation, so that the engineer carrying weekday support is not also carrying the weekend.

At current and planned team size, engineers can expect a weekday shift roughly every five weeks and a weekend shift roughly every five weeks. The two are scheduled as far apart from each other as the rotation allows, so that a weekday shift and a weekend shift do not fall close together.

Qualifications

• A bachelor's degree in computer science, software engineering, or a related field; equivalent professional experience considered.

• Minimum of 7 years of experience in Site Reliability Engineering, DevOps, Production Engineering, Platform Engineering, or a related infrastructure-focused role, with at least 4 years in a role formally titled Site Reliability Engineer or carrying explicit SLO / error-budget accountability.

• Deep, hands-on experience operating native Kubernetes. Managed distributions such as AKS and EKS are acceptable, but experience running native or self-managed Kubernetes is strongly preferred and is the primary technical requirement for this role.

• Demonstrated experience optimizing Kubernetes clusters, including right-sizing workloads and node pools, resource management, and reducing infrastructure cost without sacrificing reliability. Be prepared to walk through a specific cluster optimization project you led.

• Demonstrated experience increasing operational visibility, including building dashboards, metrics pipelines, and alerting in an environment where little or none existed before.

• Experience designing and operating workloads for safe horizontal scaling across multiple replicas, including idempotency, concurrency, and state handling considerations.

• Experience designing or operating multi-cluster high-availability architectures, including failover behavior, traffic routing, and cross-cluster service deployment.

• Experience supporting customer-facing production systems with uptime, reliability, and incident-response responsibilities.

• Experience diagnosing and resolving production incidents across application, platform and Kubernetes infrastructure layers, including workload scheduling, storage, ingress, and cluster-level failures.

• Experience operating Kubernetes outside of strictly managed cloud environments, including bare-metal, colocation, edge, or hybrid infrastructure.

• Experience with Kubernetes operational tooling and ecosystem technologies such as Rancher, Helm, autoscaling frameworks, observability stacks, or distributed storage systems.

• Strong understanding of infrastructure automation and Infrastructure as Code concepts using tools such as Terraform and Ansible.

• Experience supporting CI/CD and production deployment pipelines. GitHub Actions is used for CI/CD at Climavision.

• Experience with monitoring, logging, and observability platforms such as DataDog, Prometheus, Grafana, Loki, OpenTelemetry, or comparable technologies.

• Experience operating distributed systems and microservice-based architectures in production environments.

• Working knowledge of Microsoft Azure infrastructure.

• Strong troubleshooting skills across infrastructure, application, and platform layers.

• Demonstrated experience participating in a structured production on-call rotation supporting business-critical systems.

• Working familiarity with Jira, Confluence, and Microsoft Entra, which the team uses day to day for ticketing, documentation, and authentication.

• Strong written and verbal communication skills, including incident documentation and postmortem authoring.

• Experience working in start-up, scale-up, or other fast-moving engineering environments, and comfort with the pace and ambiguity that comes with them.

Nice to have, but not required:

• Experience operating Kubernetes platforms using RKE2 and Rancher, which is how Climavision manages its clusters.

• Experience with Octopus Deploy.

• Experience with SOC 2 or comparable security auditing and compliance work.

• Experience supporting hybrid cloud and colocation infrastructure environments.

• Experience with service mesh technologies such as Istio.

• Experience with Kubernetes-native storage platforms such as Longhorn.

• Experience operating PostgreSQL or PostGIS in Kubernetes environments.

• Experience with distributed messaging systems such as RabbitMQ or NATS.

• Experience supporting GPU-enabled workloads in Kubernetes.

• Familiarity with reliability engineering practices, including SLIs, SLOs, error budgets, and operational maturity metrics.

Physical Demands & Work Environment:

• This is a full-time, exempt position

• Fully Remote - Australia

• Candidates based in Australia should expect standard local business hours on the east coast of Australia, which provides overlap with the United States team in the early morning Eastern time.

• This job requires frequent use of a computer to complete tasks, attend meetings, and communicate via Microsoft Teams.

Once you land this position, you’ll get to enjoy:

• Benefits of a dynamic and growing organization

• A challenging, hands-on role that will have real impact on the business

• Competitive compensation

• Comprehensive benefits package

• 401(k) Savings Plan

• Medical/Dental/Vision Benefits

• Health Savings Account (HSA) and Flexible Spending Account (FSA)

• Unlimited Paid Time-off

• 11 Paid Holidays

• Paid Parental Leave

• Company Paid Short-term Disability (STD)

• Company Paid Long-term Disability (LTD)

• Company Paid Life Insurance

The salary range for this position is $130,000-170,000 annually, however Climavision considers several factors when extending an offer of employment including but not limited to, the applicant’s education, experience, the responsibilities of the role, training, knowledge, skills, and abilities, as well as internal equity and alignment with market data. Any offer of employment is contingent on completion of a background check to company standard. Please note this job description is not designed to cover or contain a comprehensive listing of activities, duties or responsibilities that are required of the employee for this job. Duties, responsibilities, and activities may change at any time with or without notice.

Climavision is an equal opportunity employer. All aspects of employment including the decision to hire, promote, discipline, or discharge, will be based on merit, competence, performance, and business needs. We do not discriminate on the basis of race, color, religion, marital status, age, national origin, ancestry, physical or mental disability, medical condition, pregnancy, genetic information, gender, sexual orientation, gender identity or expression, veteran status, or any other status protected under federal, state, or local law.

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Climavision's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Climavision's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Climavision's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.