Skip to content

Open nowPosted 59 days ago

Senior Principal Site Reliability Engineer

qfg70 open roles

Where
5700 Yonge St, North York, ON M2M 4K2, Canada
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowSenior Principal Site Reliability Engineerqfg · 5700 Yonge St, North York, ON M2M 4K2, Canada
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on qfg's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

7.8% of postings close within 7 days. Measured by our own scanner across the market.

Share of postings closed within
  1. 1.6%1 day
  2. 3.4%3 days
  3. 7.8%7 days
  4. 14.3%14 days
  5. 33.7%30 days
This job: posted 59 days ago

The posting

  What’s in it for you as an employee of QFG?

Health & wellbeing resources and programs 

Paid vacation, personal, and sick days for work-life balance

Competitive compensation and benefits packages

Work-life balance in a hybrid environment with at least 3 days in office

Career growth and development opportunities

Opportunities to contribute to community causes

Work with diverse team members in an inclusive and collaborative environment

  This job posting is for an existing vacancy   We’re looking for our next Senior Principal Site Reliability Engineer. Could It Be You?   The Senior Principal Site Reliability Engineer is directly responsible for the stability, resiliency, and scalability of business-critical brokerage back-end applications running across a hybrid on-premises and cloud architecture. This individual drives reliability engineering practices — SLOs/SLIs, observability, incident response, and capacity planning — while also making hands-on, code contributions directly into multiple applications across different technology stacks using an inner-source model. The role blends deep technical execution with cross-team influence: this person is expected to identify systemic reliability risks, fix them where they appear, and raise the operational bar for every team they touch. This position is a strong fit for a hands-on senior principal engineer who is energized by fixing production reliability at the source, is comfortable navigating multiple codebases and cloud/on-prem environments, and wants to have an outsized impact on the resiliency of a regulated, high-availability brokerage platform.   Need more details? Keep reading…   Application Stability, Reliability & Growth

Own the end-to-end reliability posture of critical brokerage back-end applications, driving measurable improvements in availability, latency, and error budgets across on-premises and cloud environments.

Define and track Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets in partnership with application teams; use them to prioritize reliability work over feature work when warranted.

Lead root cause analysis and blameless post-incident reviews for high-severity production incidents; drive remediation items to closure and identify systemic patterns across applications.

Establish and mature observability practices (metrics, logging, tracing, alerting) so that failures are detected proactively and diagnosed quickly across a heterogeneous, multi-stack estate.

Build capacity planning, load testing, and chaos/failure-injection practices to validate resilience before incidents occur.

Champion a culture of operational excellence, toil reduction, and AI and automation-first thinking across engineering teams.

  Cross-Stack Engineering & Inner Sourcing

Make hands-on, code contributions directly into multiple applications spanning different languages, frameworks, and stacks, using an inner-source model to fix reliability defects, add instrumentation, and improve resiliency patterns.

Partner with individual application teams to raise pull requests, follow their contribution standards, and pair with owning engineers so fixes land safely and are properly reviewed and owned long-term.

Identify recurring reliability anti-patterns across codebases (e.g., missing timeouts/retries, unbounded queues, improper connection pooling) and drive standardized, reusable fixes or shared libraries.

Contribute to and help govern internal reliability tooling, shared SDKs, and common patterns (circuit breakers, backoff/retry, health checks) that can be inner-sourced across teams.

  Cloud Scalability & Hybrid Architecture

Design and advise on cloud scalability strategies (auto-scaling, load balancing, multi-region/multi-AZ failover, caching, queuing) for workloads that span on-premises data centers and public cloud.

Guide capacity and cost-aware scaling decisions, balancing performance, resiliency, and cloud spend across hybrid deployments.

Evaluate and recommend cloud-native and hybrid resiliency patterns (e.g., disaster recovery, active-active/active-passive architectures, data replication strategies) appropriate for regulated brokerage workloads.

  Organizational Awareness & Risk

Bring strong organizational awareness of the operational, financial, regulatory, and reputational risk that production incidents pose to a brokerage business, and factor that into prioritization.

Participate in risk assessments related to system reliability, availability, and disaster recovery, partnering with Risk, Compliance, and Information Security as needed.

Contribute to change management and release governance practices that reduce the likelihood and blast radius of production incidents.

Promptly identify, escalate, and help remediate reliability or security-related incidents in accordance with company policy.

  Leadership & Influence

Act as a technical reference and mentor for reliability engineering practices, coaching application teams on operational excellence without formal direct reports.

Influence architecture and design decisions across multiple teams by bringing a reliability and scalability lens to reviews and planning.

Document and evangelize reliability standards, runbooks, and best practices; lead or contribute to internal tech talks and communities of practice.

Partner with engineering leadership to define the reliability roadmap and report on progress against stability goals

  So are YOU our next Senior Principal Site Reliability Engineer.? You are if you…

Bachelor's or Master's degree in Computer Science, Information Systems, Engineering, or a related field, or equivalent combination of education and experience.

8+ years of software engineering and/or site reliability engineering experience, including production ownership of business-critical applications; financial services or brokerage experience strongly preferred.

Demonstrated ability to read, debug, and make minor-to-moderate code changes across multiple languages/stacks (e.g., Java, .NET, Node.js/TypeScript, Python) in an inner-source or cross-team contribution model.

Deep experience with cloud scalability strategies on one or more major providers (AWS, Azure, GCP), including auto-scaling, load balancing, multi-region resiliency, and cost-aware capacity planning.

Experience operating and supporting hybrid architectures spanning on-premises data centers and cloud environments.

Strong background in observability tooling (e.g., Prometheus/Grafana, Datadog, Splunk, ELK, AppDynamics, Dynatrace) and building actionable alerting and dashboards.

Practical experience defining and operating against SLOs/SLIs/error budgets and running blameless post-incident reviews.

Experience with CI/CD pipelines and infrastructure-as-code (e.g., Terraform, Ansible, CloudFormation) in support of reliable, repeatable deployments.

Solid understanding of microservices architecture, distributed systems failure modes, and resiliency patterns (circuit breakers, retries/backoff, bulkheads, timeouts).

Familiarity with relational and NoSQL data stores and their operational/scaling characteristics.

Experience with incident management and on-call practices (e.g., PagerDuty, Opsgenie) including leading major incident response

Knowledge of security, audit, and regulatory considerations relevant to brokerage / financial services production systems.

Excellent communication skills, with the ability to influence engineers and stakeholders across many teams without direct authority.

Strong documentation, analytical, and problem-solving skills

Compensation Information:

Base salary range: $150,000 - $190,000

The final compensation package will be commensurate with the successful candidate's experience, skills, and geographic location (Canada). It includes a comprehensive benefits plan and a competitive incentive (bonus) program for Full-Time Permanent roles.

  Sounds like you? Click below to apply! #LI-Hybrid

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against qfg's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on qfg's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    qfg's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.