Skip to content

Open nowPosted 8 hours ago

Senior/Staff Platform Engineer

Jobgether3,797 open roles

Where
Canada
Work mode
Remote
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowSenior/Staff Platform EngineerJobgether · Canada
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Jobgether's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

7.9% of postings close within 7 days. Measured by our own scanner across the market. Jobgether postings stay open a median of 5 days.

Share of postings closed within
  1. 1.6%1 day
  2. 3.6%3 days
  3. 7.9%7 days
  4. 14.9%14 days
  5. 34.2%30 days
This job: posted 8 hours ago

Jobgether median: 5 days open

The posting

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior/Staff Platform Engineer based in Canada.

The Senior/Staff Platform Engineer will build, operate, and evolve large-scale production infrastructure with a strong focus on Kubernetes and reliability. This hands-on role spans cloud infrastructure, Linux, networking, observability, CI/CD, automation, and production operations. You will diagnose complex infrastructure challenges across multiple technical layers and build tooling that improves reliability, scalability, and developer experience. The position offers significant autonomy, with ownership of ambiguous initiatives from initial design through implementation and production operation. You will work directly with technical stakeholders, contribute to architecture and reliability decisions, and take a leading role in critical incident resolution. Your impact will extend beyond day-to-day operations as you establish stronger engineering practices, automation, observability, and operational resilience. This is an ideal opportunity for a highly experienced engineer who combines deep systems expertise with strong technical judgment and communication skills.

Accountabilities:

  • Design, build, operate, and continuously improve production Kubernetes platforms, owning cluster architecture, networking, workload isolation, resource management, security, upgrades, scaling, and reliability.
  • Troubleshoot Kubernetes and infrastructure issues beyond the application layer, including CNI and networking problems, scheduling, node behavior, resource constraints, controllers, and cluster-level failures.
  • Operate highly available infrastructure across cloud, hybrid, virtualized, and/or bare-metal environments, diagnosing complex issues across Kubernetes, containers, Linux, networking, and underlying infrastructure.
  • Develop and maintain production tooling and automation using Go, Python, or Java to improve platform operations, deployment, troubleshooting, reliability, and developer experience.
  • Build internal services, APIs, integrations, and operational tools where needed, while applying sound software engineering practices such as testing, code review, maintainability, and documentation.
  • Own the reliability and operational health of critical production infrastructure, leading or significantly contributing to incident response, root-cause analysis, durable remediation, and disaster recovery initiatives.
  • Define and improve SLOs, SLIs, alerting, and operational processes, using logs, metrics, traces, profiling tools, and system-level diagnostics to improve availability, performance, capacity, resilience, and recovery times.
  • Build and maintain infrastructure as code using Terraform or equivalent technologies, creating reusable patterns and improving automation as the platform evolves.
  • Develop and improve CI/CD and deployment workflows while balancing delivery speed with reliability, security, scalability, and operational requirements.
  • Participate in cloud and infrastructure migrations, including dependency analysis, networking, cutover planning, rollback strategies, and production validation.
  • Build and enhance monitoring, dashboards, logging, distributed tracing, and alerting to improve incident detection, diagnosis, and recovery.
  • Work directly with customers, engineers, and technical stakeholders to understand requirements, investigate issues, communicate trade-offs, and drive effective technical solutions.
  • Own complex infrastructure initiatives from problem definition through design, implementation, and production operation, contributing to RFCs, architecture discussions, design reviews, and broader technical direction.
  • Mentor engineers and help strengthen engineering and operational practices while operating independently in ambiguous situations and taking ownership when immediate direction is unavailable.
  • 10+ years of professional experience in Platform Engineering, Site Reliability Engineering, Infrastructure Engineering, DevOps, or related fields is preferred for Staff-level candidates, with significant hands-on experience operating complex production infrastructure and distributed systems.
  • Demonstrated experience building and operating production Kubernetes platforms is required, rather than experience limited to deploying applications onto existing clusters.
  • Production programming experience in Go, Python, or Java is required, with the ability to read, debug, maintain, and contribute to production codebases and automation.
  • Strong experience with production reliability, incident response, troubleshooting, root-cause analysis, and operational ownership is essential.
  • Demonstrated ability to independently own complex technical initiatives from ambiguous starting points through design, implementation, and production operation.
  • A degree in Computer Science, Engineering, or a related field is preferred, although equivalent practical experience may be considered.
  • Deep understanding of Kubernetes infrastructure, including cluster architecture, networking and CNI, NetworkPolicy, scheduling, resource management, nodes, security, RBAC, and cluster behavior.
  • Strong Linux fundamentals and hands-on experience troubleshooting production systems, combined with a solid understanding of DNS, routing, load balancing, connectivity, and cloud/Kubernetes networking.
  • Production experience with at least one major cloud platform, such as AWS, GCP, or Alicloud, along with infrastructure-as-code experience using Terraform or equivalent tooling.
  • Experience with configuration management and automation technologies such as Ansible, Puppet, or similar platforms.
  • Strong observability expertise using metrics, logs, traces, dashboards, and alerting platforms such as Prometheus, Grafana, OpenTelemetry, Datadog, or equivalent technologies.
  • Experience with CI/CD infrastructure, Docker, container tooling, modern software delivery practices, high availability, capacity planning, disaster recovery, and production resilience.
  • Experience planning and executing production cloud or infrastructure migrations, operating large-scale or multi-cluster Kubernetes environments, or working with hybrid, on-premises, virtualized, or bare-metal infrastructure is preferred.
  • Additional experience with Kubernetes controllers or operators, advanced networking, service mesh, mTLS, workload identity, multi-cloud infrastructure, security hardening, IAM, compliance, capacity planning, performance engineering, or developer platform tooling is highly valued.
  • Previous technical leadership, mentoring, Staff/Principal-level responsibilities, or experience working directly with external customers in consulting or service-delivery environments is preferred.
  • Exceptional written and verbal technical communication skills are required, with the ability to explain architecture, risks, trade-offs, and technical decisions clearly to customers, engineers, and technical leadership.
  • Strong analytical and problem-solving abilities, high ownership, sound technical judgment, and the ability to operate independently in ambiguous environments are essential.
  • The ideal candidate can influence others without formal management authority and is comfortable taking the lead on complex technical problems without continuous oversight.
  • This is a hands-on platform and reliability engineering position. Candidates should have personally built, operated, troubleshot, and improved production infrastructure rather than primarily consuming managed services or deploying applications onto Kubernetes.
  • The role requires availability during core hours aligned with Pacific Time, from 8:00 AM to 5:00 PM PST, as well as participation in an on-call rotation approximately every four to five weeks.
  • Fully remote work arrangement within Canada.
  • Opportunity to work on large-scale Kubernetes platforms and complex production infrastructure with significant technical ownership.
  • High degree of autonomy to shape infrastructure architecture, reliability initiatives, automation, and operational practices.
  • Exposure to cloud, hybrid, networking, observability, CI/CD, security, and distributed systems at scale.
  • Opportunity to collaborate directly with experienced engineers, technical stakeholders, and customers on challenging infrastructure initiatives.
  • Ability to influence technical direction, mentor other engineers, and contribute to architecture and engineering best practices.
  • Remote-first environment designed to support distributed collaboration across North America and Latin America.
  • Opportunity to work on technically challenging projects where reliability, scalability, automation, and operational excellence are core priorities.

How Jobgether works:

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Why Apply Through Jobgether?

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

#LI-CL1

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Jobgether's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Jobgether's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Jobgether's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.