Skip to content

Open nowPosted 274 days ago

Site Reliability Engineer

plenit12 open roles

Where
Oficina Plaza España
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowSite Reliability Engineerplenit · Oficina Plaza España
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on plenit's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

8.0% of postings close within 7 days. Measured by our own scanner across the market.

Share of postings closed within
  1. 1.6%1 day
  2. 3.6%3 days
  3. 8.0%7 days
  4. 15.0%14 days
  5. 34.2%30 days
This job: posted 274 days ago

The posting

About this role

The challenge As a Site Reliability Engineer, you will play a critical role in ensuring the stability, availability, resilience, and performance of our cloud platform in production. Your mission will cover two complementary areas. First, you will work proactively to prevent incidents by identifying operational risks, anticipating capacity constraints, improving production changes, resolving recurring problems, and driving reliability improvements, migrations, and corrective initiatives. Second, when incidents occur, you will help restore service quickly by coordinating technical diagnosis, controlling the situation, communicating impact, and leading recovery efforts. This is a senior, hands-on role for someone with a strong operational mindset who is comfortable working in complex production environments, investigating technical issues, making decisions under pressure, and taking ownership of platform reliability. You will be part of the Site Reliability Engineering team, working in a practical environment focused on keeping the platform stable while continuously reducing operational risk and manual effort. A key part of the role will be to identify what could affect the platform before it becomes an incident. You will assess risks, make them visible, define mitigation or correction plans, assign ownership, and follow actions through to completion. You will also contribute to capacity planning, production scaling, infrastructure improvements, migrations, observability, automation, and operational readiness. When incidents happen, you will act as a technical reference and coordination point. You will help establish control, assess impact, accelerate diagnosis, involve the right teams, communicate clearly, and restore service as quickly and safely as possible. You will work closely with engineering, systems, networking, storage, database, security, and product teams. Your operational perspective will help ensure that reliability, scalability, recoverability, and operability are considered in platform changes and evolution. Participation in an on-call rotation and response to critical incidents outside regular working hours are part of the role.

Requirements that are important for us We are looking for a senior Site Reliability Engineer with strong experience operating critical production environments and a proven ability to identify risks, troubleshoot complex issues, and drive long-term reliability improvements. Relevant experience and expected outcomes

  • Operating critical production infrastructures while ensuring availability, stability, performance, and recoverability.
  • Identifying technical and operational risks and driving mitigation, correction, or contingency plans.
  • Maintaining a prioritized backlog of risks, recurring problems, capacity constraints, and reliability improvements.
  • Leading incident response processes, including impact assessment, technical diagnosis, coordination, communication, and service restoration.
  • Performing root cause analysis and implementing corrective and preventive actions.
  • Managing recurring problems and ensuring they remain visible, owned, prioritized, and followed through to resolution.
  • Monitoring demand, system limits, growth trends, bottlenecks, and saturation points.
  • Contributing to capacity planning and executing infrastructure scaling in production environments.
  • Reviewing infrastructure changes, deployments, maintenance activities, and migrations from a reliability perspective.
  • Defining implementation, validation, rollback, and contingency plans for production changes.
  • Improving observability through metrics, logs, traces, alerts, dashboards, and service-health indicators.
  • Improving runbooks, diagnostic playbooks, recovery procedures, and operational documentation.
  • Identifying repetitive tasks, unnecessary escalations, and manual processes that should be automated or simplified.
  • Working with systems that support web applications, including a solid understanding of HTTP, DNS, TCP/IP, and load balancing.
  • Administering technologies such as NGINX, Apache, load balancers, databases, and related web infrastructure.
  • Strong experience with Linux systems and working knowledge of Windows environments.
  • Good understanding of networking, virtualization, storage, databases, and infrastructure dependencies.
  • Familiarity with Kubernetes and containerized platforms is highly valuable.
  • Experience with cloud infrastructure is highly valuable.

Key skills and expected impact

  • Strong troubleshooting capabilities and the ability to analyze metrics, logs, traces, events, and system behaviour.
  • Broad technical knowledge across infrastructure, applications, networks, storage, and databases.
  • Strong risk-based thinking, with the ability to distinguish urgent work from important preventive work.
  • Ability to turn operational concerns into concrete actions, owners, deadlines, and measurable outcomes.
  • Ability to remain calm under pressure while acting decisively and proactively.
  • Strong ownership of platform stability and service recovery.
  • Strong coordination skills during incidents, migrations, changes, and technical escalations.
  • Clear communication with technical teams, stakeholders, and affected parties.
  • Ability to challenge unsafe changes or insufficient operational preparation constructively.
  • Strong documentation habits and commitment to shared operational knowledge.
  • Continuous-improvement mindset focused on reducing incidents, recovery time, operational effort, and manual intervention.
  • Motivation to act as a key technical reference during complex and high-impact scenarios.

Tools

  • Observability tools for metrics, logs, traces, alerting, and dashboards.
  • Infrastructure and system-administration tools across Linux and Windows environments.
  • Networking, virtualization, storage, and database technologies.
  • Web infrastructure tools such as load balancers, NGINX, Apache, and proxies.
  • Cloud infrastructure and configuration-management platforms.
  • Kubernetes and container orchestration technologies.
  • Incident-management and on-call tools.
  • Operational documentation tools, runbooks, playbooks, and incident procedures.
  • Capacity-planning, performance-analysis, and infrastructure-scaling tools.
  • Automation and scripting tools.

What success looks like Success in this role means that operational risks are identified early, capacity issues are anticipated, production changes are safer, recurring problems are permanently addressed, and the platform becomes progressively more observable, resilient, scalable, and easier to operate. When incidents occur, they are controlled quickly, communicated clearly, and resolved with reduced recovery time. The objective is not only to respond better, but to reduce the number, frequency, and impact of incidents over time.

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against plenit's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on plenit's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    plenit's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.