Skip to content

Open nowPosted 18 hours ago

Site Reliability Engineer, AI Observability

Appnovation Technologies53 open roles

Where
New York, Austin, Miami, Dallas
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowSite Reliability Engineer, AI ObservabilityAppnovation Technologies · New York, Austin, Miami, Dallas
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Appnovation Technologies's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

8.1% of postings close within 7 days. Measured by our own scanner across the market. Appnovation Technologies postings stay open a median of 8 days.

Share of postings closed within
  1. 1.7%1 day
  2. 3.6%3 days
  3. 8.1%7 days
  4. 15.1%14 days
  5. 34.0%30 days
This job: posted 18 hours ago

Appnovation Technologies median: 8 days open

The posting

About us

Appnovation is a global, full-service digital partner that combines Strategy, Experience & Design, Engineering and Managed Services. We build digital solutions that deliver real impact today and serve as foundations for future growth. Bold ambition. Practical action. Endless possibilities.

We’re looking for a Site Reliability Engineer to keep a shared observability platform for LLM-based applications running for a global life sciences client. The platform is built on Langfuse and self-hosted on Kubernetes on AWS, with ClickHouse as the analytical store, PostgreSQL for metadata and Redis as the ingestion queue, all delivered through Argo CD.

Two things need building rather than maintaining. The platform has no monitoring, alerting or defined service levels today, and you will own putting them in place. Infrastructure is defined as code throughout, in Kubernetes manifests, Helm values and Argo CD applications.

Alongside the platform itself, you will own the runbooks that make an incident survivable by someone other than the author, and the onboarding and support path for the internal teams that depend on the platform.

ROLE RESPONSIBILITIES

  • Platform Operations: Diagnose and resolve failures across ClickHouse, PostgreSQL, Redis, the ingestion workers and the Kubernetes layer beneath them, including ingest backpressure from queue depth and worker drain behaviour.
  • ClickHouse Operations: Own ClickHouse under the operator model, including Keeper quorum, replication, shard and replica topology, and S3 storage tiering.
  • Provable Backup and Restore: Rehearse the restore, time it, document it and test it against its failure modes, rather than assuming a successful backup job means a recoverable system.
  • Safe Upgrades: Plan and rehearse upgrades in a lower environment, with a rollback plan that works even if a migration is only partly complete.
  • Monitoring and Service Levels: Build monitoring, alerting and service levels from scratch, so problems are found here before a user reports them.
  • Infrastructure as Code: Maintain Kubernetes manifests, Helm values and Argo CD applications so every change goes through the delivery pipeline.
  • Runbooks and SOPs: Write and maintain runbooks and SOPs that let a colleague resolve an incident without the author present.
  • Onboarding and Support: Run the onboarding and support path for internal teams that depend on the platform, and triage what they bring.
  • Automation: Turn recurring operational work into automation.
  • Upgrade Partnership: Work with the platform engineer who owns what the platform offers. They decide what to adopt and how it is configured; you own the migration and its rollback.

QUALIFICATIONS

  • Hands-on experience with Kubernetes on AWS (managed EKS), with routine work done through Helm values and Argo CD applications.
  • Hands-on experience running ClickHouse in production, including replication and Keeper quorum, shard and replica topology, and backup and restore, ideally run through an operator.
  • Experience building monitoring and alerting from scratch, including service levels that reflect what users actually experience rather than what is easy to measure.
  • Experience upgrading self-hosted software safely, including schema migrations, rehearsal in a lower environment and a rollback plan for a partly completed migration.
  • PostgreSQL and Redis operations deep enough to debug metadata-store and queue problems, including backpressure and worker drain.
  • Strong operational writing: runbooks, SOPs and post-incident reviews that a colleague can follow unaided during an incident.

PREFERRED QUALIFICATIONS

  • Observability engineering, including OpenTelemetry Collector pipelines, alerting design and Grafana dashboards.
  • OIDC or enterprise SSO integration with a corporate identity provider.
  • GitHub Actions for plan and apply pipelines with approval gates.
  • Experience running LLM observability tools such as Langfuse, LangSmith or Arize Phoenix.
  • Experience in pharma, life sciences or another regulated industry.

WHO YOU ARE

  • You don’t trust a backup until you’ve restored from it
  • You want to find problems before users do
  • You write runbooks for the person on call at 3am, not for yourself
  • You stay calm in incidents and focus on fixing the process afterward
  • You automate anything you have to do twice
  • You work well inside a client team and build trust quickly
  • You have prior experience in consulting
  • Prior experience and connections in the Life Sciences industry is preferred

Thank you for your interest in a career with Appnovation Technologies! Please note that only those selected for an interview will be contacted.

At Appnovation, we recognize that diverse teams are the strongest teams. Diversity, Equity & Inclusion is not only something that we embrace - we celebrate it! We are proud to be an Equal Opportunity Employer and we encourage applicants from all backgrounds, lived experiences and industries to apply. Come join us at Appnovation, and learn more about how we stay true to our company values as we build better lives through better digital.

Accommodations are available upon request throughout the recruitment process.

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Appnovation Technologies's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Appnovation Technologies's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Appnovation Technologies's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.