Skip to content

Open nowPosted 42 hours ago

LLM Application Engineer

Workable (global search)108,016 open roles

Where
Argentina
Work mode
Remote
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowLLM Application EngineerWorkable (global search) · Argentina
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Workable (global search)'s own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

7.9% of postings close within 7 days. Measured by our own scanner across the market. Workable (global search) postings stay open a median of 7 days.

Share of postings closed within
  1. 1.6%1 day
  2. 3.6%3 days
  3. 7.9%7 days
  4. 14.9%14 days
  5. 34.0%30 days
This job: posted 42 hours ago

Workable (global search) median: 7 days open

The posting

We're hiring a mid-senior LLM Application Engineer on a remote monthly retainer to design, ship, and harden production AI pipelines for client products at Bolder Apps. You'll own structured extraction and classification systems that turn messy real-world inputs (email, HTML, PDFs, images) into reliable product data, with measurable quality gates, evals, and cost control. You'll work on Firebase / GCP-style backends with product, mobile, and QA. We want someone who has shipped LLM apps for real users, not demos. You should be strong across modern LLMs and especially fluent with Google Gemini (multimodal prompts, structured outputs, failure modes, and cost/latency tradeoffs), with solid experience on other major providers too. If you can hit hard quality targets, keep dollars-per-run honest, and leave runbooks another engineer can pick up, we want to talk.

About Us

Bolder Apps is a product development studio that partners with US-based startups and established companies to build and scale innovative digital products. We specialize in AI-powered development, full-cycle product creation, and engineering team augmentation. Our mission is simple: build bolder, faster, and smarter.

Our Culture & Values (read before applying)

We move fast. We take ownership. We work with AI, not against it. And we expect everyone to bring ideas, not wait for instructions.

There are no daily checklists, no micromanagement, and no corporate politics. Instead, you'll have autonomy, trust, and a team that's always ready to help you grow. At Bolder Apps, impact matters more than titles, and curiosity matters more than seniority.

If you want a place where you can level up fast and actually see your work making a difference - welcome aboard.

Requirements

Responsibilities

  • Own production LLM pipelines end to end: ingestion, multimodal model calls, structured records, and storage, including confidence flags, retries, and idempotent rescans
  • Design prompt and schema strategies (including schema-aligned or constrained outputs) so results are consistent and product-ready
  • Build classification and filtering layers on top of extraction (taxonomy mapping, demographic or audience filters, deduplication, and related cleanup logic)
  • Define and run evaluation harnesses (golden sets, regression suites, online metrics) so quality does not regress when prompts, models, or parsers change
  • Hit and report against hard quality targets (precision-style gates for completeness, duplicates, incorrect inclusions, image presence, and similar product SLAs)
  • Optimize token usage, model tiering, caching, and batching to keep dollars-per-run and latency under control
  • Harden reliability for long-running async jobs (timeouts, partial recovery, memory limits, safe production deploys)
  • Partner with Flutter / mobile and QA on field contracts, review queues, and incident debugging
  • Document architecture and runbooks so ownership is shared, not a single point of failure
  • Stay current on Gemini and peer LLM APIs; recommend when to swap models, add fallbacks (e.g. document AI), or tighten schemas
  • Shipped LLM applications in production (not demos only): prompts, structured outputs, retries, observability, and real failure handling
  • Strong hands-on experience with Google Gemini, including multimodal (text + image / document-style) workflows and structured extraction
  • Practical experience with at least one other major LLM stack (OpenAI, Anthropic, or similar) and good judgment on when to use which
  • Structured extraction from messy inputs: HTML, PDFs, images, and mixed email-like content
  • Classification / taxonomy systems on top of LLM outputs
  • Evaluation discipline: offline evals, regression suites, and production quality metrics tied to clear acceptance criteria
  • Cost and latency awareness: token budgeting, cheaper tiers, caching, batching; can explain dollars-per-run tradeoffs to a PM
  • Python backend experience on serverless cloud (Cloud Functions or equivalent) and document stores (e.g. Firestore) or similar GCP patterns
  • English at C1 or above for client-adjacent debugging with a PM
  • US hours overlap through roughly 5 PM EST when live coordination is needed
  • Ownership habits: honest estimates, early blockers, finished releases

Nice to have

  • Schema-aligned LLM frameworks (BAML, Instructor, Outlines, or similar)
  • Google Document AI or other OCR / document intelligence as a fallback path
  • Gmail API / OAuth products and restricted-scope compliance familiarity
  • Computer vision for product-image quality checks
  • Firebase + GCP ops (secrets, regions, schedulers, cost monitoring)
  • Building eval corpora from real production data and iterating until contractual SLAs pass
  • Agency or multi-client studio experience

Benefits

  • Fully remote and async-friendly, with required overlap through ~5 PM EST when client or release coordination needs it
  • Monthly retainer structure with recurring AI pipeline work for engineers who keep production quality and cost honest
  • Real autonomy over how you structure prompts, schemas, evals, and deploys. We do not hand you a rigid playbook
  • Direct line to PMs, mobile engineers, and decision-makers
  • Tooling budget for the LLM and cloud tools you need to move fast
  • A peer network of product-minded builders across overlapping client projects
From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Workable (global search)'s own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Workable (global search)'s form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Workable (global search)'s answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.