Skip to content

Open nowPosted 30 days ago

Senior AI Engineer, Agents

Workable (global search)108,016 open roles

Where
Spain
Work mode
Remote
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowSenior AI Engineer, AgentsWorkable (global search) · Spain
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Workable (global search)'s own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

7.9% of postings close within 7 days. Measured by our own scanner across the market. Workable (global search) postings stay open a median of 7 days.

Share of postings closed within
  1. 1.6%1 day
  2. 3.6%3 days
  3. 7.9%7 days
  4. 14.9%14 days
  5. 34.0%30 days
This job: posted 30 days ago

Workable (global search) median: 7 days open

The posting

About EveryWatch

EveryWatch is the largest and most trusted data source in the secondary watch market. Established by a group of watch lovers, EveryWatch was created in response to the increasing popularity of luxury timepieces, with the aim of bringing unprecedented transparency and insight to the watch market. The first platform of its kind, EveryWatch combines all aspects of the watch world under one roof: a one-stop shop for watch collectors, vendors, and enthusiasts.

Position Overview

We're looking for Senior AI Engineer, Agents - an Agent Architect who will be responsible for designing, building, and scaling AI-powered solutions across EveryWatch. Working closely with our engineering, product, sales, and data teams, they will identify opportunities where AI and LLMs can automate complex processes, improve decision-making, and create new capabilities for our users and internal teams.

This is a highly hands-on role for someone who combines strong Python and LLM engineering skills with a creative, proactive mindset. You will not simply be given a roadmap. You will be expected to understand how the business operates, identify where AI can make a meaningful difference, propose new solutions, and take them from idea to production.

Our existing AI work, including WatchChat, provides a foundation to build on. The opportunity now is to expand that foundation into a broader ecosystem of intelligent agents that can work with EveryWatch's unique watch-market data and support collectors, dealers, sales teams, data operations, and engineering.

In short, we're looking for someone who doesn't just ask, “What can we build?” but “What should we build?”

Requirements

What you'll actually do

  • Invent the agent roadmap. Sit with sales, data and product, find the repetitive expert work, and come back with a ranked list of agents worth building - with a real view on feasibility, cost and impact. This is the core of the job, and it doesn't stop after the first quarter.
  • Build them yourself. You are hands-on. You design the architecture and you write the code - tools, loops, sub-agents, memory, state, evaluation. Not a spec-writer with a team underneath.
  • Own the platform under the agents. Every new agent should be cheaper to build than the last: shared tool layer over EverWatch data (pricing, auctions, listings, references, portfolios), an MCP surface over our existing backend, shared memory, tracing, and a reusable eval harness.
  • Harden WatchChat alongside us. Multi-turn state and memory, latency, cost, tool-call reliability, regression gates. It's live-bound and it has to stay right.
  • Make quality measurable. Golden multi-turn datasets, programmatic verifiers for tool/argument correctness, LLM-as-judge on held-out sets. If we can't measure an agent, we don't ship it.
  • Treat cost and latency as design constraints. Cheap models for routing and intent, strong models where they earn their keep; context budgets, caching, and - where it pays off - fine-tuning (SFT/LoRA on curated production trajectories) instead of ever-larger prompts.

What we're looking for

Three things, and we won't trade any of them away:

  1. Creative. You generate agent ideas the business hadn't thought of, and you can tell the difference between one that will work and one that demos well. You start from a use case, not a framework.
  2. Strong architect. You can design an agent platform that's still standing in five years - state, memory, tool boundaries, sub-agent decomposition, evaluation, failure modes, cost. You'll be asked to critique our current architecture in the interview, and we expect you to find things.
  3. Heavily hands-on. You ship. Deep production experience, writing the code yourself, at pace.

Must have

  • 6+ years shipping production software, of which 2+ on LLM systems that real users touched.
  • Real agentic depth: ReAct or equivalent loops, tool/function calling, planners, state & checkpointing, long-term memory, HITL steps, streaming. Not "I called the OpenAI API."
  • Hands-on LangGraph (or a strong argument for something better) plus a tracing/observability stack - LangSmith, Langfuse or similar.
  • Python in production: FastAPI, async job processing, clean service boundaries.
  • RAG done properly: retrieval + re-ranking + relevance judgement, and honest evaluation of all three.
  • Evaluation as a habit, not an afterthought (RAGAS/DeepEval/GEVAL, LLM-as-judge, pass@k).
  • Cloud production experience - AWS (Bedrock, SQS, EC2/EKS, S3) or equivalent - with Docker and CI/CD.
  • The spine pushes back on us. We explicitly want someone who asks "why did you build it like that?"

Nice to have

  • LLM post-training: SFT with LoRA/QLoRA, preference/RL methods, reward and verifier design.
  • Voice agents (LiveKit/Pipecat or similar), or vision-language work.
  • Multi-agent orchestration and MCP integrations.
  • Text-to-SQL over a real, messy production schema.
  • Report/document generation agents - structured, sourced, client-ready output.
  • Guardrails, PII handling, hallucination detection, risk scoring.
  • Interest in watches, collectibles or market data. Not required - curiosity about the domain is.

Benefits

What you get

  • Ownership of EverWatch's entire agent layer, and the roadmap for it — not a corner of someone else's.
  • A dataset that doesn't exist anywhere else, and users who care whether the answer is right.
  • Direct line to the CTO; decisions in days, not quarters.
  • Budget for models, tooling, and the coding agents you want to work with.
  • Competitive compensation, remote-first.
From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Workable (global search)'s own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Workable (global search)'s form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Workable (global search)'s answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.