Skip to content

Open nowPosted 3 days ago

Staff Data Engineer — Twenty

General Catalyst portfolio5,558 open roles

Where
New York, NY, USA
Work mode
On site
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowStaff Data Engineer — TwentyGeneral Catalyst portfolio · New York, NY, USA
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on General Catalyst portfolio's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

7.8% of postings close within 7 days. Measured by our own scanner across the market. General Catalyst portfolio postings stay open a median of 28 days.

Share of postings closed within
  1. 1.7%1 day
  2. 3.5%3 days
  3. 7.8%7 days
  4. 14.6%14 days
  5. 34.0%30 days
This job: posted 3 days ago

General Catalyst portfolio median: 28 days open

The posting

About the Company

America is under sustained cyber attack. Our adversaries infiltrate our networks, steal our IP, and degrade the digital infrastructure that modern life runs on. They’ve learned—correctly—that those attacks rarely produce consequences.

Twenty was founded to change that, by making our adversaries think twice before they attack us. Our vision is American and allied primacy in cyberspace—a future where they cannot contest us, deterrence is assured, and the free world remains secure.

Founded in 2024, Twenty Technologies (www.twenty.io) industrializes offensive cyber operations for the U.S. and its allies. Headquartered in Arlington, Virginia, Twenty has raised $168M from Khosla Ventures, Accel, Caffeinated Capital, Friends & Family Capital, Point72 Ventures, General Catalyst, and In-Q-Tel.

Mission | On Site | Full Time | U.S. Citizenship Required / No Active Clearance Required

Role Summary

You'll be based in Twenty's New York City office, building and operating the data integrations that feed Twenty's mission-critical platform. Twenty builds software for cyber operations, and the applications this role feeds are used by operators and analysts working against hard national security problems. This is a hands-on engineering role with real production ownership: you'll build the pipelines that retrieve, process, and integrate new data sources into our applications, monitor the health of the data flowing through the systems you can reach, and debug production issues when they surface. Where your work runs in environments you can't access, you'll partner with Twenty's forward deployed engineers, who operate on-site at customer facilities, to get it deployed and keep it healthy. You'll join the data platform team, reporting to its engineering manager. The data integration tooling you'll build on is young and under active development. You'll be among its first production users, and what you learn in operation will directly shape how it evolves.

If you're an experienced data engineer who wants your work in the hands of operators the same week you build it, and you'd rather own a production system end to end than ship features into a backlog, this role is for you.

What You'll Do

Data Source Integration & Pipeline Development

  • Design and build data pipelines that retrieve, parse, transform, and load new data sources into Twenty's applications, primarily as AWS Glue jobs written in PySpark.
  • Investigate unfamiliar source systems and data formats, working with Twenty's engineers to understand semantics, access patterns, and constraints before writing code.
  • Design schemas and data models that fit the platform's performance characteristics and the shape of the data, including analytical models in ClickHouse.
  • Write code that is testable, debuggable, and maintainable by engineers operating it in environments you can't access. Tests carry unusual weight here: they're how the field trusts a change you'll never see run.
  • Write integration code that meets the security standards of classified environments: no hardcoded secrets, disciplined credential handling, and audit-ready practices throughout, so your work moves through customer approval processes without friction.

Deployment & Field Support

  • Deploy, operate, and iterate on integrations against the systems reachable from the office, end to end.
  • Package pipelines and integrations destined for restricted environments so forward deployed engineers and customer deployment teams can install, verify, and operate them without you in the room.
  • Support field deployments remotely: reproduce issues, cut fixes, and move them through the pipeline quickly when an on-site engineer is blocked.
  • Operate and improve pipelines built elsewhere on the team, and feed operational learnings back into their design.

Production Health & Monitoring

  • Monitor the health of data pipelines and platform services using the LGTM stack (Grafana, Loki, Tempo, Mimir).
  • Track data quality, throughput, and freshness; identify anomalies and degradations before they become incidents.
  • Build and improve the dashboards and alerts that make pipeline health visible to the team and to the forward deployed engineers who depend on it.

Incident Response & Debugging

  • Diagnose and resolve production issues in the data path: failed ingests, malformed source data, pipeline stalls, performance degradation.
  • Own resolution of incidents within your sphere of responsibility, escalating with a clear picture of what you found, not just what triggered the alert.
  • Drive root cause analysis and post-incident reviews, and turn findings into runbooks, fixes, or upstream design changes.
  • Participate in on-call rotation using PagerDuty, with clear escalation paths across the engineering team.

Field & Platform Collaboration

  • Work closely with forward deployed engineers on-site at customer facilities, sharing context in both directions so field reality and platform design stay connected.
  • Turn operational observations, data source opportunities, and recurring patterns from the field into platform capabilities rather than one-off fixes.

Who You Are

  • You independently identify the right solution to ambiguous, open-ended problems. New data sources arrive undocumented and messy; you investigate, form a design, and deliver without waiting for a spec.
  • You respond with urgency to operational issues and own resolution within your sphere. You're unafraid to call an incident when the signal warrants it.
  • You proactively create and update runbooks and documentation for the components you own. Some of the environments your code runs in are ones you'll never touch, and what you write down is what the engineers there have to work with.
  • You make informed decisions by consulting the right stakeholders and balancing detail with the big picture, and you execute against the spirit of the requirement, not just the letter.
  • You actively seek out and eliminate sources of toil. Repetitive manual work is a design problem, and you treat it like one.
  • You understand the customer and the mission well enough to know which work has the greatest impact, and you redirect your focus when what you're doing isn't moving the needle.
  • You tailor your communication to your audience, whether that's a forward deployed engineer at a customer site, a product engineer in Arlington, or a teammate across the desk, and you proactively share information so the right people stay informed and aligned.
  • You're comfortable building for restricted environments where deployment is gated, feedback loops are longer than you'd like, and disciplined engineering practice is the price of admission.

Must Have

  • 5+ years of experience in data engineering, or software engineering with a substantial data infrastructure focus.
  • Expert Python skills, with a track record of designing and maintaining production-grade codebases that other engineers extend and depend on.
  • Hands-on experience building and operating ETL/ELT pipelines with Spark (PySpark), ideally as AWS Glue jobs.
  • Strong SQL and schema design skills, including the ability to analyze access patterns and make sound partitioning and indexing decisions.
  • Experience with column-oriented analytical databases (example technologies: ClickHouse, Redshift, BigQuery).
  • Experience debugging production data systems: reading logs, tracing failures across components, and root-causing issues under time pressure.
  • Demonstrated ability to work independently on production systems, including sound judgment about when to decide and when to escalate.
  • Ability to work on-site full-time at Twenty's New York City office, with an initial onboarding period at our Arlington, VA office and occasional travel there afterward.

Nice to Have

  • Experience with data lake and large-scale query technologies (example technologies: Apache Iceberg, Delta Lake, Trino, Presto, Athena).
  • Experience with streaming and message queue technologies (example technologies: Kafka, NATS, Kinesis, RabbitMQ).
  • Hands-on experience with observability tooling (Grafana, Datadog, Splunk, or similar).
  • Experience deploying and operating software in air-gapped, classified, or otherwise restricted environments.
  • Familiarity with containerized environments (Docker and container orchestration) and CI/CD concepts.

Tech Environment (You Might Work With)

  • Cloud: AWS (EC2, ECS, RDS, CloudWatch), including deployments into closed enclave environments
  • Data processing: AWS Glue jobs running PySpark
  • Analytics: ClickHouse, with ClickPipes for ingestion
  • Containers: Docker, orchestrated on Amazon ECS
  • Observability: Grafana, Loki, Tempo, Mimir (LGTM stack)
  • Alerting / on-call: PagerDuty
  • Messaging: NATS
  • Databases: PostgreSQL
  • Data integration: Twenty's purpose-built data platform and integration layer (in active development)
  • Languages in use across engineering: Go, TypeScript, Python

Security / Work Environment

This role requires U.S. citizenship and eligibility to obtain a U.S. Government security clearance; no active clearance is required to start. Work is performed at Twenty's New York City office. Onboarding begins with a stint of roughly two weeks at Twenty's Arlington, VA office, and travel to Arlington after that is occasional and as needed. Some of the systems this role builds for are restricted, and tooling constraints apply when working with them.

Benefits What's on the table:

  • Health. Medical, dental, and vision plan options. Life / AD&D, disability coverage options.
  • Family. Paid parental leave for eligible full-time employees. 12 weeks for birthing parents, 4 for non-birthing parents, 6 weeks for adoptive, foster, or intended parents through surrogacy.
  • Vacation. Paid holidays and flexible PTO. Take what you need.
  • Retirement. 401(k) with pre-tax and Roth options. HSA/FSA options, dependent care FSA.

Benefits vary by location, role, and eligibility. Full plan details provided during the interview and offer process.

If this role sounds like you, apply and share with us your interest

Due to U.S. government contract and security requirements, this role is limited to U.S. citizens. Some positions may also require eligibility to obtain and maintain a U.S. Government security clearance. Any active clearance requirement will be listed in the role description.

Twenty is an equal opportunity employer. We consider all qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, veteran status, disability, or any other protected status, consistent with applicable law.

If you need a reasonable accommodation during the hiring process, let us know and we will work with you.

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against General Catalyst portfolio's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on General Catalyst portfolio's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    General Catalyst portfolio's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.