Skip to content

Open nowPosted 3 hours agoWe saw it 72 min after it went up

Cluster Datacenter Manager

Cerebras120 open roles

Where
United States
Work mode
On site
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowCluster Datacenter ManagerCerebras · United States
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Cerebras's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

8.1% of postings close within 7 days. Measured by our own scanner across the market. Cerebras postings stay open a median of 32 days.

Share of postings closed within
  1. 1.8%1 day
  2. 3.5%3 days
  3. 8.1%7 days
  4. 15.1%14 days
  5. 33.9%30 days
This job: posted 3 hours ago

Cerebras median: 32 days open

The posting

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.

Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership https://openai.com/index/cerebras-partnership/ with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.

THE ROLE

Cerebras is seeking a hands-on Datacenter Cluster Manager to lead compute infrastructure operations and operational reliability across multiple data center sites. This leader manages technicians, engineers and contractors responsible for the health, availability, physical maintenance and lifecycle of high-performance compute and network infrastructure. The role requires strong technical judgment in server systems, networking, fiber connectivity and secure asset handling, alongside a working understanding of power, cooling and colocation dependencies. The Cluster Manager leads incident response, drives operational discipline, and ensures new deployments transition safely into sustained production support.

RESPONSIBILITIES

COMPUTE SYSTEMS OPERATIONS AND RELIABILITY

- Own day-to-day operational health and service readiness of compute clusters, server racks, supporting network infrastructure and associated physical systems across assigned sites.

- Lead hardware incident triage, fault isolation, break/fix, field-replaceable unit (FRU) replacement, diagnostic validation and escalation to systems engineering or vendors.

- Use monitoring, telemetry, logs and alerting to recognize degraded systems, prioritize impact, coordinate recovery and verify restoration of service.

- Oversee preventive maintenance, firmware or hardware change execution where authorized, spare parts readiness, and repeat-failure analysis.

- Partner with cluster ops , network and reliability engineering teams on root cause analysis, corrective actions and recurring fleet health issues.

NETWORK INFRASTRUCTURE AND FIBER TROUBLESHOOTING

- Apply practical understanding of data center network architecture, rack-level connectivity, switching, management networks and network redundancy to guide onsite diagnosis and escalation.

- Lead physical-layer troubleshooting of fiber and copper links, including patching, optics and transceivers, polarity, cleanliness, labeling and continuity testing using appropriate tools.

- Ensure structured cabling, fiber routing, documentation and change controls are maintained to engineering standards.

- Coordinate with network engineering to isolate physical connectivity faults versus configuration or software issues; execute approved remediation without assuming ownership of network design.

SECURE MEDIA, ASSETS AND HARDWARE LIFECYCLE

- Own physical asset accountability from inbound receiving, inspection, staging and inventory through deployment, repair, return and disposition.

- Enforce secure handling of data-bearing devices, including authorization, chain of custody, access controls, approved sanitization or destruction workflows and documented transfer.

- Ensure adherence to customer-specific controls, including two-person authorization and clean-in/clean-out procedures where applicable.

- Maintain accurate rack elevations, asset records, serial numbers, spares, repair histories and inventory reconciliation.

PEOPLE LEADERSHIP AND CLUSTER EXECUTION

- Lead, coach and evaluate technicians, engineers and contractor resources across multiple facilities; establish clear ownership, training and performance expectations.

- Plan staffing, shift coverage, on-call rotations and escalation readiness based on cluster demand, production service commitments and deployment schedules.

- Prioritize work across locations, protect production operations during concurrent builds and maintain consistent SOPs, MOPs, ticketing and change management.

- Own incident command, stakeholder communications, post-incident reviews and measurable corrective action tracking.

DEPLOYMENT, EXPANSION AND PRODUCTION HANDOVER

- Oversee operational readiness for new compute capacity, including receiving, rack integration, network and fiber validation, power-on checks, system health verification and production acceptance.

- Coordinate installation, commissioning, vendor work and handover with deployment, infrastructure, engineering and program teams.

- Confirm documentation, tooling, spares, staffing, access and escalation paths are ready before accepting operational ownership.

CRITICAL FACILITIES, VENDORS AND OPERATIONAL CONTROLS

- Understand dependencies between compute availability and electrical distribution, UPS, generators, PDUs/RPPs, HVAC, liquid cooling, CDUs and environmental monitoring.

- Coordinate with colocation providers and facilities specialists on maintenance risk, alarms, degraded cooling or power conditions and incident restoration.

- Manage vendor performance, site safety, access controls, maintenance windows, service commitments, operating expenses and contractor utilization.

REQUIRED QUALIFICATIONS

- 10+ years of experience in data center, compute infrastructure, server operations or related technical operations, including 5+ years leading technical teams.

- Demonstrated experience managing multi-site operations or complex data center environments with 24/7 production support expectations.

- Strong hands-on foundation in server hardware, rack infrastructure, hardware diagnostics, break/fix and systems health monitoring.

- Working understanding of data center network topology and physical connectivity, including Ethernet, fiber optics, transceivers and structured cabling troubleshooting.

- Experience with asset management, data-bearing media security, chain of custody and controlled equipment movement.

- Experience leading incidents, change control, technical escalations, vendor coordination and deployment-to-operations handovers.

- Ability to interpret facility power and cooling dependencies and engage subject matter experts to manage compute service risk.

- Clear communication, sound technical judgment and experience developing technicians and engineers.

PREFERRED QUALIFICATIONS

- Experience supporting large compute clusters or high-density server environments.

- Exposure to direct liquid cooling, CDU operations, thermal telemetry and the relationship between cooling performance and system health.

- Familiarity with Linux-based troubleshooting, BMC/IPMI or equivalent out-of-band management, server logs and hardware telemetry.

- Experience using fiber inspection and testing tools, network test equipment and disciplined physical-layer troubleshooting procedures.

- Experience with Grafana, Prometheus, PagerDuty, Jira, DCIM/BMS or equivalent monitoring and workflow tools.

- Knowledge of customer security requirements, colocation operating models, SLAs and relevant compliance controls.

WHAT SUCCESS LOOKS LIKE

- Healthy, available compute and network infrastructure with fast detection, structured triage and durable resolution of recurring failures.

- Consistent physical-layer troubleshooting, secure media handling and accurate hardware and asset records across the cluster.

- Skilled onsite teams with reliable coverage, clear escalation ownership and consistent operational execution.

- New compute capacity enters production with verified readiness, documented ownership and minimal impact on live services.

- Effective coordination with facilities partners on power and cooling risks without losing focus on compute operations.

WORKING EXPECTATIONS

- Regular onsite presence and travel between assigned cluster facilities, with additional travel for deployments or operational priorities.

- Availability to lead critical incidents, scheduled maintenance and deployment activities outside standard hours when needed.

- Specific location, reporting line, travel expectations and compensation to be confirmed for each requisition.

Why Join Cerebras

People who are serious about software make their own hardware. At Cerebras, we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we’ve reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras:

1. Build a breakthrough AI platform beyond the constraints of the GPU.

2. Publish and open source their cutting-edge AI research.

3. Work on one of the fastest AI supercomputers in the world.

4. Enjoy job stability with startup vitality.

5. Our simple, non-corporate work culture that respects individual beliefs.

Find out more about what it's like to work at Cerebras here https://www.cerebras.ai/join-us!

Apply today and become part of the forefront of groundbreaking advancements in AI!

Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.

This website or its third-party tools process personal data. For more details, click here https://www.cerebras.net/privacy/ to review our CCPA disclosure notice.

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Cerebras's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Cerebras's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Cerebras's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.