Skip to content

Open nowPosted today

Cyberinfrastructure Engineer

University of Chicago359 open roles

Where
Chicago, IL
Work mode
On site
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowCyberinfrastructure EngineerUniversity of Chicago · Chicago, IL
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on University of Chicago's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

8.1% of postings close within 7 days. Measured by our own scanner across the market. University of Chicago postings stay open a median of 29 days.

Share of postings closed within
  1. 1.7%1 day
  2. 3.6%3 days
  3. 8.1%7 days
  4. 15.0%14 days
  5. 33.9%30 days
This job: posted today

University of Chicago median: 29 days open

The posting

Department

PSD Enrico Fermi Institute: Gardner Group

About the Department

The MANIAC Lab, located within the Enrico Fermi Institute of the Physical Sciences Division at the University of Chicago, builds and operates advanced cyberinfrastructure for scientific instruments investigating the fundamental mysteries of nature. The Lab operates data-intensive high-throughput computing facilities that are part of a global computing grid used to reconstruct and analyze particle collisions recorded by the ATLAS detector at the Large Hadron Collider (LHC) at CERN in Geneva, Switzerland (https://atlas.cern). The Lab also serves as a gateway for U.S. researchers to national cyberinfrastructure resources through the Open Science Grid Consortium (OSG) and the Institute for Research and Innovation in Software for High Energy Physics (IRIS-HEP).

The Lab operates the NSF Shared Tier 3 Analysis Facility (https://af.uchicago.edu). The facility supports ATLAS physicists analyzing the complete LHC Run 2 and Run 3 datasets as the collaboration prepares for the High-Luminosity LHC (HL-LHC), which begins with Run 4 around 2030. It provides computing infrastructure and services for the South Pole Telescope (SPT-3G) and operates a worldwide distributed data network for the XENON dark matter search experiment at the Gran Sasso National Laboratory in Italy.

The Lab is increasingly focused on AI-enabled research computing. Its GPU platforms support machine learning training and inference, including models served through NVIDIA Triton and large language models hosted on facility hardware. The Lab has also built an agentic AI platform based on the Model Context Protocol (MCP). It gives researchers' AI assistants secure, identity-brokered access to facility services such as HTCondor, Kubernetes, Rucio, and ServiceX. The same platform supports AI-assisted facility operations, where sandboxed agents analyze operational metrics and propose corrective actions for human review.

The Lab provides the Scalable Systems Laboratory (SSL), a software testing and integration platform for IRIS-HEP. IRIS-HEP develops software and computing solutions for the HL-LHC era. Through the SSL Deployment Factory, the Lab is packaging the Kubernetes infrastructure it already operates into tested, versioned bundles that other facilities can deploy. These bundles are bootstrapped with Kubespray and managed through GitOps with Flux, and cover services such as JupyterHub, BinderHub, ServiceX, and the MCP platform. The goal is to shorten a process that currently takes months of manual work at each site.

Job Summary

As the Cyberinfrastructure Engineer at the MANIAC Lab within the Enrico Fermi Institute, you will report directly to Research Assistant Professor Giordon Stark. As part of our team (https://maniaclab.uchicago.edu), you will help build and operate advanced cyberinfrastructure for science collaborations such as the ATLAS experiment at the CERN LHC. These collaborations rely on data-intensive, distributed computing technologies such as HTCondor, Kubernetes, and Ceph.

The Cyberinfrastructure Engineer will help operate the storage, compute, and GPU infrastructure in the University's campus data centers. This infrastructure connects researchers to the national-scale OSG fabric and to emerging AI/ML service platforms. The work combines hands-on Linux systems administration and on-site hardware support with cloud-native operations using Kubernetes, Helm, containers, and GitOps tooling. It also involves distributed storage (Ceph via Rook, ZFS), configuration management (Puppet), and monitoring and alerting (Prometheus, Grafana).

The team is collaborative and partly distributed. We meet weekly over Zoom and coordinate daily in Slack. Team members are expected to take ownership of their work and function independently. The position is jointly supported by the ATLAS Midwest Tier 2 Center (MWT2), focused on production operations, and by IRIS-HEP, focused on building the infrastructure that enables AI agents to drive physics analysis and facility operations for the HL-LHC era.

Responsibilities

  • Facility operations (MWT2 and the UChicago Analysis Facility, approximately 50%)
  • Provides systems administrative services for Linux compute clusters (CPU and GPU), storage systems, and related support servers. These systems support ATLAS production and analysis at the Midwest Tier 2 Center and the Analysis Facility.
  • Operates, upgrades, and scales Ceph distributed storage (including Rook-managed Ceph on Kubernetes). The goal is to meet growing HL-LHC capacity and throughput requirements.
  • Performs on-site hardware work in campus data centers, including installing, retrofitting, and replacing servers, storage, and GPUs. Manages vendor support cases and spare-parts inventory.
  • Maintains monitoring, dashboards, and alerting (e.g., Prometheus, Grafana) for storage, compute, network, and hardware health.
  • Applies operating system and service security patches and vulnerability mitigations in coordination with University security requirements.
  • Performs network diagnostics, throughput measurement, and analysis for both LAN and WAN. Supports the facility's participation in WLCG Data Challenges and HL-LHC readiness scale tests.
  • Participates in the team's operations support rotation and maintains documentation and operational runbooks.
  • Agentic research infrastructure (IRIS-HEP, approximately 50%)
  • Deploys and operates the infrastructure for agent-driven physics analysis. This includes the Lab's MCP gateway, the MCP servers it brokers access to, and the credential services behind them. Together, these let researchers' AI assistants securely use facility resources such as HTCondor, Kubernetes, Rucio, and ServiceX.
  • Supports IRIS-HEP integration challenges and demonstrations of end-to-end agentic analysis workflows. The work spans dataset discovery, batch processing, analysis, and inference run through AI agents.
  • Operates GPU platforms for machine learning and AI workloads, including locally hosted large language models and sandboxed agent environments.
  • Helps develop and operate AI-assisted operations agents that monitor HTCondor, Kubernetes, Ceph, and related services. These agents propose corrective actions under human review.
  • Deploys and operates services using container-based approaches (Docker, Kubernetes, Helm) and GitOps workflows (e.g., Flux). Contributes reusable deployment bundles to the SSL Deployment Factory so other facilities can adopt these services.
  • Learns new distributed computing, AI/ML, and infrastructure-as-a-service technologies.
  • Maintains complex system and network administration functions. Works with moderate guidance to administer simple systems and assists in the administration of larger systems.
  • Installs, configures, and maintains operating system workstations and servers. Performs software installations and upgrades to operating systems and layered software packages. Monitors and tunes the system to achieve optimum performance levels, acquiring higher-level skills in the process.
  • Performs other related work as needed.

Minimum Qualifications

Education:

Minimum requirements include a college or university degree in related field.

Work Experience:

Minimum requirements include knowledge and skills developed through 2-5 years of work experience in a related job discipline.

Certifications:

---

Preferred Qualifications

Education:

  • Bachelor’s degree in Computer Science, Computer Engineering, Physics or related field.

Experience:

  • Strong experience managing Linux operating systems.
  • Configuration management and build systems for large numbers of computers using tools such as Puppet, Chef and Ansible.
  • Experience operating distributed storage at scale, particularly Ceph (including Rook on Kubernetes).
  • Hands-on data center hardware experience, including server diagnostics, component replacement, and working with vendor support (e.g., Dell iDRAC/warranty processes).
  • Experience with monitoring and alerting tools such as Nagios, Prometheus, Grafana, and Alertmanager.

Technical Knowledge or Skills:

  • Unix/Linux operating systems administration tools and shell scripts.
  • Distributed storage systems such as Ceph; local file systems such as ZFS.
  • Git version control, scripting (Bash, Python) and automation.
  • Knowledge and expertise in technologies such as TCP/IP and related protocols; networked file systems, including NFS.
  • Knowledge or experience with batch scheduling systems such as Slurm or HTCondor.
  • Knowledge of container technologies such as Docker, Kubernetes, Helm, OpenShift/OKD, OpenStack.
  • Familiarity with identity and access management (e.g., Keycloak, OAuth/OIDC).

Preferred Competencies

  • Strong oral and written communication skills.
  • Initiative and capacity for teamwork and creativity.
  • Ability to effectively communicate and collaborate with team members, supervisors, and researchers.
  • Ability to manage complex technical details and switch between projects.
  • Ability to work independently with minimal supervision, take ownership of issues through resolution, and keep the team informed.
  • Comfortable collaborating in a distributed team through weekly Zoom meetings and daily Slack communication.
  • Knowledge or experience with Spark, Dask or Ray.
  • Knowledge of GitOps and cluster lifecycle tooling such as Flux, Argo CD, or Kubespray.
  • Experience with GPU servers, including NVIDIA drivers, CUDA, and GPU scheduling in Kubernetes or HTCondor.
  • Familiarity with AI/ML infrastructure, such as model serving, LLM-based tooling, or MCP, is a plus.

Working Conditions

  • Able to work on-site in University data centers on a regular basis to physically install, move, and replace hardware of up to 50lbs/person.

Additional Documents

  • Resume/CV (required)
  • Cover letter (required)
  • Professional Reference Information (preferred)

The University of Chicago uses AI-assisted tools to streamline and augment some recruitment processes; however, AI is not used to make hiring decisions.

When applying, the document(s) MUST be uploaded via the My Experience page, in the section titled Application Documents of the application.

Job Family

Information Technology

Role Impact

Individual Contributor

Scheduled Weekly Hours

37.5

Drug Test Required

No

Health Screen Required

No

Motor Vehicle Record Inquiry Required

No

Pay Rate Type

Salary

​ FLSA Status

Exempt

​ Pay Range

$80,000.00 - $90,000.00

The included pay rate or range represents the University’s good faith estimate of the possible compensation offer for this role at the time of posting.

Benefits Eligible

Yes

The University of Chicago offers a wide range of benefits programs and resources for eligible employees, including health, retirement, and paid time off. Information about the benefit offerings can be found in the Benefits Guidebook.

Posting Statement

The University of Chicago is an equal opportunity employer and does not discriminate on the basis of race, color, religion, sex, sexual orientation, gender, gender identity, or expression, national or ethnic origin, shared ancestry, age, status as an individual with a disability, military or veteran status, genetic information, or other protected classes under the law. For additional information please see the University's Notice of Nondiscrimination.

Job seekers in need of a reasonable accommodation to complete the application process should call 773-702-5800 or submit a request via Applicant Inquiry Form.

All offers of employment are contingent upon a background check that includes a review of conviction history. A conviction does not automatically preclude University employment. Rather, the University considers conviction information on a case-by-case basis and assesses the nature of the offense, the circumstances surrounding it, the proximity in time of the conviction, and its relevance to the position.

The University of Chicago's Annual Security & Fire Safety Report (Report) provides information about University offices and programs that provide safety support, crime and fire statistics, emergency response and communications plans, and other policies and information. The Report can be accessed online at: http://securityreport.uchicago.edu. Paper copies of the Report are available, upon request, from the University of Chicago Police Department, 850 E. 61st Street, Chicago, IL 60637.

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against University of Chicago's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on University of Chicago's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    University of Chicago's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.