The posting
Mission: Groq is building high-performance AI infrastructure designed to make inference fast, predictable, and scalable. Our Network Engineering & Datacenter teams design and operate the systems that provide the compute, storage, and networking capabilities behind Groq's rapidly growing AI infrastructure.
We are looking for a Storage Engineer, Deployment & Support to deploy, validate, and operate the storage infrastructure behind Groq's global AI infrastructure. As a Storage Engineer, Deployment & Support, you will take approved storage designs from implementation planning through production deployment and operational handoff. You will own storage system bring-up, configuration, validation, performance acceptance, troubleshooting, expansion, upgrades, and ongoing operational health across large-scale AI/HPC environments.
This role owns the underlying storage infrastructure lifecycle and partners closely with Platform Engineering to ensure storage capabilities are reliable, production-ready, and consumable. The role requires strong storage and networking fundamentals, structured troubleshooting, performance analysis, and the ability to drive deployments and operational work across multiple teams and vendors.
Responsibilities & opportunities in this role:Deploy and validate storage infrastructure across new and existing Groq datacenters, including high-performance storage systems supporting AI/HPC workloads. Configure and bring up storage systems against approved designs and implementation standards, including cluster initialization, storage nodes, interfaces, protocols, and supporting infrastructure. Drive storage deployment execution end-to-end, directing and coordinating on-site datacenter technicians and vendors through physical rack/stack, cabling, and installation while owning configuration, validation, troubleshooting, performance acceptance, and production handoff. Develop and maintain implementation procedures, validation plans, deployment checklists, system inventories, as-built documentation, and operational runbooks. Perform storage acceptance and performance validation, including throughput, IOPS, latency, capacity, data distribution, cluster health, and recovery behavior. Monitor and maintain storage health, capacity, availability, and performance, identifying and resolving issues before they impact production workloads. Troubleshoot storage and hardware failures across storage nodes, drives, network paths, protocols, and supporting infrastructure, and perform root cause analysis for deployment and production issues. Execute and coordinate storage expansions, upgrades, hardware refreshes, and other lifecycle activities while minimizing production impact. Manage hardware readiness for deployments and expansions, including rack/stack coordination, inventory and asset record updates, RMA coordination, spares, and vendor shipments. Coordinate deployment audits and execute established QA/QC processes to ensure storage infrastructure meets Groq standards. Partner closely with storage design, network, compute and platform teams to identify and resolve blockers throughout deployment and operations. Provide operational support during and after deployments, including maintenance activities, incidents, remediation, break-fix events, and vendor escalation when required. Develop tools and scripts that improve storage deployment, configuration, validation, testing, health checks, and repeatable operational workflows. Capture lessons-learned from deployments and incidents and contribute to global implementation standards and deployment playbooks.
Ideal candidates have/are:4+ years of experience in storage engineering, systems engineering, network engineering, datacenter operations, or related infrastructure roles, with hands-on experience deploying, operating, or troubleshooting storage systems. Strong hands-on experience operating and troubleshooting large-scale distributed or high-performance storage environments. Experience with modern storage platforms and technologies such as VAST, WEKA, DDN, Lustre, Ceph, or comparable distributed storage solutions. Strong Linux systems administration and troubleshooting skills in production environments. Strong understanding of storage protocols and data paths, including NFS, NFS over RDMA, NVMe-oF, and high-throughput Ethernet or InfiniBand connectivity. Strong understanding of storage performance concepts, including throughput, IOPS, latency, capacity, and utilization. Understanding of RDMA, RoCE, InfiniBand, and other high-performance networking concepts used in AI/HPC storage environments. Ability to troubleshoot systematically across storage, Linux, network, and hardware layers and drive issues through root cause and resolution. Experience developing deployment or operational tooling using Python, Bash, Ansible, APIs, or similar technologies. Strong ownership, communication, and time-management skills, with the ability to manage multiple deployments and operational priorities under demanding timelines. Comfortable operating in fast-moving environments where deployment plans and requirements can evolve quickly. Ability to travel to global datacenter locations for deployments, maintenance activities, and other site-specific needs.
Compensation Groq is committed to providing competitive compensation through our Total Cash philosophy, which incorporates potential bonus value directly into base pay. The total cash salary range for this position, which is inclusive of the potential bonus value, is $270,400–$401,600, with individual placement determined by your geographic location, experience, skills, and alignment with internal compensation standards. This range is specific to candidates located in the United States. Compensation for international candidates will vary based on local market dynamics. Beyond cash compensation, Groq also offers a Long-Term Incentive (LTI) Program and a robust suite of employee benefits.
US Job Posting This position may require access to technology and/or information subject to U.S. export control laws and regulations, including the Export Administration Regulations (EAR). To comply with these requirements, candidates for this role must meet certain citizenship or residency criteria. Specifically, they must qualify as U.S. Persons for export control purposes (i.e., U.S. citizen, U.S. lawful permanent resident (Green Card holder), or a protected individual under 8 U.S.C. § 1324b(a)(3) such as a refugee or asylee), or otherwise be eligible for an applicable export license.



