ML Hardware Achitect
Applying for this one?
We write the CV against this exact posting — its wording, its requirements — not a template with your name in it.
$25, one-time. No subscription.
Note: By applying to this position you will have an opportunity to share your preferred working location from the following: Tel Aviv, Israel; Haifa, Israel.
Minimum qualifications:
- Bachelor's degree in Computer Engineering, Electrical Engineering, Computer Science, a related field, or equivalent practical experience.
- 15 years of experience in computer architecture, ML accelerator design, or high-performance processor architecture.
- Experience leading architectural definition and authoring architecture specifications for silicon or compute IP blocks.
- Experience with performance modeling, workload profiling, and hardware-software co-design.
Preferred qualifications:
- Master's degree or PhD in Electrical Engineering, Computer Engineering, or Computer Science with an emphasis on computer architecture or ML hardware systems.
- 5 years of experience leading the architectural definition and microarchitecture of AI/ML accelerators from concept through production.
- Deep knowledge of modern deep learning workloads (Transformers, MoE, Diffusion, Generative AI inference) and their system bottlenecks (memory capacity, KV cache bandwidth, interconnect scaling).
- Strong understanding of high-performance memory subsystems (custom SRAM architectures, high-bandwidth memory hierarchies, caching schemes).
- Experience working with modern ML frameworks (PyTorch, JAX, TensorFlow) and ML compilers/runtimes (XLA, TVM, Triton).
About the job
In this role, you’ll work to shape the future of AI/ML hardware acceleration. You will have an opportunity to drive cutting-edge TPU (Tensor Processing Unit) technology that powers Google's most demanding AI/ML applications. You’ll be part of a team that pushes boundaries, developing custom silicon solutions that power the future of Google's TPU. You'll contribute to the innovation behind products loved by millions worldwide, and leverage your design and verification expertise to verify complex digital designs, with a specific focus on TPU architecture and its integration within AI/ML-driven systems.
In this role, you will help shape the future of Google Cloud’s next-generation AI infrastructure, architecting high-performance Machine Learning silicon designed to power hyperscale AI inference. You will have an opportunity to drive accelerator technology that powers Generative AI models, large language models (LLMs), and emerging agentic workloads where throughput, latency, memory bandwidth, and energy efficiency are mission-critical.
You will be part of a silicon architecture team pushing the boundaries of custom computing. Leveraging your deep expertise in hardware-software co-design, machine learning algorithms, and computer architecture, you will define and optimize custom compute engines and memory hierarchies that accelerate the world's most advanced AI models across Google Cloud datacenters.
The AI and Infrastructure team is redefining what’s possible. We empower Google customers with breakthrough capabilities and insights by delivering AI and Infrastructure at unparalleled scale, efficiency, reliability and velocity. Our customers include Googlers, Google Cloud customers, and billions of Google users worldwide.
We're the driving force behind Google's groundbreaking innovations, empowering the development of our cutting-edge AI models, delivering unparalleled computing power to global services, and providing the essential platforms that enable developers to build the future. From software to hardware our teams are shaping the future of world-leading hyperscale computing, with key teams working on the development of our TPUs, Vertex AI for Google Cloud, Google Global Networking, Data Center operations, systems research, and much more.
Responsibilities
- Lead the architectural definition, modeling, and specification of next-generation, high-performance ML compute IP and acceleration blocks for Cloud AI silicon.
- Own the ML IP architecture specification throughout the entire product lifecycle: concept exploration, cycle-accurate modeling, implementation, silicon bring-up, and production.
- Partner closely with leading AI research and algorithm teams (e.g., Google DeepMind, Gemini research teams) and software compiler teams (XLA, PyTorch) to explore architectural trade-offs and define hardware requirements for emerging model architectures.
- Drive comprehensive architecture studies, evaluating compute dataflows, numerical formats, sparsity, and specialized acceleration mechanisms such as key-value (KV) cache optimization.
- Drive performance, latency, power efficiency, and silicon area projections across model topologies and workload configurations.
Seen 14 hours ago · Google postings close after a median of 29 days.
Original posting on Google's site ↗
Posting text belongs to the employer. Removal requests: contact us.
Nearby
Live postings like this one
Same employer first, then the same role elsewhere.
- 1h ago
- 1h ago
- 1h ago
Senior Director, Centralized Manufacturing Technical Operations
Sunnyvale, CA, USA; Austin, TX, USA; Redwood City, CA, USA
1h ago- 1h ago
Land Development Manager, Data Centers
Kirkland, WA, USA; Addison, TX, USA; New York, NY, USA; Reston, VA, USA; San Francisco, CA, USA; Sunnyvale, CA, USA; Thornton, CO, USA
1h agoBusiness Strategy and Operations Program Manager, Data Center Operations
Sunnyvale, CA, USA; Atlanta, GA, USA; Council Bluffs, IA, USA; Midlothian, TX, USA; The Dalles, OR, USA; Leesburg, VA, USA; Lenoir, NC, USA; Sparks, NV, USA; San Francisco, CA, USA
1h ago- 1h ago
One job at a time
One posting. One CV. $25.
Pick the job you actually want and we write for it.