Skip to content

Software Engineering Manager, AI/ML Infrastructure and Performance Engineering

Google

Bengaluru, Karnataka, India

Applying for this one?

We write the CV against this exact posting — its wording, its requirements — not a template with your name in it.

Get my CV for this job

$25, one-time. No subscription.

In most instances, this position requires in-person interviews as part of the hiring process.

Minimum qualifications:

  • Bachelor's degree or equivalent practical experience.
  • 8 years of experience with software engineering, machine learning infrastructure, computer architecture, distributed computing, people management, debugging tools, communication.
  • Experience with people management, and building infrastructure to improve performance of Machine Learning (ML) systems and applications.

Preferred qualifications:

  • Experience with Performance Optimization, Graphics Processing Unit (GPU) Programming, High Performance Computing, Large Language Model, Open Source Contributor.
  • Experience with the Machine Learning infra and frameworks and hands-on experience with GPU or Tensor Processing Unit (TPU) performance analysis.
  • Experience in building agentic workflows for performance debugging and optimization. Ability to generate ideas and resolve ambiguity.
  • Hands-on experience with Machine Learning (ML) frameworks such as TensorFlow, JAX, and PyTorch, Keras. Experience with ML Inference frameworks such as vLLM, SG Lang, Pathways and experience in open-source software development, including experience in releasing and supporting open-source projects.

About the job

Like Google's own ambitions, the work of a Software Engineer goes beyond just Search. Software Engineering Managers have not only the technical expertise to take on and provide technical leadership to major projects, but also manage a team of Engineers. You not only optimize your own code but make sure Engineers are able to optimize theirs. As a Software Engineering Manager you manage your project goals, contribute to product strategy and help develop your team. Teams work all across the company, in areas such as information retrieval, artificial intelligence, natural language processing, distributed computing, large-scale system design, networking, security, data compression, user interface design; the list goes on and is growing every day. Operating with scale and speed, our exceptional software engineers are just getting started -- and as a manager, you guide the way.

With technical and leadership expertise, you manage engineers across multiple teams and locations, a large product budget and oversee the deployment of large-scale projects across multiple sites internationally.

Our team is focused around driving continuous improvements to the machine learning software/hardware stacks through providing insightful performance debugging for workloads and custom kernels. We provide insights by summarizing different views of captured profile data - such as trace timelines, memory usage, Compiler profiles, ML graph summariesThe ML, Systems, & Cloud AI (MSCA) organization at Google designs, implements, and manages the hardware, software, machine learning, and systems infrastructure for all Google services (Search, YouTube, etc.) and Google Cloud. Our end users are Googlers, Cloud customers and the billions of people who use Google services around the world.

We prioritize security, efficiency, and reliability across everything we do - from developing our latest TPUs to running a global network, while driving towards shaping the future of hyperscale computing. Our global impact spans software and hardware, including Google Cloud’s Vertex AI, the leading AI platform for bringing Gemini models to enterprise customers.

Responsibilities

  • Learn and build an intuitive understanding of existing data collection, analysis, and visualization workflows with deep introspection across Frameworks, Accelerated Linear Algebra (XLA) and runtime stack.
  • Support new and exciting ML paradigms (such as horizontal scaling for upcoming TPU chips) by making contributions across the end to end stack and analysis tools. Partner with ML Stack leads to understand model optimization use cases and bring debugging to feel native in 3P environments (VSCode, Cursor, Grafana, etc).
  • Work with OSS ML inference frameworks such as vLLM, SGLang to provide insights in Xprof about performance improvement opportunities.
  • Partner with other teams that own various parts of the ML stack to understand performance optimization use cases. Work with OSS ML inference frameworks such as TorchTPU, vLLM, SGLang to provide insights into performance bottlenecks.

Seen 1 hour ago · Google postings close after a median of 29 days.

Original posting on Google's site ↗

Posting text belongs to the employer. Removal requests: contact us.

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

One job at a time

One posting. One CV. $25.

Pick the job you actually want and we write for it.

Get my CV for this job