Skip to content

Network Operations Consultant

Gruve

Pune, Maharashtra, India

About Gruve

Gruve is an innovative software services startup dedicated to transforming enterprises to AI powerhouses. We specialize in cybersecurity, customer experience, cloud infrastructure, and advanced technologies such as Large Language Models (LLMs). Our mission is to assist our customers in their business strategies utilizing their data to make more intelligent decisions. As a well-funded early-stage startup, Gruve offers a dynamic environment with strong customer and partner networks.

Position summary:

Technical lead for the NOC / Network & Device Management track and L3 for the PulseAI infrastructure layers. Major-incident commander for network and infrastructure events across the AI Fabrik estate and PulseAI customer environments (GPU servers, control-plane/infrastructure nodes, OpenShift cluster networking, front-end and RoCEv2 back-end fabrics, switches and storage), owner of network operational quality, device lifecycle and change governance, vendor engineering liaison for network and hardware vendors, and the customer-facing voice for network and infrastructure service reviews; partners with the Security Operations Consultant on cross-domain incidents and deputises for the Security Operation Manager on infrastructure matters.

Key responsibilities:

  • Command Severity 1 / P1 major incidents on the network and infrastructure side end to end — fabric, edge and firewall outages, GPU-server and node hardware faults, RoCEv2 back-end fabric degradation, OpenShift cluster-networking and node-level failures, storage outages — until management command engages; own the technical bridge and the SLA status cadence to customer authorised contacts.
  • Own root cause analysis for network and infrastructure incidents: complex RCA on EVPN-VXLAN/BGP, fabric and edge health, interface/optics and firewall issues across the in-scope Juniper data-center devices (QFX, EX, MX) and Cisco cdFMC/FMC, FTD/SRX; GPU-node, NIC/RDMA and storage-path faults; and OpenShift node lifecycle, MachineConfig and OVN-Kubernetes/CNI cluster-networking problems — with structured problem management to eliminate repeat incidents.
  • Chair change advisory for the network estate and the PulseAI infrastructure layers, aligned with the intent-based (Apstra) fabric design: approve and execute high-risk changes — fabric node additions, routing changes, firewall policy pushes, firmware/BIOS/GPU driver updates, OpenShift node-pool and MachineConfig changes, node drain/reboot cycles — enforcing maintenance windows, customer notice and the pre-change backup gate in coordination with the Security Operations Consultant.
  • Own the device and infrastructure lifecycle and preventive-maintenance roadmap: firmware/BIOS/driver currency, EOL/EOS tracking and spares posture across the Juniper fabric and edge devices, Cisco firewalls, PulseAI GPU and control-plane nodes, RoCEv2 fabric switches and storage; own GPU-cluster serviceability and capacity-governance inputs.
  • Own infrastructure observability standards for the NOC pod: fabric, GPU-cluster, OpenShift node and storage health dashboards in Grafana, streaming/SNMP/syslog telemetry coverage, alert thresholds for the SLA appendix, and the diagnostic runbooks for the GPU / switch / storage / fabric layers of the monitor / remediate / escalate matrix.
  • Lead the infrastructure side of customer onboarding: validate GPU servers, nodes, front-end and back-end fabric, switches and storage against the environment validation checklist (telemetry reachability per switch, storage throughput, access path via outbound collector, site-to-site VPN or jump host) within the 14-day Ready-for-Install window; own the network and infrastructure prerequisites for OpenShift cluster deployment.
  • Act as escalation interface and vendor engineering liaison for network and hardware: Juniper and Cisco TAC end to end (including RMA and smart-hands coordination; Cloudflare where a network issue touches that platform), NVIDIA for GPU hardware, NIC/RDMA and fabric issues, OEM/neocloud and storage vendors for hardware faults; interface to customer network engineering.
  • Own network and infrastructure SLAs/KPIs; lead the network-track input into the monthly operational reviews and quarterly business reviews alongside the Service Delivery Manager.
  • Approve runbook and authority-matrix changes for the NOC pod; develop team capability — certification plan and cross-training matrix (Juniper, Cisco, Red Hat OpenShift, Kubernetes networking, NVIDIA) for the NOC bench — and drive shift-quality audits and post-incident reviews; coordinate with the Security Operations Consultant on cross-domain incidents and shared reporting.

Mandatory Qualifications:

  • 8–11 years network operations / data-center infrastructure experience, including prior senior/lead experience in a 24×7 NOC or infrastructure-operations team.
  • Incident-command track record on network/infrastructure P1s with customer-visible outcomes.
  • Architecture-level grasp of BGP, EVPN-VXLAN fabric design and Juniper platforms (QFX, EX, MX), Cisco cdFMC/FMC, FTD/SRX; demonstrated independent RCA ownership on P1/P2/P3 network incidents.
  • Deep Kubernetes/OpenShift infrastructure operations experience in production — node lifecycle and MachineConfig, cluster networking (OVN-Kubernetes / CNI, Cilium), Services/Ingress/LoadBalancer and east-west flows, storage (CSI/NFS), oc/kubectl diagnostics — with the ability to troubleshoot fabric-to-cluster (GKE and OpenShift) connectivity end to end and separate cluster-side from network-side faults; Linux (RHEL) administration.
  • GPU-cluster infrastructure experience: NVIDIA GPU Operator and driver/firmware lifecycle, DCGM-class telemetry, RDMA/RoCEv2 NIC health, lossless-Ethernet fabric behaviour (PFC/ECN), NVLink/NVSwitch topologies and common GPU hardware failure modes on RTX PRO 6000 / HGX B300-class servers or equivalent.
  • Experience operating against contractual SLAs (acknowledgement, restoration, availability, service credits), chairing change governance, and running vendor TAC / engineering escalations through to fix.
  • Crisp written and verbal customer reporting; automation mindset

Preferred Qualifications:

  • JNCIE/JNCIP or CCIE/CCNP; ITIL Foundation/Practitioner.
  • Apstra or other intent-based networking exposure; IaC and Python/Ansible automation for operations.
  • Red Hat OpenShift certifications (EX280 / EX380) or RHCE; CKA; exposure to AI/ML workload scheduling and GPU node pools on OpenShift.
  • GPU-cluster performance troubleshooting — RoCEv2/PFC/ECN tuning, ECMP polarisation, NCCL-visible latency/jitter — and high-performance fabric telemetry.
  • Data-center/hyperscale or neocloud network operations background; presales/solutioning exposure.
  • Working knowledge of GCP networking (VPC-native GKE, Cloud Load Balancing) and Cloudflare where network issues touch those platforms.
  • JNCIE/JNCIP or CCIE/CCNP; ITIL Foundation/Practitioner.
  • Apstra or other intent-based networking exposure; IaC and Python/Ansible automation for operations.
  • Red Hat OpenShift certifications (EX280 / EX380) or RHCE; CKA; exposure to AI/ML workload scheduling and GPU node pools on OpenShift.

Why Gruve

At Gruve, we foster a culture of innovation, collaboration, and continuous learning. We are committed to building a diverse and inclusive workplace where everyone can thrive and contribute their best work. If you’re passionate about technology and eager to make an impact, we’d love to hear from you.

Gruve is an equal opportunity employer. We welcome applicants from all backgrounds and thank all who apply; however, only those selected for an interview will be contacted.

Seen 27 hours ago · Gruve postings close after a median of 13 days.

Original posting on Gruve's site ↗

Posting text belongs to the employer. Removal requests: contact us.

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

One job at a time

One posting. One CV. $25.

Pick the job you actually want and we write for it.

Get my CV for this job