The posting
Position Summary
The AI & Automation Operations Manager will lead the design, implementation, and operationalization of AI-driven workflows, intelligent automation, and observability capabilities across IT Operations and IT Service Management (ITSM). This role is responsible for advancing the organization's journey toward AI-augmented and autonomous operations by leveraging AIOps, Observability tools, Service Now and Generative AI technologies to improve service reliability, reduce operational risk, and accelerate incident resolution.
The role will also provide operational support within the Digital Command Center (DCC), supporting major incident management, executive communications, and continuous service improvement initiatives.
Key Responsibilities
Primary Responsibilities
AI Incident Management
- Design and implement AI-powered workflows for incident detection, triage, categorization, prioritization, and routing.
- Develop AI-assisted root cause analysis, incident enrichment, and predictive incident management capabilities.
- Create intelligent remediation runbooks and automated response workflows.
- Drive reduction of alert noise through event correlation and service intelligence.
AI-Driven Problem & Change Management
- Build workflows that identify recurring issues and proactively generate problem investigations.
- Implement AI-enabled trend analysis, risk scoring, and change impact assessments.
- Support automated generation of knowledge articles, PIR summaries, and operational insights.
AI-Augmented Autonomous Operations
- Develop self-healing and closed-loop remediation capabilities.
- Expand automation across observability tools and ITSM processes.
- Establish governance and operational controls for AI-enabled decision-making and autonomous actions.
- Create dashboards and metrics to demonstrate operational efficiency, reliability improvements, and business value realization.
Secondary Responsibilities
Digital Command Center (DCC) Operations
- Serve as Incident Commander or Major Incident Manager during high-severity incidents.
- Lead cross-functional incident bridges and coordination efforts.
- Develop and deliver executive communications and stakeholder updates.
- Support Incident, Problem, Change, and Knowledge Management governance activities.
- Drive continuous improvement initiatives to enhance operational maturity and resiliency.
Required Qualifications & Skills
Experience
- 5+ years of experience in monitoring, event management, AIOps, observability, production engineering, Site Reliability Engineering (SRE), or IT operations.
- 3+ years of hands-on Observability agentic workflows and Service Now Agent Orchestrator deployments for Operations Center environments.
- Proven experience implementing operational automation, orchestration, or AI-driven solutions.
- Experience supporting enterprise-scale applications and business-critical services.
- Experience leading or supporting major incident management activities.
Technical Skills
- Strong Splunk, Now Assist development and troubleshooting skills.
- Experience with data onboarding, event normalization, MCP service modeling, KPI creation, dashboard development, and analytics.
- Knowledge of Generative AI, AI Agents, Copilot technologies, and AI-assisted workflow design.
- Experience with Python, PowerShell, Bash, REST APIs, or similar automation technologies.
- Experience integrating observability, monitoring, cloud, and ITSM platforms.
- Strong understanding of Incident, Problem, Change, and Knowledge Management processes.
- Experience with ServiceNow or equivalent ITSM platforms.
Leadership & Collaboration
- Excellent communication, documentation, and stakeholder management skills.
- Ability to lead under pressure during critical incidents and service disruptions.
- Strong analytical, problem-solving, and continuous improvement mindset.
- Experience working effectively within global, cross-functional delivery teams.
Preferred Qualifications
- ITIL certification.
- Splunk Power User, Admin, Architect, or ITSI certifications.
- ServiceNow certifications.
- AI, Machine Learning, Data Analytics, or Automation certifications.
- Cloud certifications (AWS, Azure, or GCP).
- Experience implementing autonomous operations or self-healing platforms.



