The posting
Role purpose
Guided by overall objectives of operational reliability, speed, and cost efficiency, you will be responsible for the architecture, design, implementation and day-to-day operations of the SIA Group’s core Data Centre network, keeping mission-critical services highly available, performant, and secure. This engineering role requires an analytical mindset, good awareness and appreciation of operational practices, and strong expertise in automation and AI-enabled operations in order to reduce toil, improve incident response, and strengthen resilience at scale.
What you will do
1) Service Operations & Reliability
•Provide operational support for core DC networking (routing/switching, segmentation, connectivity)
•Drive Incident, Problem, Change, and Configuration MANAGEMENT to meet service targets and standards
•Serve as technical lead during major incidents, performing deep-dive troubleshooting, RCA, and driving corrective/preventive actions
•Plan and execute complex changes with risk assessment, readiness checks, back-out planning, and post-change validation
•Maintain high-quality configuration and operational documentation; continuously improve runbooks/SOPs
2) Operational Automation & AI-enabled Ops
•Design and build automation to improve BAU consistency and efficiency (e.g.,config deployment/validation, compliance, health checks, reporting)
• Use Ansible, Python, and APIs to standardize operational workflows.
• Apply AIOps/AI-assisted operations (event correlation, anomaly detection, noise reduction, and predictive alerting) to reduce MTTR
• Scale automation patterns
3) Security Operations & Zero Trust
•Operate and improve the security posture of DC network infrastructure
•Support micro-segmentation and Zero Trust-aligned controls within the datacenter network
•Proactively identify resiliency and security gaps and drive remediation through engineering changes and automation
4) Vendor & Service Delivery Management
•Collaborate with vendors and stakeholders to ensure service availability and timely issue resolution
•Track incidents/problems/tickets to ensure SLA adherence, quality updates, and timely closure
•Participate in service delivery reviews and drive actions that improve outcomes
5) Continuous Improvement & Technology Refresh (Ops-led)
•Identify operational pain points and implement improvements across process, tooling, monitoring, and automation
•Participate in POCs/lab validations to improve operability, reliability, and automation
________________________________________
What we are looking for
•~6+ years of hands-on network operations/engineering experience in a multi-vendor environment.
•Strong DC networking fundamentals (routing/switching, troubleshooting, complex change execution).
•Experience with operational practices: incident/problem/change management,RCA, and service reliability.
•Experience implementing process re-engineering and automation to improve processes.Strong problem-solvingskills with the ability to isolate and resolve issues in complex environments.
•Excellent team player; ability to manage conflicts to achieve common goals.
•Strong communication skills; comfortable engaging technical and non-technicalstakeholders.
Good to have
•Experience in two or more of: core DC networking, security, networkautomation.
•Working knowledge of Ansible/Python/APIs (or motivation to build thiscapability quickly).
•Familiarity with segmentation/micro-segmentation and operational securityhygiene.
•Certifications such as Cisco/F5/Palo Alto/Check Point (or equivalent).
•Awareness of AIOps/AI trends and ability to apply them pragmatically.



