The posting
Job Description:
- Develop, test, and maintain automation scripts, workflows, dashboards, and operational utilities to reduce manual effort.
- Configure and maintain enterprise monitoring and observability across infrastructure, applications, logs, and capacity.
- Build dashboards, alerts, thresholds, service views, reports, and support monitoring onboarding and alert tuning.
- Monitor system health and investigate alerts, incidents, failed jobs, integrations, performance, availability, and capacity issues.
- Support incident, problem, and change activities, including troubleshooting, impact analysis, validation, and post-change monitoring.
- Review operational configurations, monitoring, backup, and logging controls to ensure effective implementation and coverage.
Requirements:
- Bachelor’s degree in Computer Science, IT, Engineering, or equivalent, with 3–5 years of relevant experience.
- Strong analytical, troubleshooting, problem-solving, communication, documentation, and stakeholder coordination skills.
- Hands-on scripting experience with Python, PowerShell, or Shell and knowledge of Linux, Windows, networking, storage, and virtualisation.
- Experience with monitoring/observability tools such as BMC Helix, Remedy, Prometheus, Grafana, ELK, Dynatrace, or Splunk.
- Practical experience in dashboards, alerting, event/log analysis, integrations, and operational troubleshooting.
- Knowledge of cloud, DevOps, CI/CD, APIs, version control, or AIOps is advantageous.



