The posting
We are seeking a Platform Engineer to operate and maintain Ceph-based storage services supporting Kubernetes and OpenShift environments.
Responsibilities
- Operate and maintain Ceph and OpenShift Data Foundation (ODF) clusters.
- Monitor storage health, capacity, performance, latency, throughput, and device status.
- Perform OSD replacement, node replacement, recovery, rebalancing, backfill, and routine maintenance.
- Troubleshoot placement group issues, slow operations, quorum problems, hardware failures, and storage-related network issues.
- Plan and execute upgrades, patching, expansion, and configuration changes.
- Support Ceph CSI, RBD, and CephFS storage services.
- Troubleshoot persistent volume provisioning, attachment, mounting, expansion, and performance issues.
- Maintain monitoring dashboards, alerts, runbooks, capacity plans, and recovery procedures.
- Conduct recovery testing for disk, node, service, and network failures.
- Coordinate hardware, firmware, and infrastructure maintenance activities.
- Participate in incident response and root cause analysis.
- Develop and maintain operational automation scripts where applicable.
Requirements
- Degree or Diploma in Information Technology, Computer Science, Engineering, or related discipline.
- Hands-on experience administering and supporting Ceph environments.
- Experience with storage monitoring, capacity planning, upgrades, expansion, and hardware replacement.
- Good understanding of Linux systems administration.
- Experience benchmarking storage workloads and analysing performance bottlenecks.
- Strong troubleshooting skills across storage platforms, operating systems, networks, and hardware.
- Knowledge of HDD, SSD, NVMe, HBA, firmware, and storage networking concepts.
- Experience with Kubernetes and/or OpenShift environments.
- Knowledge of backup, snapshots, replication, and disaster recovery processes.
- Experience with automation using Ansible, Python, Shell scripting, or similar tools.
Preferred Skills
- Experience with OpenShift Data Foundation (ODF).
- Familiarity with Prometheus, Grafana, or similar monitoring tools.
- Experience supporting production infrastructure environments.



