The posting
Your Role
You will help us build, operate and industrialize the storage foundation of our on-premise cloud platform. As part of a small, experienced team, you will own the Ceph clusters behind our block, object and file storage services – essentially the storage every customer workload eventually lands on. The platform is still actively being built, which means you will have real influence on the storage architecture, hardware lifecycle, automation strategy and technical direction. The role goes beyond Ceph itself: you will regularly work across bare metal, networking, OpenStack, Kubernetes and GitOps. As a Senior Engineer, you take ownership of your area and help make storage predictable, durable and highly automated as we scale into the petabytes
Unser Tech Stack 🚀
· Ceph · OpenStack · Kubernetes · KVM · Linux · Bare Metal · Ansible · Terraform · Go · Cinder ·Neutron · Nova · RBD · Cilium · RGW/ S3 · Prometheus/ Grafana · FluxCD / ArgoCD · Git· Python · Claude Code · Cursor · Agentic Coding Tooling
Your Responsibilities
- Design, operate and evolve our Ceph clusters for block, object / S3 and file storage - from CRUSH topology and failure domains to pool design, placement groups, erasure coding, capacity headroom and performance.
- Own the storage hardware lifecycle end to end – from qualification and burn-in through firmware, controller and disk health to OSD add/drain/replace, node retirement and hardware refresh. A key part of the job is automating these processes without creating tenant-visible impact.
- Automate the provisioning of storage nodes on OpenStack-managed bare metal, working with Ironic, RAID and disk configuration, SEDs, provider networks and VLANs. You will also own the artefact path around images, package mirrors, certificates and other dependencies needed to bring up clusters in environments with tightly controlled external connectivity.
- Run upgrades, rebalancing and reconfigurations as routine production operations, evolve our Infrastructure as Code and GitOps setup with Ansible, Terraform and FluxCD / ArgoCD, and make sure storage integrates cleanly with both OpenStack and Kubernetes.
- Own capacity planning and growth forecasting, use telemetry to understand real consumption, and build out monitoring, testing and failure-injection around performance, durability and security.
- AI-assisted engineering is simply part of how we work today. You use LLMs and agentic tools where they genuinely help – across development, testing, reviews, incident triage, knowledge retrieval or automation - and help us push further into self-healing and auto-remediation.
- Act as a technical reference and sparring partner across Ceph, storage, automation and AI tooling, document technical decisions and share your knowledge with the team.
What we offer
- Exceptional team spirit across all departments and national borders; we live #OneTeam
- Exciting work in a highly innovative and international environment with cutting-edge technologies
- 32 vacation days, increasing with length of service
- Flexible working hours, home office options, and a secure permanent position with market- and performance-based compensation
- Employer-funded pension plan and an attractive insurance package
- OVHcloud covers 50% of public transportation costs
- Up to €400 annual financial contribution from OVHcloud towards sports activities (gym membership, sports classes, etc.)
- Through Corporate Benefits, you receive attractive discounts at numerous shops and companies
- We contribute to the leasing of your cargo bike
- Regular company events and free cold and hot beverages
- Several years of hands-on experience as an SRE, Storage Engineer or Platform Engineer running production storage infrastructure, with solid experience deploying, operating and debugging Ceph in production.
- You know the Ceph fundamentals from real-world operations: CRUSH, pools and placement groups, replication vs. erasure coding, BlueStore, scrubbing, recovery and backfill under load.
- You have managed storage infrastructure end to end, from the physical layer with drives, firmware, BIOS, controllers and hardware diagnostics through rebalancing, host drains, capacity expansion and hardware generation migrations.
- You are comfortable using Kubernetes as a control plane, not only as a workload runtime, and understand operators/controllers, custom resources, reconciliation and GitOps. You can also move confidently around OpenStack and debug issues involving Ironic, Neutron, Nova or Glance when needed.
- You have dealt with situations where durability really matters – such as degraded or near-full clusters, recovery under load or split-brain conditions – and you are used to making changes in a staged, evidence-based and reversible way.
- You have strong Linux and bare-metal skills, understand the block layer, filesystems and I/O behaviour, and are comfortable working with Ansible, Terraform and GitOps.
- AI-assisted engineering is already part of your day-to-day work. You have practical experience with LLMs and agentic tools and know where they can meaningfully support development, testing, reviews or operations.
- Ideally, you also bring experience with RGW / S3, RBD mirroring, CephFS or ceph-csi, deeper OpenStack storage integrations and Go and/or Python. Experience around VLAN/BGP, storage performance tuning, NVMe/BlueStore, observability, data protection, encryption, auto-remediation, security baselines or multi-site storage is also relevant for the role.
- You work autonomously, bring a strong sense of ownership, and are comfortable debugging problems where the actual root cause may sit somewhere between storage, networking, Kubernetes and OpenStack. You would rather verify what is happening from evidence than assume how the system should behave.
- We work in an international environment in English, so you should feel comfortable discussing technical topics, documenting decisions and working with the team in English.



