Backstory
We have been thinking about a “cloud without frills” for quite a while. So we bought the hardware and the required network infrastructure and got started.
It was based on Ceph Nautilus and vanilla OpenStack Stein and ran quite well. The problems:
- The network: 1 Gbit was simply not enough; it should be 10 Gbit, with bonding 20 Gbit
- HP DL360p G6 servers: simply too old
- HDDs in Ceph: simply too slow
The new initiative
You only learn sustainably from your own mistakes and experience. So the new concept looks like this:
Network:
- 10 Gbit fibre switch as backbone
- All Ceph and OpenStack nodes get 10 Gbit network cards
- With bonding we can increase to 20 Gbit (if 10 Gbit is not enough)
Ceph:
- Relaunch on Ceph Octopus (15.2)
- Switching to SSDs
OpenStack:
- Throwing out the old G6 servers
- Replacing them with newer, more power-efficient models
- Upgrading the RAM (128 GB per compute node was … nice :-))
- Fresh installation with OpenStack Victoria
The requirements
- Fully automated setup: From the operating system installation (Kickstart) through base provisioning (Ansible) and the Ceph installation (ceph-ansible) to the subsequent OpenStack installation (kolla-ansible), everything should be deployed fully automatically by a Jenkins pipeline.
- Performance must be right: After the setup, a basic performance test is carried out. A sample environment is created with terraform for this purpose. The test is then run with gatling.
- Monitoring: All of this should of course be monitored as well. Tools for this are the Elastic Stack, Prometheus, Grafana and OpenStack’s own Monasca.
- Kubernetes: We obviously also want k8s on board. We want to implement that with OpenStack’s Magnum.
… More to come …