Monitoring & High Availability
- checkmk
- Nagios
- Grafana
- Enterprise Manager
- RMAN
- Patroni
- Galera
- HAProxy
A system is only truly in operation when someone sees the outage coming before it happens. We build monitoring that measures the right things and spots problems. And we build architectures that absorb a failure: failover clusters since the Solaris and Veritas days, today Patroni on Kubernetes, MariaDB Galera, HAProxy/nginx/Apache/Squid with Keepalived and GlusterFS. Plus backup and recovery that we actually test on a regular basis.
For clients with high availability requirements we build systems that survive the loss of an entire data center (disaster recovery). We test the real thing with our clients on a regular basis.
Typical tasks
- Build system and service monitoring with checkmk or Nagios, alerting with sense and proportion
- Metrics and logs: Grafana, InfluxDB and Telegraf, ELK stack with Logstash and Filebeat
- Database monitoring with Oracle Enterprise Manager and Grid Control, introduced from 10g to 12c
- Backup and recovery with RMAN and Veritas NetBackup, recovery tests instead of hope
- Failover clusters and high availability: Veritas Cluster Server, Patroni, MariaDB Galera, HAProxy and Keepalived, GlusterFS, Samba with CTDB
- Patch management and security patching, secrets with HashiCorp Vault
From practice
For RCI Banque we built a highly available Nextcloud system with MariaDB Galera and GlusterFS and an HA Samba system with CTDB; the OpenShift clusters are monitored with checkmk. An older example: for around 120 test databases at Vodafone we designed and implemented a failover cluster concept with Veritas Cluster Server and EMC SAN between 2002 and 2007 and introduced Oracle Enterprise Manager Grid Control as the central database monitoring.