Skip to content

Backup and Restore

Starting from ClusterControl 1.7.1, you can back up the ClusterControl server and restore it (together with metadata about your managed databases) onto another server using ClusterControl CLI. It backs up the ClusterControl application as well as all of its configuration data.

There are 4 options available under the s9s backup command:

Flag Description
--save-controller Saves the state of the controller into a tarball.
--restore-controller Restores the entire controller from a previously created tarball (created by using --save-controller).
--save-cluster-info Saves the information the controller has about one cluster.
--restore-cluster-info Restores the information the controller has about a cluster from a previously created archive file.

Note

These options manage backups of the ClusterControl controller and its metadata, not backups of the databases it monitors. To back up and restore a managed database cluster, see Backup & Restore in the User Guide.

Prerequisites

Before saving or restoring, make sure of the following:

  • The source and target ClusterControl hosts run the same s9s/cmon version. Restoring across different minor versions may work for some cluster types but is not guaranteed, and increases the chance of hitting the known issues below.
  • The SSH key ClusterControl uses to reach the database nodes exists at the same path on the target host. For example, if the key file is located under /root/.ssh/id_rsa on the source, make sure the same path and file exist on the target.
  • The target host has working, passwordless SSH key-based access to every node in every cluster contained in the backup. Missing SSH access to even one node causes the restore job to fail during the grant step (see Known issues and limitations).
  • --restore-controller requires a target host with no existing cluster configuration. If /etc/cmon.d/ already contains a cmon_<id>.cnf file for one of the clusters being restored, the job fails immediately. Remove any conflicting cluster (s9s cluster --drop --cluster-id=<id>) and its leftover config file before retrying.

Back up all clusters

To back up the ClusterControl controller with all clusters together with their metadata, run the following command on the ClusterControl node as root user (or with sudo):

s9s backup \
    --save-controller \
    --backup-directory=$HOME/ccbackup \
    --output-file=controller.tar.gz \
    --log

The --output-file must be a filename or physical path (if you want to omit the --backup-directory flag), and the file must not exist beforehand. ClusterControl does not replace the output file if it already exists. By specifying the --log flag, it waits until the job is executed and shows the job logs in the terminal. The same logs can be accessed via ClusterControl GUI → Activity center → Jobs → Save Controller.

The save controller job performs the following procedures:

  1. Retrieve the controller configuration and export it to JSON.
  2. Export the CMON database as a MySQL dump file.
  3. For every database cluster:
    1. Retrieve the cluster configuration and export it to JSON.

Note

In the output, you may notice the job found is N + 1 cluster, for example Found 3 cluster(s) to save even though you only have two database clusters. This includes cluster ID 0, which carries special meaning in ClusterControl as the global initialized cluster. It does not belong to the CmonCluster component, which is the database cluster under ClusterControl management.

Back up an individual cluster

To back up an individual cluster with its metadata, use the --cluster-id option to specify the cluster ID, on the ClusterControl node as root user (or with sudo):

s9s backup \
    --save-cluster-info \
    --cluster-id=2 \
    --backup-directory=$HOME/ccbackup \
    --output-file=cc-replication-2.tar.gz \
    --log

Restore an individual cluster to another server

Restoring an individual cluster's metadata is intended for moving one cluster's management from an old ClusterControl controller to a new one, not for restoring a cluster while it is still registered elsewhere.

  1. Copy the backup file to the target host:

    scp cc-replication-2.tar.gz root@<new_server>:~/
    
  2. On the target host, run the restoration. Do not pass --cluster-id; ClusterControl allocates a new cluster ID automatically from the archive contents:

    s9s backup \
        --restore-cluster-info \
        --input-file=$HOME/cc-replication-2.tar.gz \
        --debug \
        --log
    
  3. Verify the cluster restored correctly:

    s9s cluster --list --long
    

Restore the entire controller to another server

Use this to migrate the controller and every cluster it manages onto a new host in a single operation, for example when replacing aging hardware or moving to a new operating system.

  1. Install the same ClusterControl version on the new host. See Installation.

  2. Make sure the same SSH key ClusterControl uses to reach the database nodes exists at the same path on the new host, and that it has working access to every node in every cluster contained in the backup (see Prerequisites).

  3. Perform the restoration:

    s9s backup \
        --restore-controller \
        --input-file=$HOME/controller.tar.gz \
        --debug \
        --log
    
  4. Wait a moment, then verify the clusters restored correctly:

    s9s cluster --list --long
    

    Tip

    If the cluster list comes back empty right after the restore, wait about 30 seconds and run the command again. cmon is still restarting.

Known issues and limitations

The following behavior has been confirmed when restoring onto a target running the same ClusterControl version:

  • The monitoring/Prometheus host is not repointed to the new controller. Both --restore-cluster-info and --restore-controller correctly update the controller host entry to the new server's address, but the Prometheus monitoring host entry keeps pointing at the old server's IP. Update this manually, or follow Migrating ClusterControl to Another Server, which installs Prometheus fresh on the new host as part of the procedure.
  • s9s cluster --drop is the correct option to unregister a cluster, not --remove-cluster (which does not exist and returns unrecognized option). Dropping a cluster removes it from ClusterControl only; the database itself keeps running.
  • --restore-cluster-info only completes successfully when the target controller has no other cluster already registered. Every cluster shares the same controller-side Prometheus/monitoring host entry; if the target already manages another cluster, restoring a second one fails while trying to recreate that shared monitoring host object (Host is already part of some other cluster), and the job rolls back everything it created. Always restore onto a target host with no existing clusters.
  • Dropping a cluster may log Failed to purge cluster data on dcps db: Unknown database 'dcps'. This is informational only and does not affect the drop operation; the dcps database is an optional component not present on all installations.