Disaster Recovery
How to recover SwarmCracker from node failures, Raft corruption, and cluster outages.
Overview
Section titled “Overview”SwarmCracker cluster state is stored in multiple locations:
| Component | Location | Backup Strategy |
|---|---|---|
| Raft log (cluster state) | /var/lib/swarmkit/ |
Stop the manager and copy the directory |
| VM state | /var/lib/firecracker/ |
VM snapshots |
| VM snapshots | /var/lib/firecracker/snapshots/ |
Copy to remote storage |
| Config | /etc/swarmcracker/config.yaml |
Auto-generated on join |
| Join tokens | /var/lib/swarmkit/join-tokens.txt |
Backed up with the state directory |
Scenario 1: Worker Node Failure
Section titled “Scenario 1: Worker Node Failure”Detection
Section titled “Detection”# Check the suspect node (run locally there, or over ssh)swarmcracker cluster health
# List cluster nodes (run on a manager)swarmcracker node ls# The failed worker's STATUS is shown as "Down"Recovery Steps
Section titled “Recovery Steps”-
Verify the node is truly dead:
Terminal window ssh dead-worker -- swarmcracker cluster health -
Remove the dead node from the cluster:
Terminal window swarmcracker node rm <NODE_ID>(Add
--forceif the node is still reported asReady.) -
VMs on the dead node are lost. SwarmKit will reschedule services with desired replicas. Verify:
Terminal window swarmcracker service ps <service-name> -
Replace the worker:
Terminal window # Get a new join token on the managerswarmcracker cluster token worker# On the replacement node: install binary + setup deps, then joincurl -fsSL https://raw.githubusercontent.com/restuhaqza/SwarmCracker/main/install.sh | sudo bashsudo swarmcracker setup install --download-kernel --download-rootfssudo swarmcracker setup networksudo swarmcracker setup config --non-interactivesudo swarmcracker cluster join <MANAGER_IP>:4242 --token <TOKEN>
Scenario 2: Manager Node Failure
Section titled “Scenario 2: Manager Node Failure”⚠️ Manager failure requires immediate action — the Raft consensus log is on the manager.
If you have >1 manager (recommended for production)
Section titled “If you have >1 manager (recommended for production)”SwarmKit’s Raft consensus handles this automatically. The remaining managers elect a new leader.
# Check node roles and healthswarmcracker node lsswarmcracker node inspect <node-id> # shows the node's roleIf you have only 1 manager (typical for small clusters)
Section titled “If you have only 1 manager (typical for small clusters)”You need to promote a worker or set up a new manager.
Option A: Promote a worker to manager
Section titled “Option A: Promote a worker to manager”-
On a healthy worker, stop the swarmd service:
Terminal window sudo systemctl stop swarmcracker-worker -
Clear old worker state (prevents CA conflicts):
Terminal window sudo rm -rf /var/lib/swarmkit/certificatessudo rm -rf /var/lib/swarmkit/worker -
Start as manager with force-new-cluster:
Terminal window sudo swarmd-firecracker \--manager \--force-new-cluster \--state-dir /var/lib/swarmkit \--listen-remote-api 0.0.0.0:4242 \--advertise-remote-api <THIS_NODE_IP>:4242 \--bridge-name swarm-br0 \--enable-cni -
Re-join remaining workers:
Terminal window # Get new join tokensudo cat /var/lib/swarmkit/join-tokens.txt# On each workersudo swarmcracker cluster join --token <TOKEN> <NEW_MANAGER_IP>:4242
Option B: Restore from a state-directory backup (if you have one)
Section titled “Option B: Restore from a state-directory backup (if you have one)”SwarmCracker does not expose a Raft snapshot command; back up and restore the manager state as files (see Preventive Measures).
# Stop the manager, then restore the backed-up state directorysudo systemctl stop swarmcracker-managersudo rm -rf /var/lib/swarmkitsudo cp -a /backup/swarmkit /var/lib/swarmkit
sudo swarmd-firecracker \ --manager \ --force-new-cluster \ --state-dir /var/lib/swarmkit \ --listen-remote-api 0.0.0.0:4242Scenario 3: Raft Log Corruption
Section titled “Scenario 3: Raft Log Corruption”Symptoms: Manager won’t start, or swarmctl commands return inconsistent results.
Recovery
Section titled “Recovery”-
Stop the manager:
Terminal window sudo systemctl stop swarmcracker-manager -
Try Raft recovery:
Terminal window # SwarmKit has a built-in recovery mechanismsudo swarmd-firecracker \--manager \--force-new-cluster \--state-dir /var/lib/swarmkit \--listen-remote-api 0.0.0.0:4242 \--advertise-remote-api <IP>:4242This creates a new Raft log from the most recent consistent state.
-
If recovery fails, reset and rebuild:
Terminal window sudo swarmcracker cluster resetsudo swarmcracker cluster init --advertise-addr <IP>:4242# Re-join all workers
⚠️ Resetting loses all running VM state. Use this only as a last resort.
Scenario 4: VM Snapshot Recovery
Section titled “Scenario 4: VM Snapshot Recovery”If a VM’s state is lost but you have snapshots (see the Snapshots guide for the full workflow):
# List available snapshotsswarmcracker vm snapshot list --task <VM_ID>
# Restore from snapshotswarmcracker vm snapshot restore <SNAPSHOT_ID>
# Verifyswarmcracker cluster status <VM_ID>Automated snapshot backup
Section titled “Automated snapshot backup”Set up a cron job to copy snapshots to remote storage:
0 2 * * * root rsync -avz /var/lib/firecracker/snapshots/ backup-server:/backups/swarmcracker/Scenario 5: Full Cluster Outage
Section titled “Scenario 5: Full Cluster Outage”All nodes lose power or network.
Recovery Steps
Section titled “Recovery Steps”-
Power on the manager node first.
-
Wait for it to start:
Terminal window # Manager should auto-start via systemdsudo systemctl status swarmcracker-managerswarmcracker cluster health -
Power on worker nodes one at a time:
Terminal window # On each workersudo systemctl start swarmcracker-workersudo journalctl -u swarmcracker-worker -f# Look for "Node joined" in logs -
Verify cluster:
Terminal window swarmcracker cluster health --format json | jq .swarmcracker node ls -
Verify VMs:
Terminal window swarmcracker vm listswarmcracker task ls --all -
VMs that were running before the outage will need to be recreated:
Terminal window # If using services (recommended), SwarmKit handles this automaticallyswarmcracker service update <service> --force
Preventive Measures
Section titled “Preventive Measures”1. Regular manager state backups
Section titled “1. Regular manager state backups”SwarmCracker has no Raft snapshot command. Back up the manager’s state directory as files while the manager is stopped:
# Run daily via cron on the manager node#!/bin/bashBACKUP_DIR="/backup/swarmcracker/$(date +%Y-%m-%d)"mkdir -p "$BACKUP_DIR"
# Backup manager state (Raft log, certificates, join tokens)systemctl stop swarmcracker-managercp -a /var/lib/swarmkit "$BACKUP_DIR/swarmkit"systemctl start swarmcracker-manager
# Backup configcp /etc/swarmcracker/config.yaml "$BACKUP_DIR/"2. Multiple managers (production)
Section titled “2. Multiple managers (production)”For production clusters with >3 nodes, run 3 managers. SwarmKit’s Raft consensus requires odd numbers.
# Add a second managerswarmcracker cluster token manager
# On the new node: join with the manager tokenswarmcracker cluster join <LEADER_IP>:4242 --manager --token <TOKEN>3. VM state redundancy
Section titled “3. VM state redundancy”Use SwarmKit services instead of raw VMs for stateless workloads. Services automatically reschedule tasks when a worker fails.
# Good: service with replicasswarmcracker service create --name web --image nginx:alpine --replicas 3
# Instead of: single VMswarmcracker vm create --name web-vm nginx:alpine4. Health monitoring
Section titled “4. Health monitoring”# Nagios/Icinga compatible exit codeswarmcracker cluster health --format nagios# CRITICAL: 2 checks failed | kvm=pass firecracker=pass ...# OK: all checks passed | kvm=pass ...
# JSON for scripting/monitoringswarmcracker cluster health --format json | jq '.healthy'Quick Reference
Section titled “Quick Reference”| Situation | Command |
|---|---|
| Check cluster health | swarmcracker cluster health |
| List nodes | swarmcracker node ls |
| Remove dead node | swarmcracker node rm <NODE_ID> |
| Get join token | swarmcracker cluster token worker |
| Force new cluster | swarmd-firecracker --manager --force-new-cluster ... |
| Back up manager state | stop the manager, then cp -a /var/lib/swarmkit <dest> |
| Restore VM snapshot | swarmcracker vm snapshot restore <ID> |
| Service reschedule | swarmcracker service update <name> --force |