Operations Guide — SwarmCracker
How to operate, monitor, troubleshoot, and maintain a SwarmCracker cluster in production.
Health Checks
Section titled “Health Checks”Node-Level Health
Section titled “Node-Level Health”Run the built-in doctor to verify node health:
swarmcracker doctorPassing checks:
| Check | What it verifies | Healthy Output |
|---|---|---|
| CPU virtualization (KVM) | /dev/kvm exists and the KVM module is loaded |
✓ CPU virtualization (KVM) — KVM available |
| Firecracker binary | firecracker is in PATH |
✓ Firecracker binary — found at /usr/local/bin/firecracker |
| Firecracker version | firecracker --version runs |
✓ Firecracker version — v1.15.1 |
| Kernel image | Kernel present at the configured path | ✓ Kernel image — /usr/share/firecracker/vmlinux (25.0 MB) |
| TUN/TAP device | /dev/net/tun is available |
✓ TUN/TAP device — /dev/net/tun available |
| Available memory | Free memory from /proc/meminfo |
✓ Available memory — 16.0 GB total, 7.5 GB available |
| CPU cores | Processor count | ✓ CPU cores — 8 CPU core(s) |
| Bridge interface | swarm-br0 exists |
✓ Bridge interface (swarm-br0) — swarm-br0 exists |
| Port 4242 | SwarmKit API port availability | ✓ Port 4242 — port 4242 available |
| Manager/Worker service | systemd unit is active | ✓ Worker service — worker service is active |
| Join tokens | /var/lib/swarmkit/join-tokens.txt present |
✓ Join tokens — found at /var/lib/swarmkit/join-tokens.txt |
| Firecracker processes | Running VM process count | ✓ Firecracker processes — 3 Firecracker process(es) running |
API Health Endpoint
Section titled “API Health Endpoint”The daemon exposes a health check endpoint on its health address
(--health-addr, default 127.0.0.1:8080):
# Localcurl -s http://127.0.0.1:8080/healthz
# Response{ "healthy": true, "checks": { "kvm": { "status": "ok", "message": "..." }, "bridge": { "status": "ok", "message": "..." }, "firecracker": { "status": "ok", "message": "..." } }}The same HTTP server exposes Prometheus metrics at
http://127.0.0.1:8080/metrics.
Cluster Health
Section titled “Cluster Health”# List all nodesswarmcracker node ls
# Expected output:# ID STATUS HOSTNAME AVAILABILITY# abc123def456 Ready manager-1 Active# def456abc789 Ready worker-1 Active# ghi789def012 Ready worker-2 ActiveVM Health
Section titled “VM Health”# Check specific VMswarmcracker cluster status <vm-id>
# Watch modewatch -n2 'swarmcracker cluster status <vm-id>'Monitoring
Section titled “Monitoring”Metrics
Section titled “Metrics”The node’s health server exposes Prometheus metrics at
http://127.0.0.1:8080/metrics:
# Scrape this node's metricscurl -s http://127.0.0.1:8080/metrics
# Per-task resource usageswarmctl metrics <task-id>
swarmcracker metricsstill runs but is deprecated. Use the Prometheus endpoint above (orswarmctl metrics) instead.
Key metrics (all prefixed swarmcracker_):
swarmcracker_vms_running— VMs currently running on this nodeswarmcracker_vms_total— VM lifecycle transitions by statusswarmcracker_vm_cpu_seconds— CPU time consumed per VM (task_id,servicelabels)swarmcracker_vm_memory_bytes— Memory used per VMswarmcracker_vm_net_rx_bytes/swarmcracker_vm_net_tx_bytes— Network I/O per VMswarmcracker_vm_boot_duration_seconds— VM boot timeswarmcracker_vxlan_peers/swarmcracker_vxlan_expected_peers— Overlay peersswarmcracker_manager_health/swarmcracker_raft_health— Manager/Raft healthswarmcracker_disk_usage_bytes— Disk usage of SwarmCracker directories
Logging
Section titled “Logging”Log locations:
| Log | Path | Rotation |
|---|---|---|
| swarmd-firecracker (daemon) | /var/log/swarmcracker/daemon.log |
systemd journal |
| VM console logs | /var/log/firecracker/<vm-id>.log |
Per-VM |
| dnsmasq (DHCP) | /tmp/dnsmasq.log |
Manual |
| Firecracker stderr | captured by daemon | — |
View logs:
# Daemon logssudo journalctl -u swarmcracker-worker -f
# Specific VM consoleswarmcracker vm logs --follow <vm-id>
# Logs for each task in a service (find the task IDs first)swarmcracker service ps <service-name>swarmcracker vm logs --follow <task-id>
# dnsmasq DHCP logstail -f /tmp/dnsmasq.logAttach to a VM console:
swarmcracker vm attach connects your terminal to the microVM’s serial console
(the same channel the guest kernel and ttyS0 use). For an image whose command
is an interactive shell (for example CMD ["/bin/bash"]), this gives you a shell
inside the VM.
# Find the task ID (service ps shows a 12-character prefix)swarmcracker task ls
# Attach using the full ID or any unique prefixswarmcracker vm attach 5f3a1b2c9d
# Non-default socket directory (must match the daemon's --socket-dir)swarmcracker vm attach --socket-dir /var/run/firecracker 5f3a1b2c9dPress Ctrl-P Ctrl-Q to detach; the VM keeps running. The console socket is
owned by the daemon (root, mode 0600), so run the command as root or with
sudo.
Note: console attach is available for VMs managed by the daemon (service tasks). VMs started in the foreground by
swarmcracker vm createare not attachable, because the CLI process that owns their console exits with the VM.
Log levels: Set via --log-level flag or logging.level in config.yaml.
Available: debug, info, warn, error.
Set to debug for troubleshooting (verbose, includes token operations at debug level only):
swarmd-firecracker --debugCommon Troubleshooting
Section titled “Common Troubleshooting”VM Won’t Start
Section titled “VM Won’t Start”Symptoms: vm create or service create returns error.
Checklist:
-
KVM available?
Terminal window ls -la /dev/kvm# If missing: modprobe kvm && modprobe kvm-intel (or kvm-amd) -
Firecracker binary?
Terminal window which firecrackerfirecracker --version # Should be v1.15.1+ -
Kernel image?
Terminal window ls -la /usr/share/firecracker/vmlinux# Expected: ~25MB ELF kernel -
Rootfs exists?
Terminal window ls -la /var/lib/firecracker/rootfs/# Should show .ext4 files for each pulled image -
Bridge exists?
Terminal window ip link show swarm-br0# If missing: swarmcracker cluster init (recreates infrastructure) -
Socket directory writable?
Terminal window ls -la /var/run/firecracker/# Permissions should be 0755, owned by the daemon user -
Sufficient resources?
Terminal window swarmcracker doctor # Check memory/CPU available
VM Crashes / Exits Immediately
Section titled “VM Crashes / Exits Immediately”Symptoms: VM starts then stops within seconds.
Checklist:
-
Check VM console log:
Terminal window swarmcracker vm logs <vm-id> -
Common causes:
- Kernel panic: Wrong kernel or missing modules. Check boot args.
- Rootfs not found: Verify path in config.
- Init system failure: Check if tini/dumb-init is in rootfs.
- OOM: VM has insufficient memory. Increase
--memory/-m. - Missing command: Container image doesn’t have the specified command.
-
Check Firecracker output:
Terminal window journalctl -u swarmcracker-worker | grep <vm-id>
Network Issues
Section titled “Network Issues”VMs Can’t Reach Internet
Section titled “VMs Can’t Reach Internet”-
NAT enabled?
Terminal window iptables -t nat -L POSTROUTING | grep MASQUERADE# Should show rule for 192.168.127.0/24 -
IP forwarding?
Terminal window sysctl net.ipv4.ip_forward# Should be 1 -
DHCP working?
Terminal window cat /tmp/dnsmasq.log | tail -20# Should show DHCPOFFER/DHCPACK
Cross-Node VM Communication Fails
Section titled “Cross-Node VM Communication Fails”-
VXLAN enabled on both nodes?
Terminal window ip link show | grep vxlan# Should show swarm-br0-vxlan -
VXLAN peers correct?
Terminal window swarmcracker network vxlan list# Should list all worker IPs -
UDP 4789 open between nodes?
Terminal window nc -zvu <other-node-ip> 4789 -
Bridge FDB entries correct?
Terminal window bridge fdb show dev swarm-br0-vxlan -
Consul registration?
Terminal window consul catalog services# Should show swarmcracker-worker
VM Has No IP Address
Section titled “VM Has No IP Address”-
Check TAP device:
Terminal window ip link show | grep tap- -
Check IP allocator:
Terminal window swarmcracker doctor # Checks IPAM state -
Static IP with no DHCP fallback? If using static IP mode and the VM expects DHCP, add
ip=dhcpto kernel args.
Cluster Issues
Section titled “Cluster Issues”Worker Can’t Join
Section titled “Worker Can’t Join”-
Connectivity to manager:
Terminal window nc -zv <manager-ip> 4242 -
Valid join token?
Terminal window # On manager (prints the worker and manager join tokens)swarmcracker cluster token -
Firewall? Ports needed:
4242(SwarmKit gRPC API)4789UDP (VXLAN overlay)8500(Consul, if enabled)
-
Time sync?
Terminal window timedatectl status# Clocks must be within a few seconds
Manager Lost Quorum
Section titled “Manager Lost Quorum”If you lose 2 of 3 managers:
-
On the remaining manager:
Terminal window swarmd-firecracker --manager --force-new-cluster -
Rejoin workers:
Terminal window # On each workerswarmcracker cluster join --token <new-token> <manager-ip>:4242
Performance Issues
Section titled “Performance Issues”VMs Are Slow
Section titled “VMs Are Slow”-
Check CPU steal:
Terminal window curl -s http://127.0.0.1:8080/metrics | grep swarmcracker_vm_cpu -
Check memory pressure:
Terminal window free -hswarmcracker doctor # Shows available memory -
Disk I/O bottleneck?
- Use
blockdriver instead ofdirfor database workloads - Check rootfs is on fast storage (SSD/NVMe)
- Use
-
Network throughput?
- VXLAN adds ~50 bytes overhead per packet
- For overlay networks, MTU is set to 1450 automatically
Backup and Restore
Section titled “Backup and Restore”VM Snapshots
Section titled “VM Snapshots”# Create snapshot of running VMswarmcracker vm snapshot create <task-id>
# List snapshotsswarmcracker vm snapshot list --task <task-id>
# Restore from snapshotswarmcracker vm snapshot restore <snapshot-id>Configuration Backup
Section titled “Configuration Backup”# Backup config directorytar czf swarmcracker-config-$(date +%Y%m%d).tar.gz /etc/swarmcracker/
# Backup state (includes certs, tokens, task state)tar czf swarmcracker-state-$(date +%Y%m%d).tar.gz /var/lib/swarmcracker/Volume Backup
Section titled “Volume Backup”# Backup a volumetar czf volume-<name>-$(date +%Y%m%d).tar.gz /var/lib/swarmcracker/volumes/<name>/Full Node Backup
Section titled “Full Node Backup”#!/bin/bash# Full backup scriptBACKUP_DIR="/backup/swarmcracker/$(date +%Y%m%d_%H%M%S)"mkdir -p "$BACKUP_DIR"
# Configcp -r /etc/swarmcracker "$BACKUP_DIR/config"
# State (stop the daemon first if possible)systemctl stop swarmcracker-worker # or swarmcracker-manager on a manager nodecp -r /var/lib/swarmcracker "$BACKUP_DIR/state"systemctl start swarmcracker-worker
# Rootfs imagescp -r /var/lib/firecracker/rootfs "$BACKUP_DIR/rootfs"
# Volumescp -r /var/lib/swarmcracker/volumes "$BACKUP_DIR/volumes"
# Compresstar czf "$BACKUP_DIR.tar.gz" -C "$(dirname "$BACKUP_DIR")" "$(basename "$BACKUP_DIR")"rm -rf "$BACKUP_DIR"Cluster Upgrade
Section titled “Cluster Upgrade”Rolling Upgrade (Zero Downtime)
Section titled “Rolling Upgrade (Zero Downtime)”-
Drain a worker:
Terminal window swarmcracker node drain worker-1 -
Wait for all VMs to move:
Terminal window swarmcracker task ls --node worker-1# Should show no running tasks -
Upgrade the worker:
Terminal window # On worker-1systemctl stop swarmcracker-worker# Deploy new binarycp new-swarmd-firecracker /usr/local/bin/systemctl start swarmcracker-worker -
Activate the worker:
Terminal window swarmcracker node activate worker-1 -
Verify health:
Terminal window swarmcracker node ls | grep worker-1# Should show Ready, Active -
Repeat for remaining workers, then managers.
Manager Upgrade
Section titled “Manager Upgrade”Managers must be upgraded one at a time:
-
Verify quorum:
Terminal window swarmcracker node ls | grep manager# Need 2+ managers healthy for quorum -
Drain and upgrade:
Terminal window # Same as worker upgrade, but re-initialize if neededswarmcracker cluster leave# Upgrade binary, then re-joinswarmcracker cluster join --token <token> <leader-ip>:4242
Resource Management
Section titled “Resource Management”Capacity Planning
Section titled “Capacity Planning”| Per-VM Overhead | Value |
|---|---|
| Firecracker process | ~50 MB RSS |
| TAP device | Negligible |
| Bridge + VXLAN | ~5 MB kernel memory |
| dnsmasq (shared) | ~10 MB RSS |
| State tracking | ~5 MB per VM |
Formula:
Available VMs = (Total_RAM - System_Reserved - 200MB) / (VM_Size + 50MB)Example for a 16GB worker with 512MB VMs:
(16GB - 2GB - 0.2GB) / (0.512GB + 0.05GB) = 24.5 → 24 VMs maxScaling Up
Section titled “Scaling Up”# Add workers# 1. Provision new node with Firecracker + kernel# 2. Get join token from managerswarmcracker cluster token worker
# 3. On new workerswarmcracker cluster join <manager-ip>:4242 --token SWMTKN-1-xxx
# 4. Verifyswarmcracker node ls# Should show new node as Ready, ActiveScaling Down
Section titled “Scaling Down”# Drain workerswarmcracker node drain worker-3
# Wait for VMs to rescheduleswarmcracker task ls --node worker-3
# Leave clusterswarmcracker cluster leave
# Clean upswarmcracker cluster resetSecurity Operations
Section titled “Security Operations”Certificate Rotation
Section titled “Certificate Rotation”SwarmCracker uses SwarmKit’s built-in mutual TLS, and node certificates are
rotated automatically by SwarmKit. There is no CLI command to force a
rotation: swarmcracker cluster token only displays join tokens.
Secret Management
Section titled “Secret Management”SwarmCracker does not currently expose secret management through its CLI:
neither swarmcracker nor swarmctl has a secret command, and
swarmcracker service create has no --secret flag. Distribute sensitive
data through your image build or a mounted volume until secret support lands.
Firewall Rules
Section titled “Firewall Rules”Minimum required ports:
| Port | Protocol | Source | Destination | Purpose |
|---|---|---|---|---|
| 4242 | TCP | All nodes | Managers | SwarmKit gRPC |
| 4789 | UDP | All workers | All workers | VXLAN overlay |
| 8500 | TCP | All nodes | Consul nodes | Service discovery |
| 4242/tcp | TCP | Admin | Managers | CLI access |
# Example iptables rulesiptables -A INPUT -p tcp --dport 4242 -s 192.168.1.0/24 -j ACCEPTiptables -A INPUT -p udp --dport 4789 -s 192.168.1.0/24 -j ACCEPTRoutine Maintenance
Section titled “Routine Maintenance”- Check node health:
swarmcracker doctor - Check running VMs:
swarmcracker vm list - Check disk space:
df -h /var/lib/firecracker/rootfs
Weekly
Section titled “Weekly”- Run image cleanup: check
/var/log/swarmcracker/daemon.logfor “Periodic cleanup completed” - Check for orphaned VMs: review daemon logs for “Found orphaned VM”
- Review metrics trends
- Check available disk for snapshots:
du -sh /var/lib/firecracker/snapshots/
Monthly
Section titled “Monthly”- Test backup restoration
- Review and update firewall rules
- Check for SwarmCracker updates
- Rotate Consul tokens (if using ACLs)
- Verify VXLAN cross-node connectivity
Emergency Procedures
Section titled “Emergency Procedures”Node Failure
Section titled “Node Failure”Worker failure:
- Drain the failed node:
swarmcracker node drain <worker> - SwarmKit reschedules VMs to other workers
- Replace hardware, reprovision, rejoin
Manager failure (1 of 3):
- Cluster continues operating. Replace the failed manager.
Manager failure (2 of 3 — loss of quorum):
- On the surviving manager:
swarmd-firecracker --manager --force-new-cluster - Replace failed managers, join as followers
- Rejoin workers
Full Cluster Recovery
Section titled “Full Cluster Recovery”- Stop all
swarmd-firecrackerprocesses - Restore from backup on the designated manager node
- Start with
--force-new-cluster - Restore workers from backups, rejoin one at a time
- Verify:
swarmcracker node lsshould show all nodes Ready
Disk Full
Section titled “Disk Full”-
Immediate: Delete old snapshots
Terminal window swarmcracker vm snapshot cleanup --max-age 168h -
Short-term: Manually trigger image cleanup
Terminal window # Set max_image_age_days to 1 in config, restart daemon -
Long-term: Configure auto-cleanup in config:
images:max_cache_size_mb: 10240snapshot:max_snapshots: 10max_age: 168h # 7 days
Configuration Reference
Section titled “Configuration Reference”See Configuration Guide for all config keys and defaults.
Quick Reference Card
Section titled “Quick Reference Card”# Healthswarmcracker doctor # Node health checkcurl 127.0.0.1:8080/healthz # API health check
# Clusterswarmcracker node ls # List nodesswarmcracker node inspect <node> # Node detailsswarmcracker cluster token worker # Get a worker join token
# VMsswarmcracker vm list # List VMs (CLI-created + service tasks)swarmcracker vm status <vm-id> # VM details (service task IDs work too)swarmcracker vm logs -f <vm-id> # Follow VM logs (CLI-created VMs)swarmcracker vm attach <vm-id> # Attach to the VM serial consoleswarmcracker vm stop <vm-id> # Stop a CLI-created VMswarmcracker cluster status <vm-id> # VM status
# Servicesswarmcracker service ls # List servicesswarmcracker service ps <service> # Service tasks
# Snapshotsswarmcracker vm snapshot create <vm-id> # Snapshot VMswarmcracker vm snapshot list # List snapshotsswarmcracker vm snapshot restore <snap> # Restore from snapshot
# Networkswarmcracker network vxlan ls # VXLAN peers
# Recoveryswarmcracker cluster reset --hard # Reset nodeswarmcracker cluster leave # Leave clusterSee Also: Architecture Overview | Configuration Guide | Networking Guide | Security Guide