End-to-End Testing — The Full Journey
From bare metal to running production workloads, step by step.
What This Covers
Section titled “What This Covers”We’ll walk through the entire lifecycle of a SwarmCracker deployment:
Bare Metal → Install → Set Up Cluster → Deploy Services → Scale → Update → Snapshot → Rollback → Monitor → Clean UpTarget Setup
Section titled “Target Setup”┌──────────────────────────────────────────────────────────────┐│ A 3-Node Cluster ││ ││ ┌─────────────┐ ┌──────────────┐ ┌──────────────┐ ││ │ Manager-1 │ │ Worker-1 │ │ Worker-2 │ ││ │ 192.168.1.10│ │ 192.168.1.11 │ │ 192.168.1.12 │ ││ │ │ │ │ │ │ ││ │ SwarmKit API │ │ 3x nginx VMs │ │ 2x nginx VMs │ ││ │ swarmctl │ │ 1x redis VM │ │ 2x redis VMs │ ││ └──────────────┘ │ swarm-br0 │ │ swarm-br0 │ ││ │ gRPC │ VXLAN │ │ VXLAN │ ││ └──────────┴──────┬───────┘────────────────┘ │└──────────────────────────┴──────────────────────────────────┘Phase 0: Get Your Environment Ready
Section titled “Phase 0: Get Your Environment Ready”What You Need (Per Node)
Section titled “What You Need (Per Node)”| Role | vCPU | RAM | Disk |
|---|---|---|---|
| Manager | 2+ | 2 GB | 20 GB SSD |
| Worker | 4+ | 8 GB | 40 GB SSD |
OS and Kernel
Section titled “OS and Kernel”# Ubuntu 22.04 LTS (recommended) or Debian 12+cat /etc/os-release
# You'll need kernel 5.15+ for Firecrackeruname -rCheck KVM Access
Section titled “Check KVM Access”# See if KVM is therels -la /dev/kvm
# Verify your CPU supports virtualizationlscpu | grep Virtualization# You should see: VT-x (Intel) or AMD-V (AMD)
# If KVM isn't loaded yetsudo modprobe kvm_intel # Intel# orsudo modprobe kvm_amd # AMD
# Make it persistent across rebootsecho "kvm_intel" | sudo tee /etc/modules-load.d/kvm.confNetwork Setup
Section titled “Network Setup”# Each node needs a static or reserved DHCP IP# Make sure nodes can reach each otherping -c 3 192.168.1.11 # from managerping -c 3 192.168.1.12 # from manager
# Open the ports we needsudo ufw allow 4242/tcp # SwarmKit gRPCsudo ufw allow 7946/tcp # SwarmKit controlsudo ufw allow 7946/udp # SwarmKit gossipsudo ufw allow 4789/udp # VXLAN overlayPhase 1: Install SwarmCracker
Section titled “Phase 1: Install SwarmCracker”One-Line Install (On All Nodes)
Section titled “One-Line Install (On All Nodes)”curl -fsSL https://raw.githubusercontent.com/restuhaqza/SwarmCracker/main/install.sh | sudo bashsudo swarmcracker setup install --download-kernel --download-rootfssudo swarmcracker setup networksudo swarmcracker setup config --non-interactiveThis sets up:
- swarmcracker binary →
/usr/local/bin/swarmcracker - swarmd-firecracker →
/usr/local/bin/swarmd-firecracker - swarmcracker-agent →
/usr/local/bin/swarmcracker-agent - Firecracker v1.15.1 + jailer →
/usr/local/bin/(viasetup install) - Kernel →
/usr/share/firecracker/vmlinux(viasetup install --download-kernel) - Rootfs →
/var/lib/firecracker/rootfs/bionic.rootfs.ext4(viasetup install --download-rootfs) - Bridge + NAT →
swarm-br0(viasetup network) - Default config →
/etc/swarmcracker/config.yaml(viasetup config) - Data directory →
/var/lib/swarmkit/
Build From Source (Alternative)
Section titled “Build From Source (Alternative)”git clone https://github.com/restuhaqza/SwarmCracker.gitcd SwarmCracker
# Install build toolsmake install-tools
# Build everythingmake all
# Install to systemsudo cp build/swarmcracker /usr/local/bin/sudo cp build/swarmd-firecracker /usr/local/bin/sudo cp build/swarmcracker-agent /usr/local/bin/Make Sure It Works
Section titled “Make Sure It Works”swarmcracker version# SwarmCracker v0.10.0# Build Time: <RFC3339 timestamp># Git Commit: <short SHA># Go Version: go1.26.x (linux/amd64)
swarmcracker --helpPhase 2: Get Your Cluster Running
Section titled “Phase 2: Get Your Cluster Running”Start the Manager
Section titled “Start the Manager”# On manager-1 (192.168.1.10)sudo swarmcracker cluster init \ --advertise-addr 192.168.1.10:4242 \ --listen-addr 0.0.0.0:4242You should see:
✓ SwarmKit manager initialized✓ Control socket: /var/run/swarmkit/swarm.sock✓ TLS certificates generated✓ Join tokens saved to /var/lib/swarmkit/join-tokens.txt✓ Node ID: abc123def456Get Your Join Tokens
Section titled “Get Your Join Tokens”# Option A: Read the saved tokenssudo cat /var/lib/swarmkit/join-tokens.txt
# Option B: Ask the CLIsudo swarmcracker cluster token workerAdd Workers to the Cluster
Section titled “Add Workers to the Cluster”# On worker-1 (192.168.1.11)sudo swarmcracker cluster join 192.168.1.10:4242 \ --hostname worker-1 \ --token SWMTKN-1-<worker-token>
# On worker-2 (192.168.1.12)sudo swarmcracker cluster join 192.168.1.10:4242 \ --hostname worker-2 \ --token SWMTKN-1-<worker-token>Check That Everything’s Healthy
Section titled “Check That Everything’s Healthy”# List all nodesswarmctl ls-nodes# ID STATUS HOSTNAME AVAILABILITY ROLE# abc123 READY manager-1 ACTIVE MANAGER# def456 READY worker-1 ACTIVE WORKER# ghi789 READY worker-2 ACTIVE WORKER
# Check cluster statusswarmcracker cluster health✅ Phase 2 done: You’ve got a working 3-node cluster
Phase 3: Deploy Some Services
Section titled “Phase 3: Deploy Some Services”Launch an Nginx Web Service
Section titled “Launch an Nginx Web Service”swarmctl create-service nginx:latest# Service created: svc-nginx-143022# Image: nginx:latestScale It Up
Section titled “Scale It Up”# Scale to 5 replicas across workersswarmctl scale svc-nginx-143022 5See Your Tasks (MicroVMs)
Section titled “See Your Tasks (MicroVMs)”swarmctl ls-tasks# ID SERVICE STATUS NODE STATE# task-001 nginx RUNNING worker-1 RUNNING# task-002 nginx RUNNING worker-1 RUNNING# task-003 nginx RUNNING worker-1 RUNNING# task-004 nginx RUNNING worker-2 RUNNING# task-005 nginx RUNNING worker-2 RUNNINGEach task is its own Firecracker microVM with its own kernel.
Add a Redis Backend
Section titled “Add a Redis Backend”swarmctl create-service redis:7-alpineswarmctl scale svc-redis-<id> 3Inspect Your Running VMs
Section titled “Inspect Your Running VMs”swarmcracker vm list# VM ID SERVICE NODE STATUS MEMORY VCPUS# vm-nginx-001 nginx worker-1 RUNNING 128MB 1# vm-nginx-002 nginx worker-1 RUNNING 128MB 1# vm-nginx-003 nginx worker-1 RUNNING 128MB 1# vm-nginx-004 nginx worker-2 RUNNING 128MB 1# vm-redis-001 redis worker-1 RUNNING 256MB 1# vm-redis-002 redis worker-2 RUNNING 256MB 1Test Connectivity
Section titled “Test Connectivity”# Get the VM IP from the bridge# Each VM gets an IP on swarm-br0 (like 172.17.0.x)# Try pinging between VMs to verify networking✅ Phase 3 done: 8 microVMs running (5 nginx + 3 redis)
Phase 4: Update Services and Roll Back
Section titled “Phase 4: Update Services and Roll Back”Rolling Update
Section titled “Rolling Update”# Update to a new nginx versionswarmctl update svc-nginx-143022 --image nginx:1.25
# SwarmKit does a rolling update:# 1. Start new VM with nginx:1.25# 2. Run health checks# 3. Shift traffic away from old VM# 4. Stop old VM# 5. Repeat for each replicaUpdate With Environment Variables
Section titled “Update With Environment Variables”swarmctl update svc-nginx-143022 \ --env NGINX_PORT=8080 \ --env WORKER_PROCESSES=autoWatch the Update
Section titled “Watch the Update”# Watch tasks during the updateswarmctl ls-tasks
# Check logsswarmcracker vm logs <task-id>Rollback (Using Snapshots — See Phase 5)
Section titled “Rollback (Using Snapshots — See Phase 5)”If something goes wrong, restore from a snapshot you made before the update.
✅ Phase 4 done: Zero-downtime rolling update works
Phase 5: Snapshots and Recovery
Section titled “Phase 5: Snapshots and Recovery”Make a Pre-Update Snapshot
Section titled “Make a Pre-Update Snapshot”# Before changing things, snapshot a VMswarmctl snapshot create <task-id> pre-update-v1List Your Snapshots
Section titled “List Your Snapshots”swarmctl snapshot list# NAME CREATED TASK_ID# pre-update-v1 2026-04-11T13:00:00Z <task-id>Restore a Snapshot
Section titled “Restore a Snapshot”# Something broke? Inspect the snapshot metadata to roll back:swarmctl snapshot restore pre-update-v1Delete Old Snapshots
Section titled “Delete Old Snapshots”swarmctl snapshot rm old-backup✅ Phase 5 done: Snapshots let you undo mistakes
Phase 6: Manage Nodes
Section titled “Phase 6: Manage Nodes”Drain a Worker (Maintenance Mode)
Section titled “Drain a Worker (Maintenance Mode)”# Drain worker-1 for a kernel upgradeswarmctl drain worker-1
# Tasks get rescheduled to worker-2swarmctl ls-tasks# All tasks now on worker-2Bring the Worker Back
Section titled “Bring the Worker Back”swarmctl activate worker-1# Tasks rebalance back to worker-1Promote a Worker to Manager
Section titled “Promote a Worker to Manager”# Add a second manager for high availabilityswarmctl promote worker-2
swarmctl ls-nodes# worker-2 now shows ROLE: MANAGERDemote a Manager
Section titled “Demote a Manager”swarmctl demote worker-2✅ Phase 6 done: Node lifecycle operations work
Phase 7: Monitor and Debug
Section titled “Phase 7: Monitor and Debug”Check Cluster Status
Section titled “Check Cluster Status”swarmcracker cluster healthView Service Logs
Section titled “View Service Logs”swarmcracker vm logs <task-id>
# Follow logs in real-timeswarmcracker vm logs --follow <task-id>Inspect a Specific Task/VM
Section titled “Inspect a Specific Task/VM”swarmctl inspect <task-id># Full JSON with VM config, network, resourcesCheck the Systemd Service
Section titled “Check the Systemd Service”# If running as a systemd service (manager nodes use swarmcracker-manager)sudo systemctl status swarmcracker-workersudo journalctl -u swarmcracker-worker -fDebug Networking
Section titled “Debug Networking”# Check the bridgeip link show swarm-br0
# Check VXLAN (interface is <bridge>-vxlan; default bridge is swarm-br0)bridge fdb show dev swarm-br0-vxlan
# Check TAP devices for each VMip link show tap-*✅ Phase 7 done: You can see what’s going on
Phase 8: Manage Storage
Section titled “Phase 8: Manage Storage”Create a Volume
Section titled “Create a Volume”# --size is an integer number of megabytesswarmcracker volume create app-data --size 1024Attach Volume to a Service
Section titled “Attach Volume to a Service”swarmctl update accepts only --image, --replicas, and --env; it has no
--volume flag. Volume mounts are defined in the service/task spec.
List Volumes
Section titled “List Volumes”swarmcracker volume ls✅ Phase 8 done: Persistent storage works
Phase 9: Clean Up
Section titled “Phase 9: Clean Up”Remove Services
Section titled “Remove Services”swarmctl rm-service svc-nginx-143022swarmctl rm-service svc-redis-<id>Verify All VMs Stopped
Section titled “Verify All VMs Stopped”swarmcracker vm list# (empty)Leave the Cluster (Workers)
Section titled “Leave the Cluster (Workers)”# On worker-1 and worker-2sudo swarmcracker cluster leaveTear Down the Manager
Section titled “Tear Down the Manager”# On manager-1sudo swarmcracker cluster leave --forceClean Up
Section titled “Clean Up”sudo rm -rf /var/lib/swarmkit/sudo rm -rf /var/run/swarmkit/sudo rm -rf /etc/swarmcracker/✅ Phase 9 done: Clean teardown verified
Validation Checklist
Section titled “Validation Checklist”Use this to make sure everything works:
Infrastructure
Section titled “Infrastructure”- KVM available on all nodes
- Nodes can reach each other
- Required ports open (4242, 7946, 4789)
Cluster
Section titled “Cluster”- Manager initialized
- Workers joined
- All nodes show READY
- TLS certificates generated
- gRPC communication works
Workloads
Section titled “Workloads”- Service created and running
- Service scaled to multiple replicas
- Tasks spread across workers
- Each task = 1 running microVM
- Cross-VM networking works (VXLAN)
Lifecycle
Section titled “Lifecycle”- Rolling update (zero downtime)
- Environment variables injected
- Snapshot created and verified
- Snapshot restore works
- Service removal and VM cleanup
Node Operations
Section titled “Node Operations”- Worker drain → tasks rescheduled
- Worker activate → tasks rebalanced
- Worker promoted to manager
- Manager demoted to worker
Monitoring
Section titled “Monitoring”- Cluster status command works
- Service logs accessible
- Task inspection returns valid JSON
- Network debug commands work
Cleanup
Section titled “Cleanup”- All services removed
- All VMs stopped
- Workers left cluster
- Manager torn down
- Data directories cleaned
Automate It
Section titled “Automate It”The production example has a deploy script that automates phases 1-3:
cd examples/production-cluster/./deploy.sh --manager 192.168.1.10 --workers 192.168.1.11,192.168.1.12See examples/production-cluster/README.md for details.
Local Development Alternative
Section titled “Local Development Alternative”If you want to test without multi-node hardware:
cd examples/local-dev/./start.shThis runs a single-node cluster (manager + worker on the same machine). See examples/local-dev/README.md.
Quick Troubleshooting
Section titled “Quick Troubleshooting”| Problem | Check This |
|---|---|
| KVM not found | ls /dev/kvm → modprobe kvm_intel |
| Worker won’t join | Ping manager, check token, check port 4242 |
| VMs not starting | journalctl -u swarmcracker-worker -f, check kernel images |
| No cross-node networking | Check VXLAN: bridge fdb show dev swarm-br0-vxlan |
| Service stuck updating | swarmctl ls-tasks, check health check config |
| Snapshot fails | Check disk space, VM must be RUNNING |
This covers the complete SwarmCracker lifecycle. Each phase can be tested independently.