Golden Image Recipes — Analysis & Design
Related: Snapshots, pkg/image/, pkg/translator/, pkg/swarmkit/translator.go, recipes/
Implementation status (2026-09-28)
Section titled “Implementation status (2026-09-28)”Landed: the recipe loader/validator (pkg/golden/recipe.go), the builder
(pkg/golden/builder.go, builder_real.go), the exported pkg/image building
blocks (PullImage, ExtractImageToDir, CreateExt4FromDir, ParseDiskSize),
and the swarmcracker image build|list|inspect CLI. Work items 1, 2 and 9 below
are done; the rest remain.
Verified on a real KVM node (amd64, Firecracker v1.15.1). Each recipe was
built fresh, the runtime checked inside the ext4 with debugfs, then booted
under Firecracker with a tap NIC and its serial log inspected:
| Recipe | Build | Runtime in image | Guest boot |
|---|---|---|---|
almalinux-9-docker |
✅ | ✅ | ✅ containerd + docker, multi-user |
ubuntu-24.04-docker |
✅ | ✅ | ✅ containerd + docker, multi-user |
debian-12-docker |
✅ | ✅ | ✅ containerd + docker, multi-user |
alpine-3.20-docker |
✅ | ✅ | ✅ Starting Docker Daemon ... [ ok ] |
go test ./pkg/golden/ and go test ./cmd/swarmcracker/ -run Image pass and
go vet is clean. The reusable harness is
test-automation/scripts/golden-matrix-test.sh.
Issues the matrix exposed and the fixes applied to the builder/recipes:
- No
/tmpin minimal OCI rootfs →apt/apt-keyfailed. The chroot runner now creates/tmp,/var/tmp,/run,/var/log,/rootfirst. - apt/dpkg tried to start services through a non-running init → added a
policy-rc.d(exit 101) during provisioning, removed afterwards. - RHEL clones ship
curl-minimal, conflicting withcurl→ RPM recipes dropcurland usednf --allowerasing. - Missing
/etc/sysctl.don AlmaLinux → recipesmkdir -pit. - OpenRC
networkingneeds/etc/network/interfacesanddockerdepends on it → Alpine configureseth0from the kernelip=parameter via an init script. - 90 s
dev-ttyS0.devicestall withoutudev→ systemd recipes now installudev/systemd-udev.
systemd recipes still rely on the kernel ip= parameter for addressing; wiring
networkd/NetworkManager for systemd guests is a follow-up.
Booting through the service path is also wired: a service can set the
swarmcracker.golden label (or use swarmcracker service create --golden <name[@version]>). Controller.Prepare resolves the artifact from the golden
store (--golden-dir, default /var/lib/firecracker/golden), skips OCI image
preparation, and records the rootfs and kernel profile as task annotations so
the SwarmKit translator boots the golden guest with the kernel the recipe
pinned. Each task first materializes its own writable copy of the template
under the rootfs dir (Artifact.Materialize, reflink/sparse-aware), because
replicas must not share a read/write rootfs and a hardened daemon may not be
able to write to the golden store at all. Removal deletes that copy and never
the golden artifact. A missing artifact fails the task; there is no silent OCI
fallback.
Verified on the KVM node: a goldsvc service (label
[email protected]) booted to multi-user.target
from /usr/share/firecracker/vmlinux-runtime with the recipe boot args, and
service rm removed the per-task copy while the template’s md5 stayed
identical.
Booting through the CLI is now wired: swarmcracker vm create --golden <name[@version]> resolves the artifact, marks the task with a prebuilt rootfs,
and the executor/translator boot the recipe’s own init with its boot args.
Verified on the node: the Ubuntu 24.04 golden VM reached multi-user with
dockerd 29.1.3 listening on /run/docker.sock, and vm list shows it as
golden:[email protected].
1. Problem statement
Section titled “1. Problem statement”Today SwarmCracker turns one OCI image into one microVM that runs exactly one process tree. The flow is:
OCI image ──pull/extract──> dir ──inject tini + /init wrapper──> mkfs.ext4 ──> boot │ kernel: ... nomodules init=/sbin/init one writable ext4 root deviceTwo requests are not served by this model:
- “Golden images” — a curated, versioned, pre-built rootfs per distro that boots fast and is reused across tasks instead of being re-derived from layers every time.
- “A microVM that can run a container runtime inside (Docker/containerd)” — a full guest OS where the VM is the host and Docker runs inside it.
These overlap: a golden image is the natural delivery vehicle for a Docker-capable VM. This document analyses feasibility and defines a recipe format for building both.
2. Current-state analysis (grounded in code)
Section titled “2. Current-state analysis (grounded in code)”2.1 What the preparer does
Section titled “2.1 What the preparer does”pkg/image/preparer.go:
- Pulls the OCI image with
go-containerregistry(daemon-free), falls back todocker/podmanCLI. - Extracts the flattened filesystem into a temp dir.
DetectInitType(pkg/image/detector.go) classifies the image asscratch | systemd | openrc | sysvinit | tini | dumb-init | none.systemdis treated asincompatibleand fails preparation.- Injects an init wrapper (
pkg/image/init.go,wrapper.go): writes/sbin/tini, generates a/sbin/initshell script that mounts proc/sys/dev, configures eth0 from the kernelip=arg, exports OCIENV, drops to OCIUSER, andexecs the OCIENTRYPOINT/CMDunder tini./init -> /sbin/init. - Injects essentials (
essentials.go):resolv.conf,hosts,nsswitch,machine-id, and creates/tmp,/run,/var/log,/root. - Creates an ext4 with
mkfs.ext4 -dat content-size + 50 % overhead, minimum 100 MiB, honouring a per-serviceswarmcracker.disklabel.
2.2 What the translator/boot path does
Section titled “2.2 What the translator/boot path does”pkg/translator/translator.goandpkg/swarmkit/translator.gohardcode boot args:console=ttyS0 reboot=k panic=1 pci=off nomodules init=/sbin/init(init=/initin one path), plusip=<ip>::<gw>:<mask>::eth0:off.- Exactly one drive is created:
task.ID-> the rootfs.ext4,is_root_device: true,is_read_only: false. machine-configis sized from SwarmKitResources.Reservations(min 1 vCPU / 512 MiB).
2.3 What the snapshot package adds
Section titled “2.3 What the snapshot package adds”pkg/snapshot can pause a VM and write vm.state + vm.mem, restoring in
2–3× less than cold boot. A booted golden VM snapshot is the fastest
possible golden-image delivery — but it pins the exact kernel/rootfs/boot-args,
so it must be versioned with the recipe.
2.4 The hard constraints for “Docker inside”
Section titled “2.4 The hard constraints for “Docker inside””| Constraint | Where it lives today | Impact |
|---|---|---|
init=/sbin/init tini wrapper replaces the real init |
image/init.go, wrapper.go |
Must be bypassed for a “VM host” golden image; systemd/OpenRC must own PID 1 |
| systemd rejected as incompatible | image/detector.go |
Must become allow-listed per recipe |
nomodules boot arg |
translators | Fine if every needed driver is built-in (=y), which we control via the kernel profile |
| one root drive | translators | Docker needs an image store; either enlarge root or add a second writable data disk |
| 100 MiB default rootfs | preparer.go |
Full distro + Docker needs GiBs; recipe must set disk size |
pci=off |
translators | OK — Firecracker uses virtio-mmio, not PCI |
| one process = one task | executor/SwarmKit model | A Docker-capable VM is a host; mapping a SwarmKit task to work inside it needs a guest agent (see §6) |
3. Kernel feasibility — evidence
Section titled “3. Kernel feasibility — evidence”I downloaded the Firecracker CI guest kernel config
(resources/guest_configs/microvm-kernel-ci-x86_64-6.1.config, 3563 lines) and
grepped the options a container runtime needs:
| Capability | Option | Value |
|---|---|---|
| Overlay storage driver | CONFIG_OVERLAY_FS |
=y |
| Bridge + veth | CONFIG_BRIDGE, CONFIG_VETH |
=y |
| Netfilter core | CONFIG_NETFILTER, CONFIG_NF_CONNTRACK, CONFIG_NF_TABLES |
=y |
| NAT / masquerade | CONFIG_NF_NAT, CONFIG_NETFILTER_XT_TARGET_MASQUERADE, CONFIG_IP_NF_NAT, CONFIG_IP_NF_TARGET_MASQUERADE |
=y |
| iptables filter/mangle | CONFIG_IP_NF_IPTABLES, CONFIG_IP_NF_FILTER, CONFIG_IP_NF_MANGLE, CONFIG_NETFILTER_XTABLES |
=y |
| addrtype / conntrack match | CONFIG_NETFILTER_XT_MATCH_ADDRTYPE, ..._CONNTRACK |
=y |
| cgroup controllers | CONFIG_CGROUPS, MEMCG, CGROUP_SCHED, CFS_BANDWIDTH, CPUSETS, BLK_CGROUP, CGROUP_PIDS/DEVICE/FREEZER, CGROUP_BPF |
=y |
| namespaces | CONFIG_NAMESPACES, NET_NS, PID_NS, USER_NS, IPC_NS, UTS_NS |
=y |
| misc runtime | BINFMT_MISC, POSIX_MQUEUE, KEYS, SECCOMP, BPF_SYSCALL, VXLAN, EXT4_FS |
=y |
| Firecracker devices | VIRTIO_BLK, VIRTIO_NET, VIRTIO_MMIO, VIRTIO_PCI, SERIAL_8250_CONSOLE, PRINTK, ACPI, PCI |
=y |
| Bridge filtering for k8s/Docker | CONFIG_BRIDGE_NETFILTER |
=y |
Conclusion: the stock Firecracker CI 6.1 kernel is close to being able to run
Docker. Everything essential is built in, so nomodules is not fatal.
Gaps to close in a dedicated guest-runtime kernel profile:
| Option | Current | Why |
|---|---|---|
CONFIG_NETFILTER_XT_MATCH_COMMENT |
not set | Docker/CNI insert -m comment --comment …; insertion fails without it |
CONFIG_MACVLAN |
not set | Docker macvlan networks; some CNI plugins |
CONFIG_IP_VS |
not set | Docker Swarm ingress / ipvs load-balancing mode |
CONFIG_IP_NF_RAW, CONFIG_IP_NF_ARPTABLES, CONFIG_BRIDGE_NF_EBTABLES |
not set | iptables-legacy raw/arp/ebtables tables used by some CNIs and legacy Docker paths |
CONFIG_NETFILTER_XT_MATCH_BPF |
not set | Cilium/eBPF CNIs |
CONFIG_SECURITY_APPARMOR |
not set | Docker AppArmor confinement (optional; document that guests run unconfined) |
CONFIG_SECURITY_SELINUX |
=y |
Fedora/Rocky/AL2023 ship SELinux; keep enabled, ship policy in rootfs |
CONFIG_ZRAM |
not set | Optional swap/compressed memory under pressure |
CONFIG_DM_THIN_PROVISIONING |
not set | Only if devicemapper storage driver is required (avoid; use overlay2) |
We keep every option built-in (=y) so the existing nomodules boot arg stays
valid and we avoid building an initramfs.
4. Recipe model
Section titled “4. Recipe model”4.1 Concept
Section titled “4.1 Concept”A recipe is a declarative build input that produces one immutable artifact:
recipe.yaml ──> builder ──> golden-<name>-<version>.ext4 └─> golden-<name>-<version>.json (metadata/checksums) └─> kernel profile → vmlinux-runtime-<ver>Artifacts are content-addressed, versioned, and cached. A sealed image is
registered by name so tasks can request golden:[email protected].
4.2 Schema (proposed, recipes/<name>.yaml)
Section titled “4.2 Schema (proposed, recipes/<name>.yaml)”apiVersion: swarmcracker.io/v1alpha1kind: GoldenImagemetadata: name: ubuntu-24.04-docker version: 1.0.0 description: Ubuntu 24.04 with Docker Engine, systemd, cgroup v2
spec: arch: [amd64, arm64]
source: # how to obtain the base userspace type: oci # oci | rootfs-tar | cloud-image ref: docker.io/library/ubuntu:24.04
init: system: systemd # systemd | openrc | sysvinit | custom bootArgs: # appended to the translator's base args - systemd.unified_cgroup_hierarchy=1 - systemd.journald.forward_to_console=1 # init= is derived from system (or set explicitly for custom)
kernel: profile: guest-runtime-6.1 # selects the vmlinux variant mustBoot: true
runtime: # the container runtime installed *inside* name: docker # docker | containerd | podman | none version: "27.5" storageDriver: overlay2 cgroupVersion: v2 install: script # script | packages | mirror | none
disk: rootMinSize: 4GiB # mkfs.ext4 size floor dataDisk: # optional second writable drive for the image store size: 20GiB mount: /var/lib/docker fs: ext4
network: guestCIDR: 172.17.0.0/16 # docker0 subnet placeholder
provision: | # run inside a chroot (qemu-user for cross-arch) set -eux export DEBIAN_FRONTEND=noninteractive apt-get update apt-get install -y ca-certificates curl iproute2 iptables nftables install -m0755 /tmp/get-docker.sh /root/get-docker.sh /root/get-docker.sh systemctl enable docker
seal: # hygiene before mkfs - truncate -s 0 /etc/machine-id - rm -f /etc/ssh/ssh_host_* - rm -rf /var/cache/apt /var/lib/apt/lists/*
health: - command: systemctl is-system-running --wait - command: docker info timeout: 60s
verify: # boot-time assertions run by the builder/CI - name: kernel-has-overlayfs guest: grep -qw overlay /proc/filesystems - name: docker-bridge guest: ip link show docker0Fields are intentionally close to the existing config vocabulary so the Go
implementation can reuse ImagesConfig/ExecutorConfig validation.
4.3 Recipe catalogue
Section titled “4.3 Recipe catalogue”| Recipe | Init | Runtime | Notes |
|---|---|---|---|
alpine-3.20-docker |
OpenRC | Docker | Smallest; musl; OpenRC must mount cgroups |
debian-12-docker |
systemd | Docker | Conservative, glibc baseline |
ubuntu-22.04-docker |
systemd | Docker | Matches Firecracker CI rootfs |
ubuntu-24.04-docker |
systemd | Docker | Current LTS, cgroup v2 default |
fedora-42-docker |
systemd | Docker | cgroup v2, SELinux enforcing |
rocky-9-docker |
systemd | Docker | RHEL-compatible, SELinux |
almalinux-9-docker |
systemd | Docker | RHEL-compatible, SELinux |
amazonlinux-2023-docker |
systemd | Docker | Firecracker-tested platform |
opensuse-leap-15-docker |
systemd | Docker | Btrfs-friendly, SELinux/AppArmor variants |
arch-docker |
systemd | Docker | Rolling; reproducible snapshot date required |
distroless-minimal |
tini | none | The current single-workload model, as a recipe |
busybox-scratch |
tini | none | Current scratch handling, as a recipe |
See recipes/ for the concrete YAML and recipes/README.md for the schema
reference and build instructions.
5. Building a golden image
Section titled “5. Building a golden image”5.1 Pipeline
Section titled “5.1 Pipeline”- Resolve base (
ociref pinned by digest, or tarball). - Extract to a working dir (reuse
preparer.extractOCIImage). - Provision in a chroot:
- native arch:
chrootdirectly; - cross arch: register
qemu-user-staticviabinfmt_miscand chroot (tonistiigi/binfmtpattern) — the host already supportsbinfmt_misc. - mount
/proc,/sys,/dev,/dev/ptsinto the chroot for package scripts.
- native arch:
- Install the runtime per recipe (
runtime.install). - Configure init: enable units/services, mount cgroups for OpenRC,
set
docker0CIDR, enablenet.ipv4.ip_forward=1. - Seal (recipe
sealsteps); optionally set the rootfs read-only baseline. - Size & format:
mkfs.ext4atmax(content*1.5, rootMinSize)(reusecreateExt4ImageWithOverhead). Record checksum + recipe digest. - Smoke-boot with the
guest-runtimekernel and runverifyassertions (machine-readable JSON result). - Register: write
<name>@<version>.json; optionally take a snapshot of the booted VM as the fast-path artifact. - Publish: copy to each worker’s rootfs dir (or an object store the fetcher understands).
5.2 Kernel build
Section titled “5.2 Kernel build”Start from the Firecracker CI config, apply
recipes/kernel/guest-runtime-6.1.fragment, build vmlinux with make vmlinux
(x86_64) / make Image (aarch64), install as
/usr/share/firecracker/vmlinux-runtime. Keep options =y. The fragment must
be validated by a CI job that boots a recipe and runs the verify commands.
6. Execution model — how a task uses a Docker-capable VM
Section titled “6. Execution model — how a task uses a Docker-capable VM”This is the part that needs a product decision. Three options:
Option A — “runtime available, still one workload” (smallest change)
Section titled “Option A — “runtime available, still one workload” (smallest change)”Golden image boots its real init, but SwarmCracker still injects the task command
as the workload. The image simply has docker on PATH. The service can shell
out to Docker, but there is no dockerd-by-default and no orchestration inside.
Cheap, but doesn’t really deliver “VM that runs Docker”.
Option B — “VM as host + guest agent” (recommended first target)
Section titled “Option B — “VM as host + guest agent” (recommended first target)”The golden VM boots systemd + dockerd and runs a small SwarmCracker guest
agent on PID 1’s supervision. The agent:
- receives the service’s OCI image ref + command/env/labels over vsock (preferred; Firecracker has a vsock device) or serial,
docker pull+docker runinside the guest,- streams logs/health/exit status back,
- applies the task’s mounts/configs/secrets by materialising them as guest
bind mounts or
docker runflags.
SwarmCracker’s executor keeps its task state machine but, for
runtime_mode: vm, it provisions the golden VM and drives the agent instead of
Initialize/Start/Delete of a single process. This delivers real Docker-in-VM
with one host-facing task. Cost: a guest-agent protocol and lifecycle work.
Option C — “nested swarm” (future)
Section titled “Option C — “nested swarm” (future)”Run swarmcracker-agent inside the golden VM and join the same (or a child)
cluster. The inner VM becomes a schedulable node; the outer executor is just a
VM lifecycle provider. Most elegant long-term, biggest blast radius (nested
networking, two schedulers, token distribution). Defer.
Recommendation: ship recipes + guest-runtime kernel + Option A first
(unblocks images and proves the kernel), then Option B.
7. Networking
Section titled “7. Networking”Inside the guest:
eth0 (virtio-net, 192.168.127.2/24) ── docker0 (172.17.0.0/16) ── containers- systemd images: enable
net.ipv4.ip_forward=1(/etc/sysctl.d/99-docker.conf); Docker setsFORWARDpolicy itself. - OpenRC/Alpine: an init script must mount cgroup v2 at
/sys/fs/cgroupwith-o rw,relatime(andcgroup2fstype), then startdockerd. - The host already NATs the guest (
network.nat_enabled); nested container traffic NATs twice (container → docker0 → eth0 → host MASQUERADE). Correct but worth documenting for MTU and conntrack limits. BRIDGE_NETFILTERis built in; ensurenet.bridge.bridge-nf-call-iptables=1only if a workload needs strict bridge filtering (k8s); Docker itself warns when it’s absent.
8. Storage
Section titled “8. Storage”- Root: ext4, recipe-sized (4–8 GiB floor). Everything writable.
- Data disk (recommended): second Firecracker drive, ext4, mounted at
/var/lib/docker. Requires the translator to emit >1 drive and the rootfs/metadata to name the mount. Protects the base image from image-store growth and lets the base be made read-only later. - Read-only base + overlay (firecracker-containerd style,
overlay-init): strongest immutability; defer until Option B lands. - Snapshots: after first successful boot, snapshot the golden VM; restore is the fast path. Snapshot artifacts must record the recipe+kernel digest.
9. Security
Section titled “9. Security”- Docker-in-VM is not a new host-privilege boundary: the guest kernel is still the isolation boundary; nested containers are namespaces inside the guest. Keep the jailer and seccomp on the Firecracker process.
- Never pass the host Docker socket into a guest.
- AppArmor is not enabled in the FC kernel; recipes must document that guest workloads are not AppArmor-confined. SELinux-capable distros keep SELinux enforcing with the distro’s container policy.
- Golden images must be scanned and signed; the registry/fetcher should
verify a digest before boot.
sealremoves host keys, machine-id, and caches. nomodulesstays: no module loading in the guest means noinsmodsurface.
10. Gaps / work items in the Go codebase
Section titled “10. Gaps / work items in the Go codebase”| # | Item | Package | Notes |
|---|---|---|---|
| 1 | GoldenImage recipe type + YAML loader/validator |
new pkg/golden |
✅ done |
| 2 | Recipe builder (extract → chroot provision → seal → mkfs.ext4) | pkg/golden, pkg/image |
✅ done (reuses ExtractImageToDir, CreateExt4FromDir) |
| 3 | Allow systemd/openrc as first-class init for vm mode |
pkg/image/detector.go, init.go |
✅ prebuilt-rootfs tasks boot the recipe init; preparer detector untouched |
| 4 | Make init= and boot args recipe-driven |
pkg/translator, pkg/swarmkit/translator.go |
✅ via task annotations in pkg/translator; swarmkit translator pending |
| 5 | Multi-drive support (root + data disk) | translators | drives[] already a slice |
| 6 | executor.runtime_mode: container | vm config + task annotation |
pkg/config, pkg/executor |
✅ annotation-based (swarmcracker.prebuilt_rootfs); no config field needed |
| 7 | Guest agent + vsock protocol | new pkg/guestagent |
Option B |
| 8 | guest-runtime kernel build + CI boot validation |
scripts/, CI |
kconfig fragment in recipes/kernel/ |
| 9 | swarmcracker image build/list/inspect CLI |
cmd/swarmcracker |
✅ done |
| 10 | Register golden images and resolve golden:<name>@<ver> |
pkg/image, pkg/executor |
content-addressed cache |
| 11 | Snapshot a booted golden VM as fast-path artifact | pkg/snapshot |
tie snapshot to recipe digest |
11. Risks & open questions
Section titled “11. Risks & open questions”- Kubernetes vs Docker. Recipes target Docker; if k8s-in-VM is a goal,
CONFIG_MEMCG_SWAP,IP_VS,MACVLAN,BRIDGE_NF_EBTABLESand kubeadm preflight become mandatory. Scope now, expand later. - Cross-arch builds need
binfmt_misc+qemu-user-staticon the builder; document the dependency. - Image size is the enemy of fast boot. Option: two artifacts per recipe — a fat build image and a stripped runtime image.
- systemd in Firecracker is proven (firecracker-containerd boots systemd),
but their config used
systemd.unified_cgroup_hierarchy=0. Modern distros default to cgroup v2 (=1); verify per distro. - Guest clock can drift; systemd images want
CONFIG_KVM_GUEST/KVM clock (present) and possibly chrony/NTP from the host. - Registry credentials for private base images must flow to the builder and to dockerd inside the guest.
12. Next step
Section titled “12. Next step”The recipe loader/builder and CLI are done (see implementation status above).
Next: work item 8 (build and validate the guest-runtime kernel), then work
items 3–5 to let a task boot a golden image with its own init and multi-drive
layout, and finally Option B (pkg/guestagent) so a SwarmKit task can drive
docker run inside the golden VM.