k7d: Forking a Live Kubernetes Cluster in 105 ms for RL Rollouts — the Fine Print
Sources
- Katakate k7d — repository and README
- k7d LATENCY_BUDGETS.md — the CI-enforced latency table
- k7d v0.8.0 release notes — coordinated cluster keyframes
- DeepSeekMath — the paper that introduced GRPO (arXiv 2402.03300)
- Show HN: K7d — Fork live Kubernetes clusters (Aug 2026 thread)
- k7d client reference — examples/cluster-tree-search/k7d_client.py
Every platform team building AI-agent training infrastructure eventually hits the same wall: reinforcement learning on real infrastructure — Kubernetes, charts, admission controllers, CNIs — needs thousands of isolated, resettable, and above all identical worlds. The standard answer is a cold kind or k3s cluster per trial: about 30 seconds of boot time and a full RAM bill per copy. When your GRPO loop wants a group of 16 rollouts from the same starting state, that tax compounds into the dominant cost of the entire training run.
k7d — a Rust VMM and containerd shim from Katakate (docs at docs.katakate.org) — takes the opposite bet: boot the cluster once, then fork the live machine. The headline: a running 3-VM Kubernetes cluster forks in ~105 ms on the minimal CI guest, a Ubuntu k3s+Cilium+Tetragon cluster in ~1.1 s, and 50 three-VM cluster forks land in ~4.1 s on one 64 GB bare-metal box. Yesterday's v0.8.0 release added coordinated cluster keyframes that cut the source pause down to the residual dirty pages via live pre-copy — the most architecturally significant change since the project went public in August.
Methodology disclosure, because we owe you that: our analysis box has no /dev/kvm, so nothing in this piece is hands-on. What we did instead is audit the source of truth most VMM projects don't have: k7d wires every latency claim on its README into CI-enforced budgets (LATENCY_BUDGETS.md, constants imported from crates/k7d/src/latency.rs into crates/k7d/tests/test_latency.rs). Every number below is traced to a repo file. That is k7d's single most credible property — and, as we'll see, its own budget file is refreshingly honest about where the numbers are loose.
The structural shift: fork the universe, don't reset it
Group-relative policy optimization — GRPO, introduced in the DeepSeekMath paper — compares rewards within a group of rollouts from a shared prompt. The group-relative baseline is what makes it sample-efficient, and it is also what makes it brutally sensitive to environment state. If member A starts from a cold page cache, a different etcd revision, or a half-ready Deployment while member B starts warm, the reward gap between them is noise, not signal. The paper's authors controlled this with math on the model side; on the infrastructure side, nobody controls it — because a "reset" script that re-applies manifests never reproduces the same kernel timers, the same TLS session state, or the same in-flight reconcile loop.
k7d's answer is to make the copies byte-identical by construction: a fork is a copy-on-write clone of the live machine state — RAM, disk, processes, established TCP between member VMs, in-cluster TLS, kube-apiserver state, everything. From inside the fork, nothing moved. The README's phrasing is right: that is the difference between "we reset the env" and "we cloned the universe."
The second shift is economic. CoW forks share the parent's memory pages until someone writes: 50 cluster copies cost 50× the dirty pages, not 50× the guest RAM. That is what makes a 64 GB box a plausible RL gym. The daemon exposes the fork loop as a budgeted tree — you protect winners, prune losers, and let the daemon auto-evict under RAM and disk caps while the loop runs.
The third shift is the one nobody else does: k7d owns the guest clock. Because the VMM controls kvmclock, per-vCPU TSC offsets, and timer re-arms, the same pause that makes a fork correct can also move every member's clock by the same delta. An 8.7-minute guest-time episode that is mostly kubelet waiting on probes compresses 1.83× with lockstep clock jumps at zero fidelity cost — the parts of an RL episode that are pure waiting stop costing wall-clock time.
Architectural blueprint: what actually happens in 105 ms
The design is one daemon (k7d, one process, one address space, many VMs) owning a KVM VMM, a snapshot-tree manager, and the guest clock. Your trainer talks JSON-lines over a Unix socket at /run/k7d/k7d.sock. A fork, end to end, is a short critical section on the source:
sequenceDiagram
participant T as Trainer (JSON-lines client)
participant D as k7d daemon
participant S as Source VMs (3-node cluster)
participant C as Child VMs (forks)
T->>D: tree_fork_batch(n)
D->>S: pause every vCPU, drain in-flight virtio I/O
D->>D: dirty bitmap = KVM dirty log ∪ device-written pages
D->>C: map parent RAM copy-on-write, copy only dirty pages
D->>C: new bridge per child, replay guest IPs / MACs
D->>C: restore vCPUs, re-arm timers, KVM_SET_CLOCK (never backwards)
D->>S: resume
D-->>T: n children + latency_ns
Three implementation details decide whether any of this is real, and the project documents all three in its CHALLENGES.md (nearly 200 recorded footguns):
- Device writes are invisible to the hypervisor's dirty log. KVM's dirty bitmap only sees CPU writes. k7d's block devices write guest memory from userspace, so a naive fork taken after disk I/O would resurrect pre-I/O bytes — silent corruption that no test at the application layer would catch. The fix: track device-dirty pages in the VMM, merge them into the fork bitmap, and drain in-flight I/O before the bitmap is read.
- A restored guest with a blank timer chip freezes time. An early fork left the interval timer unprogrammed — no timer interrupts,
CLOCK_REALTIMEstuck, and Kubernetes quietly parking pods. The fix: re-arm the timer and reset the paravirtual clock on every restore. This is the class of bug that separates a demo from something an RL loop can trust for ten thousand resets. - Identical IPs only work if the L2 domains are separate. Each fork gets a fresh Linux bridge with the same guest IPs and MACs as the source. Inside the guest, nothing moved — same addresses, same ARP cache, same established TCP — so kubelet, the CNI, and the control plane keep running. Separate bridges mean forks can't see each other. Without this trick, every fork would force a kubelet restart and the sub-second claims would collapse.
The tree the daemon manages underneath all of this is the part platform engineers should internalize — it is a resource model, not just a data structure:
[base cluster — root, 48 s bring-up]
| | |
[fork A] [fork B] [fork C]
| (pruned: | \
[fork A1] low reward) [fork C1] [fork C']
(protected: (evicted: (tree_rollback
winner) RAM budget) from C)
RAM: one mapping, parent pages shared until a fork writes
disk: reflink keyframes (v0.8.0) — children CoW from the keyframe
budget: max_live_vms / max_live_ram_bytes / max_disk_bytes / max_chain_depth
The v0.8.0 change that matters: keyframes cut the pause to the residual
Fork latency has two components: the pause you impose on the source (all vCPUs frozen while dirty pages are read and child mappings are created) and the materialization of children. Until v0.7.x, the pause scaled with guest RAM and dirtiness. The v0.8.0 release notes describe the fix with unusual precision: coordinated cluster keyframes (tree_create_cluster, tree_adopt_cluster, and in-tree children) now "cut every native member at one instant. Live pre-copy writes RAM while the guest runs, so the pause is the residual rather than guest RAM."
That is live migration's iterative-copy trick applied to forking: copy most of RAM while the guest runs, then pause only long enough to sweep the residual delta. The reply now reports pause_ns and max_pause_ns per TreeCreated, and an adopted root keyframe can be refreshed into a new generation — meaning a long-lived gym can re-baseline its root without paying the 48 s Ubuntu bring-up again. The same release adds an opt-in per-tree dilation policy (fixed / down / track, with caller knobs for window, cadence, and ladder) and opt-in per-tree KSM (ksm: true, off by default, inherited by forks, same-tenant trees only — the daemon deliberately refuses to start ksmd for you).
Read the release cadence as a warning as much as a feature: eight minor releases in six weeks (v0.2.1 on Aug 17 → v0.8.0 on Sep 28). The wire API is moving. Pin versions in any gym you build.
The numbers, audited
Here is what we mean by "CI-enforced": budgets are set at roughly 2× the typical observed value on the reference bare-metal node (an AMD Ryzen 5 3600, 64 GB, NVMe — about €40/month of metal), and an integration test fails if the daemon's latency_ns drifts past budget. Selected rows from the budgets file, with the honesty flags it documents itself:
| Operation | Enforced budget | Typical | What the file admits |
|---|---|---|---|
| Warm fork, <25% dirty (1 VM) | 50 ms | ~5 ms | Warm 2-vCPU source: budget 75 ms, typical "TBD — start generous, tighten after measuring" |
| Cluster-tree fork, 3 VMs | 750 ms | ~75 ms quiet / 188–494 ms loaded | "Intentionally generous" vs the ~75 ms design projection |
| Inner-k3s fork under load churn | 1 s | ~104 ms | The README's headline ~105 ms is this churn-loaded row, not a quiet-host best case |
| 50× 3-VM batch fork | 20 s | ~3.85 s (~77 ms/cluster) | Full net jail costs ~2 ms/cluster because all member TAPs pin in ONE nft transaction; one nft process per TAP measured 12.7 s — 3× worse |
| Ubuntu VM boot → agent on vsock | 120 s | ~3.2 s | "Deliberately loose… a regression that doubles Ubuntu boot time passes" |
Two audit findings worth your attention. First, the file documents that under full-suite serial runs, test_cluster_fork_warm_3vm_under_50ms "swings up to ~500 ms" versus ~25 ms run alone — a 20× swing between isolated and suite-mode measurement. Your quiet-host numbers will not survive a loaded CI box; the project budgets for that, and so should you. Second, the README quotes VM boot→agent at ~163 ms while the budgets file's typical is ~194–205 ms. The gap is small but it is exactly the kind of drift you get when prose and tests live in different files — k7d at least has the tests; most VMM marketing has neither.
What a fork costs as your in-cluster surface grows — per-profile deltas on the 3-VM inner-k3s fixture, all from the budgets file:
| Profile | Typical fork | Delta vs lean baseline |
|---|---|---|
| Lean baseline (add-ons off, load churn) | ~104 ms | — |
| + CoreDNS | ~175 ms | +71 ms |
| + ServiceLB / Traefik ingress | ~206 ms | +102 ms |
| + local-path PVC workload | ~200 ms | +96 ms |
| + metrics-server | ~196 ms | +92 ms |
| + NetworkPolicy (default-deny + allow) | ~195 ms | +91 ms |
The pattern: every extra reconciling controller adds roughly 90–100 ms to the fork because more vCPUs are mid-tick when the pause lands, and more memory is dirty. A "realistic" cluster with all of these is a ~2× fork versus the lean fixture — still well under the 1 s budget, but a long way from the marketing number.
Time warp, read skeptically
The clock tricks are the part of k7d with no equivalent anywhere else, and also the part where the README's own table rewards careful reading. The 8.7-minute guest-time episode, 3-node k3s:
| Mechanism | Wall time | Fidelity cost |
|---|---|---|
| Stock clock | 522 s | — |
| Lockstep jump (auto-warp, no kernel patch): when every vCPU is idle, pause ~0.2 ms at ~20 Hz, add δ to kvmclock + every TSC offset | 286 s (1.83×) | Zero — same probe failures, restarts, leases as stock |
| Continuous dilation ×N (~200-line carried KVM patch): guest clocks run N× wall | 79 s | Real and bounded: guest-visible fsync 2.3 → 14 ms, pod RTT 0.69 → 0.92 ms, a guest-paced churn loop does 37 rounds instead of 86 |
| Both combined | 76 s | Dilation consumes the idle windows that jumps need — they don't stack |
Now the skeptical footnote the table invites: the dilation row is labeled "(8.00×) at N=8," but 522 s stock → 79 s is a 6.6× wall speedup. The "(8.00×)" is the dilation factor the guest clock runs at, not the speedup your episode gets — because "only waiting is compressible." Work — CPU, I/O with real device latency — does not dilate; an episode that is ~97% waiting is what yields 6.6× at N=8. Read every vendor ×-label as a factor, never as a speedup, and the README's own fidelity column (fsync 6× slower guest-visible) tells you exactly what you trade: any RL reward signal that correlates with real I/O latency is now measured against a distorted clock. The zero-cost option is the lockstep jump, and the daemon exposes it as auto_warp(cluster_id, mode) with off / jump / micro — a declined warp is always correct, which is the right default posture.
Runnable: a GRPO group loop against the real API
The reference client is a single-file, no-dependency JSON-lines client (examples/cluster-tree-search/k7d_client.py) — the project explicitly says to treat it as a template for wiring your trainer, not an SDK. Setup first: you need a Linux amd64 host with /dev/kvm; the quickstart doctor-checks /dev/kvm, /dev/vhost-vsock, /dev/net/tun, and cgroup v2. No arm64 build exists.
git clone https://github.com/Katakate/k7d && cd k7d
sudo ./scripts/quickstart.sh # boots a busybox 3-node tree, forks 4 branches, prints wall-clocks
k7d doctor # host checks; does not start a daemon
k7d quickstart # 3-node busybox + 4 forks (daemon must be up)
The loop below uses only verbs and fields that exist in the reference client and demo, with the demo's own budget values (64 GiB RAM / 256 GiB disk, headroom for root + branches + a rollback fork). This is the shape of a GRPO group step:
from k7d_client import K7dClient # the reference client; K7D_SOCKET env overrides the path
client = K7dClient() # default: /run/k7d/k7d.sock, 600 s timeout
TREE = "grpo-run-01"
GROUP = 16 # rollouts per generation (the G in GRPO)
# 1) Root the tree at a warm checkpoint. Budgets are enforced by the daemon,
# not by your discipline: auto_evict prunes unprotected branches under pressure.
budget = {
"max_live_vms": 3 * (GROUP + 2) + 16, # demo formula: root + branches + rollback headroom
"max_live_ram_bytes": 64 * 1024**3,
"max_disk_bytes": 256 * 1024**3,
"max_chain_depth": 32,
}
created = client.tree_create_cluster(
3, # vm_count — N-node, no hard-coded 3
{"kernel": KERNEL, "initrd": INITRD,
"memory_mb": 256, "vcpus": 1, "vsock": True},
tree_id=TREE, budget=budget,
)
root = created["root_id"]
# 2) Fork the whole group from the root — byte-identical starts by construction.
# dilation=8 is opt-in (v0.8.0 policy); drop it for fidelity-sensitive rewards.
batch = client.tree_fork_batch(TREE, root, [f"gen0-{i}" for i in range(GROUP)], dilation=8)
forks = batch["forks"]
for f in forks:
assert not any(f.get("fork_memory_full_copy") or []) # CoW broke if this fires — stop the run
# 3) Drive each rollout over vsock by CID (the daemon does NOT proxy your exec):
# keep f["guest_cids"] from the fork reply, or re-read tree_nodes -> live_guest_cids.
# Score with YOUR reward model; k7d owns environments and budgets only.
scores = sorted(((reward(f), f) for f in forks), key=lambda x: x[0], reverse=True)
winner = scores[0][1]
# 4) Keep the winner, prune the losers — budgets make the tree safe to leave running.
client.tree_protect(TREE, winner["node_id"]) # budget pressure can no longer evict it
for _, f in scores[1:]:
client.tree_prune(TREE, f["node_id"])
# 5) Roll forward (fork from the winner) or roll back — rollback is non-destructive:
# it is another fork from an earlier node, the earlier node survives.
nxt = client.tree_fork_batch(TREE, winner["node_id"], [f"gen1-{i}" for i in range(GROUP)])
client.tree_auto_evict(TREE) # enforce RAM/disk caps now
# 6) Teardown: stops guests, releases broker pins — caller-supplied disks stay YOURS to unlink.
client.tree_drop(TREE)
Fields worth knowing on the wire, from the demo and release notes:
{
"tree_id": "grpo-run-01",
"forks": [
{
"node_id": "n-a1b2",
"guest_cids": [4, 5, 6],
"fork_memory_full_copy": [false, false, false]
}
]
}
Three API-contract footguns, all documented but all easy to trip on the first week:
list_vmsis only the sandbox registry. Tree members (created or adopted) do not appear in it, and the oldcreate_vmvm_iddoes not reach a tree fork. Exec goes vsock-to-CID; keepguest_cidsfrom the fork reply or re-readtree_nodes.stop_vmon an id adopted into a tree errors by design and does not free tree disks — teardown istree_prune(root) ortree_drop. The v0.8.0 release notes also confirmtree_dropdoes not unlink caller-supplied block devices: disk cleanup is your job.dump_serialafter a fork reads a fresh COM1 buffer that starts empty. Dump the source console before forking, or address the child by its new CID orvm_idof the formtree_id/node_id.
Two backends, two isolation stories — pick with your eyes open
One daemon, two VMMs: backend: native (k7d's in-process rust-vmm) or backend: firecracker (stock Firecracker plus its stock jailer). A tree is one backend, never mixed. The trade is blunt and the README states it plainly:
| native (k7d) | firecracker (k7d-fc) | |
|---|---|---|
| Warm fork of a running VM | ~5 ms | ~76–140 ms per child (jailer spawn each) |
| Per-VM jail (own uid, chroot, seccomp) | No — CoW fork requires parent and child memory in one address space | Yes — stock jailer |
| Guest time warp / dilation | Yes | No — Create fails loudly, nothing silently drops |
| virtiofs (hostPath) | Yes | No |
| Built for | RL gyms, fleets you control | Hostile / multi-tenant guests |
The security consequence of native is the line every evaluator should underline: because live CoW fork requires sibling memory mappings in one process, isolation between sibling forks of the same tenant is weaker than a jailed VMM, and the daemon is by design shared across branches of one tree. For RL training that is an acceptable trade — the branches are all yours. For anything multi-tenant, use the Firecracker backend, where a guest→VMM bug dies in the jail. The daemon also ships privilege separation (K7D_PRIVSEP=on: a root node broker with no listening socket plus a uid-k7d, empty-cap VM broker; k7d doctor reports it), and selected critical paths carry machine-checked proofs — Kani on unsafe memory arithmetic, Aeneas→Lean on the tree budget/eviction model (eviction never frees a page a live descendant references). Selected paths, not the whole runtime — and no external security audit exists yet.
Context from the neighborhood, since you'll be asked "isn't this just X": Kata Containers and Firecracker can snapshot-and-restore, but not live-fork a running VM; E2B-style sandbox fan-out tools restore N copies of one sandbox (~220 ms class); the closest sandbox-side neighbor is forkd, which CoW-forks Firecracker guest RAM (56–150 ms per branch, live UFFD-WP branch in v0.4) but forks one sandbox VM, not a multi-node cluster, and runs a vendored Firecracker patch without the jailer. k7d's differentiator is the whole-cluster unit, the budgeted tree API over it, and the clock ownership. The sibling project Katakate k7 (810 stars, self-hosted sandbox infra with CLI/API/Python SDK) is the k8s-orchestration layer you pair with k7d when sandboxes — not clusters — are your unit.
Day-2 operational warnings
- Metrics to watch are response fields, not a Prometheus endpoint. k7d does not document a daemon metrics surface; the honest telemetry is in the replies — v0.8.0's
pause_ns/max_pause_nsper keyframe, andlatency_nson everyTreeFork(the same fields CI enforces). Log them per call, alert whenpause_nstrends toward your tolerance, and watch host-levelfreeplus KSMpages_shared/pages_sharingif you opt into KSM. CoW means forked clusters slowly materialize pages as they diverge — a 50-fork gym that fit at hour zero can pressure RAM at hour three.tree_auto_evictis the mitigation;tree_protectis how your winners survive it. - The daemon is a single point of failure by architecture. Native VMs are threads in the VM broker; a broker crash kills every native VM on the host, and the escape hatch (
K7D_VM_BROKER_RESTART) ships off. Treat every k7d tree as disposable lab state: the durable artifacts are your manifests, charts, and the protected winners — re-materialize from a refreshed keyframe, don't host anything on k7d you cannot lose. - Blast radius: never point this at production. Forks preserve in-cluster state but not TCP to the outside world — the far end never forked. Any forked workload holding external connections (databases, registries, cloud APIs) sees them break post-fork. This is a training-gym primitive, full stop.
- Net jail costs are real and measurable. The full net jail (egress+input chains, per-TAP anti-spoof) costs ~2 ms/cluster on the 50-fork batch because all member TAPs pin in one nft transaction — 3.85 s versus 2.44 s with
K7D_NET_JAIL=off. One nft process per TAP instead measured 12.7 s. If your fork numbers look 3× worse than the README, check whether your environment splits the nft transaction before you file a bug. - Single host, x86_64, KVM only. No arm64, no macOS/Windows, cross-node fork is roadmap-only. N-node clusters are supported (any
vm_count), but the limit is host RAM: the README's own math says ~3.2 GiB/node, so a 20-node base alone is ~64 GiB. Your gym topology is RAM-shaped; plan the box before the experiment matrix. - Adoption reality check. The repo went public Aug 11, 2026; the Show HN thread drew 3 points; the project runs on one primary test machine and 5 stars as of this writing. The engineering culture (CI-enforced latency claims, a 200-entry challenge ledger, formal methods on the eviction path, ≤30k auditable LOC) is several grades above its traction. Our read: study it now, prototype on disposable metal, and let the release cadence — currently weekly — settle before anything touches a workflow your team depends on.
Who should skip it
- Multi-tenant sandbox hosters: the native backend's shared-address-space CoW is the wrong isolation model for hostile guests; you want the k7d-fc profile at best, plain Firecracker or Kata at worst.
- Anyone without bare-metal KVM: nested virtualization murders the latency story, and there is no arm64 build. If your lab is cloud VMs, the ~105 ms number is not yours.
- GPU-heavy RL: the GPU fixture is explicitly a mock — DCGM-shaped series,
NVIDIA_VISIBLE_DEVICESset tofake-gpu-NIDs, Kueue flavors namedfake-h100/fake-l40s. Not CUDA, not NVML, and the README says so in bold. Real GPU fleet work still lives where our GPU capacity planning guide and vLLM review put it: on real cards with real DCGM. - Anyone needing cross-node or distributed forks today: host-local trees only. If your scenarios span racks, wait for the roadmap.
The verdict
k7d is the first project we've audited where the README's most quotable number is a CI assertion rather than a benchmark paste — and where the same project volunteers the conditions under which its own budgets are loose (suite-mode swings, deliberately-generous ceilings, a 20% prose-vs-tests drift it doesn't bother to reconcile). The design is genuinely novel where it counts: dirty-bitmap union with device-written pages, per-fork L2 replay, clock ownership as a first-class primitive, and now keyframes that make the fork pause proportional to residual dirt instead of RAM. The caveats are equally structural: one host, one address space, one test machine, no audit, an API that moves weekly, and a security trade that is only acceptable because the intended tenant is your own RL loop. That is not a "skip" — that is a "watch closely, prototype on €40 of metal, and steal the ideas even if you never run the daemon." The reset-vs-fork question it raises is the right question for every team training agents against infrastructure, and its answer — byte-identical worlds as a budgeted tree operation — is where this whole niche is going. For the runtime layer your agents actually serve from, see how containerd 2.4 landed and our AI fleet architecture patterns; for the training side, this space — k7d, forkd, and whatever forks them both — is worth a quarterly re-read.
References & further reading
- Katakate k7d — repository and README — the primary source: fork mechanics, two backends, time warp, security model
- k7d LATENCY_BUDGETS.md — the CI-enforced latency table and its own honesty flags (crates/k7d/tests/test_latency.rs imports the constants)
- k7d v0.8.0 release notes — coordinated cluster keyframes, dilation policy, opt-in KSM, pause_ns reporting
- k7d CHALLENGES.md — ~200 recorded implementation footguns; the three we cite are #40, #43
- k7d_client.py — reference JSON-lines client — the tree API surface used in our GRPO loop
- docs.katakate.org — project documentation site
- DeepSeekMath (arXiv 2402.03300) — the paper that introduced GRPO and the group-relative baseline
- forkd — nearest sandbox-side neighbor: Firecracker CoW branch (56–150 ms), UFFD-WP live branch
- Firecracker — the jailed-VMM baseline k7d-fc wraps and native k7d trades away
- Show HN: K7d (Aug 18, 2026) — the author's launch thread with the design rationale in the comments