k7d: Forking a Live Kubernetes Cluster in 105 ms for RL Rollouts — the Fine Print

Sources

Every platform team building AI-agent training infrastructure eventually hits the same wall: reinforcement learning on real infrastructure — Kubernetes, charts, admission controllers, CNIs — needs thousands of isolated, resettable, and above all identical worlds. The standard answer is a cold kind or k3s cluster per trial: about 30 seconds of boot time and a full RAM bill per copy. When your GRPO loop wants a group of 16 rollouts from the same starting state, that tax compounds into the dominant cost of the entire training run.

k7d — a Rust VMM and containerd shim from Katakate (docs at docs.katakate.org) — takes the opposite bet: boot the cluster once, then fork the live machine. The headline: a running 3-VM Kubernetes cluster forks in ~105 ms on the minimal CI guest, a Ubuntu k3s+Cilium+Tetragon cluster in ~1.1 s, and 50 three-VM cluster forks land in ~4.1 s on one 64 GB bare-metal box. Yesterday's v0.8.0 release added coordinated cluster keyframes that cut the source pause down to the residual dirty pages via live pre-copy — the most architecturally significant change since the project went public in August.

Methodology disclosure, because we owe you that: our analysis box has no /dev/kvm, so nothing in this piece is hands-on. What we did instead is audit the source of truth most VMM projects don't have: k7d wires every latency claim on its README into CI-enforced budgets (LATENCY_BUDGETS.md, constants imported from crates/k7d/src/latency.rs into crates/k7d/tests/test_latency.rs). Every number below is traced to a repo file. That is k7d's single most credible property — and, as we'll see, its own budget file is refreshingly honest about where the numbers are loose.

The structural shift: fork the universe, don't reset it

Group-relative policy optimization — GRPO, introduced in the DeepSeekMath paper — compares rewards within a group of rollouts from a shared prompt. The group-relative baseline is what makes it sample-efficient, and it is also what makes it brutally sensitive to environment state. If member A starts from a cold page cache, a different etcd revision, or a half-ready Deployment while member B starts warm, the reward gap between them is noise, not signal. The paper's authors controlled this with math on the model side; on the infrastructure side, nobody controls it — because a "reset" script that re-applies manifests never reproduces the same kernel timers, the same TLS session state, or the same in-flight reconcile loop.

k7d's answer is to make the copies byte-identical by construction: a fork is a copy-on-write clone of the live machine state — RAM, disk, processes, established TCP between member VMs, in-cluster TLS, kube-apiserver state, everything. From inside the fork, nothing moved. The README's phrasing is right: that is the difference between "we reset the env" and "we cloned the universe."

The second shift is economic. CoW forks share the parent's memory pages until someone writes: 50 cluster copies cost 50× the dirty pages, not 50× the guest RAM. That is what makes a 64 GB box a plausible RL gym. The daemon exposes the fork loop as a budgeted tree — you protect winners, prune losers, and let the daemon auto-evict under RAM and disk caps while the loop runs.

The third shift is the one nobody else does: k7d owns the guest clock. Because the VMM controls kvmclock, per-vCPU TSC offsets, and timer re-arms, the same pause that makes a fork correct can also move every member's clock by the same delta. An 8.7-minute guest-time episode that is mostly kubelet waiting on probes compresses 1.83× with lockstep clock jumps at zero fidelity cost — the parts of an RL episode that are pure waiting stop costing wall-clock time.

Architectural blueprint: what actually happens in 105 ms

The design is one daemon (k7d, one process, one address space, many VMs) owning a KVM VMM, a snapshot-tree manager, and the guest clock. Your trainer talks JSON-lines over a Unix socket at /run/k7d/k7d.sock. A fork, end to end, is a short critical section on the source:

sequenceDiagram
    participant T as Trainer (JSON-lines client)
    participant D as k7d daemon
    participant S as Source VMs (3-node cluster)
    participant C as Child VMs (forks)
    T->>D: tree_fork_batch(n)
    D->>S: pause every vCPU, drain in-flight virtio I/O
    D->>D: dirty bitmap = KVM dirty log ∪ device-written pages
    D->>C: map parent RAM copy-on-write, copy only dirty pages
    D->>C: new bridge per child, replay guest IPs / MACs
    D->>C: restore vCPUs, re-arm timers, KVM_SET_CLOCK (never backwards)
    D->>S: resume
    D-->>T: n children + latency_ns

Three implementation details decide whether any of this is real, and the project documents all three in its CHALLENGES.md (nearly 200 recorded footguns):

The tree the daemon manages underneath all of this is the part platform engineers should internalize — it is a resource model, not just a data structure:

                 [base cluster — root, 48 s bring-up]
                    |            |            |
              [fork A]      [fork B]      [fork C]
                    |       (pruned:       |       \
              [fork A1]      low reward)  [fork C1] [fork C']
            (protected:                 (evicted:   (tree_rollback
             winner)                      RAM budget)  from C)

RAM:  one mapping, parent pages shared until a fork writes
disk:  reflink keyframes (v0.8.0) — children CoW from the keyframe
budget: max_live_vms / max_live_ram_bytes / max_disk_bytes / max_chain_depth

The v0.8.0 change that matters: keyframes cut the pause to the residual

Fork latency has two components: the pause you impose on the source (all vCPUs frozen while dirty pages are read and child mappings are created) and the materialization of children. Until v0.7.x, the pause scaled with guest RAM and dirtiness. The v0.8.0 release notes describe the fix with unusual precision: coordinated cluster keyframes (tree_create_cluster, tree_adopt_cluster, and in-tree children) now "cut every native member at one instant. Live pre-copy writes RAM while the guest runs, so the pause is the residual rather than guest RAM."

That is live migration's iterative-copy trick applied to forking: copy most of RAM while the guest runs, then pause only long enough to sweep the residual delta. The reply now reports pause_ns and max_pause_ns per TreeCreated, and an adopted root keyframe can be refreshed into a new generation — meaning a long-lived gym can re-baseline its root without paying the 48 s Ubuntu bring-up again. The same release adds an opt-in per-tree dilation policy (fixed / down / track, with caller knobs for window, cadence, and ladder) and opt-in per-tree KSM (ksm: true, off by default, inherited by forks, same-tenant trees only — the daemon deliberately refuses to start ksmd for you).

Read the release cadence as a warning as much as a feature: eight minor releases in six weeks (v0.2.1 on Aug 17 → v0.8.0 on Sep 28). The wire API is moving. Pin versions in any gym you build.

The numbers, audited

Here is what we mean by "CI-enforced": budgets are set at roughly 2× the typical observed value on the reference bare-metal node (an AMD Ryzen 5 3600, 64 GB, NVMe — about €40/month of metal), and an integration test fails if the daemon's latency_ns drifts past budget. Selected rows from the budgets file, with the honesty flags it documents itself:

OperationEnforced budgetTypicalWhat the file admits
Warm fork, <25% dirty (1 VM)50 ms~5 msWarm 2-vCPU source: budget 75 ms, typical "TBD — start generous, tighten after measuring"
Cluster-tree fork, 3 VMs750 ms~75 ms quiet / 188–494 ms loaded"Intentionally generous" vs the ~75 ms design projection
Inner-k3s fork under load churn1 s~104 msThe README's headline ~105 ms is this churn-loaded row, not a quiet-host best case
50× 3-VM batch fork20 s~3.85 s (~77 ms/cluster)Full net jail costs ~2 ms/cluster because all member TAPs pin in ONE nft transaction; one nft process per TAP measured 12.7 s — 3× worse
Ubuntu VM boot → agent on vsock120 s~3.2 s"Deliberately loose… a regression that doubles Ubuntu boot time passes"

Two audit findings worth your attention. First, the file documents that under full-suite serial runs, test_cluster_fork_warm_3vm_under_50ms "swings up to ~500 ms" versus ~25 ms run alone — a 20× swing between isolated and suite-mode measurement. Your quiet-host numbers will not survive a loaded CI box; the project budgets for that, and so should you. Second, the README quotes VM boot→agent at ~163 ms while the budgets file's typical is ~194–205 ms. The gap is small but it is exactly the kind of drift you get when prose and tests live in different files — k7d at least has the tests; most VMM marketing has neither.

What a fork costs as your in-cluster surface grows — per-profile deltas on the 3-VM inner-k3s fixture, all from the budgets file:

ProfileTypical forkDelta vs lean baseline
Lean baseline (add-ons off, load churn)~104 ms—
+ CoreDNS~175 ms+71 ms
+ ServiceLB / Traefik ingress~206 ms+102 ms
+ local-path PVC workload~200 ms+96 ms
+ metrics-server~196 ms+92 ms
+ NetworkPolicy (default-deny + allow)~195 ms+91 ms

The pattern: every extra reconciling controller adds roughly 90–100 ms to the fork because more vCPUs are mid-tick when the pause lands, and more memory is dirty. A "realistic" cluster with all of these is a ~2× fork versus the lean fixture — still well under the 1 s budget, but a long way from the marketing number.

Time warp, read skeptically

The clock tricks are the part of k7d with no equivalent anywhere else, and also the part where the README's own table rewards careful reading. The 8.7-minute guest-time episode, 3-node k3s:

MechanismWall timeFidelity cost
Stock clock522 s—
Lockstep jump (auto-warp, no kernel patch): when every vCPU is idle, pause ~0.2 ms at ~20 Hz, add δ to kvmclock + every TSC offset286 s (1.83×)Zero — same probe failures, restarts, leases as stock
Continuous dilation ×N (~200-line carried KVM patch): guest clocks run N× wall79 sReal and bounded: guest-visible fsync 2.3 → 14 ms, pod RTT 0.69 → 0.92 ms, a guest-paced churn loop does 37 rounds instead of 86
Both combined76 sDilation consumes the idle windows that jumps need — they don't stack

Now the skeptical footnote the table invites: the dilation row is labeled "(8.00×) at N=8," but 522 s stock → 79 s is a 6.6× wall speedup. The "(8.00×)" is the dilation factor the guest clock runs at, not the speedup your episode gets — because "only waiting is compressible." Work — CPU, I/O with real device latency — does not dilate; an episode that is ~97% waiting is what yields 6.6× at N=8. Read every vendor ×-label as a factor, never as a speedup, and the README's own fidelity column (fsync 6× slower guest-visible) tells you exactly what you trade: any RL reward signal that correlates with real I/O latency is now measured against a distorted clock. The zero-cost option is the lockstep jump, and the daemon exposes it as auto_warp(cluster_id, mode) with off / jump / micro — a declined warp is always correct, which is the right default posture.

Runnable: a GRPO group loop against the real API

The reference client is a single-file, no-dependency JSON-lines client (examples/cluster-tree-search/k7d_client.py) — the project explicitly says to treat it as a template for wiring your trainer, not an SDK. Setup first: you need a Linux amd64 host with /dev/kvm; the quickstart doctor-checks /dev/kvm, /dev/vhost-vsock, /dev/net/tun, and cgroup v2. No arm64 build exists.

git clone https://github.com/Katakate/k7d && cd k7d
sudo ./scripts/quickstart.sh   # boots a busybox 3-node tree, forks 4 branches, prints wall-clocks

k7d doctor                     # host checks; does not start a daemon
k7d quickstart                 # 3-node busybox + 4 forks (daemon must be up)

The loop below uses only verbs and fields that exist in the reference client and demo, with the demo's own budget values (64 GiB RAM / 256 GiB disk, headroom for root + branches + a rollback fork). This is the shape of a GRPO group step:

from k7d_client import K7dClient          # the reference client; K7D_SOCKET env overrides the path

client = K7dClient()                      # default: /run/k7d/k7d.sock, 600 s timeout
TREE = "grpo-run-01"
GROUP = 16                                # rollouts per generation (the G in GRPO)

# 1) Root the tree at a warm checkpoint. Budgets are enforced by the daemon,
#    not by your discipline: auto_evict prunes unprotected branches under pressure.
budget = {
    "max_live_vms": 3 * (GROUP + 2) + 16,             # demo formula: root + branches + rollback headroom
    "max_live_ram_bytes": 64 * 1024**3,
    "max_disk_bytes": 256 * 1024**3,
    "max_chain_depth": 32,
}
created = client.tree_create_cluster(
    3,                                                  # vm_count — N-node, no hard-coded 3
    {"kernel": KERNEL, "initrd": INITRD,
     "memory_mb": 256, "vcpus": 1, "vsock": True},
    tree_id=TREE, budget=budget,
)
root = created["root_id"]

# 2) Fork the whole group from the root — byte-identical starts by construction.
#    dilation=8 is opt-in (v0.8.0 policy); drop it for fidelity-sensitive rewards.
batch = client.tree_fork_batch(TREE, root, [f"gen0-{i}" for i in range(GROUP)], dilation=8)
forks = batch["forks"]
for f in forks:
    assert not any(f.get("fork_memory_full_copy") or [])   # CoW broke if this fires — stop the run

# 3) Drive each rollout over vsock by CID (the daemon does NOT proxy your exec):
#    keep f["guest_cids"] from the fork reply, or re-read tree_nodes -> live_guest_cids.
#    Score with YOUR reward model; k7d owns environments and budgets only.
scores = sorted(((reward(f), f) for f in forks), key=lambda x: x[0], reverse=True)
winner = scores[0][1]

# 4) Keep the winner, prune the losers — budgets make the tree safe to leave running.
client.tree_protect(TREE, winner["node_id"])               # budget pressure can no longer evict it
for _, f in scores[1:]:
    client.tree_prune(TREE, f["node_id"])

# 5) Roll forward (fork from the winner) or roll back — rollback is non-destructive:
#    it is another fork from an earlier node, the earlier node survives.
nxt = client.tree_fork_batch(TREE, winner["node_id"], [f"gen1-{i}" for i in range(GROUP)])
client.tree_auto_evict(TREE)                               # enforce RAM/disk caps now

# 6) Teardown: stops guests, releases broker pins — caller-supplied disks stay YOURS to unlink.
client.tree_drop(TREE)

Fields worth knowing on the wire, from the demo and release notes:

{
  "tree_id": "grpo-run-01",
  "forks": [
    {
      "node_id": "n-a1b2",
      "guest_cids": [4, 5, 6],
      "fork_memory_full_copy": [false, false, false]
    }
  ]
}

Three API-contract footguns, all documented but all easy to trip on the first week:

Two backends, two isolation stories — pick with your eyes open

One daemon, two VMMs: backend: native (k7d's in-process rust-vmm) or backend: firecracker (stock Firecracker plus its stock jailer). A tree is one backend, never mixed. The trade is blunt and the README states it plainly:

native (k7d)firecracker (k7d-fc)
Warm fork of a running VM~5 ms~76–140 ms per child (jailer spawn each)
Per-VM jail (own uid, chroot, seccomp)No — CoW fork requires parent and child memory in one address spaceYes — stock jailer
Guest time warp / dilationYesNo — Create fails loudly, nothing silently drops
virtiofs (hostPath)YesNo
Built forRL gyms, fleets you controlHostile / multi-tenant guests

The security consequence of native is the line every evaluator should underline: because live CoW fork requires sibling memory mappings in one process, isolation between sibling forks of the same tenant is weaker than a jailed VMM, and the daemon is by design shared across branches of one tree. For RL training that is an acceptable trade — the branches are all yours. For anything multi-tenant, use the Firecracker backend, where a guest→VMM bug dies in the jail. The daemon also ships privilege separation (K7D_PRIVSEP=on: a root node broker with no listening socket plus a uid-k7d, empty-cap VM broker; k7d doctor reports it), and selected critical paths carry machine-checked proofs — Kani on unsafe memory arithmetic, Aeneas→Lean on the tree budget/eviction model (eviction never frees a page a live descendant references). Selected paths, not the whole runtime — and no external security audit exists yet.

Context from the neighborhood, since you'll be asked "isn't this just X": Kata Containers and Firecracker can snapshot-and-restore, but not live-fork a running VM; E2B-style sandbox fan-out tools restore N copies of one sandbox (~220 ms class); the closest sandbox-side neighbor is forkd, which CoW-forks Firecracker guest RAM (56–150 ms per branch, live UFFD-WP branch in v0.4) but forks one sandbox VM, not a multi-node cluster, and runs a vendored Firecracker patch without the jailer. k7d's differentiator is the whole-cluster unit, the budgeted tree API over it, and the clock ownership. The sibling project Katakate k7 (810 stars, self-hosted sandbox infra with CLI/API/Python SDK) is the k8s-orchestration layer you pair with k7d when sandboxes — not clusters — are your unit.

Day-2 operational warnings

Who should skip it

The verdict

k7d is the first project we've audited where the README's most quotable number is a CI assertion rather than a benchmark paste — and where the same project volunteers the conditions under which its own budgets are loose (suite-mode swings, deliberately-generous ceilings, a 20% prose-vs-tests drift it doesn't bother to reconcile). The design is genuinely novel where it counts: dirty-bitmap union with device-written pages, per-fork L2 replay, clock ownership as a first-class primitive, and now keyframes that make the fork pause proportional to residual dirt instead of RAM. The caveats are equally structural: one host, one address space, one test machine, no audit, an API that moves weekly, and a security trade that is only acceptable because the intended tenant is your own RL loop. That is not a "skip" — that is a "watch closely, prototype on €40 of metal, and steal the ideas even if you never run the daemon." The reset-vs-fork question it raises is the right question for every team training agents against infrastructure, and its answer — byte-identical worlds as a budgeted tree operation — is where this whole niche is going. For the runtime layer your agents actually serve from, see how containerd 2.4 landed and our AI fleet architecture patterns; for the training side, this space — k7d, forkd, and whatever forks them both — is worth a quarterly re-read.

References & further reading