Microsoft MXC Review: OS-Native Sandboxing for Agentic Tool Calls, Tested on Ubuntu

Sources

Microsoft shipped the first stable release of MXC (Microsoft eXecution Containers) on October 7, 2026 — an SDK-first, MIT-licensed sandboxing layer for running untrusted code: model output, plugins, agentic tool calls. It puts policy-driven containment behind typed APIs in Rust, .NET, and Node, with OS-native backends — Bubblewrap on Linux, Seatbelt on macOS, AppContainer/Windows Sandbox/WSL containers on Windows, and experimental Hyperlight micro-VMs. The repo crossed 1,900 stars within months of going public in February 2026. We pulled the actual @microsoft/mxc-sdk 1.0.0 tarball from the npm registry, ran its native executor on Ubuntu 24.04, and tried to break every promise in the policy model — including the one where the vendor's own README says your first week will be access-denied failures. The same V1 surface ships on NuGet as Microsoft.Mxc.Sdk and crates.io as mxc-sdk, and the release follows semver 2.0 for the 1.x line, per the release notes.

The short version: the policy engine is unusually rigorous — refusals fail closed with precise, actionable errors, the network model is directional deny-by-default with an honest, documented v6 blind spot, and the schema rejects legacy fields instead of silently accepting them. The engineering visible in the binary's diagnostics (enforcement probes for every declared policy bit, capability-bit enumeration tests, unprivileged iptables in the sandbox's own namespace) is well above the average for agent-infrastructure releases. But our hands-on battery on a stock Ubuntu 24.04 host hit the central production trap: MXC's availability probe reports the bubblewrap backend as available while every single launch dies with bwrap: loopback: Failed RTM_NEWADDR: Operation not permitted, because Ubuntu 24.04 restricts unprivileged user namespaces via AppArmor. The executor has a precise remediation string for exactly this condition sitting in its binary — the ruleless default path never prints it. If your fleet is Ubuntu 23.10+, budget for sysctl/AppArmor remediation before you write a line of integration code.

Executive Scorecard

DimensionScoreWhy
Reliability6/10Deterministic refusals, structured errors, clean exit codes. But the availability probe's isSupported:true does not mean launchable — on AppArmor-restricted hosts every run fails at bwrap's first netlink call.
DX7/10Typed SDKs, JSON schema with exact contract versions, dry-run validation, capture-denials tooling. The native Node FFI chain (koffi) needs cmake when prebuilds miss; executor binary named lxc-exec is a naming landmine.
Cost9/10MIT license, no telemetry outside Windows, no commercial tier dependency. The entire stack — bubblewrap, slirp4netns, iptables — is free and unprivileged.
Security8/10Deny-by-default filesystem, directional network policy enforced in-namespace, capability bits probe-tested, secrets scrubbed from env. The AppArmor gap is a host config issue, not a design flaw — but it fails with a wrong diagnostic on the default path.
Linux maturity (self-tested)5/10v1.0.0 is two days old; the schema registers a 1.1.0-alpha contract already. Zero contained executions succeeded on our stock Ubuntu host without root. Windows backends and lifecycle APIs remain unverified here.

Verdict up front: platform teams building agentic tool-execution infrastructure on Linux should treat MXC as the strongest open-source policy layer currently shipping — and the AppArmor finding below as a mandatory pre-adoption checklist item, not a blocker. Windows-first agent platforms (the repo's own center of gravity: AppContainer, Windows Sandbox, WSLC) get a richer backend set today. Skip MXC entirely if you need containers-as-a-service, multi-tenant SaaS sandboxing with billing, or a network model that can express hostname-level egress rules — it cannot, by design.

Architecture Mechanics

MXC is not a runtime you install; it is an SDK dependency that builds into your application. Your code constructs a ContainerRequest (command + filesystem policy + network policy + containment backend), the SDK lowers it to the backend's native machinery, and the workload runs contained. The Linux default backend is Bubblewrap — the same unprivileged user-namespace sandbox Flatpak uses — which means no daemon, no root, no container runtime:

your app (Rust / .NET / Node)
    │  ContainerRequest{ command, filesystem, network, containment }
    ▼
MXC SDK (in-process) ── validates against exact schema contract
    │                    (registered: 0.9.0-alpha, 1.0.0, 1.1.0-alpha)
    ▼
lxc-exec native executor   ← yes, "lxc" — it runs bubblewrap
    │
    ├─ ruleless egress deny ──► bwrap --unshare-net --clearenv
    │                              (private netns, lo only)
    ├─ egress rules / proxy ──► supervisor userns + slirp4netns
    │                              + nsenter iptables-restore
    │                              (MXC_EGRESS / MXC_INGRESS chains)
    ▼
filesystem baseline: --ro-bind-try /bin /sbin /lib* /usr/* /etc
                     + DNS stub dirs under /run; everything else invisible
    ▼
sh -c "your command" ── dies with the workload; no lifecycle to manage

Three mechanics matter for platform teams:

Deny-by-default filesystem. The sandbox sees system binaries, libraries, /etc, and DNS resolver directories — nothing else. $HOME (credentials, SSH keys, browser cookies), /opt, /var, /run/user/<uid> are invisible until you bind them with readonlyPaths / readwritePaths. Denied paths are masked: directories with an empty tmpfs, files with /dev/null — classified by on-disk type via symlink_metadata, and symlinked paths are canonicalized before masking so the mask lands on the real object. The docs are explicit that root-owned secrets in /etc (/etc/shadow, ssh_host_*) remain unreadable because user-namespace UID mapping does not bypass kernel DAC.

Directional network policy, enforced unprivileged. network.egress / network.ingress sections default to deny. Ruleless deny isolates the netns entirely; rules put the sandbox behind slirp4netns with an unprivileged supervisor holding CAP_NET_ADMIN inside a user namespace — the sandbox drops the capability before the workload starts, so it cannot flush its own chain. DNS names in rules are rejected at validation (IP literals/CIDRs only — the sandbox resolves names itself, and a lookup that disagreed with the rule's resolution would bypass the chain). The except field is lowered by CIDR subtraction, not accept-then-hope. An IPv6 allow is warned as ineffective (slirp runs without --enable-ipv6, so the namespace has no v6 route — tracked in issue #955, closed), and an IPv6 deny block shorter than /96 that contains the IPv4-mapped range is rejected outright rather than programmed as a fail-open rule.

Environment hygiene. --clearenv is absolute: host env never leaks. From schema 0.9 the child gets a default block (PATH, TERM, HOME only when process.cwd resolves), and the docs call out the trap that HOME equals the working directory — dotfiles inside it are read as user-level tool config, so point HOME elsewhere for untrusted workspaces.

Hands-On: The Refusal Matrix Works

We ran 20+ configurations through the real executor (lxc-exec from the 1.0.0 npm tarball) against a JSON config battery. The validation layer refused every dangerous or unsatisfiable combination exactly as the backend guide documents, with error text precise enough to paste into a runbook. Selected results, verbatim from our terminal:

# ingress "allow" — cannot be delivered, refused:
$ lxc-exec --config t8.json
Bubblewrap: network.ingress.default='allow' is not supported. The sandbox
runs in a private network namespace reached through slirp, which has no
route in until a host port is forwarded to it, and the schema carries no
port list to forward. A listener inside the sandbox is reachable from the
sandbox only. Use network.ingress.default='deny'.

# hostname in a cidr field — refused at parse, not resolved:
Configuration parse error: network.egress.allow[0].to[0].cidr must be a
valid network CIDR: couldn't parse address in network: invalid IP address syntax

# egress rules + proxy — combination would discard the rules:
Configuration parse error: runtimeConfig.networkProxy requires
network.egress.default='deny' with no direct allow or deny rules

# legacy field — exact-contract rejection, no silent acceptance:
Configuration parse error: Invalid configuration at `network.defaultPolicy`:
unknown field `defaultPolicy`, expected `egress` or `ingress` at line 5
column 30; schema 0.9 migration: use network.egress.default

# old contract version — refused, with the registered list:
Configuration parse error: Unsupported contract version. Registered
versions are: 0.9.0-alpha, 1.0.0, 1.1.0-alpha

# IPv6 mapped-range trap — refused rather than fail-open:
Bubblewrap: network rule '::/0' is an IPv6 block shorter than /96 that
contains the IPv4-mapped range '::ffff:0:0/96'. Addresses in that range
travel as IPv4 packets, so an ip6tables rule would not match them and the
rule would be silently unenforced. Write the IPv4 side explicitly instead.

Two structural observations from the battery. First, the error taxonomy is two-tier and consistent: configuration problems exit 1 with Configuration parse error:; backend refusals exit 255 with a JSON error envelope ({"error":{"code":"backend_error",...}}). That separation is genuinely useful for CI gating. Second, the registered-contract list inside the shipping 1.0.0 binary already includes 1.1.0-alpha — the Hyperlight micro-VM backend we could not launch (no /dev/kvm on the test host; the executor refused with error: Hyperlight requires KVM: /dev/kvm must be readable and writable by this user). The hyperlight containment value is only reachable under the 1.1.0-alpha contract plus --experimental; under 1.0.0 the enum rejects it. Expect the stable schema to move again soon after this review ships.

The Node SDK's capture layer is honest about failure: our run() call returned {stdout:"", stderr:"bwrap: loopback: ...", exitCode:1, timedOut:false, warnings:[]} — the diagnostic is captured in the result's stderr field even though the SDK does not classify it. The raw bwrap error never reaches your logs unstructured via the SDK path; you just have to grep your own results object.

Critical Failure Modes

1. The availability probe lies on Ubuntu 24.04 (the headline finding)

This is the failure that will bite every Linux-adopting team. MXC's documented prerequisite check is cat /proc/sys/kernel/unprivileged_userns_clone — ours printed 1, the passing value. The CLI probe (lxc-exec --available-backends) reported bubblewrap available, with a warning only about the missing slirp4netns (which proxy mode needs; the ruleless path we tested does not). The Node SDK's getPlatformSupport() returned {isSupported: true, availableMethods: ["bubblewrap"]}. Then every launch — all 20+ configs, every containment value including the generic process default — died in 8 ms with:

bwrap: loopback: Failed RTM_NEWADDR: Operation not permitted

The cause: Ubuntu 23.10+ ships kernel.apparmor_restrict_unprivileged_userns=1 (Ubuntu's own announcement covers the change), which blocks the uid_map write inside user namespaces. Our host: apparmor_restrict_unprivileged_userns = 1, max_user_namespaces = 15309, Ubuntu 24.04.5. Plain bwrap --unshare-net outside MXC fails identically — this is a host restriction, not an MXC bug. What is an MXC gap: the engine contains a precise remediation string for exactly this condition (visible in the binary: "On Ubuntu 24.04+ this is usually AppArmor's 'kernel.apparmor_restrict_unprivileged_userns=1': give the caller an AppArmor profile that permits userns, or clear that sysctl on an ephemeral host...") — but it is wired to the proxy-mode user-namespace path, not the ruleless default. The default path surfaces raw bwrap output. A team that trusts getPlatformSupport() as a gate will ship integrations that pass pre-flight and fail every execution. Treat probe results as advisory until you have run one real workload on every Linux fleet variant.

2. Node SDK dependency chain requires a build toolchain on the edge

The Node SDK binds to the native engine via koffi (FFI). koffi ships prebuilt platform binaries as optional dependencies (@koromix/koffi-linux-x64 et al.), and with those present the SDK loads cleanly — our manual tarball install verified the full chain: koffi 3.2.1 + @koromix/koffi-linux-x64 3.2.1 + Node 24.21. But in environments where optional dependencies are stripped (offline mirrors, strict internal registries, air-gapped CI), koffi falls back to building from source and demands cmake — a hard stop on locked-down hosts. The Rust path avoids this entirely (static build into your binary), which is the safer embedding choice for appliance-style deployments.

3. No resource limits on the Linux one-shot path

The 1.0.0 schema exposes cpuCount and memoryMb fields — only for the WSLC (Windows) backend. The bubblewrap/LXC/ProcessContainer one-shot paths have no CPU, memory, or cgroup controls in the stable contract. An agent-generated while true loop in a contained process competes for host resources constrained only by your node-level limits. Platform teams must wrap MXC workloads in their own cgroup/scheduler policy (systemd slices, K8s pod resources if wrapped in a container). We could not verify timeout-kill behavior end-to-end on this host (the sandbox dies pre-workload), so we make no claim about timeoutMs enforcement beyond the field's existence.

4. Hostname-level egress is delegated, not expressed

Rules accept IP literals and CIDRs only. Hostname-based egress control — the thing most agent-platform teams actually want ("allow the agent to reach api.github.com only") — belongs to an external proxy via runtimeConfig.networkProxy, and MXC is explicit that it does not forward host lists to that proxy. The design is coherent (DNS answers shift; a stale resolution would bypass the chain), but the operational consequence is real: hostname policy means standing up and operating your own proxy fleet, and the bundled proxy is a testing-only deliberately-permissive server behind --allow-testing-features — never production.

5. Stateful lifecycle is backend-conditional

The provision/start/exec/stop/deprovision lifecycle is where MXC differs from one-shot bwrap: the lifecycle contract supports it, but bubblewrap implements one-shot only (its docs say it is a ScriptRunner, not a StatefulSandboxBackend). Linux persistent containers today mean LXC (root) — the unprivileged default has no stateful mode. Windows gets IsolationSession and WSLC for stateful operation. If your agentic pattern is "provision a warm sandbox, exec many tool calls, deprovision" on Linux without root, MXC 1.0.0 does not serve it yet.

The Node SDK, for Reference

The versioned V1 surface is a clean break from the root package: import from @microsoft/mxc-sdk/v1, and the root exports nothing. Our probe script against the real SDK:

import { getPlatformSupport, getAvailableBackends, run, spawn } from '@microsoft/mxc-sdk/v1';

// Probe (memoized for the module lifetime — restart after host remediation):
getPlatformSupport();
// → { isSupported: true, availableMethods: ["bubblewrap"],
//     bubblewrapNetwork: { proxyEnforcement: "unsupported", warnings: [...] } }

// Captured one-shot:
const out = await run({
  command: 'node -e "console.log(\'hi\')"',
  network: { egress: { default: 'deny' } },
  timeoutMs: 30_000,
});
// → { stdout, stderr, exitCode, timedOut, warnings }

// Streaming with SDK-owned pipes (untaken streams drained internally):
const p = await spawn({ command: 'uname -a', timeoutMs: 10_000 });
p.standardOutput?.on('data', c => process.stdout.write(c));
await p.wait(); p.dispose();

One more DX wart worth naming: the Linux executor binary is lxc-exec — named after the LXC container system it optionally drives, while the default Linux path runs bubblewrap. In incident triage, an executor named after the wrong technology is a real source of misdiagnosis ("why is my bubblewrap sandbox running under LXC?" — it is not). The Windows binary is wxc-exec.exe. Naming asymmetry between platforms plus an engine that already registers an experimental contract version inside a stable release are small things; they are also the kind of thing you notice when a project is two days into its stable life.

Who Should Skip This

Who it is for: agent-platform and AI-infra teams shipping products that execute model-written code on user or customer hosts (the Windows backends — AppContainer, Windows Sandbox, WSLC — are the repo's center of gravity and its most mature surface); platform teams building tool-execution layers who want a policy engine that fails closed and refuses unsatisfiable configs; and anyone who needs the same containment model expressed across three SDKs over OS-native isolation rather than another containerd dependency.

Verdict

MXC is the rare agent-infrastructure release where the internals are more impressive than the announcement. The policy engine's discipline — every declared capability bit probe-tested against rendered iptables output, unsatisfiable postures refused instead of half-enforced, legacy config rejected with migration guidance, the IPv4-mapped-range trap caught at validation — reflects a security team that has shipped sandbox escapes before. The bubblewrap backend documentation alone is among the most honest isolation docs we have read, including a paragraph admitting its own ingress chain is defense-in-depth for a path that cannot exist yet. That is the culture you want under agentic execution.

The gap between that engineering and operational readiness is the availability story. A stock Ubuntu 24.04 host — the single most common Linux for platform teams — passes every documented prerequisite and fails every execution, with the precise remediation string present in the binary but unreachable from the default path. That is a v1.0.0 sin: forgivable, and precisely the kind of thing this review exists to catch before it bites your on-call. The remediation is known and small (an AppArmor profile or a sysctl), but until the probe-then-fail gap closes, run one real workload per fleet variant before trusting any MXC pre-flight. For Windows-centric agent platforms the calculus is simpler: this is the most credible open-source containment layer for agentic tool calls shipping today, and the v1 SDK surface is the right integration point to bet on.

References & Further Reading