Milestones
List view
## TL;DR Track the three GitHub changes implementing POL-706's infra-local deployment decision. Linear: https://linear.app/polygala/issue/POL-706/refine-infra-local-deploy
No due date•0/4 issues closed## TL;DR Keep MITM secret substitution working across expiry and restarts, with safe guest trust replacement.
No due date•3/3 issues closed## TL;DR Persist public tunnel declarations and enforce them at both proxy ingress paths. Parent issue: #1685. API slice: #1686. Proxy slice: #1687.
No due date•3/3 issues closed**Priority: P1.** New capability, built after P0; the M0 design leaves room for it. Replace gvproxy with our own network backend in Rust, behind the virtio-net device we already write in M2. Today the backend is a Go c-archive (`gvisor-tap-vsock` plus gVisor's netstack) linked into the shim. **Why** - Seccomp is disabled for every network-enabled box today, because the Go runtime runs inside the shim process. - Removes Go from the shim, and with it the c-archive linking workarounds. - One packet path we own: no unixgram hop, no VFKIT framing, less per-box overhead. **Tasks, staged** - Packet path: terminate the guest's frames in the VMM; DHCP and DNS for the box; outbound TCP and UDP through host sockets (translating flows avoids writing a full TCP stack); port forwarding with the expose and unexpose API - Allow-list filtering and the DNS sinkhole - TLS MITM: SNI peek, secret substitution, WebSocket support, using the existing per-box CA - Traffic shaping: the per-box bandwidth cap - Stats, DHCP leases and the forwards listing that the runtime reads today - Fuzzing for anything parsing guest or network input; throughput and latency at or above the gvproxy baseline - Re-enable per-thread seccomp for network-enabled boxes - Remove the Go bridge and `libgvproxy-sys` **Done when** the network suites pass on the new backend on all three hosts, seccomp is on for network-enabled boxes, throughput and latency are at or above the gvproxy baseline, and no Go code is left in the shim. **Blocked by** VMM M2 **Cheaper alternative to consider first:** run gvproxy as a child process instead of in-process. That restores seccomp for networked boxes without a rewrite.
No due date**Priority: P2 (lowest).** Windows boxes already run through WSL2. Run boxes on Windows without WSL (#69). **Tasks** - Windows Hypervisor Platform backend for x86_64 - Shim process model, isolation and logging on Windows - Host transports on Windows: vsock bridges and the gvproxy connection - virtio-fs passthrough on NTFS - Windows builds of the SDKs and CLI, and CI on Windows hosts **Done when** a box runs on Windows x86_64 without WSL, and the core integration suites pass there. **Blocked by** VMM M5
No due date**Priority: P1.** New capability, built after P0; the M0 design leaves room for it. Give boxes GPU acceleration. **Tasks** - Decide the approach per platform, and the API: - macOS arm64: virtio-gpu with Vulkan (Venus), rendered on the host through MoltenVK - Linux: whole-GPU passthrough with VFIO - Linux: PCI transport, VFIO passthrough and IOMMU setup, used only for passed-through devices - macOS: virtio-gpu device with the Venus backend and shared-memory mapping - Guest kernel config and guest user-space drivers - GPU assignment and isolation: one passed-through GPU per box on Linux - Integration tests: a GPU compute workload runs in a box **Done when** a GPU compute workload runs in a box on macOS arm64, and on Linux hosts with a supported GPU. **Blocked by** VMM M2
No due date**Priority: P1.** New capability, built after P0; the M0 design leaves room for it. Save a running box's full state (memory, vCPUs, devices) and restore it later or as a new box. **Tasks** - API design: naming, and how it relates to the disk snapshot API in #205; then Python, Node, Go and C SDKs, REST API and CLI - Pause the VM and save vCPU registers, interrupt-controller state and device state (Hypervisor.framework and KVM) - Save guest memory to a file; restore by mapping it copy-on-write, so restores and clones start fast - Restore in a new shim process: reconnect vsock and gvproxy, and resync the guest clock - Define what happens across a restore to open virtio-fs files and to network connections - Use it for an AutoPause/AutoResume that keeps memory (#1003) - Integration tests: snapshot, restore and clone, with running processes surviving **Done when** a running box can be snapshotted and restored with its processes intact on all three hosts, and a restore is faster than a cold boot. **Blocked by** VMM M4
No due date**Priority: P1.** New capability, built after P0; the M0 design leaves room for it. Change a running box's CPU, memory and disk without restarting it. **Tasks** - API design and naming survey, then Python, Node, Go and C SDKs, REST API and CLI - vCPUs: boot with the maximum vCPU count, and bring vCPUs online or offline in the guest - Memory: virtio-mem to add or remove memory within a maximum reserved at boot - Disk: grow the container disk, notify the guest of the new capacity, and grow ext4 online - Host limits on Linux: update the box cgroup's `cpu.max` and `memory.max` together with the guest size - Network bandwidth cap (#1435) changeable at runtime - Integration tests: grow and shrink under load **Done when** CPU count, memory and disk size of a running box can be changed through every SDK on all three hosts, and the guest sees the new values without a restart. **Blocked by** VMM M2
No due date**Priority: P1.** New capability, built after P0; the M0 design leaves room for it. Mount and unmount host directories in a running box without restarting it. **Tasks** - API design and naming survey, then Python, Node, Go and C SDKs, REST API and CLI - The virtio-fs server adds and removes exports while the VM runs; no new device per mount - The jailed shim gets access to the new host directory; for example, the runtime opens it and passes the descriptor to the shim - The guest agent mounts the export at the requested path in the running container, and unmounts it on detach - Read-only and read-write mounts - Integration tests: attach, use and detach a mount while commands run **Done when** a host directory can be attached to and detached from a running box on all three hosts, through every SDK, without a restart. **Blocked by** VMM M3
No due date**Priority: P0 (highest).** Boxes work on the new VMM exactly as they do today. Every box runs on our stack, and libkrun is gone. **Tasks** - Make the new VMM the default on Linux, with libkrun as a fallback for one release - Make the new VMM the default on macOS, with libkrun as a fallback for one release - Remove libkrun and libkrunfw: engine module, `libkrun-sys`, submodules, runtime bundling, dlopen workarounds, seccomp rules, docs **Done when** releases ship only the new VMM, our kernel build and `boxlite-guest` as init. That also retires the libkrunfw loading bugs such as #1509. **Blocked by** VMM M3, VMM M4
No due date**Priority: P0 (highest).** Boxes work on the new VMM exactly as they do today. The new VMM is ready for production. **Tasks** - Fault-injection tests for hypervisor errors - No panics reachable from guest-controlled input - Fuzz virtqueue descriptor chains and vsock packets in CI - Per-thread seccomp profiles for VMM threads - Record libkrun baselines: boot time, fs/block/net throughput, per-box memory - Performance at or above the libkrun baselines - Soak test on the e2e runner - Decide on nested virtualization (#1054) and custom kernels (#1041) **Done when** - Security suites pass. - Fuzzers run in CI. - Performance meets the libkrun baselines. - Soak runs are clean. **Blocked by** VMM M2. Can run alongside M3.
No due date**Priority: P0 (highest).** Boxes work on the new VMM exactly as they do today. User volumes work on the new VMM. **Tasks** - virtio-fs passthrough on Linux: ownership, modes, xattrs, special files, rename and unlink edge cases - virtio-fs passthrough on macOS: Linux file semantics on APFS - Read-only shares enforced on the host (#454), and single-file volumes - Any number of volumes without adding devices (fixes the #935 failure class) - In-guest filesystem conformance tests (pjdfstest subset) in CI - Fuzz FUSE request parsing in CI **Done when** copy_ownership, mount_security, secret_substitution, security_enforcement and network_spec pass on the new VMM on all three hosts, and a box with many directory volumes starts. **Blocked by** VMM M2
No due date**Priority: P0 (highest).** Boxes work on the new VMM exactly as they do today. A real box runs on the new VMM. **Tasks** - virtio-mmio transport and one split-virtqueue implementation, unit-tested without a hypervisor - virtio-blk: raw and qcow2, read-only, flush and discard - virtio-vsock: host Unix-socket bridges in both directions - virtio-net: unixgram to gvproxy with the VFKIT handshake - virtio-console log output and virtio-rng - virtio-balloon with free page reporting, which libkrun already runs - virtio-fs core for the SHARED share and file copy; one virtio-fs device serves all shares - `boxlite-guest` as PID 1 on the ext4 root - Guest clock resync after host sleep - Shim integration: engine registration, jailer profiles, per-thread seccomp on Linux - Make the engine selectable for tests; today it is hard-coded in `vmm_spawn.rs` - Tag each integration test with the devices it needs **Done when** every integration suite that needs no user volumes passes on the new VMM on all three hosts: lifecycle, shutdown, detach, recovery, exec_options, exec_user, run_main_command, resize_tty, zygote_integration, copy, snapshot, clone_export_import, network, gvproxy_backend. **Blocked by** VMM M1
No due date**Priority: P0 (highest).** Boxes work on the new VMM exactly as they do today. Our VMM boots our kernel on macOS arm64, Linux x86_64 and Linux arm64. **Tasks** - Hypervisor.framework backend (macOS arm64, macOS 15 or later): guest memory, vCPU run loop, exits returned as values, the in-kernel GICv3 - KVM backend (Linux x86_64 and arm64) - arm64 platform: kernel Image, FDT, GICv3, PL011 serial, PL031 RTC, PSCI - x86_64 platform: bzImage/ELF, boot params, CPUID, irqchip, 8250 serial, CMOS RTC - Guest kernel: pinned Linux LTS plus BoxLite config, reproducible CI build, shipped as a file - Lifecycle API: `run()` returns the guest exit status or a typed error; stop handle; nothing in the VMM exits the process - CI that boots VMs on macOS arm64, Linux x86_64 and Linux arm64 - CI smoke test that boots a test initramfs and checks serial output and exit code **Done when** CI boots our kernel into a test initramfs on all three hosts, and the exit status reaches the caller. **Blocked by** VMM M0
No due date•4/16 issues closed