Two defects found while getting a GPU through Proxmox to a nested guest.
Passthrough gave every address its own pcie-root-port, which splits a GPU from
its own HDMI audio: 05:00.0 and 05:00.1 arrived in the guest as two devices on
two buses instead of functions 0 and 1 of one device. Navi needs both halves on
one device to reset or power-manage either, so the guest got a card stuck in D3
and a reset that could not be performed. Addresses are now grouped by
everything left of the function digit, and each group goes behind one root port
at one slot with multifunction=on on function 0 -- which is also where the
VBIOS and the VGA route belong.
The data-disk setup used Initialize-Disk/New-Partition/Format-Volume. Only the
first of those works that early in specialize; the rest need services that are
not up yet, and with ErrorActionPreference=Stop the script gave up straight
after writing a GPT header. The result was a 50G disk carrying 24KB of nothing
and a ProfilesDirectory pointing at a volume that never existed. diskpart works
at that stage. It also tries assigning the letter before laying the disk out,
so a disk that already holds a profile is lettered rather than cleaned.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0117qMyjpuXsjpVAcpJbFD8g
Image builds have been dying with "Argument list too long" from sed, mktemp
and timeout alike -- commands whose argv is trivial, which is the tell that it
was the environment that had grown, not the arguments.
The X11 forwarding hint is looked up with `ls -t /tmp/.vmix-display-* | head -1`.
Under nullglob a non-matching pattern is removed from the command line rather
than passed through literally, so `ls -t` runs with no arguments at all and
lists the working directory instead. In a nix build that directory is the
build tree, whose newest file is nix's own env-vars dump. The result is that
VMIX_DF becomes "env-vars", the SDL branch is taken on a machine with no X at
all, and DISPLAY is exported with a slice of the env dump inside it. From that
line onward every exec in the build fails with E2BIG.
find does the same lookup without depending on how the shell treats an
unmatched pattern, and -type f keeps a stray directory out of it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0117qMyjpuXsjpVAcpJbFD8g
A disk-mode overlay defaults to C:\uwfswap.sys -- on the very volume being
protected, which is the opposite of the point. uwfmgr grew a create-swapfile
subcommand for exactly this, and it only accepts the call while the filter is
off and the overlay is already in disk mode, so the ordering in the generated
script is forced rather than stylistic.
Configuration is deferred to the target's first boot through RunOnce instead
of running in the build. Two reasons: uwfmgr does not exist until the DISM
feature has been through a reboot, and the swapfile belongs on the real data
volume rather than on the throwaway copy the build attaches. Enabling the
filter needs one more restart after that, which the script asks for itself.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0117qMyjpuXsjpVAcpJbFD8g
Three things, all in service of putting Proxmox in a VM that can still hand a
GPU to its own guests, and of a Windows VM whose profile survives its OS disk.
pci.viommu.enable emits `-device intel-iommu,intremap=on,caching-mode=on` and
forces kernel-irqchip=split, which interrupt remapping requires. Without an
IOMMU of its own a guest cannot bind a passed-through device to vfio-pci, so
it can never forward one on. The device leads the command line because QEMU
realizes devices in order and intel-iommu must precede what it translates.
pci.vgaPassthrough (default true, so nothing changes for existing VMs) makes
x-vga=on optional. It was forced on the first passthrough device, which is
wrong for a card the guest only forwards onward: it claims the VGA path the
emulated console adapter needs.
customizeImage gains extraDisk, a blank disk attached for the Audit Mode boot
and emitted as the derivation's `data` output. generalize uses it for dataDisk
and profilesDirectory, so the disk is partitioned and the profile relocated
under OOBE in the build VM. That is what removes the need for delayOobeRun --
previously the volume ProfilesDirectory names could not exist until the image
reached real hardware. A second output rather than a directory keeps ${image}
meaning the OS qcow2 for every existing consumer.
generalize also picks up staticIP and profilesDirectory.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0117qMyjpuXsjpVAcpJbFD8g
After= was placed in [Service], which systemd ignores ('Unknown key
After'), so the interfaces.d merge + ifreload raced networking.service
on every boot. On losing boots vmbr0 never ran DHCP and the proxmox
guest came up without its LAN IP (bridging still worked, so inner VMs
stayed reachable while the PVE host itself was not).
Move After= to [Unit] ordering against networking.service, and mkdir
/run/network in ExecStartPre: the image-build workaround for
ifupdown2#276 doesn't survive boots since /run is tmpfs.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
All wan.net.vmix@* instances created their veth pair with the same
temporary peer name 'vhost' in the host namespace before moving it into
their netns. Parallel starts at boot could steal each other's peer ends,
pairing a host-side vn-<ns> with another namespace's vhost (mismatched
/30s, dead links) or leaving the pair stranded in the host namespace.
Use a per-namespace temporary name (vh-<ns>) and rename to vhost only
after the move, making concurrent creation collision-free.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
NixOS firewall sets conf.all.forwarding=false via mkDefault, which
overrides ip_forward=1. Use normal priority to beat mkDefault.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The NixOS module was importing lib directly with the host's pkgs,
causing image customization to use the host's guestfs-tools instead
of vmix's locked version. guestfs-tools 1.52.2 (from host nixpkgs)
has a bug that overwrites /boot/grub/grub.cfg with resolv.conf
content, breaking VM boot.
Now vmixLib is built once in flake.nix with vmix's own nixpkgs and
passed through the overlay to pkgs.vmixLib. Removes overlay.nix and
module.nix as the logic is inlined in flake.nix.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Switch MAS from /HWID to /Z-Windows (TSforge ZeroCID) which is
hardware-independent and survives VM migration
- Re-install product key and restart SPP service before TSforge
to restore licensing state after sysprep
- Add nicModel option to customizeImage and generalize for images
without VirtIO drivers
- Update MAS activation script to latest version
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Allows overriding the QEMU NIC model during builds (e.g. e1000 for
images without VirtIO drivers). Enables MAS activation on upstream
images that lack VirtIO network drivers.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Win11 LTSC 2024 RDP works with MAS. The edition switch issue was
specific to Win10 LTSC 2021.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
New-NetFirewallRule with -Profile Any is more reliable than
Enable-NetFirewallRule (predefined rules may not exist or be
profile-scoped). Set UserAuthentication=1 (NLA) per standard
RDP configuration. Settings take effect after reboot.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
MAS HWID switches Enterprise LTSC to IoT Enterprise S which lacks
the RDP server listener. Skip activation to preserve the edition.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Add nicModel option (default: virtio-net-pci) to allow e1000 for
images without VirtIO drivers
- Restore MAS activation with slmgr /ipk to switch back from IoT
Enterprise S to Enterprise LTSC (which has native RDP server)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
When spice.vgamem is set (e.g. 64), uses -device qxl-vga,vgamem_mb=N
instead of -vga qxl (which defaults to 16MB). When null (default),
uses -vga qxl for backwards compatibility.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
MAS HWID activation switches the edition from Enterprise LTSC to IoT
Enterprise LTSC (which lacks the RDP server listener). Re-apply the
Enterprise LTSC product key after activation to restore RDP capability.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
sc config fails silently for these services. Use reg add to set
Start=2 (automatic) directly in the registry instead.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
TermService alone doesn't create the RDP listener — SessionEnv (Remote
Desktop Configuration) and UmRdpService (Port Redirector) must also be
running. Use PowerShell Enable-NetFirewallRule to enable the built-in
Remote Desktop firewall rules for all network profiles instead of
creating custom netsh rules.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- generalize.nix: add enableRDP option that re-enables RDP in
post-oobe.cmd after sysprep resets registry (firewall rules,
TermService auto-start, disable NLA)
- Fix OOBE AutoLogon: create user with blank password (Windows
ignores unattend passwords), set real password via net user in
post-oobe.cmd, and explicitly set AutoAdminLogon registry values
- Add LogonCount=999 for persistent AutoLogon across reboots
- Remove unused rdpEntries import from registry/default.nix
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Copy Xauthority to a world-readable temp file so nix build users
(nixbld*) can authenticate to X11. Add --option sandbox relaxed so
__noChroot derivations can access the X11 socket and xauth file.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Only set sandbox = "relaxed" when vmix.namespaces is non-empty.
Safe to import as a default module on all hosts.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Images:
- laptopUpstream: bare OS install with AHCI, no templates
- laptopSlim: essentials only (debloat, registry tweaks)
- laptop: full (essentials + all apps)
- win10/win11 images use rec for self-references
CLI:
- preserve recovery partition (4) during disk copy
- expand partition 3 up to partition 4 boundary
- remove VNC CLI flag (use vncDisplay in nix configs instead)
Flake:
- add devShell with vmix alias and PS1 prompt
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Extract all vmix CLI logic (build, copy, run) from flake.nix into
cli.nix. flake.nix is now 30 lines — just wiring.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Laptop images now use AHCI storage + e1000 network instead of VirtIO.
This fixes "inaccessible boot device" on real hardware — the AHCI→NVMe
driver transition is handled by Windows, unlike VirtIO→NVMe which isn't.
- makeImage: useAHCI flag switches disk to ide-hd and network to e1000
- customizeImage: auto-detects useAHCI from original image, propagates it
- win10/win11 laptop images: useAHCI = true
- vmix run: --ahci flag for running laptop images in QEMU
- generalize: PlainText password tags in OOBE unattend XML
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
SDL display:
- try SDL, auto-fallback to headless if it fails (no crash)
- SDL_VIDEODRIVER=x11 to avoid wayland socket path issues
- suppress XDG_RUNTIME_DIR warnings
Disk copy:
- zap-all before writing to clear old partition tables
- delete recovery partition (4) before resizing partition 3
- use parted resizepart (preserves partition GUID for BCD)
- remote: nix-shell for sgdisk/parted/ntfsresize on target
- remote: lz4 compression for faster streaming
- remote: pv progress bar with disk size
- -y/--yes flag to skip confirmation prompt
Generalize:
- delay-oobe-run=true defers OOBE + activation to real hardware
- clean cached Autounattend from Windows\Panther before sysprep
- taskkill sysprep.exe on first login (CopyProfile artifact)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
CLI:
- `vmix run <qcow2>` boots image with QEMU (SDL if DISPLAY, snapshot mode)
- --generalize supports delay-oobe-run=true to defer OOBE + activation
to first boot on real hardware (for physical disk deployments)
Templates:
- essentials.virtioDrivers: installs VirtIO drivers only (no guest agent)
used in laptop bundle for network access during Office download
- generalize: delayOobeRun flag controls sysprep /shutdown vs /reboot
delays OOBE, user creation and HWID activation to target device
Build:
- suppress XDG_RUNTIME_DIR and homeless-shelter warnings in SDL mode
- remove invalid ICH9-LMB global properties
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- macvtaps working
- only 1 dnsmasq service per namespace
- vms binds to networking services
- lans with domains
- vms no longer assigned same ip (machine id issues)
-