sysprep /generalize regenerates the machine SID on every build, so an account
built by one image does not match a profile left on a persistent disk by an
earlier one -- different SID, so file ACLs, the NTUSER.DAT hive and ProfileList
all mismatch, and the profile will not load.
keepMachineSid drops /generalize and uses /oobe alone. The SID is then
inherited from the cached base install derivation, which is content-addressed
and so identical across every rebuild of the layers above it; the account,
always RID 1000, comes out the same each time. A profile kept on a data disk
then matches exactly, with no ownership or ProfileList fixups.
Without /generalize the specialize pass does not run, so the profile relocation
cannot ride the unattend there. It is written to the registry offline instead,
before the build's OOBE, which virt-win-reg applies ahead of the Audit Mode
boot. And because /generalize is also what strips MountedDevices, dropping it
means the data disk keeps its drive letter into the shipped image -- the
letterless-first-boot race that sent profiles temporary goes away at the root.
Default is unchanged (/generalize), correct for an image deployed to many
hosts; keepMachineSid is for an image that is always the same one machine.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0117qMyjpuXsjpVAcpJbFD8g
It worked on the first boot and was gone after a later internal reboot -- no
IPv4 at all, not even DHCP. DHCP is turned off before the address is set, so
anything that stops the set mid-way leaves the interface with nothing. A stale
ARP entry for the address, left on the network by the previous instance,
tripped duplicate-address detection and made New-NetIPAddress throw; with
-ErrorAction Stop that aborted the script with DHCP already off.
DadTransmits 0 turns that detection off so the static binds regardless of what
the network remembers, and the assignment is now retried a few times rather
than fatal on the first throw. Moved to its own .ps1 -- a wait loop and a retry
are not worth keeping correct inside a cmd one-liner.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0117qMyjpuXsjpVAcpJbFD8g
The relocated profile went temporary on the target and stayed that way. The
cause was not the profile: the image ships with ProfilesDirectory set to
D:\Users but with no drive letter for the data disk. generalize strips
MountedDevices, and the specialize pass that re-asserts the letter runs only in
the build VM, never on the target -- so the zvol boots letterless, the first
autologon cannot find D:\Users\sagar (event 1511), and Windows falls back to a
temporary profile, renames the real ProfileList key to .bak, and the fault
sticks on every later logon.
Confirmed by reading the shipped image offline: ProfileList has the SID at
D:\Users\TEMP with a .bak sibling at D:\Users\sagar, MountedDevices carries no
\DosDevices\D:, and the sagar hive on the zvol is intact -- so nothing was
wrong but the letter.
An onstart SYSTEM task now runs the existing (idempotent) data-disk init, whose
diskpart assign writes MountedDevices and so makes D: persistent for every
later boot. Only the first boot is exposed to the race; if it left a .bak, a
small PowerShell heal puts the key back, drops the temp profile, and reboots
once -- after which D: is persistent and the real profile loads. Same onstart /
SYSTEM mechanism the static address already uses.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0117qMyjpuXsjpVAcpJbFD8g
Setting it from post-oobe.cmd could never have worked, and the reason is worth
writing down. Without delayOobeRun, OOBE runs inside the build VM -- whose NIC
is qemu user networking, on a different subnet, with a different MAC. The
address was being applied to an adapter that does not exist on the real host.
Windows then meets the target's NIC as new hardware and defaults to DHCP.
RDP came through the same script unharmed because its settings are
registry-wide rather than per-adapter, which is why one worked and the other
did not despite sitting a few lines apart.
So the script is now registered as an onstart scheduled task running as SYSTEM.
It was already idempotent, and per-boot also survives the adapter being
replaced again later.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0117qMyjpuXsjpVAcpJbFD8g
The VM came up holding a DHCP lease rather than the address it was told to
take, while RDP -- configured a few lines earlier in the same script -- worked
fine. So post-oobe.cmd was running; only the addressing failed.
Two reasons, both fixed. The interface arrives DHCP-managed and nothing turned
DHCP off, so New-NetIPAddress had no lasting effect. And FirstLogonCommands can
run before the adapter is up, so it is now waited for rather than assumed.
Moved out of post-oobe.cmd into its own file. The command is long and full of
quotes and pipes, which is not a thing to leave at the mercy of cmd's parsing.
It also logs, so the next failure can be read off the disk instead of inferred.
Pings are now allowed too. Windows blocks ICMP by default, which makes a box
at a fixed address look dead to everything that checks it the obvious way --
including me, for a while.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0117qMyjpuXsjpVAcpJbFD8g
The data volume comes out correctly partitioned, formatted and labelled, and
containing nothing but $RECYCLE.BIN and System Volume Information -- so the
disk work lands and the relocation does not. Two candidates, both cheap to
address together.
CopyProfile is now dropped whenever profilesDirectory is set. Sysprep choosing
a profile to copy into Default while the profile root is being moved is the
likelier of the two, and a profile that persists is worth more than the Audit
Mode customizations that CopyProfile preserves.
FolderLocations is now named in oobeSystem as well as specialize. Which pass
honours it is not something the documentation is crisp about, and saying it
twice costs nothing.
This matters more than it looks: with UWF protecting C: and the profile still
on C:, every profile write lands in the overlay and is discarded on reboot,
which makes the whole VM stateless.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0117qMyjpuXsjpVAcpJbFD8g
Formatting it from a specialize RunSynchronousCommand is not enough on its
own. FolderLocations is applied by Shell-Setup while the disk is prepared by
Deployment, and component order within a pass is not guaranteed -- so the
relocation can be evaluated before the volume it names exists, which fails
silently and leaves profiles on C:. That is what a correctly formatted data
disk carrying nothing but NTFS metadata was telling us.
Audit Mode is a fully booted OS with the disk already attached, so doing it
before sysprep makes the volume unconditionally present by the time any pass
looks for it. The specialize copy stays, now purely to re-assert the drive
letter after generalize clears MountedDevices.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0117qMyjpuXsjpVAcpJbFD8g
Two defects found while getting a GPU through Proxmox to a nested guest.
Passthrough gave every address its own pcie-root-port, which splits a GPU from
its own HDMI audio: 05:00.0 and 05:00.1 arrived in the guest as two devices on
two buses instead of functions 0 and 1 of one device. Navi needs both halves on
one device to reset or power-manage either, so the guest got a card stuck in D3
and a reset that could not be performed. Addresses are now grouped by
everything left of the function digit, and each group goes behind one root port
at one slot with multifunction=on on function 0 -- which is also where the
VBIOS and the VGA route belong.
The data-disk setup used Initialize-Disk/New-Partition/Format-Volume. Only the
first of those works that early in specialize; the rest need services that are
not up yet, and with ErrorActionPreference=Stop the script gave up straight
after writing a GPT header. The result was a 50G disk carrying 24KB of nothing
and a ProfilesDirectory pointing at a volume that never existed. diskpart works
at that stage. It also tries assigning the letter before laying the disk out,
so a disk that already holds a profile is lettered rather than cleaned.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0117qMyjpuXsjpVAcpJbFD8g
Image builds have been dying with "Argument list too long" from sed, mktemp
and timeout alike -- commands whose argv is trivial, which is the tell that it
was the environment that had grown, not the arguments.
The X11 forwarding hint is looked up with `ls -t /tmp/.vmix-display-* | head -1`.
Under nullglob a non-matching pattern is removed from the command line rather
than passed through literally, so `ls -t` runs with no arguments at all and
lists the working directory instead. In a nix build that directory is the
build tree, whose newest file is nix's own env-vars dump. The result is that
VMIX_DF becomes "env-vars", the SDL branch is taken on a machine with no X at
all, and DISPLAY is exported with a slice of the env dump inside it. From that
line onward every exec in the build fails with E2BIG.
find does the same lookup without depending on how the shell treats an
unmatched pattern, and -type f keeps a stray directory out of it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0117qMyjpuXsjpVAcpJbFD8g
A disk-mode overlay defaults to C:\uwfswap.sys -- on the very volume being
protected, which is the opposite of the point. uwfmgr grew a create-swapfile
subcommand for exactly this, and it only accepts the call while the filter is
off and the overlay is already in disk mode, so the ordering in the generated
script is forced rather than stylistic.
Configuration is deferred to the target's first boot through RunOnce instead
of running in the build. Two reasons: uwfmgr does not exist until the DISM
feature has been through a reboot, and the swapfile belongs on the real data
volume rather than on the throwaway copy the build attaches. Enabling the
filter needs one more restart after that, which the script asks for itself.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0117qMyjpuXsjpVAcpJbFD8g
Three things, all in service of putting Proxmox in a VM that can still hand a
GPU to its own guests, and of a Windows VM whose profile survives its OS disk.
pci.viommu.enable emits `-device intel-iommu,intremap=on,caching-mode=on` and
forces kernel-irqchip=split, which interrupt remapping requires. Without an
IOMMU of its own a guest cannot bind a passed-through device to vfio-pci, so
it can never forward one on. The device leads the command line because QEMU
realizes devices in order and intel-iommu must precede what it translates.
pci.vgaPassthrough (default true, so nothing changes for existing VMs) makes
x-vga=on optional. It was forced on the first passthrough device, which is
wrong for a card the guest only forwards onward: it claims the VGA path the
emulated console adapter needs.
customizeImage gains extraDisk, a blank disk attached for the Audit Mode boot
and emitted as the derivation's `data` output. generalize uses it for dataDisk
and profilesDirectory, so the disk is partitioned and the profile relocated
under OOBE in the build VM. That is what removes the need for delayOobeRun --
previously the volume ProfilesDirectory names could not exist until the image
reached real hardware. A second output rather than a directory keeps ${image}
meaning the OS qcow2 for every existing consumer.
generalize also picks up staticIP and profilesDirectory.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0117qMyjpuXsjpVAcpJbFD8g
After= was placed in [Service], which systemd ignores ('Unknown key
After'), so the interfaces.d merge + ifreload raced networking.service
on every boot. On losing boots vmbr0 never ran DHCP and the proxmox
guest came up without its LAN IP (bridging still worked, so inner VMs
stayed reachable while the PVE host itself was not).
Move After= to [Unit] ordering against networking.service, and mkdir
/run/network in ExecStartPre: the image-build workaround for
ifupdown2#276 doesn't survive boots since /run is tmpfs.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
All wan.net.vmix@* instances created their veth pair with the same
temporary peer name 'vhost' in the host namespace before moving it into
their netns. Parallel starts at boot could steal each other's peer ends,
pairing a host-side vn-<ns> with another namespace's vhost (mismatched
/30s, dead links) or leaving the pair stranded in the host namespace.
Use a per-namespace temporary name (vh-<ns>) and rename to vhost only
after the move, making concurrent creation collision-free.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
NixOS firewall sets conf.all.forwarding=false via mkDefault, which
overrides ip_forward=1. Use normal priority to beat mkDefault.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The NixOS module was importing lib directly with the host's pkgs,
causing image customization to use the host's guestfs-tools instead
of vmix's locked version. guestfs-tools 1.52.2 (from host nixpkgs)
has a bug that overwrites /boot/grub/grub.cfg with resolv.conf
content, breaking VM boot.
Now vmixLib is built once in flake.nix with vmix's own nixpkgs and
passed through the overlay to pkgs.vmixLib. Removes overlay.nix and
module.nix as the logic is inlined in flake.nix.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Switch MAS from /HWID to /Z-Windows (TSforge ZeroCID) which is
hardware-independent and survives VM migration
- Re-install product key and restart SPP service before TSforge
to restore licensing state after sysprep
- Add nicModel option to customizeImage and generalize for images
without VirtIO drivers
- Update MAS activation script to latest version
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Allows overriding the QEMU NIC model during builds (e.g. e1000 for
images without VirtIO drivers). Enables MAS activation on upstream
images that lack VirtIO network drivers.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Win11 LTSC 2024 RDP works with MAS. The edition switch issue was
specific to Win10 LTSC 2021.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
New-NetFirewallRule with -Profile Any is more reliable than
Enable-NetFirewallRule (predefined rules may not exist or be
profile-scoped). Set UserAuthentication=1 (NLA) per standard
RDP configuration. Settings take effect after reboot.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
MAS HWID switches Enterprise LTSC to IoT Enterprise S which lacks
the RDP server listener. Skip activation to preserve the edition.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Add nicModel option (default: virtio-net-pci) to allow e1000 for
images without VirtIO drivers
- Restore MAS activation with slmgr /ipk to switch back from IoT
Enterprise S to Enterprise LTSC (which has native RDP server)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
When spice.vgamem is set (e.g. 64), uses -device qxl-vga,vgamem_mb=N
instead of -vga qxl (which defaults to 16MB). When null (default),
uses -vga qxl for backwards compatibility.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
MAS HWID activation switches the edition from Enterprise LTSC to IoT
Enterprise LTSC (which lacks the RDP server listener). Re-apply the
Enterprise LTSC product key after activation to restore RDP capability.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
sc config fails silently for these services. Use reg add to set
Start=2 (automatic) directly in the registry instead.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
TermService alone doesn't create the RDP listener — SessionEnv (Remote
Desktop Configuration) and UmRdpService (Port Redirector) must also be
running. Use PowerShell Enable-NetFirewallRule to enable the built-in
Remote Desktop firewall rules for all network profiles instead of
creating custom netsh rules.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- generalize.nix: add enableRDP option that re-enables RDP in
post-oobe.cmd after sysprep resets registry (firewall rules,
TermService auto-start, disable NLA)
- Fix OOBE AutoLogon: create user with blank password (Windows
ignores unattend passwords), set real password via net user in
post-oobe.cmd, and explicitly set AutoAdminLogon registry values
- Add LogonCount=999 for persistent AutoLogon across reboots
- Remove unused rdpEntries import from registry/default.nix
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Copy Xauthority to a world-readable temp file so nix build users
(nixbld*) can authenticate to X11. Add --option sandbox relaxed so
__noChroot derivations can access the X11 socket and xauth file.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Only set sandbox = "relaxed" when vmix.namespaces is non-empty.
Safe to import as a default module on all hosts.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Images:
- laptopUpstream: bare OS install with AHCI, no templates
- laptopSlim: essentials only (debloat, registry tweaks)
- laptop: full (essentials + all apps)
- win10/win11 images use rec for self-references
CLI:
- preserve recovery partition (4) during disk copy
- expand partition 3 up to partition 4 boundary
- remove VNC CLI flag (use vncDisplay in nix configs instead)
Flake:
- add devShell with vmix alias and PS1 prompt
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Extract all vmix CLI logic (build, copy, run) from flake.nix into
cli.nix. flake.nix is now 30 lines — just wiring.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Laptop images now use AHCI storage + e1000 network instead of VirtIO.
This fixes "inaccessible boot device" on real hardware — the AHCI→NVMe
driver transition is handled by Windows, unlike VirtIO→NVMe which isn't.
- makeImage: useAHCI flag switches disk to ide-hd and network to e1000
- customizeImage: auto-detects useAHCI from original image, propagates it
- win10/win11 laptop images: useAHCI = true
- vmix run: --ahci flag for running laptop images in QEMU
- generalize: PlainText password tags in OOBE unattend XML
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
SDL display:
- try SDL, auto-fallback to headless if it fails (no crash)
- SDL_VIDEODRIVER=x11 to avoid wayland socket path issues
- suppress XDG_RUNTIME_DIR warnings
Disk copy:
- zap-all before writing to clear old partition tables
- delete recovery partition (4) before resizing partition 3
- use parted resizepart (preserves partition GUID for BCD)
- remote: nix-shell for sgdisk/parted/ntfsresize on target
- remote: lz4 compression for faster streaming
- remote: pv progress bar with disk size
- -y/--yes flag to skip confirmation prompt
Generalize:
- delay-oobe-run=true defers OOBE + activation to real hardware
- clean cached Autounattend from Windows\Panther before sysprep
- taskkill sysprep.exe on first login (CopyProfile artifact)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
CLI:
- `vmix run <qcow2>` boots image with QEMU (SDL if DISPLAY, snapshot mode)
- --generalize supports delay-oobe-run=true to defer OOBE + activation
to first boot on real hardware (for physical disk deployments)
Templates:
- essentials.virtioDrivers: installs VirtIO drivers only (no guest agent)
used in laptop bundle for network access during Office download
- generalize: delayOobeRun flag controls sysprep /shutdown vs /reboot
delays OOBE, user creation and HWID activation to target device
Build:
- suppress XDG_RUNTIME_DIR and homeless-shelter warnings in SDL mode
- remove invalid ICH9-LMB global properties
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- macvtaps working
- only 1 dnsmasq service per namespace
- vms binds to networking services
- lans with domains
- vms no longer assigned same ip (machine id issues)
-