Most of a VM host, and the boot we stopped short of
4 October 2026 · Windows internals · Secure Kernel from the root, part 8 of 10
To stop an initialized Secure Kernel we first tried to become the process that owns the VM. Four gates later we had six of Hyper-V's own device models running in a partition of our own, the exact Windows memory map, and UEFI firmware state that reached DXE — and then we retired the whole route.
Part 6 ended on a wall with a precise shape. A resumable stop needs the VID message that reports the stop and the completion that releases it, that stream belongs to one process per partition, Vid.sys enforces it as a process-identity check at [partition+0x3780], and on a running Hyper-V VM the process holding it is vmwp.exe. Consuming from that slot takes the owner's traffic: the run that tried drained 64 of vmwp's own messages and the guest reset inside ten seconds.
So the probe in part 6 created its own partition instead, and held a VTL1 kernel-mode stop on real securekernel.exe code there. What it could not do is boot Windows in that partition — and without a booted Windows there is no initialized Secure Kernel to stop, only a mapped image being called as a library. Part 6 priced the difference in one clause: the 19-module, roughly 10 MB device-and-firmware cliff is what it costs to boot Windows in a partition you own.
This part is that climb. It is also the part where a route gets retired on purpose, which is why it is written as four gates with a fatal one first rather than as a schedule.
The question, and why the first gate had to be the one that could kill it
The plan this replaced was a schedule written against the obstacle in front of it: own the boot, reproduce the completion, then stop. Its risk is not in the work it lists but in one unvalidated assumption underneath — that Hyper-V's device model can be driven by a host process rather than by vmwp.exe and VMMS specifically. If that is false, every week spent on firmware and storage is spent on nothing.
So the gates are ordered to collect what does not depend on that assumption, and then to falsify it before anything is built on it. Four of them, each with a pass condition, a control, and a stop condition written before the work:
| gate | question | kind of answer it can give |
|---|---|---|
| K1.0 | do a VBS guest and a VBS-off twin boot identically in the minimum shape, and what is their physical memory map? | the control pair, and an input K1.2 needs |
| K1.1 | can each in-box device model initialize outside vmwp? |
fatal — a no retires the route |
| K1.2 | can the six share one partition, one service graph and one unwind? | composition, not initialization |
| K1.3 | can the in-box BIOS build and apply a bootable UEFI state? | firmware, still with no VP started |
Part 6 had already answered the cheaper half of K1.1 — the registry census of 24 in-box CLSIDs and the CoCreateInstance run that got 20 S_OK, 4 CLASS_E_CLASSNOTAVAILABLE and no faults, under a restricted token and at genuine medium integrity. Activation is not initialization. A class factory handing back an object proves registration; it says nothing about whether IVirtualDevice::Initialize can succeed without the partition object, VMBus channel manager and repository state that vmwp.exe supplies. That distinction is the whole of K1.1, and part 10 records what it cost to learn that the earlier record had flattened it.
K1.0: the same boot, twice, in the minimum shape
The lab discipline is part 3's: two guests differing in one variable. What K1.0 adds is that both source chains were flattened into private immutable bases and used only through disposable differencing children, so that three cold boots of each are three boots of the same bits rather than three boots of a drifting disk.
The shape is deliberately minimal and fixed: one VP, 4 GiB of RAM, Secure Boot off, no TPM, no network, no DVD, and no Guest Service Interface. Each base cold-booted three times. Paired saved-state reads — part 5's route, which needs no driver — then separate them directly:
| reading | VBS base | VBS-off control |
|---|---|---|
| VTL masks | 3 |
1 |
| VTL1 page-table root | readable, long mode | refused |
| Secure Kernel module list | validated, contains skci.dll |
none |
That is the positive and negative arm of the same instrument, which is the only reason the later gates' passes mean anything. It also produced a reading the next gate needed: both captures expose the same two RAM spans, which is where K1.2's memory map comes from rather than from a guess about what Windows wants.
K1.1: six devices, six children, one fatal question
tools/vdev_initialization_probe.py asks whether the six in-box device models in a minimum Windows boot will initialize in an ordinary owner process. It runs each one in a separate child process with a 30-second deadline, so a faulting device model costs one result rather than the run, and a timed-out child takes its ownership boundary down with it.
The recovered common interface is three vtable slots:
IVirtualDevice slot |
method |
|---|---|
| 3 | GetDependencies(repository, count, services, required_count) |
| 4 | Initialize(repository, uint64_argument, services) |
| 5 | Teardown() |
with IID_IVirtualDevice pinned to {0693ED7D-8A8A-4D87-A468-1103B8C63D9C}, the repository to {355AC5A8-8A94-44A9-BE14-ECF7FB8F7C3B} and the service map to {20BEEF08-C3AB-44D8-92C3-03EC0CF398DC}. The probe refuses to make a VID call at all unless every in-box binary matches a guarded size and SHA-256 — plus its embedded CodeView identity, for the four that carry one — and unless slots 3 through 5 resolve inside the guarded module at the recovered RVAs — the same build-lock discipline as part 6's native probe, for the same reason: these are undocumented, build-specific interfaces and a mismatched one is a crash rather than an error.
All six initialize and tear down, and each releases every dependency it was handed. GuestEmulationDevice, BiosVdev, RtcVdev, IoApicVdev, VmbusVdev and SynthStor, elevated, 2026-10-02:
| device | dependencies | what Initialize actually called |
|---|---|---|
GuestEmulationDevice |
13 required, 3 optional | ISecurityManager slot 11 returns zero; repository /generation_id returns the zero GUID |
BiosVdev |
16 required, 4 optional | repository /generation_id returns the zero GUID; ISecurityManager slots 12, 13 and 10 return zero |
RtcVdev |
6 required | no service method; missing exported configuration is accepted |
IoApicVdev |
4 required, 1 optional | no service method |
VmbusVdev |
4 required | handle-broker slot 3 returns E_NOTIMPL |
SynthStor |
4 required | runtime configuration absent; repository /PreallocatedResources returns false |
Five of the six run against a fresh process-local direct-VID partition that the child creates and deletes; RTC needs no partition at all. The recorded run opened IDs 0x77 through 0x7B, and partition IDs and GUIDs are fresh on every run by construction.
VMBus is the gate's discriminator, and it is the one device whose pass is about ownership rather than about tolerance. Its handle-broker lookup for VmbusVdevHandle deliberately returns E_NOTIMPL, after which VMBus goes and opens \\.\VMBus\vdev\{vm-id} itself. That open succeeds only while the probe owns a direct VID partition carrying the same bare GUID. So the inbox device is not merely indifferent to who hosts it; it finds its own partition through the VM GUID, and ours is the one it finds.
Two smaller things worth recording because they are the difference between a measurement and a story that fits. The optional IID order returned by GetDependencies is the reverse of the order in which the guest and BIOS templates query those optional services, so the probe asserts both the returned order and the observed behaviour rather than inferring one from the other. And no supplied service method other than the ones in that table is called by Initialize or Teardown — the service collection is a recording object, so "nothing else was asked for" is an observation rather than an absence of instrumentation.
The whole contract is generated and committed as vdev-contract-26100.8457.json, and the probe refuses to run against a stale manifest. That matters more than it sounds: the manifest is the thing that makes a later run on a patched host fail loudly instead of quietly measuring a different Windows.
What the pass disproves is the stop condition this gate was written to find: none of those six independent initialization paths requires vmwp process identity, an identity-bearing VMMS repository object, managed-VM state, or a second VID receive loop. What it does not prove is that the six can share one graph — which is the next gate, and was written down as a separate one before this one ran.
K1.2: one partition, real interfaces, and three clean unwinds
tools/vdev_graph_probe.py composes the six in one owner-created partition. The initialization order is VmbusVdev, IoApicVdev, BiosVdev, RtcVdev, GuestEmulationDevice, SynthStor; teardown is the exact reverse. Three of the recording stubs are replaced by interfaces the in-box objects themselves provide:
| provider | interface | IID | consumers |
|---|---|---|---|
VmbusVdev |
IVmbusServices |
{ECE3F556-F87F-4120-9E37-AAA55E5E0CA9} |
BIOS, guest emulation, SynthStor |
IoApicVdev |
IVmIoApic |
{9D33829B-58BE-4BBF-AB6E-3B16DBCEF954} |
BIOS, RTC |
BiosVdev |
IVmBios |
{9BE0B79F-68DF-4C59-9D88-4BFC1BF7A73D} |
RTC, guest emulation |
All three answer S_OK from QueryInterface, and the devices' initialization calls stay exactly the K1.1 set — so swapping a stub for the real object changed what the graph is without changing what it asks for.
The memory map is K1.0's reading rather than a round number:
| span | start | size | pages |
|---|---|---|---|
| low RAM | 0x0 |
0xF8000000 |
0xF8000 |
| PCI/MMIO hole | 0xF8000000 |
0x08000000 |
not RAM |
| high RAM | 0x100000000 |
0x08000000 |
0x8000 |
Two VSM-capable VA-backed memory blocks, one per span, bound to the partition's notification queue, with GPA ranges created under default VTL protections; one marked page per block is mapped, read back and cleared before anything depends on it. The pre-power lifecycle then follows what was recovered from vmwp.exe rather than what seemed reasonable: slot 6 StartReservingResources in graph order, the RAM-complete phase once both ranges exist, slot 7 FinishReservingResources with rollback clear, then slot 8 FreeReservedResources and teardown in reverse. Each of those slots is checked against a per-device RVA before it is called, and a partial failure calls FinishReservingResources with rollback set and frees every reservation it started.
The RAM-complete phase is the one that would have been easy to fake. VirtualMotherboard::NotifyAllDevicesRamConstructionComplete sits at vmwp.exe RVA 0x218780; decompiled, it walks every device, queries IID_IVirtualDeviceMemoryInfo {2E223C59-62C4-4D03-93E4-05674B3B94EB}, calls interface slot 4 with a 32-bit argument where it is supported, and logs a failed result without stopping the walk. Its only direct caller on this path, RVA 0x97928, passes 0. The probe mirrors that exactly — and all six devices answer E_NOINTERFACE, before and after initialization, so on this minimum graph the phase is a measured no-op. It is still issued and still asserted on every run, because "we skipped it and nothing broke" and "we ran it and nothing was there to call" are different facts.
Three consecutive runs passed on 2026-10-02, in fresh partitions 0x1A, 0x1B and 0x1C: all 18 resource calls S_OK, every initialize and every teardown S_OK, both ranges and blocks destroyed, every partition deleted.
One detail is a real finding about object lifetime rather than bookkeeping. BiosVdev retains its repository after Teardown and releases it when the COM object is destroyed, so the probe checks repository ownership only after it has closed the graph-held provider interfaces and released every device object. Checking earlier would have reported a leak; checking at process exit would have reported nothing at all.
K1.3: the firmware builds a bootable state, and we do not press start
tools/vdev_firmware_probe.py takes the graph to the edge of execution. It supplies the firmware-time services the diskless cold-power path needs — a checksummed 80-byte MADT and 144-byte SRAT carrying the recovered VRTUAL and MICROSOFT OEM fields, PCI and MMIO ranges, one package and one thread, empty SLIT and PPTT, one VP, memory mode zero, isolation type zero, no TPM — and deliberately supplies nothing optional: a service nothing requires receives E_NOINTERFACE rather than a permissive stub. Every callback catches its own errors at the native boundary and returns an HRESULT, because a Python exception unwinding through in-box code is not a test result.
Then BiosVdev::PowerOnCold imports its boot state in five non-overlapping page ranges:
| GPA page | pages | source bytes | role |
|---|---|---|---|
0x100 |
0x600 |
0x600000 |
in-box UEFI image |
0x700 |
6 | 0x6000 |
initial firmware data |
0x706 |
1 | 24 | loader data |
0x707 |
2 | 0 | zero-filled pages |
0x709 |
1 | 704 | final loader data |
Each is written through VID in at most 16-page chunks, zero-filled over the unused part of the requested range, read back whole, and required to compare byte-for-byte before the probe returns S_OK to the BIOS. Including the two pages whose source is zero bytes — a range that was asked for and contains nothing is still a range, and "we wrote nothing" and "nothing is there" are the same observation unless you read it back.
The importer then takes 19 VP0 state records in this exact order:
70001 60003 60000 60004 60005 60002 60001
40000 40002 40003 80001 80004
20011 20005 20010 20008 20009 2000A 2000B
whose measured scalars are CR0=0x80000023, CR3=0x700000, CR4=0x660, EFER=0xD00, PAT=0x7040600070406, RFLAGS=0x2, RBP=0x6E0000 and RIP=0x6E1474. RIP is required to land inside the imported UEFI image rather than merely to match that number, which is the difference between a check and a transcription. All 19 apply in one VidSetVirtualProcessorStateEx and read back in one VidGetVirtualProcessorStateEx; VID's only normalization is setting the architecturally fixed CR0.ET bit, so CR0 reads back as 0x80000033 and that is the only changed record. The probe requires that exact canonicalization, which turns a known quirk into an assertion instead of an allowance.
Of the six devices, five cold-power to S_OK — VMBus, IOAPIC, BIOS, RTC and guest emulation. SynthStor initializes and reserves and is then deliberately not powered, because this gate supplies no LUN. Three runs passed on 2026-10-03 in partitions 0x4D, 0x4E and 0x4F, each with five exact page imports, 19 applied records, five powered devices, no VP started, and the partition deleted.
And one ABI detail decided whether any of it was real. IVmBootMemoryTopology returns byte addresses and lengths, while VID's memory-block APIs take page numbers and page counts. A private execution check found it: page-scaled values made PEI call InstallPeiMemory(0, 0), which reads exactly like firmware refusing a machine with no usable RAM; byte-scaled values made the same call InstallPeiMemory(0x70A000, 0x4081000), after which it advanced through PEI into DXE and an idle HLT, and resuming the five powered devices carried it on into later timer work. That execution check is diagnostic, not this gate's acceptance result — the acceptance runs never start VP0 — but it is the evidence that the state being built is a state that boots.
Retired, deliberately, one boundary short
What remained after K1.3 was a real SynthStor LUN and an owner-side VID completion dispatcher before a Windows boot could be attempted at all, and the original costing put the rest at 4 to 8 engineering weeks to a first repeatable initialized boot and 6 to 12 through the stop and release.
It was not spent. Part 9 is the narrower route that passed on 2026-10-03: leave Windows' own vmwp responsible for firmware, storage, devices, VID receipt and completion, and debug that process instead of replacing it. Once that worked, the owner-built boot was retired as the selected route. Nothing in this part is load-bearing for the result the series ends on.
That is the right trade and it is worth being explicit about why it is not a write-off.
- The fatal gate came back negative, and that is a published fact about Hyper-V rather than about this project. Six in-box device models initialize, compose, reserve resources, accept the RAM-complete walk and tear down in a process that is not
vmwp.exe, with no VMMS object, no managed-VM state and no second receive loop. The route was abandoned because a cheaper one appeared, not because it was blocked. - No irreducible managed service ever turned up, which is also what keeps the contingency closed: adopting a third-party host user-mode VMM would have added a VMM stack without exercising the in-box VID receive path this work depends on, and stock OpenVMM's Windows path is VTL0-only, so it does not constrain the guest's CPL at all.
- The gates were ordered so that retirement was cheap. K1.1 — the one that could have killed the route — ran before the firmware work rather than after it. Had it failed, the stop would have cost one gate instead of eight weeks.
What this part does not claim: no Windows booted in a partition we own, no VP ever started in an acceptance run, no synthetic disk I/O, and no Secure Kernel initialized by any of it. The three probes are checked in and build-locked; the bench, the disk lineage and the host component detail are not, and stay out of the public record on purpose.
The stop itself is part 9.