Stopping an initialized Secure Kernel
4 October 2026 · Windows internals · Secure Kernel from the root, part 9 of 10
The word carrying the whole result is initialized. A normally booted VBS guest's Secure Kernel was stopped at an instruction we chose, held, read twice, stepped sixteen times, restored and resumed — with its text unchanged and vmwp.exe still running the VM afterwards.
Everything before this part either reads Secure Kernel or stops something that is not quite it. Part 5 decodes a checkpoint, part 7 decodes a running guest, and part 6 holds a VTL1 kernel-mode breakpoint on real securekernel.exe code — in a partition we created, with no Windows booted in it. Part 8 is the attempt to close that gap by becoming the VM host, retired one boundary short of a boot.
This is the route that closed it instead, and the move is small enough to state in one line: do not try to be vmwp.exe; attach a debugger to it.
The shape of the result in one line. The VID completion stream a resumable stop needs belongs to exactly one process per partition, and on a managed VM that process is Windows' own
vmwp.exe. So leave it holding the VM — firmware, storage, devices, message receipt, native completion — and drive one of its threads through its own internal registration call for a temporary exception handler. The privileged primitive then only has to read and write VTL1 VP registers for an exact partition. Nothing is injected intovmwp: no helper thread, no DLL, and the controller consumes no VID queue of its own.
Two measurements had to land before that was worth attempting, and both are about what a stop is allowed to touch.
The stop no longer has to patch Secure Kernel's text
A software breakpoint is a byte of memory plus a trap that arrives somewhere. Part 4 established that the byte can be written; part 6 established that the trap, in an owned partition, arrives at VID and names the faulting instruction. What neither establishes is that you may patch an initialized Secure Kernel — which is exactly the thing whose integrity the rest of the system is built on, and which HVCI and skci.dll are in the business of checking.
You do not have to. On the guarded Vid.sys, VidHandleExceptionIntercept reads the vector's claim slot, constructs mapped message type 0x01000002 with the vector in its 16-byte payload, and enqueues VidExceptionInterceptReturnCallback. That callback tests one byte — exchange-buffer offset +0x148: nonzero calls VidInterceptAdvanceInstructionPointer, zero skips it and goes straight to the common completion.
So if the advance byte is cleared, a RIP written while the message is pending survives completion. Part 6's own probe was extended to prove it dynamically rather than infer it from the decompilation, on the tiny long-mode image whose entire state the probe controls:
owner partition 0x4c created
pending breakpoint state: vtl=1 cpl=0 rip=0x10008 rsp=0x1cffd8
held pending intercept for 265 ms
pending VTL1 RIP write verified: rip=0x10180
completed the first breakpoint without advancing the redirected VTL1 RIP
pending breakpoint state: vtl=1 cpl=0 rip=0x10180 rsp=0x1cffd8
received redirected VTL1 breakpoint: rip=0x10180 rsp=0x1cffd8
pending VTL1 RIP write verified: rip=0x10009
restored the original VTL1 continuation while the redirected breakpoint was pending
pending VTL1 state write passed: redirected, restored, and resumed without instruction advance
deleted owner partition
Read that as four separate claims, because it was written to be falsifiable at each: the write takes effect (rip=0x10180 reads back), it affects subsequent execution (the next marked trap arrives at exactly 0x10180, with the same RSP, so the stack was not disturbed), completion neither overwrites nor advances it, and the original continuation can be restored the same way with the loop's own resume witness proving it ran. A successful return from VidSetVirtualProcessorStateEx would have claimed none of that.
And the address to redirect to is published by the target itself. SkdInitDebuggerDataBlock stores &DbgBreakPointWithStatus into KdDebuggerDataBlock+0x20 — measured at securekernel+0xAC210 on 26100.9457, and present on 28000.2952 and 29617.1000 at their own RVAs. An initialized Secure Kernel hands over a breakpoint instruction of its own, at an address it computed, which is the one place a redirect can land without modifying a byte of it.
The first stop inside a running Secure Kernel
2026-10-03, on an allowlisted disposable one-VP VBS boot. The guest stopped at its existing securekernel.exe!DbgBreakPointWithStatus, stayed held while its state was read twice and the two reads agreed, had its original state restored, was completed through vmwp's native dispatcher, and the temporary handler was removed after its deferred cleanup. Afterwards the VM kept its heartbeat and a monotonic uptime for more than 60 seconds.
The two agreeing reads are not belt-and-braces. A single read of a state you are about to make decisions about cannot distinguish "the VP is held" from "the VP is running and you sampled it" — which is the same failure part 7 measured in the reading direction, where a page of .data moved inside twenty seconds.
Then the same lifecycle without a breakpoint instruction anywhere. DR0 and DR7 stopped VTL1 CPL0 before a selected five-byte instruction; TF produced the second vector-1 event at exactly the decoded successor address; DR6 classified the two stops as B0 and then BS, so each stop says which mechanism produced it rather than being assumed; and the original RIP, RSP, RFLAGS, debug registers and unchanged text were all verified before the event was released.
Sixteen steps through Secure Kernel's own control flow
A redirected stop proves the mechanism. It does not prove you can watch Secure Kernel do its job, because the instruction being executed is one you sent it to.
So the natural-flow gate leaves RIP untouched and arms securekernel!KiTimerInterrupt. The guest reaches it through its own execution, and from there the session made 16 guarded steps through register, stack, memory and conditional-branch instructions. Two rules make that sequence evidence rather than a sequence of hopes:
- Every step after the first re-proves the exact current instruction bytes before anything is mutated. The first step may reuse the hardware-stop instruction it already verified; nothing else inherits a guard.
- A step declares where it may land. One default fall-through, or at most four explicit destinations for a branch — and the conditional branch in that run admitted only its two decoded destinations. A single-step is accepted only with
DR6.BSset,TFstill set in the held state,RFclear, andRIPinside that bounded set.
Continue then preserved the guest's execution progress — the advanced RIP, RSP and ordinary flags — while restoring the saved debug registers and the baseline TF/RF bits. That asymmetry is the point of having two modes: redirect mode puts everything back including RIP, because it moved the guest; natural mode must not, because the guest got there itself and rewinding it would be the one unrecoverable thing in the whole run.
Then the same session on VP1 of a two-vCPU boot. The exact-build profile reads the native vector event's VP field, and the dispatcher ignores any event not reporting the selected VP — otherwise "a stop arrived" and "our stop arrived" are the same observation on a multiprocessor guest. VP1 stopped and stepped over the guarded Secure Kernel NOP while the adapter's VM-wide pause kept both VPs stable, and the independent audit afterwards found VP0's debug registers untouched, VP1's baseline restored, TF/RF clear on both, guest text unchanged, the same healthy vmwp, and no scoped crash record.
That gate also produced the run's cheapest lesson. A preceding natural-flow attempt on VP1 timed out, because that interrupt was not scheduled there. Bounded recovery restored the baseline and resumed the VM before faulting the session — so an event that never arrives is a fault with the guest intact, and not a broken breakpoint.
Two further runs from fresh differencing children repeated the whole 16-instruction lifecycle at new partition IDs and new Secure Kernel load bases, each with its own 60-second audit, which is what makes this a three-run result rather than one run told three ways.
The contract, which exists so that the privileged half stays out of this repository
windbg-mcp ships no driver, no private VID layout and no provider DLL, for the same reason the live-kernel tier ships no KDNET wiring: the privileged component is the operator's. What the repository owns is the boundary.
The provider is a child process speaking JSON lines. It announces itself, and every request and response afterwards repeats the full target and the current opaque epoch:
windbg-mcp-sk-control/1
{"protocol":1,"target":{...},"epoch":"running-..."}
The target is the VM GUID, hypervisor partition ID, VP, VTL and expected CR3. Addresses are fixed-width hex strings, so that a JSON implementation cannot round a 64-bit value and hand back a neighbouring page. Provider stdout is read on a dedicated non-DbgEng thread with a 10-second deadline per complete line — including the banner — so a provider that stays alive and stops answering cannot pin the worker's engine thread; EOF is reported immediately.
The epochs are what make an answer mean anything:
| transition | from | to |
|---|---|---|
begin_arm |
running | arming |
finish_arm |
arming | running |
publish_stop |
running | stopped |
release |
stopped | running |
Register access is legal only in arming or stopped. arming is entered while the debugger has already paused the VM and is installing or clearing DR state; stopped requires a held dispatcher event. A provider must refuse register access in running, on a stale epoch, for a different target, for an unsupported register, and for any compare guard that no longer matches — writes are compare-and-write, every named register checked before any replacement is applied. Every transition must issue a token never seen before in that session; returning to an older non-adjacent token is a protocol fault rather than a valid rotation, which is how a replayed epoch is caught instead of honoured.
Those provider-local epochs stay inside the worker. What a caller sees is a separate monotonic epoch the controller issues for each public running or stopped transition, and each controller instance carries a 256-bit nonce from the Windows system RNG — so two provider processes, or two concurrent MCP sessions, cannot produce equal epoch strings that would authorize a replay against another VP or another VM. A token being unguessable is not the point; a token being unambiguous across sessions is.
The readable bank is rip, rsp, rflags, cr3, cs, dr0 through dr3, dr6, dr7 and vsm_vp_status, of which cr3, cs and vsm_vp_status are read-only. A held event records message type, vector, VP, VTL, CPL, dispatcher context, advance flag, and an unclassified debug-exception reason; revision 2 accepts only mapped exception type 0x01000002, vector 1, a selected VP at VTL1 CPL0, a nonzero dispatcher context, and native advance clear.
That word unclassified is the wire change, and it was made deliberately incompatible. Revision 1 had the provider publish an already-classified cause — this was a hardware breakpoint in slot n, that was a single step. Revision 2 has it publish debug_exception and nothing more, and the worker reads DR6 after publication and reports the validated slot or step cause in the stop record. The privileged half of this boundary is the half that should be saying least: a classification is an interpretation, and the interpretation belongs to the side that armed the breakpoint and knows which slots it owns. A revision 1 provider is therefore refused at its ready banner rather than being handed a value outside the schema it declared — a wire break taken on purpose, because the alternative is a provider silently reporting a cause the debugger no longer asked it for.
One ordering rule in that list is load-bearing and easy to get backwards. release consumes the held epoch immediately before the worker completes the event natively — not after. Once native completion runs, the next vector-1 event may arrive at once, so the provider has to be ready to accept a new publish_stop already. The worker stays in its own releasing phase across that window, and a failure inside it faults the session and runs bounded recovery.
What refuses, and what happens when something dies
The adapter that drives vmwp is build-guarded rather than clever. Its profile argument is one JSON file or a directory of at most 128 entries, of which exactly one must declare a vmwp.exe path and SHA-256 matching the current local image — and the bound is enforced during the iteration, counting non-JSON entries too, so selecting a profile never first materializes an unbounded directory. The identity is then checked again against DbgEng's loaded module before any mutation, because a file on disk agreeing with a hash is not the same claim as the process under the debugger being that image. Each profile carries an absolute vmwp.exe path, SHA-256 and SizeOfImage; function and stop-site RVAs; the original bytes of every software-breakpoint site; exact-build event-context, event-VP and native-advance offsets; bounded scratch offsets — and no debugger command text at all. Every debugger command is a fixed internal operation with validated numeric substitutions, breakpoints are created and removed through the typed DbgEng API with their original bytes read back after removal, and before any mutation the adapter checks the VM GUID against vmwp's command line, the image identity, every guarded site, the current VTL1 CR3 and the selected Secure Kernel instruction. It refuses an existing mapping at the requested scratch base, and a pre-existing breakpoint at a site it would own.
Two bench facts are baked into it. It detaches with .detach /h at the pending native breakpoint, because a measured unhandled detach terminates vmwp — which ends the VM. And the delayed Suspend-VM completion kick runs on a host helper thread that makes no DbgEng call, because the engine never leaves its owner thread.
The release path is where those two facts turn into an ordering, and the order is the whole of it. On a matching event the adapter records the exact vmwp system thread before removing its owned breakpoint; release then reselects that thread, completes the registered and native callbacks through guarded return boundaries, gives native completion one watchdog-bounded run slice, detaches handled — and only then joins the delayed Hyper-V helper, because that helper's pause request would otherwise deadlock against a vmwp the debugger has stopped.
Containment got sharper in the same pass, and the distinction is worth stating because it is the one a cleanup routine gets wrong. A conservatively held pause means a resume may be owed; it does not mean provider writes are authorized. So recovery restores the provider baseline only while that exact native event is still retained — or when the pause completed before any event was outstanding at all — and otherwise leaves the baseline untouched and contains the session with the native event incomplete. A retained event becomes release-eligible only after its raw advance field and its owned breakpoint cleanup both pass. "We are paused, so it is safe to write" is the assumption that sentence exists to refuse.
Both failure injections ran live.
- Wrong build. Changing the profiled
vmwpSHA made the adapter refuse before provider mutation, resume the guest, and leave the debug registers,TF/RFand guest text unchanged. - Dead provider. A provider exiting immediately before acknowledging
publish_stopmade the wait return a terminal fault withtarget_left_paused=true;end_sessionreportedrecovery_required=trueandreleased=false, and retained the exact worker rather than cleaning up. An independent provider read then foundDR0still armed at the ownedKiTimerInterruptaddress and the guarded text unchanged.
The second one is the honest half of the result, and the record says so in those terms: because the dead provider could no longer restore that state, the dedicated failure child was discarded after its evidence was captured, and ending the unresolved debugger replaced vmwp — which is what an unhandled DbgEng detach does. So that run proves containment. It does not prove recovery, and it does not prove the guest stays healthy after a provider dies. A fail-closed teardown that keeps the worker and admits only another teardown attempt is worth having precisely because the alternative is a silent claim of cleanup.
The fan-out that passed offline and cannot run
The controller underneath all of this is written for one provider per selected VP, and it has offline tests for the fan-out — including a VP1 win with VP0's baseline restored. The concrete build-guarded adapter now refuses that fan-out before it starts a provider, and the restriction is a live reading rather than caution.
Two things were measured on the two-provider attempt:
Suspend-VMdid not return while the winning VID event was outstanding — not after the event thread had been redirected to owned scratch, and not after DbgEng performed a handled detach. The pause a fan-out would need in order to touch the losing VP is the one thing the winning event prevents.- Holding that first callback and pumping the remaining
vmwpthreads produced no second selected-VP callback inside the bounded deadline. At this point the native dispatcher path is serialized, so there is no second event to catch while the first is held.
So the narrow path treats one selected VP's retained event as the provider-write barrier, which is what the ordering above is for. And the fan-out is split into gates that each need live evidence before any of it is claimed again: a two-phase dispatcher release that can finish the winning native event and prove a post-event VM pause before any unheld provider write; then restoring providers that never generated an event under that pause; then handling and restoring a losing provider that generates a breakpoint event while the winner is stepping; and only then the two-provider, four-slot, 16-step scheduler-selected lifecycle with its own health audit.
The generic controller keeps its offline fan-out tests. They are a model of a machine that, measured, does not behave that way — which is part 10's subject rather than this one's.
Seven tools, and why this one is a session when the live source was not
Part 7 ended by arguing against shipping its live source as a tool: gate S5x watched a page of securekernel.exe's .data change inside twenty seconds, so a live read cannot keep the promise a capture session's decode-on-open makes. That argument does not transfer here, and the reason is the epoch. A live-control answer is not "the state of a running guest", it is "the state of a VP held at an event this session owns" — and every mutating call has to present that exact opaque epoch, so a stale or replayed step is rejected before it mutates anything.
The securekernel group therefore grew from four tools to eleven, and this server's openers from seven to eight — open_sk_live_control being the eighth, after the six debugger openers and the capture that holds no debuggee:
| tool | what it does |
|---|---|
open_sk_live_control |
records the requested VM, partition, selected VP, CR3, vmwp PID, dispatcher pointer, profile and provider commands; validates their shape, loads the profile and checks the register provider's declared identity and capabilities. It does not pause, attach to or inspect the VM |
sk_live_arm |
the first one pauses the VM and validates the live target — the VM-to-vmwp binding, live memory against the requested CR3, each guarded instruction, the attach, and vmwp's exact build and dispatcher sites — and only then saves the selected VP's baseline and installs one to four explicitly slotted execution breakpoints; redirect mode requires a single address, natural mode arms the set |
sk_live_wait |
pumps vmwp until the selected VP reaches any armed address; retains that exact callback thread and returns the winning DR slot with two matching register snapshots, the guarded bytes, the bounded accepted next addresses, the arm mode, and a fresh controller epoch |
sk_live_registers |
the retained stop record, read without consuming its epoch |
sk_live_read_memory |
guest virtual memory through the live VTL1 page-table walk, bound CR3, tied to the held epoch and carrying its starting GPA; running, stale and unmapped reads are refused whole rather than returned short |
sk_live_step |
consumes the stopped epoch and arms one trap-flag step, with the current instruction and the bounded destination set |
sk_live_continue |
consumes the epoch, restores and verifies the complete baseline, clears execution control, completes the owned event; the session can be armed again |
Which call does the proving is itself a design decision, and it moved. Opening is cheap and claims nothing: it validates a request, a profile and a provider's declared identity, and the summary it returns names the vmwp debug target the caller asked for — explicitly not a claim that DbgEng attached to anything, let alone that a guest Secure Kernel is there. Everything that can only be known by touching the machine is done by the first sk_live_arm, under the pause, before a single byte is mutated. An opener that reports success for work it has not done is the same defect as a field that looks like evidence and is only a request.
Ordinary debugger tools are refused on this session, and the live operations are refused on every other session kind — because this session's DbgEng target is the host vmwp while its answers describe guest VTL1, and a tool that answered about the wrong one of those would be worse than no tool. end_session restores the selected VP, completes any owned event through the native dispatcher, removes the temporary handler and the scratch allocation, detaches from vmwp and leaves the VM running. If any of that is unconfirmed the session becomes live_control_unresolved, and what survives then is the interesting part: the worker, the session slot and an exact VM-plus-vmwp reservation outlive idle and capacity reclamation, lease expiry, shutdown and supervisor loss. The session refuses further control work and cannot be treated as released — someone has to inspect or discard that disposable VM out of band before another controller is allowed to own it. This server already had one such exception, for unresolved remote kernel controllers; a half-released VTL1 stop is the second, and for the same reason: an automatic cleanup that cannot prove what it did is worse than a session nobody reclaimed.
A further run on 2026-10-04 exercised the widened arming end to end: it selected the exact profile out of a two-entry catalog whose other entry was a wrong build, armed DR0 through DR3 at four guarded securekernel!KiTimerInterrupt instructions, classified the natural stop as slot 3, completed all 16 guarded steps including the conditional branch, restored both vCPUs' debug baselines and TF/RF state, preserved the guarded text, released the worker, kept the same healthy vmwp through the independent 60-second audit, and left the disposable VM Off.
On 2026-10-03 the opt-in smoke test drove that whole lifecycle — bind, arm, stop, inspect, step, second stop, continue, close — through the built MCP binary against the allowlisted disposable VM, and its wrapper then observed the same vmwp PID with an advancing heartbeat for 60 seconds, no scoped crash record, unchanged guarded bytes, and the VM finally Off. The test's config is read from one environment variable whose JSON must explicitly say "disposable": true, and neither the profile, the provider, nor the evidence is in version control.
The surface cost is worth stating because this server charges every caller for tools they may never use: the default surface went from 67 tools to 74, and the definitions the model is handed before it asks anything from 105,588 to 113,516 bytes — about 28k tokens — of which the securekernel group is now 15,573. That is what --tools exists for, and a client that will only ever open a capture should not be paying 13.7% of its context for a control session it cannot drive: four of those eleven tools need a checkpoint, the SDK's provider and the guest's own securekernel.exe, and the other seven need a disposable VBS VM, two operator-supplied providers and a build-matched vmwp profile. Almost nobody driving a crash dump has either setup. Those are measurements of 2026-10-04, taken from the tables a surface test checks against a running server rather than from the prose beside them, and any edit to a tool's description moves them — this one moved three times on the day it was written.
What this establishes, and what it does not
Established. An initialized, normally booted Secure Kernel can be stopped at a chosen instruction from outside the guest, held while its state is read twice and agreed, inspected, single-stepped through its own control flow 16 times with every instruction re-proved, restored, and resumed — three times, on fresh boots, at new partition IDs and new load bases; on a selected VP of a multiprocessor guest without disturbing the other; with Secure Kernel's text unchanged throughout and not one of its bytes patched — the software breakpoints are host-side, in vmwp, created and removed through the typed API with their original bytes read back — nothing injected into vmwp, no VID queue consumed by the controller, and a 60-second independent survival audit after each. Up to four slots can be armed at once, with the winning one classified by the debugger rather than by the privileged provider. The route is a documented contract with the privileged half outside the repository, and it is reachable from seven epoch-bound MCP tools.
Not established. Recovery from a provider death — that is containment, with the guest discarded. Fan-out across VPs, which has offline tests and a live refusal, and four named gates before any of it is claimed. Live debugger loss after restoration, and VM-reset identity, which exist only as offline injections. Exact-build profile coverage beyond this one vmwp build. Anything at all about a host that is not deliberately weakened, allowlisted and disposable. And none of this makes Secure Kernel inspection harder to get: that still runs off a checkpoint with no driver, no test-signing and no hypercall, which is part 5.
Nine parts, from a ret that explained a silent debugger to a held breakpoint in a live one. The part that transfers is still the last one: how to measure nothing.