A live source is not a snapshot
4 October 2026 · Windows internals · Secure Kernel from the root, part 7 of 10
The decode in part 5 was built over a source seam so that a live source could join it later. One finally did — and produced a refusal no fixture had managed, plus a two-byte measurement that stopped a tool from shipping.
Everything in part 5 reads a file. That was the point: a checkpoint needs no driver, no test-signing and no weakened host, and it can be copied off the Hyper-V host and read anywhere. But a file is a guest that stopped, and the live route from part 4 remains the only way to read one as it runs.
The decode was written for both from the beginning. Its source seam is three methods — read(gpa, len), root() and max_read() — and its doc comment said, in terms, "a future driver-backed live source joins here and changes nothing above it." Two things in it existed for a live source and had never been exercised by one:
max_read(), which exists becauseHvCallReadGpamoves at most sixteen bytes. A capture has no such limit.ReadFailure::Refused, a variant distinct from a failed read, which exists because part 4 measured the hypervisor answeringHV_STATUS_SUCCESSwith zeros and a per-accessReadIntercept. Produced by synthetic fixtures only.
An unexercised guard is a hypothesis. This part is what happened when a real source drove it.
The shape of the result in one line. The whole decode runs against a running guest through an operator-supplied transport, and reproduces every landmark. A second transport reads through the refusing hypercall and makes the decode decline to claim a negative it has not earned. And measuring what a running Secure Kernel's pages do in twenty seconds is why the live route shipped as a command-line role rather than as a fifth MCP tool.
What this repository ships, and what it does not
src/livesrc.rs is the live source, and what it adds is a client: the transport is a child process speaking a line protocol on its stdio, and the repository ships none of them. Reading another partition's VTL1 live needs a kernel component this project will not distribute — the same standing decision as the live-kernel tier's KDNET wiring. The operator supplies the transport; windbg-mcp --sk-live --transport "<command>" --image <path> drives the decode through it, and refuses to start without one.
A DLL with an agreed export was the alternative and is worse for a reason part 5 measured: it would put unsafe FFI and a vendor ABI in the crate, and a provider that __fastfails takes the whole process down with it — which is precisely what vmsavedstatedumpprovider.dll does on an encrypted capture. The process boundary that the SDK provider's robustness defect argued for is the same boundary a line protocol gets for free.
So the crate gains no privileged code, links nothing new, and contains no unsafe FFI for this. Fifteen unit tests pin the framing against in-memory pipes — including a provider banner being skipped to a sentinel, a misspelled key being refused, and a transport announcing more bytes than were asked for.
Direct mode: the whole decode, live
The bench's first transport reads physical memory through the driver interface from part 4, which reads VTL1 and never refuses. Against the running VBS guest:
| run of 1 October | re-run, 2 October | |
|---|---|---|
| root | 0x1201000, 26 present entries, self-map at [290] |
identical |
| walk | 11,819 leaf mappings over 4,229 distinct pages, 169 table reads, 0 malformed, complete | 11,820 leaves over 4,318 pages, same 169 table reads, still complete |
| scan | 11,819 pages, 6 PE headers, 1 matching the image on disk | 11,820 pages, same 6, same 1 |
| identified | securekernel.exe at 0xFFFFF8024278A000 (GPA 0xCD1000), 4 KDBG hits, 3 other PE headers rejected on KernBase |
identical |
KdDebuggerDataBlock |
0xFFFFF802428BD5E0 = base +0x1335E0, Size 0x3A0 |
identical |
SkLoadedModuleList |
0xFFFFF802428B1770 = base +0x127770 |
identical |
| image | 1,527,808 bytes gathered, 0 pages unreadable | identical |
| modules | 6, complete: securekernel.exe, skci.dll, symcryptk.dll, cng.sys, vmsvc.dll, vmsvcext.sys |
identical |
| reads | 12,375 attempted, 0 failed, 0 refused, 50,688,000 bytes | 12,376 reads, 50,692,096 bytes |
So every figure part 5 reports from a capture, the same code now reports from a running guest, through a source this repository does not contain.
And the second column is a result rather than noise. Same command, same guest, same boot, a day apart. What moved is the mapping: 89 more distinct pages behind one more leaf. Every landmark, the identified base, the debugger data block, the loader list and the module names are byte-identical.
That is the stronger form of this part's own finding, and it arrived before the measurement that was meant to establish it: the churn arm below measures page contents changing, and this is the page table changing. A capture cannot do this. It is why these numbers carry a date, and why the right instruction in the record is do not "correct" one column against the other.
Hypercall mode: the refusal, at last from a real source
The second transport drives HvCallReadGpa instead, at its real max_read of 16 — not a configuration choice but the width the seam's max_read() method exists because of.
Part 4 measured that call answering HV_STATUS_SUCCESS with zeros and a per-access ReadIntercept(2). A transport reading the top-level status alone would hand the zeros back as data. This one reads the access result:
root 0x1201000 (read from the capture, never remembered)
rootpage unreadable: Refused { detail: "ReadIntercept(2)" }
walk 0 leaf mapping(s) ... 1 unreadable table(s)
complete: false (UnreadableTables(1))
identified none
reads 2 attempted, 2 failed (2 refused), 0 byte(s)
this run identified nothing while 2 read(s) failed, so its negative has not been earned
refused is above zero for the first time outside a fixture, and the last line is the whole reason that variant is separate from the others. The decode does not report this guest has no VTL1. It reports that it could not earn the negative.
That is the one mistake the seam was shaped to make impossible. A source collapsing a refusal into success — zeros handed back as bytes — would have produced part 3's VBS-off control's answer from a VBS-on guest: no Secure Kernel found, no KDBG tag, a clean and completely wrong negative. Now demonstrated rather than argued.
One correction, because the first draft of this claim overreached. It said that collapsing a refusal into a plain failure was the same mistake. It is not — it is a different and lesser one. Read the message two lines above: a failed read is exactly what makes the decode withhold the negative. What a failure-shaped refusal loses is the classification naming the cause, not the verdict. The verdict survives either way, and saying otherwise credited the Refused variant with work the failed counter was already doing.
Two things the protocol had to learn from the bench
A transport prints before it is a transport, so the protocol opens with a sentinel. The bench's provider library prints a partition menu with an ordinary print(), so the client's first read took [ 0 ] Lab Guest Hyper-V as the protocol's shape reply. The transport now announces windbg-mcp-gpa/1, and everything before that line is echoed to stderr for the operator.
The alternative — having the transport redirect its own stdout around the provider — was tried and abandoned: duplicating fd 1 and pointing fd 1 elsewhere crashed the Python host inside its allocator, every run.
And a transport must not fill in registers it cannot read. The bench's library exposes no VTL1 control registers, so its shape reply declares cr0, cr4 and efer as unreadable with a reason rather than inventing plausible values. Inventing them would make the decode's long-mode consistency check vacuous, because that check answers unknown from a missing register by design — the honest answer defeats the guard far less than a plausible one.
A misspelled shape key is refused for the same family of reason. Silently ignored, a mistyped cr3 reads as a guest with no page-table root — which is a far more plausible-looking result than a parse error, and therefore far more dangerous.
Twenty seconds, 373 pages, two bytes
Two consequences of a live source were named when the capture tools shipped, and neither had been measured:
- The session stops being a fixed snapshot, which is the premise the capture opener's decode-on-open rests on.
sk_read_memoryacquires a target that can change between two reads.
Both are plausible. Plausible is not true — if a running Secure Kernel's pages were static in practice, the session model would need nothing. So: read pages twice with a gap and count what differs. Nothing is written.
| sample, 20 s apart | pages | changed |
|---|---|---|
inside securekernel.exe (every page of the image, translated individually) |
373 | 1, and 2 bytes |
| elsewhere in VTL1 (stride-sampled from 5,364 walked leaves) | 1,200 | 0 |
The page that moved is base +0x12B000, which the image's section table puts in .data — not .text, not .rdata, and not in any landmark the decode reports.
Two things about that sample, because its first two runs were wrong in the instructive direction. The leaf walk descends only 4 KB leaves, and this guest maps the image with large pages — so the image arm printed "0 of 0 page(s) changed", which reads as no image page changed and means no image page was looked at. It now translates every page of the image range explicitly. And the first sample took the first 400 leaves rather than a stride, which biased it to the lowest PML4 indices; the stride is what lets 1,200 pages stand for 5,364.
Which is why it is not a fifth tool
Both consequences are real, and the magnitude is what bounds them.
Decode-on-open is unsound in principle for a live source. A figure computed at the open can go stale while the session is held, and the day-apart walk above shows the page tables doing exactly that. What actually moved in twenty seconds, though, is two bytes of writable data, while the root, the identified base, the debugger data block, the loader list and all 1,200 sampled non-image pages were byte-identical.
So the honest shape of the thing is: a capture session's whole decode travels with the opener because the capture cannot change — that was the argument for putting it there — and a live source would quietly turn those fields from facts into timestamps. Keeping the live source out of that opener is a decision with a measurement behind it rather than an omission, and --sk-live is a command-line role for that reason. A reader who wants a tool against a running Secure Kernel is being told no on evidence.
And the shipped guidance now says so in both directions, which it had said nothing about in either. The plugin skill gained a playbook for the four capture tools: what the host needs, the three ways of naming a capture, why the image is required, that a capture session refuses the tools that answer about a debugger target while session teardown and interrupt go on working, that the public PDB's missing type records mean structures are hand-decoded, and that a VBS-off guest's capture opening with no VTL1 is the answer rather than a failure. The live route is stated as what it is: the operator supplies the transport, this repository ships none, and there is no MCP call for it — so a request to debug a running Secure Kernel is answered with a checkpoint or handed back.
One detail there is worth more than the playbook. Five places enumerate what this project debugs, and not one of them named Secure Kernel — not since the four tools shipped. Including the skill's own description, which is the only part of a skill in context before it is loaded. The playbook was unreachable for the exact request it exists to serve.
Where the series stands
Over three parts: a checkpoint route that needs no driver and four tools on it; a stop that can be held in VTL1 kernel mode in a partition we own, and every clause of the published trustlet-debugging capability reproduced from the root; and a live source that reads a running guest, refuses honestly when the hypercall refuses it, and is not a tool because of two bytes in .data.
What is still open is narrow and named. A stop inside an initialized Secure Kernel — the owner-partition probe boots no Windows, and a managed VM's completion stream belongs to vmwp.exe. Whether VTL can be retrofitted onto an existing partition, which remains an inference and not a measurement. And driving DbgEng itself at a live Secure Kernel target, which stays parked behind a two-part condition written down so it is checkable rather than remembered: the EXDI activation stall resolved and the engine's Secure Kernel record turning out reachable after all. Either alone buys nothing.
Closed since, 4 October 2026. The first of those three — a stop inside an initialized Secure Kernel — is now part 9, reached by debugging the managed VM's own vmwp.exe rather than by replacing it; part 8 is the replacement route, built to the edge of a Windows boot and then retired. The other two are still open exactly as written.
Four weeks, roughly thirty gates, two dead guests, one crashed trustlet, one reported SDK defect, and a running tally of measurements that were wrong before they were right. That last category turned out to be the most portable thing here, and it has its own part: how to measure nothing.