← writing

One branch, several compares: reading A64 conditional-compare chains

4 October 2026 · Windows internals · A windbg-mcp note, outside the Secure Kernel series

A driver's IOCTL map said seventeen control codes. The driver accepts twenty-one. The four it never named were folded into ccmp chains — and the instruction hiding them is also the one that says whether reading them is sound.

windbg-mcp's ioctl_map walks a driver's dispatch routine and answers with the control codes it accepts, each with the block it routes to, plus an untracked list of the sites where the walk lost the code. That second list is the whole contract: a map that is silently short is worse than a map that says it is short, because the only thing anyone does with a complete-looking case list is treat it as the attack surface.

On x64, if (code == A || code == B || code == C) is three compares and three branches, and a walk that reads one comparison per branch gets all three. A64 has ccmp, so the same source is one branch fed by several compares. This walk read the instruction before the branch, filed the rest in untracked, and lost every other code in the chain. Measured on the live ARM64 target against windbg-mcp 0.18.0+g30c4af94:

rdyboost+0xef00  mov    w11,#0xC008 / movk w11,#0x56,lsl #0x10   ; w11 = 0x0056c008
rdyboost+0xef08  mov    w10,#0xA0   / movk w10,#7,lsl #0x10      ; w10 = 0x000700a0
rdyboost+0xef10  cmp    w8,w11
rdyboost+0xef14  ccmpne w8,w12,#4
rdyboost+0xef18  ccmpne w8,w10,#4
rdyboost+0xef1c  beq    rdyboost+0xef7c

Three codes, one handler, and the map named none of them: untracked carried 0xef18. The driver's second chain is the same instruction with the opposite meaning — cmp w8,#0 / ccmpne w8,w10,#0 / bne, with w10 = 0x00224194.

The nzcv immediate is the whole thing

ccmp Wn, Wm, #nzcv, cond does one of two things. If cond holds it performs the comparison and writes real flags. If it does not, it forces the flag bits to the immediate. So the immediate is not a detail of the encoding; it is the chain's operator.

With ne as the condition — which is what a compiler writes for a disjunction — and reading nzcv as the encoding holds it, with N at bit 3 and Z at bit 2:

link immediate what an earlier match forces so the branch below and a case means
ccmpne Wn,Wm,#4 Z set equality set b.eq is taken by a match at any link every readable link is a case, at its own site
ccmpne Wn,Wm,#0 all clear equality clear b.ne is taken away from the case only this link's own operand can be a case

Which makes the two chains above the same shape with inverted meanings, and makes the failure mode specific: read the first shape onto the second and you publish the arm the routine rejects. For rdyboost that arm is 0x00000000, which is exactly what the Binary Ninja companion's map carries for this driver (binja-windbg-mcp#14). So each shape is pinned by a test of its own rather than by one test and an argument.

The remedy the filed item proposed was the wrong one

The item said reading a chain means "folding several comparisons into one branch's condition". It does not. This walk already carries a set of live comparisons and already makes a case per site; what a chain needs is for that set to survive a flag write, and the nzcv bit is what decides whether it does. Nothing about the derived Condition changed.

So the fold is two functions. chained_compare recognises the instruction and asks the ordinary comparison reader — through a probe, so the width guard, the shifted-index rule and the code arithmetic are the ones every other compare gets — and absorb is the rule:

A flag write replaces the set of live compares with the one comparison it leaves. A conditional compare continues the set instead, and the nzcv immediate says whether it does.

That is the entire mechanism, and it was 800 lines. Everything below is what it took to make it safe, which was nearly three times that again.

A kept reading is good for an equality and nothing else

#4 forces Z alone — so C is left clear, which on A64 is borrow. A b.lo after the chain is therefore taken where the comparison it stands in for would not have been. The forced flags are not the flags that compare would have left, so a chained reading carried onto an edge would be evaluated against a state the target never reached, and a case built there is a code the routine sent somewhere else.

So chained readings are marked as such, and the marking costs them everything except the one question they can answer:

That last rule costs a real code on the b.lo shape. It is the right way round, and it is pinned as a characterisation rather than left implied, because the alternative is a map whose completeness depends on a flag nobody checked.

Then fourteen review rounds, and the arithmetic of them

The change was 800 lines. The rounds that followed added 2,331 more — measured as a diff of src/ioctl.rs from the first commit to the last, rather than by adding the rounds up, which gives a different number. (The record's own figure of 820 is the same measurement taken four rounds in, which is worth keeping beside this one.) Nineteen findings in all: seventeen from Codex (eleven P1, six P2) and two from CodeRabbit on the final head. Five were declined, and one of those declines was reversed a round later. Two of the rounds were against a previous round's own fix.

And the answer never moved. rdyboost reported 21 codes with nothing untracked at the first commit and reported the same 21 after the last rule, re-measured rather than recalled, at rounds 1, 9, 10, 13 and 14.

Four of those rounds are worth the space.

Count the arms instead of patching the ones you were handed

Round 1's P1: the fold needs a predecessor this walk read. ccmpne …,#4 says "if the comparison before me did not match, compare; otherwise call it equal" — so where that earlier comparison is not one of the walk's own readings, as in cmp w2,#0 / ccmpne w9,w10,#4 / b.eq, the branch below is decided by a condition about some other value and the forced arm sends every control code to the handler. Publishing w10's code there is not false; publishing it with an empty untracked reads as the whole set, which is the one thing this module must not do. So a link with no readable predecessor is dropped and its site becomes the loss.

Round 2 brought two more P1s that were the same class — a forced arm read over flags the walk cannot evaluate. Patching the two named shapes would have been two more commits and no reason to think there were not more. So the arms were counted instead: the pair (does the nzcv keep the earlier links, were they read) against (was this link's own comparison read) has six cases, and three of them need the site marked.

earlier links this link's compare what it owes
forced set, not read — the link cannot be folded at all; its site is the loss
forced set, read blind the read links are still true cases, so the site is a loss beside them
forced clear, a predecessor established the code blind the case is unknown; the site is recorded
forced set, read read nothing — this is the chain the item was filed about
forced clear read nothing — only this link's code can be a case
forced clear, no predecessor blind nothing — nothing here is about the control code at all

The third row was named by no finding. Only the count found it. That is the argument for enumerating a class rather than working a queue: two findings of one class are evidence that the class has members nobody has reported yet, and the cheapest way to find them is to write the cases down.

A class fix closes the class only if it can express the whole rule

Round 3's P1 was filed against round 2's own fix, and it is the shape worth keeping. Round 2 recorded a dropped chain by searching the readings that survived it — and the forced-clear blind arm it added in the same commit clears those readings, leaving the chain represented by its loss alone. So cmp code,A / ccmpne w2,w3,#0 / b.lo was still a complete-looking map with an accepted code in it.

What the check could not express was a chain with no reading left. So the chain became a fact the block keeps in its own right — the site of the conditional compare whose flags are live, cleared by an ordinary flag write and by a call exactly as the readings are — and the terminator asks that, then commits whichever of the reading or the loss the chain left.

A decline is a price, and a price can be wrong

Round 9 declined a P1 and recorded the limit instead: a ccmn compares the control code against the negation of its operand, and a forced-clear link after one replaced the loss it had left, so cmp code,A / ccmnne code,C,#4 / ccmpne code,B,#0 / b.eq published B without saying B is rejected where B == -C. The remedy wanted a fifth distinction on that seam, for a shape needing a negated compare against a control code. Declined as machinery, pinned as a known limit, and that round's commit subject says "the rounds stop here".

Round 10 reached the same gap through tst code,#1 / ccmpne code,B,#0 / b.eq, which admits only an odd code — so an even B is an invented case, and a bit test before a conditional compare is ordinary codegen rather than an exotic shape. The decline was wrong on price, not on fact.

And once repriced it was not a fifth distinction at all, but what three findings on that line had been about: a loss about the control code and a forced arm that admits every code shared one slot and want opposite treatment. A forced-clear link resolves the second — an earlier match takes the branch away from the case — and resolves nothing about a code nobody could name. Separated, each direction is pinned by its own mutation, which is the evidence that they really are two things.

The decline that held is worth as much. Round 6 argued that the published case is conditional on an unread comparison, which it is — and what settles it is a control: the same guard written without a conditional compare (cmp w2,w3 / b.ne skip / cmp code,B / b.eq) publishes the code and marks nothing, because the guard tests a value that is not the control code. Marking the folded form would make the map's completeness depend on the compiler's instruction selection rather than on the driver, and would make an untracked entry of every input-dependent guard a compiler happened to fold. a_guard_on_something_else_is_not_a_loss_however_it_is_written is that control, and it sits beside the ccmp fixture it controls for so that a later round cannot "fix" it.

When you separate two conflated things, every site that writes either is in scope

Round 10's separation was applied to the arm the finding named and to no other. So a forced-set link with no predecessor still assigned the loss slot — dropping a tst's site, after which the forced-clear arm had nothing left to preserve — and a readable link still overwrote the forced slot with nothing, erasing the site a blind link before it had left. Both published an invented case as definitive. Both came back as P1s in the next round.

The remedy was to enumerate all four sites that write either slot and say what each owes: an ordinary flag write replaces the loss and clears the forced arm, because the flags it replaces are what both were about; a forced-set link keeps both; a forced-clear link keeps the loss and resolves the forced arm. The defect is not at the site where it was seen.

And a join unions the readings

Two findings, two rounds apart, from one fact: the walk's join unions pending comparisons deliberately, because a comparison is a claim about the path that made it. So a non-empty set at a conditional compare proves that one incoming path compared the control code, and proves nothing about the others — with cmp code,A on one edge and cmp w2,#0 on another, a Z-forcing link admits every code along the second. The fold therefore asks for flags this block wrote, which costs a chain whose first link is in another block the codes it used to name, and gives it the site instead.

The same union sharpened the opposite rule, twice. A forced-clear link whose predecessor matched the same code is dead code — cmp code,A / ccmpne code,A,#0 / b.ne rejects every value — so the case is dropped; but finding that code on one predecessor does not make the case unreachable on every path, so the case goes only where every incoming reading named a code and every one of them is this link's. Then round 13: every path is the right test only for readings that arrived from other paths. Within one chain, any earlier link matching the code excludes it, because those links are on the same path — cmp code,A / ccmpne code,B,#4 / ccmpne code,A,#0 / b.ne makes equality mean A || B by the time the last condition is read, so every value takes the branch and the fall-through is dead. The marking that distinguishes a chained reading from a carried-in one is what tells the two apart.

Both of those were P2s found by reconciling the review, not by a notification: counting the comments on the pull request by head came to seventeen against the fifteen that had been worked, because two were anchored to a lint commit reviewed instead of the round-9 commit beside it.

The measurement, and why it was taken twice

End to end through the tool surface, against the debugger guest's own rdyboost.sys — 10.0.26100.1, SHA-256 D872CFF7…, opened as an image target, dispatch rdyboost+0xf6a0, 2026-10-03:

build codes untracked
0.21.0+gdaa11a3d (without the fold) 17 rdyboost+0xf800, rdyboost+0xf708
0.21.0+gc02152aa (with it) 21 none

The four gained are exactly 0x0056c008, 0x00070000, 0x000700a0 and 0x00224194, and nothing was lost.

Both builds are named and clean, and that is a re-take. The first pair of readings was 0.20.0+g3b8d6fee against a -dirty tree, which is a measurement nobody can reproduce; and by the time this merged, main had moved seventeen commits, one of which had touched ioctl_map's own rendering. Re-taking the baseline after that rebase is what makes the four codes attributable to this change rather than to the release — it came back identical, sites included.

And this is not the build the item was filed against. Its chains are at rdyboost+0xef00 and +0xf00c; this build's are at +0xf708 and +0xf800, and the file hashes differ. On this build the link whose register the live-target reading could not resolve is readable, so the fold recovered four codes where the filed item named three. Saying so costs nothing and is the only thing that stops the next reader treating the two as one measurement.

Twenty-six tests landed with it, most mutation-verified against the mutation they exist for — three of them deliberately controls, so that a later round does not "fix" something already right.

What is still open

The other half of the original disagreement. A Binary Ninja companion runs the same question through a different engine, and it publishes the chain's first operand while missing the ccmp operands — so the cross-tool agreement lane now disagrees about rdyboost in the other direction (binja-windbg-mcp#14). That lane has not been re-run against this fix, which needs a fresh companion capture of the build the debuggee is actually running.

The one line worth keeping

The instruction that was hiding the codes is the same instruction that says whether you may read them. #4 and #0 differ by one bit of an immediate most disassembler output prints without comment, and that bit decides whether a link is a case the driver accepts or the arm it rejects. Reading the chain is one rule, and it is the sentence in the block quote above. Knowing which of the six ways it can be unreadable is what took fourteen rounds — and it is the only half of the work that was ever going to cost somebody a wrong attack surface.

The walk is ioctl_map in windbg-mcp, the fold is chained_compare and absorb in src/ioctl.rs, and the review is #439.