Skip to content

hypervisor: kvm: save and restore the guest shadow stack pointer - #4

Closed
tonicmuroq wants to merge 1 commit into
mainfrom
fix/x86-guest-ssp
Closed

tonicmuroq wants to merge 1 commit into
mainfrom
fix/x86-guest-ssp

Conversation

@tonicmuroq

Copy link
Copy Markdown

Why

On a host whose KVM exposes CET shadow stacks to the guest (CPUID.(EAX=7,ECX=0):ECX[7]; seen with Linux 7.0 on an AMD EPYC 7443P / Zen 3), a Windows guest bugchecks 0x50 PAGE_FAULT_IN_NONPAGED_AREA right after being restored from some snapshots. It happened to about one in five of our snapshots, and every restore of a bad snapshot failed the same way. !analyze -v on the minidumps:

  • FAILURE_BUCKET_ID: AV_(null)_nt!KePopulateContinuationContext, process svchost.exe
  • Arg1 0xfffffffffffffff8 (0 − 8), Arg2 0x44 (user-mode, shadow-stack access)
  • the faulting user RIP is a few bytes past the RIP the snapshot saved for the vCPU that was in user mode

A vCPU paused while running a user thread with shadow stacks enabled keeps its live SSP in the SSP register, not in MSR_IA32_PL3_SSP. KVM exposes that register only as KVM_REG_GUEST_SSP (0x2030000300000000) through KVM_{GET,SET}_ONE_REG, and Cloud Hypervisor never saved it. The thread resumed with SSP=0, took a shadow-stack #PF on its next CALL/RET, and the kernel faulted again at 0 − 8 while building the exception's continuation context. Snapshots taken while every vCPU was in the kernel, or running a thread without shadow stacks, were unaffected — hence "some snapshots".

What

  • VcpuKvmState gains guest_ssp: Option<u64> (#[serde(default)], so older snapshots still deserialize).
  • state() reads KVM_REG_GUEST_SSP when the guest CPUID exposes shadow stacks; set_state() writes it back after the MSRs, before the vCPU events — the same place QEMU restores it (i386/kvm: Add save/restore support for KVM_REG_GUEST_SSP).
  • kvm-ioctls 0.25 only exposes get_one_reg/set_one_reg for aarch64 and riscv64, so x86 issues the ioctls directly, like the existing KVM_*_DEVICE_ATTR ones.

Guests whose CPUID has no shadow stacks (every current Intel host we run, and any host whose KVM does not virtualize CET) take the old path unchanged.

Testing

  • cargo fmt --check, and cargo clippy -p hypervisor --features kvm --target x86_64-unknown-linux-gnu -- -D warnings.
  • On the affected host (EPYC 7443P, Linux 7.0.0-31), release builds of this commit and of its parent. A Windows 10 guest ran busy headless Chrome under the patched binary and was saved repeatedly until a snapshot caught a vCPU at CPL3 with U_CET=1: save 9, where the vCPU had PL3_SSP MSR 0xc9513fef40 and guest_ssp register 0xc9513fef78. That one snapshot, restored:
    • with the parent: bugcheck, guest rebooted (agent back after 42.7 s, new LastBootUpTime, BugCheck 1001 logged)
    • with this commit: resumed in 1.3 s, LastBootUpTime unchanged, no bugcheck

On a host whose KVM exposes CET shadow stacks to the guest
(CPUID.(EAX=7,ECX=0):ECX[7]; e.g. Linux 7.0 on AMD Zen 3), a Windows
guest bugchecks 0x50 PAGE_FAULT_IN_NONPAGED_AREA at 0xfffffffffffffff8,
from nt!KePopulateContinuationContext, right after being restored from
some snapshots — about one in five on our node.

A vCPU paused while running a user thread with shadow stacks enabled
keeps its live SSP in the SSP register, not in MSR_IA32_PL3_SSP. KVM
exposes that register only as KVM_REG_GUEST_SSP through
KVM_{GET,SET}_ONE_REG, which Cloud Hypervisor never saved. The thread
therefore resumed with SSP=0, took a shadow-stack #PF on its next
CALL/RET, and the kernel faulted again reading 0 - 8 while building the
exception's continuation context. Snapshots taken while every vCPU was
in the kernel or running a thread without shadow stacks were unaffected,
which is why only some of them failed.

Save the register in VcpuKvmState when the guest CPUID exposes shadow
stacks and restore it after the MSRs, as QEMU does. The field is
optional, so snapshots taken before this change still restore, with the
old behaviour.

Signed-off-by: tonic <tonicbupt@gmail.com>
@CMGS

CMGS commented Sep 23, 2026

Copy link
Copy Markdown
Collaborator

Closing in favour of the upstream PR cloud-hypervisor#8924, sent from the x86-guest-ssp branch. Same fix, rebased on current upstream main, with the private-fn doc comments trimmed and unit tests added for the CPUID check and for restoring a vCPU state saved without guest_ssp. dev already carries it as fbda568 (dev = upstream + CI + diff snapshots + arm64 + this). Fork main stays a pure upstream mirror and dev is rebuilt by rebase, so neither of these is merged directly.

@CMGS CMGS closed this Sep 23, 2026
@tonicmuroq
tonicmuroq deleted the fix/x86-guest-ssp branch September 23, 2026 16:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants