Skip to content

[dev] hypervisor: kvm: save and restore the guest shadow stack pointer - #5

Closed
tonicmuroq wants to merge 4 commits into
devfrom
dev-x86-guest-ssp
Closed

tonicmuroq wants to merge 4 commits into
devfrom
dev-x86-guest-ssp

Conversation

@tonicmuroq

@tonicmuroq tonicmuroq commented Sep 23, 2026 •

Copy link
Copy Markdown

Cherry-pick of #4 onto dev, so the fleet build carries it. Same commit; see #4 for the root cause (guest SSP register lost across snapshot/restore on hosts that expose CET shadow stacks → Windows 0x50 after restore) and the A/B test on the affected AMD host.

CMGS and others added 4 commits September 23, 2026 17:46
Build cloud-hypervisor on the dev branch for x86_64 and aarch64, then
publish both static binaries to a rolling dev prerelease with checksums
and build provenance.

Signed-off-by: CMGS <ilskdw@gmail.com>
Every snapshot currently dumps the entire guest RAM, so the pause a
snapshot imposes grows with the memory size regardless of how little
the guest has changed. Iterative checkpointing and warm pre-copy (take
a snapshot, let the guest run, take another) repeat that full-memory
cost every round.

Add a diff snapshot: `vm.snapshot` accepts a `snapshot_type` of `full`
(default) or `diff`, exposed as `ch-remote snapshot --diff`. The first
diff of a series writes a full baseline and enables dirty-page
tracking; each subsequent diff dumps only the pages dirtied since the
previous one, so its pause is proportional to the dirtied memory
rather than to the guest RAM size. Dirty pages are harvested after the
device snapshot, so pages touched by snapshot side effects are
included.

The whole series lives in one directory, and every diff request of a
series must target that directory. A delta is a sparse
`memory-ranges.diff.N` whose extents sit at the same offsets as in the
baseline `memory-ranges`, written under a temporary name and renamed
into place so an interrupted diff never invalidates a consistent
directory; `config.json` and `state.json` are replaced the same way
and always describe the newest delta. A full dump removes leftover
delta files, since restore discovers them by name. A full snapshot, a
memory layout change, a migration, restore, or deleting the VM ends
the series; the next diff starts a new one.

Restore needs no new parameter: when `source_url` points at a
directory containing delta files, memory fills from the baseline and
each delta's dirty extents replay in sequence order through the same
SEEK_DATA walk the eager restore already uses; device state and config
come from the newest delta as before. A gap in the sequence or a delta
whose length does not match the restored layout is rejected before any
page is touched, a filesystem that cannot report extents fails the
restore rather than zeroing undirtied pages, and a diff-snapshot
series cannot be combined with `memory_restore_mode=ondemand`.

Dirty tracking is enabled lazily on the first diff rather than at
boot, so a VM that never takes a diff snapshot pays no runtime cost.

Validated: an integration test writes tmpfs markers before and after
the baseline and reads both back through a restore of the series
directory; unit tests cover extent replay over holes, the
layout-length check, and delta discovery (ordering, gap detection,
ignoring temporary files). On a three-phase workload that dirtied one
phase between snapshots, the diff pause measured ~17x shorter than the
full one.

Signed-off-by: CMGS <ilskdw@gmail.com>
A direct-kernel-booted aarch64 guest only ever saw the FDT, so every
ACPI table create_acpi_tables() had already written went unused. PCI
hotplug is ACPI-only (GED -> PHPR.PSCN -> PCNT -> DVNT/B0EJ), which left
vm.add-net and vm.remove-device returning success while the guest
noticed neither: a hot-added NIC never appeared, and an ejected one
stayed on the bus forever along with its tap.

The arm64 kernel does not need real firmware for this. It takes the RSDP
from the EFI configuration table, which it finds through
`linux,uefi-system-table` and the `linux,uefi-mmap-*` pointers in the
device tree's /chosen node, and it does not care who produced those
structures. So synthesize them:

  - arch/src/aarch64/efi.rs builds an EFI System Table, a configuration
    table and an EFI memory map.
  - fdt::create_stub_fdt() emits a device tree holding only /chosen.
    dt_is_stub() treats any other depth-1 node as proof of a real DTB
    and disables ACPI, so describing no hardware is what lets the guest
    take its hardware from ACPI — and is why acpi=force is not needed.

Selected by `--platform acpi_boot=on|off`, defaulting on for aarch64. It
applies to direct kernel boot only. A firmware boot keeps the full device
tree whatever the option says: the firmware reads its memory and devices
from that tree and publishes ACPI by itself.

The stub tree also carries `linux,uefi-secure-boot = 2`: Ubuntu's kernel
makes it a required property in efi_get_fdt_params(), and without it the
EFI handoff is abandoned, the memory map is never installed, and — with
no /memory node — memblock comes up empty and paging_init panics with
"Failed to allocate page table page". The value is
efi_secureboot_mode_disabled; 0 is efi_secureboot_mode_unset, which the
same kernel reports as "Secure boot could not be determined (mode 0)" at
warning level. Mainline ignores the property.

Four configuration tables are published. ACPI 2.0 carries the RSDP. EFI
RT Properties declares no runtime services, so the guest never calls
into firmware that does not exist. SMBIOS3 is how aarch64 Linux locates
DMI at all, which makes the tables setup_smbios() writes reachable for
the first time on this path. LINUX_EFI_MEMRESERVE lets the GICv3 ITS
persist its LPI property and pending tables through
efi_mem_reserve_persistent(); without it the guest warns twice per boot
out of irq-gic-v3-its.c.

The EFI and SMBIOS ranges are both typed as runtime-services data.
is_usable_memory() hands boot-services ranges back to memblock as System
RAM, and both of these outlive early boot: the kernel keeps appending to
the memreserve table after boot services end, and arm64 reads DMI from
arm_dmi_init(), a core_initcall that runs once the allocator is already
up.

Verified on a c4a-highmem-96-metal host, Ubuntu 24.04 arm64 guest:

  - /sys/firmware/{acpi,efi,dmi} present, ACPI0013 GED bound to a GIC
    SPI, GICv3/ITS, PSCI, PMU, the SPCR console and DMI all taken from
    ACPI
  - zero kernel warnings, and no unavailable ranges in the memory map
  - vm.add-net: GED interrupt fires, the NIC appears with no guest-side
    rescan
  - vm.remove-device: the guest ejects it and it leaves CH's device tree
  - clone: the new MAC appears by itself and DHCPs a distinct lease, the
    snapshot's NIC is gone, and no restore tap is left behind
  - acpi_boot=off still boots the same guest through the device tree

Signed-off-by: CMGS <ilskdw@gmail.com>
On a host whose KVM exposes CET shadow stacks to the guest
(CPUID.(EAX=7,ECX=0):ECX[7]; e.g. Linux 7.0 on AMD Zen 3), a Windows
guest bugchecks 0x50 PAGE_FAULT_IN_NONPAGED_AREA at 0xfffffffffffffff8,
from nt!KePopulateContinuationContext, right after being restored from
some snapshots — about one in five on our node.

A vCPU paused while running a user thread with shadow stacks enabled
keeps its live SSP in the SSP register, not in MSR_IA32_PL3_SSP. KVM
exposes that register only as KVM_REG_GUEST_SSP through
KVM_{GET,SET}_ONE_REG, which Cloud Hypervisor never saved. The thread
therefore resumed with SSP=0, took a shadow-stack #PF on its next
CALL/RET, and the kernel faulted again reading 0 - 8 while building the
exception's continuation context. Snapshots taken while every vCPU was
in the kernel or running a thread without shadow stacks were unaffected,
which is why only some of them failed.

Save the register in VcpuKvmState when the guest CPUID exposes shadow
stacks and restore it after the MSRs, as QEMU does. The field is
optional, so snapshots taken before this change still restore, with the
old behaviour.

Signed-off-by: tonic <tonicbupt@gmail.com>
@CMGS

CMGS commented Sep 23, 2026

Copy link
Copy Markdown
Collaborator

Closing in favour of the upstream PR cloud-hypervisor#8924, sent from the x86-guest-ssp branch. Same fix, rebased on current upstream main, with the private-fn doc comments trimmed and unit tests added for the CPUID check and for restoring a vCPU state saved without guest_ssp. dev already carries it as fbda568 (dev = upstream + CI + diff snapshots + arm64 + this). Fork main stays a pure upstream mirror and dev is rebuilt by rebase, so neither of these is merged directly.

@CMGS CMGS closed this Sep 23, 2026
@tonicmuroq
tonicmuroq deleted the dev-x86-guest-ssp branch September 23, 2026 16:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants