[dev] hypervisor: kvm: save and restore the guest shadow stack pointer - #5
Closed
tonicmuroq wants to merge 4 commits into
Closed
tonicmuroq wants to merge 4 commits into
tonicmuroq wants to merge 4 commits into
Conversation
Build cloud-hypervisor on the dev branch for x86_64 and aarch64, then publish both static binaries to a rolling dev prerelease with checksums and build provenance. Signed-off-by: CMGS <ilskdw@gmail.com>
Every snapshot currently dumps the entire guest RAM, so the pause a snapshot imposes grows with the memory size regardless of how little the guest has changed. Iterative checkpointing and warm pre-copy (take a snapshot, let the guest run, take another) repeat that full-memory cost every round. Add a diff snapshot: `vm.snapshot` accepts a `snapshot_type` of `full` (default) or `diff`, exposed as `ch-remote snapshot --diff`. The first diff of a series writes a full baseline and enables dirty-page tracking; each subsequent diff dumps only the pages dirtied since the previous one, so its pause is proportional to the dirtied memory rather than to the guest RAM size. Dirty pages are harvested after the device snapshot, so pages touched by snapshot side effects are included. The whole series lives in one directory, and every diff request of a series must target that directory. A delta is a sparse `memory-ranges.diff.N` whose extents sit at the same offsets as in the baseline `memory-ranges`, written under a temporary name and renamed into place so an interrupted diff never invalidates a consistent directory; `config.json` and `state.json` are replaced the same way and always describe the newest delta. A full dump removes leftover delta files, since restore discovers them by name. A full snapshot, a memory layout change, a migration, restore, or deleting the VM ends the series; the next diff starts a new one. Restore needs no new parameter: when `source_url` points at a directory containing delta files, memory fills from the baseline and each delta's dirty extents replay in sequence order through the same SEEK_DATA walk the eager restore already uses; device state and config come from the newest delta as before. A gap in the sequence or a delta whose length does not match the restored layout is rejected before any page is touched, a filesystem that cannot report extents fails the restore rather than zeroing undirtied pages, and a diff-snapshot series cannot be combined with `memory_restore_mode=ondemand`. Dirty tracking is enabled lazily on the first diff rather than at boot, so a VM that never takes a diff snapshot pays no runtime cost. Validated: an integration test writes tmpfs markers before and after the baseline and reads both back through a restore of the series directory; unit tests cover extent replay over holes, the layout-length check, and delta discovery (ordering, gap detection, ignoring temporary files). On a three-phase workload that dirtied one phase between snapshots, the diff pause measured ~17x shorter than the full one. Signed-off-by: CMGS <ilskdw@gmail.com>
A direct-kernel-booted aarch64 guest only ever saw the FDT, so every
ACPI table create_acpi_tables() had already written went unused. PCI
hotplug is ACPI-only (GED -> PHPR.PSCN -> PCNT -> DVNT/B0EJ), which left
vm.add-net and vm.remove-device returning success while the guest
noticed neither: a hot-added NIC never appeared, and an ejected one
stayed on the bus forever along with its tap.
The arm64 kernel does not need real firmware for this. It takes the RSDP
from the EFI configuration table, which it finds through
`linux,uefi-system-table` and the `linux,uefi-mmap-*` pointers in the
device tree's /chosen node, and it does not care who produced those
structures. So synthesize them:
- arch/src/aarch64/efi.rs builds an EFI System Table, a configuration
table and an EFI memory map.
- fdt::create_stub_fdt() emits a device tree holding only /chosen.
dt_is_stub() treats any other depth-1 node as proof of a real DTB
and disables ACPI, so describing no hardware is what lets the guest
take its hardware from ACPI — and is why acpi=force is not needed.
Selected by `--platform acpi_boot=on|off`, defaulting on for aarch64. It
applies to direct kernel boot only. A firmware boot keeps the full device
tree whatever the option says: the firmware reads its memory and devices
from that tree and publishes ACPI by itself.
The stub tree also carries `linux,uefi-secure-boot = 2`: Ubuntu's kernel
makes it a required property in efi_get_fdt_params(), and without it the
EFI handoff is abandoned, the memory map is never installed, and — with
no /memory node — memblock comes up empty and paging_init panics with
"Failed to allocate page table page". The value is
efi_secureboot_mode_disabled; 0 is efi_secureboot_mode_unset, which the
same kernel reports as "Secure boot could not be determined (mode 0)" at
warning level. Mainline ignores the property.
Four configuration tables are published. ACPI 2.0 carries the RSDP. EFI
RT Properties declares no runtime services, so the guest never calls
into firmware that does not exist. SMBIOS3 is how aarch64 Linux locates
DMI at all, which makes the tables setup_smbios() writes reachable for
the first time on this path. LINUX_EFI_MEMRESERVE lets the GICv3 ITS
persist its LPI property and pending tables through
efi_mem_reserve_persistent(); without it the guest warns twice per boot
out of irq-gic-v3-its.c.
The EFI and SMBIOS ranges are both typed as runtime-services data.
is_usable_memory() hands boot-services ranges back to memblock as System
RAM, and both of these outlive early boot: the kernel keeps appending to
the memreserve table after boot services end, and arm64 reads DMI from
arm_dmi_init(), a core_initcall that runs once the allocator is already
up.
Verified on a c4a-highmem-96-metal host, Ubuntu 24.04 arm64 guest:
- /sys/firmware/{acpi,efi,dmi} present, ACPI0013 GED bound to a GIC
SPI, GICv3/ITS, PSCI, PMU, the SPCR console and DMI all taken from
ACPI
- zero kernel warnings, and no unavailable ranges in the memory map
- vm.add-net: GED interrupt fires, the NIC appears with no guest-side
rescan
- vm.remove-device: the guest ejects it and it leaves CH's device tree
- clone: the new MAC appears by itself and DHCPs a distinct lease, the
snapshot's NIC is gone, and no restore tap is left behind
- acpi_boot=off still boots the same guest through the device tree
Signed-off-by: CMGS <ilskdw@gmail.com>
On a host whose KVM exposes CET shadow stacks to the guest
(CPUID.(EAX=7,ECX=0):ECX[7]; e.g. Linux 7.0 on AMD Zen 3), a Windows
guest bugchecks 0x50 PAGE_FAULT_IN_NONPAGED_AREA at 0xfffffffffffffff8,
from nt!KePopulateContinuationContext, right after being restored from
some snapshots — about one in five on our node.
A vCPU paused while running a user thread with shadow stacks enabled
keeps its live SSP in the SSP register, not in MSR_IA32_PL3_SSP. KVM
exposes that register only as KVM_REG_GUEST_SSP through
KVM_{GET,SET}_ONE_REG, which Cloud Hypervisor never saved. The thread
therefore resumed with SSP=0, took a shadow-stack #PF on its next
CALL/RET, and the kernel faulted again reading 0 - 8 while building the
exception's continuation context. Snapshots taken while every vCPU was
in the kernel or running a thread without shadow stacks were unaffected,
which is why only some of them failed.
Save the register in VcpuKvmState when the guest CPUID exposes shadow
stacks and restore it after the MSRs, as QEMU does. The field is
optional, so snapshots taken before this change still restore, with the
old behaviour.
Signed-off-by: tonic <tonicbupt@gmail.com>
tonicmuroq
force-pushed
the
dev-x86-guest-ssp
branch
from
September 23, 2026 14:58
c8a12a5 to
31ac241
Compare
Collaborator
|
Closing in favour of the upstream PR cloud-hypervisor#8924, sent from the |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Cherry-pick of #4 onto dev, so the fleet build carries it. Same commit; see #4 for the root cause (guest SSP register lost across snapshot/restore on hosts that expose CET shadow stacks → Windows 0x50 after restore) and the A/B test on the affected AMD host.