Skip to content

MacBookPro15,2: fixed ~22.6 s stall on every S3 suspend/resume after early CPU offlining (linux-t2 7.2.4 / 7.2.6) #30

Description

@maxyharr

Summary

On a MacBookPro15,2, every S3 (deep) suspend/resume cycle has a fixed ~22.6 s stall that starts right after the early CPU offlining added in t2linux/linux-t2-patches#52 ("ACPI: x86: Apple T2 systems need early CPU offlining"). From the user's side, it takes ~22 s after opening the lid for the screen to come back, regardless of how long the machine slept. Suspend/resume otherwise works.

cc @deqrocks, since the patch notes list MacBookPro15,1 as tested but not 15,2.

Environment

  • Model: MacBookPro15,2 (13", 2018/2019, 4 cores / 8 threads)
  • Firmware: 1916.60.2.0.0 (iBridge: 20.16.2059.0.0,0)
  • Distro: Arch (Omarchy), linux-t2 from arch-mact2
  • Kernels: 7.2.4-arch1-Watanare-T2-3-t2 and 7.2.6-arch2-Watanare-T2-4-t2 (same behavior)
  • mem_sleep: s2idle [deep]
  • Relevant cmdline: intel_iommu=on iommu=pt pm_async=off mem_sleep_default=deep
  • No custom system-sleep hooks unloading modules; t2bce_* suspend/resume natively.

What the logs show

Kernel monotonic timestamps from one cycle on 7.2.6 (lid close → lid open):

[   32.686015] PM: suspend entry (deep)
[   32.786044] Filesystems sync: 0.099 seconds
[   32.796019] smpboot: CPU 1 is now offline
   ... CPUs 2-6 ...
[   32.860319] smpboot: CPU 7 is now offline
               <-- 22.6 s gap -->
[   55.459457] Freezing user space processes
[   55.470275] ACPI: PM: Waking up from system sleep state S3
[   56.115265] smpboot: Booting Node 0 Processor 7 APIC 0x7
[   56.117072] PM: suspend exit

The gap between CPU 7 is now offline and Freezing user space processes was the same in every cycle I measured, whether the machine slept for ~25 s or ~34 min:

CPU 7 offline Freezing user space Gap
2379.363 2401.961 22.60 s
3876.842 3899.452 22.61 s
4266.719 4289.331 22.61 s
4304.514 4327.094 22.58 s
4703.989 4726.574 22.59 s
4892.997 4915.595 22.60 s

With pm_debug_messages=1 and pm_print_times=1, every device suspend/resume callback completes in well under a second. The largest is wiphy_suspend at ~0.5 s. So the time isn't going to device PM callbacks, and the CPU restore after resume is fast.

Since monotonic time doesn't advance during S3, I can't tell from the timestamps alone whether the 22.6 s is spent before entering S3 or during the resume. From the user's side, it shows up as a ~22 s wait after opening the lid. That's close to the "several seconds per CPU" × 7 secondary CPUs that the patch was meant to eliminate. So on this model, the slow bring-up may just be moving elsewhere instead of going away.

Things tried

  • Tracing with function_graph on pm_notifier_call_chain_robust (tracing_thresh=100ms) plus rtcwake -m mem -s 30 hung the machine hard during suspend, which needed a forced power-off, so no trace was captured.

Happy to test a patch or collect more data if you can suggest a way to find what's waiting during that window.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions