From 7b0bfa959990943082948daf0e9d7999bcadfc15 Mon Sep 17 00:00:00 2001 From: Mark Shannon Date: Mon, 28 Sep 2026 12:31:36 +0100 Subject: [PATCH] PEP 848: Clarifications and bug fixes Fix bugs in the Python code for the algorithm and clarify that code Remove from future work, work that has already been done. Improve explanation of why performance is better --- peps/pep-0848.rst | 68 +++++++++++++++++++++++------------------------ 1 file changed, 33 insertions(+), 35 deletions(-) diff --git a/peps/pep-0848.rst b/peps/pep-0848.rst index 80dc3d16c80..1a0d93d2b50 100644 --- a/peps/pep-0848.rst +++ b/peps/pep-0848.rst @@ -201,7 +201,7 @@ When the nursery is full, the unreachable cycles in the oldest aging space reachable to visited. They are reachable and cannot be garbage. """ moved_to_visited = 0 - while reachable: + while reachable and moved_to_visited < limit: root = reachable.pop() visited.append(root) moved_to_visited += 1 @@ -209,8 +209,6 @@ When the nursery is full, the unreachable cycles in the oldest aging space if obj in pending: pending.remove(obj) reachable.append(obj) - if moved_to_visited >= limit: - return moved_to_visited return moved_to_visited def old_collection(): @@ -230,14 +228,14 @@ When the nursery is full, the unreachable cycles in the oldest aging space # form transitive closure starting at obj, taking objects from pending increment = form_transitive_closure(obj, pending) candidates = len(increment) + work_to_do -= candidates survivors = collect_cycles(increment) - work_to_do -= survivors - collected = candidates - survivors + collected = candidates - len(survivors) + visited_space.extend(survivors) # If we are collecting lots of objects, that means - # there is a lot of cycle garbage and we need to - # sweep the heap faster. + # there is a lot of cycle garbage and we should sweep + # the heap faster to keep the amount of garbage down work_to_do += 2 * collected - visited_space.extend(survivors) The legacy collector -------------------- @@ -283,11 +281,17 @@ of half spaces is always rounded up to an even number. Performance ----------- -Performance is improved relative to the current collector. -The performance improvements come from doing less work in the young -generations (one collection per object, not two) and doing less work -in the old generation due to the lower survivor rate from the young -generations. +The new collector reduces the overhead of cyclic garbage collection by almost +half, although the exact amount depends on the application. + +By allowing objects longer to die, the effectiveness of the collector is +improved. This allows it to collect the same amount of garbage for less work. +Performance is further improved by scanning fewer objects during collections: + +* In the young generation: each object is only scanned once, instead of twice + in the generational GC +* In the old generation: objects are scanned at a lower rate, only increasing + that rate when necessary to collect excess garbage Peak Memory Consumption ----------------------- @@ -322,10 +326,10 @@ Calling ``gc.collect()`` is equivalent to calling ``gc.collect(2)``. ================== ============================================== Argument Effect ================== ============================================== - 0 Perform a young collection, - collecting the oldest aging space - 1 Collect an increment of the old generation - 2 or no argument Collect the whole heap + 0 Perform a young collection, + collecting the oldest aging space + 1 Collect an increment of the old generation + 2 or no argument Collect the whole heap ================== ============================================== Choosing the legacy generational collector @@ -398,24 +402,6 @@ collections will usually mean shorter pauses per collection. Future work =========== -Further reducing pause times ----------------------------- - -While the reference implementation is 1-2% faster than main (with the -generational GC), it can still have long pause times on large object graphs. -Many of the benchmarks have a single large tree as their object graph, and this -can result in long pauses, as an increment starting at the root of the tree -will contain almost the whole heap. - -This could be improved in a few ways: - -* Sorting the increments as they are either created or sent to the old - generation, so that the objects farthest from the root are picked first in - the next collection. -* Traversing the stack prior to increment formation to skip reachable objects. - This will complicate the algorithm, but could save significant amounts of - work in some cases. - Porting to the free-threaded build ---------------------------------- @@ -431,6 +417,18 @@ Porting the incremental GC to the free-threaded build will need a few changes: external arrays for the young generation and the increments. The free-threaded GC already needs to do this to partition garbage and survivors. +Further reducing pause times +---------------------------- + +While the reference implementation generally has shorter pause times, it +can still have long pause times if the transitive closure needed for an +increment is large. + +It may be possible to scan increments over multiple collections, keeping +each pause short. This would be challenging as the program may transform the +object graph of the increment between collections, but garbage cycles cannot be +modified by the program so this might be possible. There is extensive research +on concurrent collectors which also have to handle similar problems. Reference Implementation ========================