diff --git a/peps/pep-0848.rst b/peps/pep-0848.rst index 80dc3d16c80..1a0d93d2b50 100644 --- a/peps/pep-0848.rst +++ b/peps/pep-0848.rst @@ -201,7 +201,7 @@ When the nursery is full, the unreachable cycles in the oldest aging space reachable to visited. They are reachable and cannot be garbage. """ moved_to_visited = 0 - while reachable: + while reachable and moved_to_visited < limit: root = reachable.pop() visited.append(root) moved_to_visited += 1 @@ -209,8 +209,6 @@ When the nursery is full, the unreachable cycles in the oldest aging space if obj in pending: pending.remove(obj) reachable.append(obj) - if moved_to_visited >= limit: - return moved_to_visited return moved_to_visited def old_collection(): @@ -230,14 +228,14 @@ When the nursery is full, the unreachable cycles in the oldest aging space # form transitive closure starting at obj, taking objects from pending increment = form_transitive_closure(obj, pending) candidates = len(increment) + work_to_do -= candidates survivors = collect_cycles(increment) - work_to_do -= survivors - collected = candidates - survivors + collected = candidates - len(survivors) + visited_space.extend(survivors) # If we are collecting lots of objects, that means - # there is a lot of cycle garbage and we need to - # sweep the heap faster. + # there is a lot of cycle garbage and we should sweep + # the heap faster to keep the amount of garbage down work_to_do += 2 * collected - visited_space.extend(survivors) The legacy collector -------------------- @@ -283,11 +281,17 @@ of half spaces is always rounded up to an even number. Performance ----------- -Performance is improved relative to the current collector. -The performance improvements come from doing less work in the young -generations (one collection per object, not two) and doing less work -in the old generation due to the lower survivor rate from the young -generations. +The new collector reduces the overhead of cyclic garbage collection by almost +half, although the exact amount depends on the application. + +By allowing objects longer to die, the effectiveness of the collector is +improved. This allows it to collect the same amount of garbage for less work. +Performance is further improved by scanning fewer objects during collections: + +* In the young generation: each object is only scanned once, instead of twice + in the generational GC +* In the old generation: objects are scanned at a lower rate, only increasing + that rate when necessary to collect excess garbage Peak Memory Consumption ----------------------- @@ -322,10 +326,10 @@ Calling ``gc.collect()`` is equivalent to calling ``gc.collect(2)``. ================== ============================================== Argument Effect ================== ============================================== - 0 Perform a young collection, - collecting the oldest aging space - 1 Collect an increment of the old generation - 2 or no argument Collect the whole heap + 0 Perform a young collection, + collecting the oldest aging space + 1 Collect an increment of the old generation + 2 or no argument Collect the whole heap ================== ============================================== Choosing the legacy generational collector @@ -398,24 +402,6 @@ collections will usually mean shorter pauses per collection. Future work =========== -Further reducing pause times ----------------------------- - -While the reference implementation is 1-2% faster than main (with the -generational GC), it can still have long pause times on large object graphs. -Many of the benchmarks have a single large tree as their object graph, and this -can result in long pauses, as an increment starting at the root of the tree -will contain almost the whole heap. - -This could be improved in a few ways: - -* Sorting the increments as they are either created or sent to the old - generation, so that the objects farthest from the root are picked first in - the next collection. -* Traversing the stack prior to increment formation to skip reachable objects. - This will complicate the algorithm, but could save significant amounts of - work in some cases. - Porting to the free-threaded build ---------------------------------- @@ -431,6 +417,18 @@ Porting the incremental GC to the free-threaded build will need a few changes: external arrays for the young generation and the increments. The free-threaded GC already needs to do this to partition garbage and survivors. +Further reducing pause times +---------------------------- + +While the reference implementation generally has shorter pause times, it +can still have long pause times if the transitive closure needed for an +increment is large. + +It may be possible to scan increments over multiple collections, keeping +each pause short. This would be challenging as the program may transform the +object graph of the increment between collections, but garbage cycles cannot be +modified by the program so this might be possible. There is extensive research +on concurrent collectors which also have to handle similar problems. Reference Implementation ========================