Skip to content

Performance regression in anext() after moving its implementation from C to Python #158407

Description

@leandrodamascena

Bug report

Bug description:

I have a project where I build CPython from individual commits into OCI images and run them on AWS Lambda to detect performance changes before they reach a runtime release.

While comparing merged commits on main, I found that #157362, commit 0a6c1ed34118b091230ee38fc047bcc2df8e5c5e, makes explicit calls to anext() slower. I compared that commit directly with its parent, 8df6f077115484361b40952751fb7d703843c192.

The change is visible with a regular GIL build and does not require free-threading.

Reproducer:

import asyncio
import statistics
import time

ITEMS = 1_000_000
REPEATS = 7


async def items():
    for value in range(ITEMS):
        yield value


async def consume_with_anext():
    stream = items()
    started = time.perf_counter_ns()
    for _ in range(ITEMS):
        await anext(stream)
    return (time.perf_counter_ns() - started) / ITEMS


async def consume_with_async_for():
    started = time.perf_counter_ns()
    async for _ in items():
        pass
    return (time.perf_counter_ns() - started) / ITEMS


async def main():
    for name, benchmark in (
        ("anext", consume_with_anext),
        ("async for", consume_with_async_for),
    ):
        samples = [await benchmark() for _ in range(REPEATS)]
        print(f"{name}: {statistics.median(samples):.1f} ns/item")


asyncio.run(main())

On macOS arm64, using identically configured builds and seven interleaved outer repetitions, I measured await anext(async_generator) as:

Revision Median
parent 8df6f0771154 141.8 ns/item
0a6c1ed34118 228.4 ns/item

This is a 61.1% regression. The async for control changed by 0.7%.

I also isolated the builtin from the async generator itself:

Revision await generator.__anext__() await anext(generator)
parent 170.7 ns/item 137.0 ns/item
commit 167.3 ns/item 215.4 ns/item

The direct method did not regress. The additional time comes from the new Python implementation of the builtin, which performs type() and __anext__ lookup followed by a Python call, while the previous C implementation called the async iteration slot directly.

I reproduced the result in two independent AWS Lambda experiments using optimized x86_64 CPython builds with LTO. Each experiment used 12 environments and 24 paired warm samples:

Experiment Parent median Commit median Control-adjusted regression
1 180.0 ms 248.1 ms 39.8%
2 179.3 ms 245.0 ms 37.0%

This affects code that manually consumes ready or buffered async iterators with await anext(stream). This benchmark isolates CPU overhead and does not model network latency. Code using async for does not use this path.

CPython versions tested on:

CPython main branch

Operating systems tested on:

Linux, macOS

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    interpreter-core(Objects, Python, Grammar, and Parser dirs)performancePerformance or resource usagetype-featureA feature request or enhancement

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions