Bug report
Bug description:
I have a project where I build CPython from individual commits into OCI images and run them on AWS Lambda to detect performance changes before they reach a runtime release.
While comparing merged commits on main, I found that #157362, commit 0a6c1ed34118b091230ee38fc047bcc2df8e5c5e, makes explicit calls to anext() slower. I compared that commit directly with its parent, 8df6f077115484361b40952751fb7d703843c192.
The change is visible with a regular GIL build and does not require free-threading.
Reproducer:
import asyncio
import statistics
import time
ITEMS = 1_000_000
REPEATS = 7
async def items():
for value in range(ITEMS):
yield value
async def consume_with_anext():
stream = items()
started = time.perf_counter_ns()
for _ in range(ITEMS):
await anext(stream)
return (time.perf_counter_ns() - started) / ITEMS
async def consume_with_async_for():
started = time.perf_counter_ns()
async for _ in items():
pass
return (time.perf_counter_ns() - started) / ITEMS
async def main():
for name, benchmark in (
("anext", consume_with_anext),
("async for", consume_with_async_for),
):
samples = [await benchmark() for _ in range(REPEATS)]
print(f"{name}: {statistics.median(samples):.1f} ns/item")
asyncio.run(main())
On macOS arm64, using identically configured builds and seven interleaved outer repetitions, I measured await anext(async_generator) as:
| Revision |
Median |
parent 8df6f0771154 |
141.8 ns/item |
0a6c1ed34118 |
228.4 ns/item |
This is a 61.1% regression. The async for control changed by 0.7%.
I also isolated the builtin from the async generator itself:
| Revision |
await generator.__anext__() |
await anext(generator) |
| parent |
170.7 ns/item |
137.0 ns/item |
| commit |
167.3 ns/item |
215.4 ns/item |
The direct method did not regress. The additional time comes from the new Python implementation of the builtin, which performs type() and __anext__ lookup followed by a Python call, while the previous C implementation called the async iteration slot directly.
I reproduced the result in two independent AWS Lambda experiments using optimized x86_64 CPython builds with LTO. Each experiment used 12 environments and 24 paired warm samples:
| Experiment |
Parent median |
Commit median |
Control-adjusted regression |
| 1 |
180.0 ms |
248.1 ms |
39.8% |
| 2 |
179.3 ms |
245.0 ms |
37.0% |
This affects code that manually consumes ready or buffered async iterators with await anext(stream). This benchmark isolates CPU overhead and does not model network latency. Code using async for does not use this path.
CPython versions tested on:
CPython main branch
Operating systems tested on:
Linux, macOS
Bug report
Bug description:
I have a project where I build CPython from individual commits into OCI images and run them on AWS Lambda to detect performance changes before they reach a runtime release.
While comparing merged commits on
main, I found that #157362, commit0a6c1ed34118b091230ee38fc047bcc2df8e5c5e, makes explicit calls toanext()slower. I compared that commit directly with its parent,8df6f077115484361b40952751fb7d703843c192.The change is visible with a regular GIL build and does not require free-threading.
Reproducer:
On macOS arm64, using identically configured builds and seven interleaved outer repetitions, I measured
await anext(async_generator)as:8df6f07711540a6c1ed34118This is a 61.1% regression. The
async forcontrol changed by 0.7%.I also isolated the builtin from the async generator itself:
await generator.__anext__()await anext(generator)The direct method did not regress. The additional time comes from the new Python implementation of the builtin, which performs
type()and__anext__lookup followed by a Python call, while the previous C implementation called the async iteration slot directly.I reproduced the result in two independent AWS Lambda experiments using optimized x86_64 CPython builds with LTO. Each experiment used 12 environments and 24 paired warm samples:
This affects code that manually consumes ready or buffered async iterators with
await anext(stream). This benchmark isolates CPU overhead and does not model network latency. Code usingasync fordoes not use this path.CPython versions tested on:
CPython main branch
Operating systems tested on:
Linux, macOS