gh-158219: Fix bytearray constructor to not use hashed object - #158329
Conversation
* Add _PyBytes_GET_CACHED_HASH() static inline function. * _PyBytes_IsMutable() makes sure that the hash value is not cached yet. * bytearray_reinit_from_bytes() checks that the bytes object is mutable. Co-Authored-by: Cody Maloney <cmaloney@theoreticalchaos.com>
|
cc @cmaloney Another approach would be to remove the bytearray optimization to reuse a bytes object if its refcount is 1: always create a copy. Note: |
There was a problem hiding this comment.
Overall looks good. How I found this, and forward looking a bit, my hope is to add more of these cases. The Argument Clinic vectorcall work is in that same line. Hopefully make it so that bytes() and bytearray() constructed on temporaries can avoid the copy work they currently do which should make code faster without having to write very specific patterns.
That is for 3.16+ though. As I add more cases will look for simpler ways to do the cached hash + mutable check.
| if (_PyObject_IsUniquelyReferenced(encoded) | ||
| && PyBytes_CheckExact(encoded)) | ||
| && PyBytes_CheckExact(encoded) | ||
| && _PyBytes_GET_CACHED_HASH((PyBytesObject*)encoded) == -1) |
There was a problem hiding this comment.
the check exact + get cached for "can this be taken" feel like very internal APIs. For now bytearray and bytes are really close to eachother and only have one instance of this (_PyBytes_Resize takes care of generally).
|
Thanks @vstinner for the PR 🌮🎉.. I'm working now to backport this PR to: 3.15. |
|
Sorry, @vstinner, I could not cleanly backport this to |
|
GH-158397 is a backport of this pull request to the 3.15 branch. |
|
I created a small 3.15 backport which omits _PyBytes_IsMutable(): PR gh-158397, since _PyBytes_IsMutable() was not backported to 3.15 yet. |
Uh oh!
There was an error while loading. Please reload this page.