Skip to content

gh-158585: Fix _PyBytes_Resize() if hash value is already computed - #158589

Open
vstinner wants to merge 2 commits into
python:mainfrom
vstinner:bytes_resize
Open

vstinner wants to merge 2 commits into
python:mainfrom
vstinner:bytes_resize

Conversation

@vstinner

@vstinner vstinner commented Oct 2, 2026 •

Copy link
Copy Markdown
Member

_PyBytes_Resize() now creates a new bytes object if the hash value was already computed.

bytes_resize_inplace() no longer sets the hash value to -1, since this function must not be called if the hash value was already computed.

_PyBytes_Resize() now creates a new bytes object if the hash value
was already computed.

bytes_resize_inplace() no longer sets the hash value to -1, since
this function must not be called if the hash value was already
computed.
@vstinner

vstinner commented Oct 2, 2026

Copy link
Copy Markdown
Member Author

Recently, issue gh-158219 was found in bytearray(str, encoding) if a codec computes somehow the hash value of the created bytes object. I wrote PR gh-158329 to not use a bytes object (as the internal bytearray buffer) if it's the hash value was already computed. Instead, the bytes string is copied.

I propose to make a similar fix for _PyBytes_Resize(). The change should not impact any project in practice. It's more to be safer in theory (if it happens) :-)

cc @cmaloney @serhiy-storchaka


Using the new PyBytesWriter API, it should not be possible to compute the hash value of the internal bytes object.

Using _PyBytes_Resize(), even if it's unlikely, the bytes object can be used in a dictionary or exposed in Python by accident, and so its hash value can be computed. In Python 3.15 and older, _PyBytes_Resize() simply clears the cached hash value in this case, as nothing happened, and wish that everything will be fine.

I propose to change _PyBytes_Resize() behavior to be safer: create a copy if the hash value was already computed.

In Python 3.15, commit 32c2649 modified _PyBytes_Resize() last year, to replace Py_REFCNT(v) != 1 test with !_PyObject_IsUniquelyReferenced(v). (The change was also backported to Python 3.14.1.) _PyBytes_Resize() creates a copy if _PyObject_IsUniquelyReferenced() is false. In Python 3.14.0 and older, the function returns as copy if Py_REFCNT(v) is not 1. I propose to do the same if the hash value was computed.


Currently, calling _PyBytes_Resize() on a bytes object with a hash value already computes failed with an assertion error in debug mode. bytes_resize_inplace() calls assert(_PyBytes_IsMutable(v)); before resizing the object in-place and this function checks:

    // gh-158219: The hash value must not be cached yet. Otherwise, it means
    // that the bytes object was already used in Python somehow (ex: as a
    // dictionary key).
    assert(get_ob_shash((PyBytesObject *)self) == -1);

I added this assertion when I fixed the bytearray bug: commit 969af80.

@vstinner

vstinner commented Oct 2, 2026

Copy link
Copy Markdown
Member Author

Sadly, there is also a PyUnicode_Resize() function for Unicode strings (I proposed to soft deprecate it). This function already creates a copy if the hash value was already computed. It checks _PyUnicode_IsModifiable() which checks:

    (...)
    if (!_PyObject_IsUniquelyReferenced(unicode))
        return 0;
    if (PyUnicode_HASH(unicode) != -1)
        return 0;
    (...)

Enhance also _PyBytes_Resize() tests.

@cmaloney cmaloney left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

High level I think this is reasonable. On vacation currently / can't review fully.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants