Skip to content

Extend BINARY_OP_SUBSCR_USTR_INT to support latin1 characters #158649

Description

@johng

Feature or enhancement

Currently in _BINARY_OP_SUBSCR_USTR_INT specialisation we look up just ascii tables. This results in specialisation fails when a character is latin1 encoded which might be common in many European texts.

A small extension to support latin1 would enable the specialisation to support these characters such as

PyObject *res_o = (c < 128)
                ? (PyObject*)&_Py_SINGLETON(strings).ascii[c]
                : (PyObject*)&_Py_SINGLETON(strings).latin1[c - 128];

Scanning some french text results in a 14% speedup with the test done on MacOS. This avoids any fallback to the generic path for the latin characters.

def scan(text):
    i = 0
    n = len(text)
    while i < n:
        c = text[i]          # the only str[int] site
        i += 1

ASCII  = "The quick brown fox jumps over the lazy dog. " * 20
FRENCH = "Le garçon préfère les crêpes à la fraîcheur du matin. " * 20

Has this already been discussed elsewhere?

This is a minor feature, which does not need previous discussion elsewhere

Links to previous discussion of this feature:

No response

Linked PRs

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    interpreter-core(Objects, Python, Grammar, and Parser dirs)performancePerformance or resource usagetype-featureA feature request or enhancement

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions