Bug report
Bug description:
This is a self-contained reproducer of a real issue I encountered (see below):
import numbers
from concurrent.futures import ThreadPoolExecutor
from time import time
mylist = [1.0] * 100_000
def check(_):
for x in mylist:
isinstance(x, numbers.Integral)
for cores in [1, 2, 4, 8]:
start = time()
with ThreadPoolExecutor(cores) as pool:
list(pool.map(check, range(cores)))
print(cores, time() - start)
I ran this on my computer, which has 12 cores. I would expect this to the same amount of time regardless of number of threads, since they ought to be run in parallel. In fact, the output looks like this:
Python 3.14t:
1 0.030498981475830078
2 0.07532048225402832
4 0.17853212356567383
8 0.5576419830322266
Python 3.15t (3.15rc2):
1 0.032598018646240234
2 0.05810952186584473
4 0.14814376831054688
8 0.31418490409851074
Original real-world issue
I discovered this issue while benchmarking some scikit-learn code. In particular it's caused by code that uses https://github.com/scikit-learn/scikit-learn/blob/d6f188097e255822d99633f13f4d6304cc76a7c8/sklearn/utils/_missing.py#L40 on all the values in a data structure.
In practice I have optimized away much of the usage of this function, so in future versions of scikit-learn (post-1.9) it hopefully won't be a bottleneck in practice. But it's definitely a real bottleneck in sklearn 1.9, and presumably other people may encounter it in other code.
Potential source of bottleneck
Looking at the profile output of samply suggests _in_weak_set's critical section, maybe (thanks to @ngoldbaum for the link:
|
int incache = _in_weak_set(impl, &impl->_abc_cache, subclass); |
).
CPython versions tested on:
3.14, 3.15
Operating systems tested on:
Linux
Linked PRs
Bug report
Bug description:
This is a self-contained reproducer of a real issue I encountered (see below):
I ran this on my computer, which has 12 cores. I would expect this to the same amount of time regardless of number of threads, since they ought to be run in parallel. In fact, the output looks like this:
Python 3.14t:
Python 3.15t (3.15rc2):
Original real-world issue
I discovered this issue while benchmarking some scikit-learn code. In particular it's caused by code that uses https://github.com/scikit-learn/scikit-learn/blob/d6f188097e255822d99633f13f4d6304cc76a7c8/sklearn/utils/_missing.py#L40 on all the values in a data structure.
In practice I have optimized away much of the usage of this function, so in future versions of scikit-learn (post-1.9) it hopefully won't be a bottleneck in practice. But it's definitely a real bottleneck in sklearn 1.9, and presumably other people may encounter it in other code.
Potential source of bottleneck
Looking at the profile output of
samplysuggests_in_weak_set's critical section, maybe (thanks to @ngoldbaum for the link:cpython/Modules/_abc.c
Line 642 in e5fbabb
CPython versions tested on:
3.14, 3.15
Operating systems tested on:
Linux
Linked PRs
__instancecheck__and__subclasscheck__without creating a bound method #157670