Repository navigation
Conversation
11eac6e to
0b26e33
Compare
|
What about lists with mixed elements? and lists of very small sizes? and non-matches? |
|
My comment: #159092 (comment) |
Thanks for pointing this out! I’ve added benchmarks for those. The results show no measurable regression for mixed-element lists or custom objects. It improves performance for small lists of integers as well as full scans that don’t find a match. There is a small regression (around 1–2%) for some types that don’t use the fast path, such as tuple, bool, and bytes. This is likely because the three additional type checks performed for each element when none of the fast paths applies. We could potentially reduce this overhead by checking the target type before entering the loop and selecting the appropriate comparison path upfront, though mixed-element lists would still need to handle different element types. |
The reason why I did not use specialization is that I think the key difference is that the overhead here is in the comparison for each element, rather than the bytecode-level dispatch.
|
Use specialized comparisons for exact int, float and str objects to avoid unnecessary rich-comparison dispatch in list operations.
0b26e33 to
8dfd0f0
Compare
|
|
I am not against the change but this adds some maintennce burden and complicate the code. The speedupa are attractive enough but I wonder whether the operations are commom enough though. Finding a string in a list of string is relevant. Likewise for ints. But floats... not really IMO. |
Sorry for the force push. I’m running the benchmarks on a GIL-enabled build with PGO/LTO and will update the results once they’re ready. |
I’m also considering adding tests and assertions to verify that the fast paths preserve the semantics of I agree that finding strings and ints in lists is a more compelling use case than finding floats. I kept the float fast path because it gives the largest speedup in the microbenchmarks, and the implementation is small. However, I think it’s reasonable to leave floats out for now and keep only the int and str fast paths. |
Assert that the int, float and str fast paths agree with PyObject_RichCompareBool().
|
Could you run pyperformance benchmark? (with whole lists) |
Sure. I will run the full benchmark once the PGO/LTO benchmark was done. Thanks. |
|
An issue that needs to talk about is that the fast paths can cause regressions in PGO+LTO builds. The original generic comparison chain gets inlined into One fix is to move the original generic scan into a separate I think optimizing |
|
In general when we want optimizations, we look at their effect under PGO/LTO builds and not just -O3 builds. How much would we gain on those builds though? |
Notice: these are from a local fix that I have not pushed yet. The PR for now as it stands regresses these builds — GIL+PGO+LTO: Numbers from the PGO/LTO builds (each build trained on its own profile; median of 8 alternating rounds against the unmodified tree, A/A <= 1.8%):
The fix keeps the generic scan of each entry point in its own Also the full benchmark haven't been done yet. |
…ns__ Keep the generic scan in a separate non-inlinable function so that PyObject_RichCompareBool() is still inlined with LTO and PGO; drop the fast paths for list.index(), count(), remove() and the rich comparison, which are much rarer operations and do not gain enough to be worth the extra code.
|
I ran the full pyperformance. Since only the |
list.__contains__()callsPyObject_RichCompareBool()for each element. For exactint,floatandstrneedles, we can inline the comparison and avoid the generic dispatch, following the approach used bylist.sort().The fast path is limited to
list.__contains__()for now. Extending it toindex(),count(),remove()andlist.__eq__()would add more code and the benefits are limited.Benchmarks
Measured on macOS arm64 with Clang 17,
-O3, and JIT disabled. Lists contain 20,000 elements; misses scan the entire list. Results are medians of alternating runs against unmodified main. Speedups are calculated as baseline time / modified time.int, missint, hit at endstr, missfloat, missint, missThe geometric mean covers the three exact-type full-scan miss benchmarks (
int,str,float).In
pyperformance,bm_hexiomimproves by 1.07x in the GIL build. The geometric mean across 49 benchmarks remains within ±2% of 1.00x. PGO results are noisy: independent builds of the same source differ by up to ~8% on individual benchmarks.Implementation notes
With ThinLTO+PGO, LLVM inlines the generic rich-comparison path into
list_contains(). Adding the fast-path type check changes its inlining decisions and can regress the generic path.Keeping the original loop in a
Py_NO_INLINEhelper prevents ThinLTO from inlining it back into the dispatcher, preserving the generic path's performance.Applying the same approach to
index(),count()andremove()would add roughly 80 lines of specialized loops, so they are left for a separate change.Correctness
The fast path preserves
PyObject_RichCompareBool()'s identity shortcut, including NaN identity semantics. Other type combinations retain the existing comparison path.Tested with
test_list,test_sort,test_operator,test_compare,test_capi.test_abstract,test_bytesandtest_tuple, plus differential checks for subclasses, NaNs, custom__eq__, large integers, string representations and mutation during comparison.The existing
test_deopt_from_append_listenvironment failure also reproduces on unmodified main.