Skip to content

gh-159092: Optimize comparisons in list operations - #159093

Open
hetaozdh wants to merge 3 commits into
python:mainfrom
hetaozdh:list-fast-eq
Open

hetaozdh wants to merge 3 commits into
python:mainfrom
hetaozdh:list-fast-eq

Conversation

@hetaozdh

@hetaozdh hetaozdh commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

list.__contains__() calls PyObject_RichCompareBool() for each element. For exact int, float and str needles, we can inline the comparison and avoid the generic dispatch, following the approach used by list.sort().

The fast path is limited to list.__contains__() for now. Extending it to index(), count(), remove() and list.__eq__() would add more code and the benefits are limited.

Benchmarks

Measured on macOS arm64 with Clang 17, -O3, and JIT disabled. Lists contain 20,000 elements; misses scan the entire list. Results are medians of alternating runs against unmodified main. Speedups are calculated as baseline time / modified time.

Workload GIL GIL + PGO + LTO No-GIL No-GIL + PGO + LTO
int, miss 4.67x 3.73x 2.86x 2.79x
int, hit at end 4.68x 3.60x 2.73x 2.77x
str, miss 2.92x 3.44x 2.03x 2.32x
float, miss 6.40x 3.58x 3.30x 2.26x
3-element int, miss 1.50x 1.22x 1.49x 1.29x
Geometric mean (full-scan misses) 4.4x 3.6x 2.7x 2.4x

The geometric mean covers the three exact-type full-scan miss benchmarks (int, str, float).

In pyperformance, bm_hexiom improves by 1.07x in the GIL build. The geometric mean across 49 benchmarks remains within ±2% of 1.00x. PGO results are noisy: independent builds of the same source differ by up to ~8% on individual benchmarks.

Implementation notes

With ThinLTO+PGO, LLVM inlines the generic rich-comparison path into list_contains(). Adding the fast-path type check changes its inlining decisions and can regress the generic path.

Keeping the original loop in a Py_NO_INLINE helper prevents ThinLTO from inlining it back into the dispatcher, preserving the generic path's performance.

Applying the same approach to index(), count() and remove() would add roughly 80 lines of specialized loops, so they are left for a separate change.

Correctness

The fast path preserves PyObject_RichCompareBool()'s identity shortcut, including NaN identity semantics. Other type combinations retain the existing comparison path.

Tested with test_list, test_sort, test_operator, test_compare, test_capi.test_abstract, test_bytes and test_tuple, plus differential checks for subclasses, NaNs, custom __eq__, large integers, string representations and mutation during comparison.

The existing test_deopt_from_append_list environment failure also reproduces on unmodified main.

@bedevere-app bedevere-app Bot added the type-feature A feature request or enhancement label Oct 9, 2026
@hetaozdh
hetaozdh force-pushed the list-fast-eq branch 3 times, most recently from 11eac6e to 0b26e33 Compare October 9, 2026 23:08
@picnixz

picnixz commented Oct 10, 2026

Copy link
Copy Markdown
Member

What about lists with mixed elements? and lists of very small sizes? and non-matches?

@corona10

Copy link
Copy Markdown
Member

My comment: #159092 (comment)

@hetaozdh

Copy link
Copy Markdown
Contributor Author

What about lists with mixed elements? and lists of very small sizes? and non-matches?

Thanks for pointing this out! I’ve added benchmarks for those. The results show no measurable regression for mixed-element lists or custom objects. It improves performance for small lists of integers as well as full scans that don’t find a match.

There is a small regression (around 1–2%) for some types that don’t use the fast path, such as tuple, bool, and bytes. This is likely because the three additional type checks performed for each element when none of the fast paths applies. We could potentially reduce this overhead by checking the target type before entering the loop and selecting the appropriate comparison path upfront, though mixed-element lists would still need to handle different element types.

@hetaozdh

Copy link
Copy Markdown
Contributor Author

My comment: #159092 (comment)

The reason why I did not use specialization is that I think the key difference is that the overhead here is in the comparison for each element, rather than the bytecode-level dispatch.

list.index(), count(), and remove() are C methods. So specialization cannot optimize their inner loops. list.__contains__() has the same issue, and only list == list goes through COMPARE_OP. So specialization cannot cover all entry points.

PyObject_RichCompareBool() is called directly from the list's C-level loop, so the per-element comparison does not go through specialization. Specializing CONTAINS_OP could reduce some per-call dispatch overhead, but the cost addressed here is incurred for every element scanned. These are different levels of overhead, and I don't think specialization is the right place for this optimization.

Use specialized comparisons for exact int, float and str objects to avoid unnecessary rich-comparison dispatch in list operations.
@picnixz

picnixz commented Oct 10, 2026

Copy link
Copy Markdown
Member
  1. Please do not force push, we cannot review what changed otherwise.
  2. What about non-gil builds with PGO/LTO? your benchmarks are only for FT AFAIU.

@picnixz

picnixz commented Oct 10, 2026

Copy link
Copy Markdown
Member

I am not against the change but this adds some maintennce burden and complicate the code. The speedupa are attractive enough but I wonder whether the operations are commom enough though.

Finding a string in a list of string is relevant. Likewise for ints. But floats... not really IMO.

@hetaozdh

Copy link
Copy Markdown
Contributor Author
  1. Please do not force push, we cannot review what changed otherwise.
  2. What about non-gil builds with PGO/LTO? your benchmarks are only for FT AFAIU.

Sorry for the force push. I’m running the benchmarks on a GIL-enabled build with PGO/LTO and will update the results once they’re ready.

@hetaozdh

Copy link
Copy Markdown
Contributor Author

I am not against the change but this adds some maintennce burden and complicate the code. The speedupa are attractive enough but I wonder whether the operations are commom enough though.

Finding a string in a list of string is relevant. Likewise for ints. But floats... not really IMO.

I’m also considering adding tests and assertions to verify that the fast paths preserve the semantics of PyObject_RichCompareBool(), which may help with the maintenance burden.

I agree that finding strings and ints in lists is a more compelling use case than finding floats. I kept the float fast path because it gives the largest speedup in the microbenchmarks, and the implementation is small. However, I think it’s reasonable to leave floats out for now and keep only the int and str fast paths.

Assert that the int, float and str fast paths agree with PyObject_RichCompareBool().
@corona10

corona10 commented Oct 11, 2026 •

Copy link
Copy Markdown
Member

Could you run pyperformance benchmark? (with whole lists)

@hetaozdh

Copy link
Copy Markdown
Contributor Author

Could you run pyperformance benchmark? (with whole lists)

Sure. I will run the full benchmark once the PGO/LTO benchmark was done. Thanks.

@hetaozdh

Copy link
Copy Markdown
Contributor Author

An issue that needs to talk about is that the fast paths can cause regressions in PGO+LTO builds. The original generic comparison chain gets inlined into list_contains() under PGO+LTO, but adding the fast-path dispatch changes the hot path and prevents the same optimization from applying.

One fix is to move the original generic scan into a separate Py_NO_INLINE helper and dispatch to it before entering the loop. This eliminates the regressions, but applying this approach to contains, index, count, and remove adds around 150 lines of code.

I think optimizing contains is worth it, but for the others I am not sure.

@picnixz

picnixz commented Oct 11, 2026

Copy link
Copy Markdown
Member

In general when we want optimizations, we look at their effect under PGO/LTO builds and not just -O3 builds. How much would we gain on those builds though?

@hetaozdh

hetaozdh commented Oct 11, 2026 •

Copy link
Copy Markdown
Contributor Author

In general when we want optimizations, we look at their effect under PGO/LTO builds and not just -O3 builds. How much would we gain on those builds though?

Notice: these are from a local fix that I have not pushed yet. The PR for now as it stands regresses these builds — GIL+PGO+LTO: True in list[bool] +8.8%, b'x' in list[bytes] +6.4%, object() in list +9.6%; FT+PGO+LTO: list[tuple].index(x) +11.3%, list[tuple].count(x) +6.5%. The cause is that the fast-path dispatch moves the PyObject_RichCompareBool() call off the hot path, so with LTO+PGO it is no longer inlined into the loop.

Numbers from the PGO/LTO builds (each build trained on its own profile; median of 8 alternating rounds against the unmodified tree, A/A <= 1.8%):

microbenchmark GIL+PGO+LTO FT+PGO+LTO
x in list[int] (full scan) -73% -64%
list[int].index(x) -78% -62%
list[int].count(x) -79% -65%
x in list[str] -71% -57%
x in list[float] -72% -56%
list[int].remove(x) -65% -73%

The fix keeps the generic scan of each entry point in its own Py_NO_INLINE function and dispatches to it once, before the loop. It then compiles to exactly the same code as before (in the GIL+PGO+LTO build _list_contains_generic is 612 B with the same call profile as the old _list_contains), and the regression is gone in all four configurations I built (GIL / GIL+PGO+LTO / FT / FT+PGO+LTO). I haven't pushed that yet because I'm not sure the performance makes the additional codes worth.

Also the full benchmark haven't been done yet.

…ns__

Keep the generic scan in a separate non-inlinable function so that PyObject_RichCompareBool() is still inlined with LTO and PGO; drop the fast paths for list.index(), count(), remove() and the rich comparison, which are much rarer operations and do not gain enough to be worth the extra code.
@hetaozdh

Copy link
Copy Markdown
Contributor Author

I ran the full pyperformance. Since only the list.__contains__() optimization showed a measurable improvement in pyperformance, I've limited this change to that operation for now.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting review type-feature A feature request or enhancement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants