Skip to content

SOLR-18424: add oversampling + raw dense vector reranking - #4919

Open
liangkaiwen wants to merge 6 commits into
apache:mainfrom
liangkaiwen:jira/SOLR-18424
Open

liangkaiwen wants to merge 6 commits into
apache:mainfrom
liangkaiwen:jira/SOLR-18424

Conversation

@liangkaiwen

@liangkaiwen liangkaiwen commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

https://issues.apache.org/jira/browse/SOLR-18424

Description

Solr 10 introduced Binary Quantization, however the recall drop-off is fairly sharp when used as-is. This can be partially offset by oversampling (retrieving more total documents from HNSW graph) and reranking (using raw non-quantized vector values from .vec). While reranking is possible through the standard Solr reranking feature, it is cumbersome syntactically and easy to get wrong. This aims to offer a cleaner and performant oversampling/reranking natively in the KNN query parser

Solution

Code mostly written by Claude Opus 5

  • Add a param 'rerankOversample' to the KNN query parser. This param accepts a positive integer and is optional (default value=1 when not provided). Behaviorally, this value is a multiplier on the existing topK * efSearch, and if it is > 1 do a rerank stage after search down to topK only.
  • Validation of rerankOversample. The provided value should be >= 1 and oversample should not be applied to BYTE encoded vectors
  • Uses generic RescoreTopNQuery and passes the similarity function explicitly, because the lucene rerank vector function can hit a NullPointerException when similarity function is not explicitly defined (as a binary quantized field definition might not).
    • Additional bug to avoid with Lucene's Rescore API is that an empty result set from the initial query will cause an ArrayIndexOutOfBoundsException. This change includes a Solr level check for this, but Lucene may want to harden against this as well

Tests

  • Manually querying
  • Unit tests

Checklist

Please review the following and check all that apply:

  • I have reviewed the guidelines for How to Contribute and my code conforms to the standards described there to the best of my ability.
  • I have created a Jira issue and added the issue ID to my pull request title.
  • I have given Solr maintainers access to contribute to my PR branch. (optional but recommended, not available for branches on forks living under an organisation)
  • I have developed this patch against the main branch.
  • I have run ./gradlew check.
  • I have added tests for my changes.
  • I have added documentation for the Reference Guide
  • I have added a changelog entry for my change

@github-actions github-actions Bot added the documentation Improvements or additions to documentation label Sep 18, 2026
@liangkaiwen

Copy link
Copy Markdown
Contributor Author

Exploring a bug where Lucene's RescoreTopNQuery can get an index out of bounds exception when the inner query returns 0 documents

@alessandrobenedetti alessandrobenedetti left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I spent some time digging various aspects of this pull request, that is fine overall.
I have doubts on the explicit check conditions on the vectorEncoding, specifically the 'BYTE' is not supported. So possibly some of my previous comments lose sense overall if we remove that part

Generally speaking both the binary and scalar-quantised dense vector field don't support vectorEncoding='BYTE ' (logically) and from some investigation they will silently skip quantisation entirely (in this class there are various checks on the vectorEncoding and when not FLOAT32 it skips various instructions):
only FLOAT32 fields get the quantizing writer. This is the key line. lucene104/Lucene104ScalarQuantizedVectorsWriter.java, lines 127–137: }

So, my conclusion and ideal outcome is:

  • oversampling is useful only with binary or scalar quantised vector fields, so we log a warning/throw exception when used with anything different
    -on a separate ticket maybe we can add a check for both the scalar and binary quantised fields that reject the 'BYTE' encoding telling the user this is logically incompatible with the quantisation itself (and the underlying lucene quantisation engine).
    Some may argue (including myself) that binary quantisation should be compatible with a 'BYTE' vectorEncoding, but it's a different discussion and investigation, I would say (because from what I quickly saw, this is not permitted by Lucene implementation)

vectorBuilder.getByteVector(),
acceptedChildren,
topK,
candidateTopK,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

here candidateTopK could be misleading as BYTE encoding won't be supported in this contribution?

fieldName,
vectorBuilder.getByteVector(),
topK,
candidateTopK,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

same comment as below, isn't this a bit confusing given the fact in this PR BYTE encoding won't be supported?

if (rewrittenInner instanceof MatchNoDocsQuery) {
return rewrittenInner;
}
// Delegate using the already rewritten inner query rather than calling super.rewrite(), which

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I suspect this is a 'if... else...' maybe better to make it explicit?

"{!knn f=" + quantizedField + " topK=5 rerankOversample=5}[1.0, 2.0, 3.0, 4.0]",
"fl",
"id"),
EXPECTED_EXACT_TOP_5);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

should we add here an assertion with the same query but no oversampling? showing the difference?

"//result/doc[4]/str[@name='id'][.='10']",
"//result/doc[5]/str[@name='id'][.='3']");
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What's the differentiator between some of these tests and the dedicated oversampling KNN tests? I get the throw/not throw exception, but the others feel similar to the ones in the other class.

|Optional |Default: 1
|===
+
If provided, the query will retrieve a candidate pool of documents equal to topK multiplied by the rerankOversample value provided. The candidate set of documents will then be rescored with their raw vector values, and reranked down to a topK result set.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I would specify at the moment that is compatible only with scalar quantised or binary quantised vector fields

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

cat:schema cat:search documentation Improvements or additions to documentation tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants