feat: prepared-query inner product estimate, O(dim) per code - #10
Merged
Merged
Conversation
QuantizerProd::estimate_inner_product did two O(dim^2) products for every code: S * y in QJL, and Pi^T * y_hat in the MSE dequantize. Both depend only on the query once the MSE term is rewritten as <y, Pi^T y_hat> = <Pi y, y_hat> (Pi is orthogonal), so scoring a query against N codes paid N * O(dim^2). - QuantizerProd::prepare_query(y) computes Pi*y and S*y once. - QuantizerProd::estimate_inner_product(PreparedQuery, q) is O(dim) per code. - QJL::project / estimate_inner_product_projected and QuantizerMSE::inner_product_rotated are the building blocks. - The existing estimate_inner_product(y, q) now prepares and delegates, so the two entry points produce identical values. turboquant_example output is byte-identical to main (all estimates, 6 d.p.). Measured through recall's MediaIndex (100 entries, d=1536, -c opt): search 23.1 ms -> 0.56 ms on Apple Silicon, 204 ms -> 3.0 ms on Jetson Orin Nano. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
amikhail48
force-pushed
the
feat/prepared-query
branch
from
September 14, 2026 02:37
7dc18de to
4d99b00
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
QuantizerProd::estimate_inner_product(y, q)did two O(dim²) matrix-vector products for every code:S * yinQJL, andPi^T * y_hatinsideQuantizerMSE::dequantize. Both depend only on the query once the MSE term is written as ⟨y, Πᵀŷ⟩ = ⟨Πy, ŷ⟩ (Π is orthogonal). Scoring one query against N codes therefore cost N · O(dim²).This adds a prepared-query path: the O(dim²) work runs once per query, then each code is O(dim).
API (additive; nothing removed)
QuantizerProd::PreparedQuery { Vec rotated; Vec projected; }— Πy and Sy.QuantizerProd::prepare_query(const Vec& y)QuantizerProd::estimate_inner_product(const PreparedQuery&, const QuantizedProd&)— O(dim).QJL::project,QJL::estimate_inner_product_projected,QuantizerMSE::inner_product_rotated.The existing
estimate_inner_product(y, q)now prepares and delegates, so both entry points compute identical values.Verification
bazel build //...passes;//:turboquant_exampleoutput is byte-identical tomain(all 38 lines, every estimate to 6 d.p.).TurboQuantCodecadapter: batch scores equal single-call scores bit-exactly and still match ⟨query, dequantize(code)⟩ within 1e-3; clean under ASan and UBSan.Performance
Measured through recall's
MediaIndex::search(one query vs. 100 stored codes, bitwidth 3,-c opt):A standalone
estimate_inner_product(y, q)call is unchanged in complexity (it still prepares per call); the win comes from preparing once and scoring many codes.Not in this PR
Construction is still O(dim³) (Householder QR, materializing Q, and
determinant()for the sign inmake_rotation_matrix) — ~0.74 s at d=1536 on the Orin at-c opt. Separate follow-up.🤖 Generated with Claude Code