Skip to content

feat: prepared-query inner product estimate, O(dim) per code - #10

Merged
amikhail48 merged 1 commit into
mainfrom
feat/prepared-query
Sep 14, 2026
Merged

amikhail48 merged 1 commit into
mainfrom
feat/prepared-query

Conversation

@amikhail48

Copy link
Copy Markdown
Member

Summary

QuantizerProd::estimate_inner_product(y, q) did two O(dim²) matrix-vector products for every code:
S * y in QJL, and Pi^T * y_hat inside QuantizerMSE::dequantize. Both depend only on the query once the MSE term is written as ⟨y, Πᵀŷ⟩ = ⟨Πy, ŷ⟩ (Π is orthogonal). Scoring one query against N codes therefore cost N · O(dim²).

This adds a prepared-query path: the O(dim²) work runs once per query, then each code is O(dim).

API (additive; nothing removed)

  • QuantizerProd::PreparedQuery { Vec rotated; Vec projected; } — Πy and Sy.
  • QuantizerProd::prepare_query(const Vec& y)
  • QuantizerProd::estimate_inner_product(const PreparedQuery&, const QuantizedProd&) — O(dim).
  • Building blocks: QJL::project, QJL::estimate_inner_product_projected, QuantizerMSE::inner_product_rotated.

The existing estimate_inner_product(y, q) now prepares and delegates, so both entry points compute identical values.

Verification

  • bazel build //... passes; //:turboquant_example output is byte-identical to main (all 38 lines, every estimate to 6 d.p.).
  • Exercised through recall's TurboQuantCodec adapter: batch scores equal single-call scores bit-exactly and still match ⟨query, dequantize(code)⟩ within 1e-3; clean under ASan and UBSan.
  • This repo has no unit-test target, so correctness coverage lives in recall's adapter tests.

Performance

Measured through recall's MediaIndex::search (one query vs. 100 stored codes, bitwidth 3, -c opt):

dim Apple Silicon Jetson Orin Nano Super (MAXN_SUPER)
512 2.3 → 0.13 ms 13.0 → 0.39 ms
1024 8.9 → 0.31 ms 73.0 → 1.3 ms
1536 23.1 → 0.56 ms 204 → 3.0 ms

A standalone estimate_inner_product(y, q) call is unchanged in complexity (it still prepares per call); the win comes from preparing once and scoring many codes.

Not in this PR

Construction is still O(dim³) (Householder QR, materializing Q, and determinant() for the sign in make_rotation_matrix) — ~0.74 s at d=1536 on the Orin at -c opt. Separate follow-up.

🤖 Generated with Claude Code

QuantizerProd::estimate_inner_product did two O(dim^2) products for every
code: S * y in QJL, and Pi^T * y_hat in the MSE dequantize. Both depend only
on the query once the MSE term is rewritten as <y, Pi^T y_hat> = <Pi y, y_hat>
(Pi is orthogonal), so scoring a query against N codes paid N * O(dim^2).

- QuantizerProd::prepare_query(y) computes Pi*y and S*y once.
- QuantizerProd::estimate_inner_product(PreparedQuery, q) is O(dim) per code.
- QJL::project / estimate_inner_product_projected and
  QuantizerMSE::inner_product_rotated are the building blocks.
- The existing estimate_inner_product(y, q) now prepares and delegates, so the
  two entry points produce identical values.

turboquant_example output is byte-identical to main (all estimates, 6 d.p.).
Measured through recall's MediaIndex (100 entries, d=1536, -c opt): search
23.1 ms -> 0.56 ms on Apple Silicon, 204 ms -> 3.0 ms on Jetson Orin Nano.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@amikhail48
amikhail48 merged commit 9e6451f into main Sep 14, 2026
5 checks passed
@amikhail48
amikhail48 deleted the feat/prepared-query branch September 14, 2026 02:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant