Conversation
Includes previously-uncommitted sigmoid normalization found in working tree (EXPERIMENTS_LOG 2026-09-27: llama.cpp /v1/rerank returns raw logits ~[-11,+11], not Cohere [0,1]; 1/(1+e^-x) in llama_cpp branch, MIN_RERANK_SCORE stays 0.3; 7 sigmoid tests). New: MAX_RERANKER_TOPN env tumbler (PerformanceConfig.reranker_topn_keep, default 0 = current behavior) - union of threshold-passers with top-N by score, so uncalibrated cross-encoder scores (P3 target 0.271<0.3) can survive as recall floor without lowering the absolute cut. calibrate_threshold() (F1-max on labeled holdout) refuses eval sources with ValueError - the anti-overfit rule (no sweep on the 16 frozen eval rules) is encoded as code. Split protocol + sizes (>=10 holdout queries disjoint from eval-16) in threshold_calibration.py docstring. Default 0.3 unchanged pending a real holdout measurement.
…use) Root cause (verified): 3-way RRF rewards multi-tier consensus, so a target found by ONE tier only (P2: BM25 rank 0 for src/core/search/engine.py scores 1/(60+1)) loses to junk present in 2-3 tiers (2-3x the score) and is amputated by rrf_results[:limit] before the reranker ever sees it. MMR is innocent (reorder-only, no drops); bucket weights favour the target (.py 1.0 vs .txt/.md 0.5); query expansion keeps the verbatim query as variants[0]. This resolves the open contradiction in KNOWN_ISSUES (standalone BM25 rank 0 vs hybrid loss): hybrid never returns raw BM25 order. Fix: anchor_tier_winners() in scoring.py - pool = MMR-ordered RRF top-limit (order preserved for the no-reranker path) + missing per-tier top-1 winners appended from the full RRF list (reconstructed as fused entries when outside it), capped at MAX_RERANKER_INPUT. Reranker top_n (=limit) unchanged. Regression test failed before (target cut at rank ~13) / passes after. KNOWN_ISSUES P2 entry updated: root cause established, live validation + holdout calibration remain open.
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID: Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
Superseded: P2 anchor part merged via #54 (_anchor_identifier_chunks_async). Non-conflicting remainder salvaged as #63 (sigmoid scale, top-N floor, holdout calibration; tier-winner anchors deliberately skipped - would duplicate #54 at the same pool site). My two commits from this branch went to #62. Closing as superseded; branch kept as reference. |
Root cause P2 (verified, was Open)
3-way RRF rewards multi-tier consensus: the P2 target is found by ONE tier only (BM25 rank 0 for \src/core/search/engine.py\ -> 1/(60+1)≈0.0164) while junk present in 2-3 tiers at mediocre ranks accumulates 2-3x that.
rf_results[:limit]\ (\engine.py:746) then amputates the single-tier winner before the reranker ever sees it.
Ruled out by reading: MMR is reorder-only (no drops); bucket weights favour the target (.py 1.0 vs .txt/.md 0.5); query expansion keeps the verbatim query as variants[0]. This resolves the open contradiction in KNOWN_ISSUES (standalone BM25 rank 0 vs hybrid loss): hybrid never returns raw BM25 order. Regression test failed before (target cut, pool = 10x experiments artifacts) / passes after.
Design
Numbers
Caveats (open)
DO NOT MERGE yet — review requested.