Repository navigation
Conversation
BenchmarkDB.recall never passed wrrf_k to pg_recall, so the benchmark path always used the pg_recall default of 60 while the production handler (handlers/recall.py) passes settings.WRRF_K. CORTEX_MEMORY_WRRF_K now moves the benchmark exactly as it moves a live recall. Two tests pin the contract (env override reaches pg_recall; default equals the production setting). Measured scope, 2026-09-24, pgvector/pgvector:pg16, 12 memories, rerank off: on the PostgreSQL path this changes no score. The recall_memories() SQL function fuses max-normalised signal scores with a weighted sum (no rank transform); p_wrrf_k only enters the agent-topic boost term. Top-5 scores were identical to six decimals at k=10, 60 and 1000. Rank-based RRF with k runs only on the SQLite path (sqlite_store_search.py). A k sweep through reproduce.sh therefore cannot measure first-stage RRF sensitivity. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Mkvy5MtGG89cNLuQSU8xLe Signed-off-by: cdeust <cdeust@icloud.com>
… output BenchmarkDB was limited to PostgreSQL, so the default SQLite backend and its rank-based RRF were never measured through the production recall pipeline. The new mode is opt-in (backend="sqlite" or CORTEX_BENCH_BACKEND=sqlite), uses a throwaway database file removed on close, and keeps PostgreSQL as the default. The LongMemEval harness now reports the backend it ran on. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Mkvy5MtGG89cNLuQSU8xLe Signed-off-by: cdeust <cdeust@icloud.com>
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Mkvy5MtGG89cNLuQSU8xLe Signed-off-by: cdeust <cdeust@icloud.com>
The SQLite recall path summed weight/(k+rank) per signal, while PL/pgSQL recall_memories sums weighted scores divided by each signal's pool maximum (vector shifted by +1). Port the PostgreSQL semantics to the SQLite path: cosine from the sqlite-vec L2 distance, BM25 from FTS5, heat_base, and an exp(-0.01 * age_days) recency. The trigram signal has no SQLite counterpart and stays absent. wrrf_k now only scales the agent-topic bonus, as on PostgreSQL. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Mkvy5MtGG89cNLuQSU8xLe Signed-off-by: cdeust <cdeust@icloud.com>
The pilot ingests each LongMemEval-S question once, restores the store from an in-memory snapshot before every arm (a recall reinforces the memories it returns), and recalls under the previous rank-based fusion and the score fusion, with and without the reranker. The analysis script runs a paired percentile bootstrap. Moving the fusion into _fused_scores keeps recall_memories under the method-size cap without changing its result. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Mkvy5MtGG89cNLuQSU8xLe Signed-off-by: cdeust <cdeust@icloud.com>
paired_n500.jsonl holds one line per question with the top-10 session ids and first-hit rank for four arms (RRF k=60 and score fusion, each with and without the reranker); summary.json holds the paired bootstrap (10 000 resamples, seed 20260924), overall, on the first 50 questions, and per category. The runs used the code of commit 410ca2f; the later refactor reproduces the recorded retrievals on questions 100 to 103. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Mkvy5MtGG89cNLuQSU8xLe Signed-off-by: cdeust <cdeust@icloud.com>
|
ZETETIC-REVIEW: REQUEST_CHANGES Owner's draft ("do not merge"); I pushed nothing to it. This is a read of why CI is red and what it would take; the decision is the owner's. State: draft, mergeStateStatus BLOCKED, base 0037cea, 12 commits behind origin/main (d574219). The only CI run is from 2026-09-26 (run 36268350406); it predates those 12 commits. What fails (run 36268350406,
Green on that run: Lint, Type Check, Build, Windows SQLite, Docker Smoke, Release dependency set, HOL, CodeQL, Fuzz. What it would take: rebase onto current main (CI reruns on the new base), split |
Summary
Draft experiment, not for merge. It ports the PostgreSQL max-normalised score fusion into the SQLite recall path and adds a SQLite mode to the benchmark database plus a paired benchmark, to test whether the fusion explains the gap between the SQLite and PostgreSQL backends. It does not: on LongMemEval-S (500 questions, production pipeline with reranker) the paired MRR difference is -0.009 [-0.022, +0.005] and recall@10 is -0.018 [-0.032, -0.006] against the existing RRF (k=60). The branch is kept for further work on the SQLite backend.
No issue is closed by this draft.
Type of change
Test plan
check_craftsmanship.pypass.test_benchmark_db_still_reachable_lazily_and_fails_loudly_without_psycopg(psycopg is installed on the test machine).Audit notes
Results, per-question data and the analysis script are in
benchmarks/results/sqlite_score_fusion/. Signal differences left as they are: no trigram signal on SQLite, BM25 instead ofts_rank_cd, rawheat_base, nopost_tool_capturefilter. A recall reinforces heat, so the paired driver restores the store from a snapshot before each arm.Reviewer checklist
🤖 Generated with Claude Code
https://claude.ai/code/session_01Mkvy5MtGG89cNLuQSU8xLe