perf: warm-start the positive-only NNLS solve from the previous evaluation's passive set (#498) - #501
Merged
Conversation
…ation's passive set (#498) Phase 3a of the numba CPU likelihood speed restoration. fnnls_cholesky gains solver diagnostics (stats dict), accepts a mask or index warm start, reuses one Cholesky factorisation for the warm start instead of a dense solve plus a from-scratch refactorisation, and repairs an infeasible warm passive set before the outer loop (an all-True seed previously bypassed the loop and returned the clipped vector). A process-local FIFO memo (nnls_memo.py) seeds each solve from the previous evaluation's final passive set, keyed on the index space (mesh/data shapes + edge-zeroed subset); Settings.nnls_warm_start_memo, default on, AUTOARRAY_NNLS_WARM_START=0 disables; a failed seeded solve retries once from the dense-sign start. Delaunay-1250 fiducial, random-walk sequence: median active-set iterations 69.5 -> 7.0 (euclid), 32 -> 8 (hst); solve 0.54 -> 0.05 s / 0.25 -> 0.07 s; log-likelihood parity 3e-14. Solve time is never worse than the dense-sign start, including on uncorrelated i.i.d. jumps. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XySQZP6Npj4adZMkYZ96Bo
Each memo entry now also carries the error fraction of the most recent dense-sign-started solve for its key; a seeded solve whose own error fraction exceeds Settings.nnls_warm_start_error_tolerance (config default 1.5) times that reference is dropped, so the next solve restarts from the dense-sign start and refreshes the reference. Relative, not absolute: the 32-cell robustness matrix showed the absolute seed error fraction does not separate seeds that save iterations from seeds that cost them (0.048-0.138 overlap), while the seed/dense ratio does (helpful <= 0.89, worst 1.42). Non-finite or non-positive tolerance disables the guard. reconstruction_positive_only_from records seed_source and warm_start_fallback in the stats dict. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XySQZP6Npj4adZMkYZ96Bo
5 tasks
Collaborator
Author
|
Workspace leg: PyAutoLabs/autolens_profiling#184 (merges after this PR — library-first gate). |
This was referenced Aug 28, 2026
Closed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Phase 3a of the numba CPU sparse-operator likelihood speed restoration (epic
numba-cpu-likelihood, closes #498). The positive-only reconstruction solve (fnnls_cholesky, Bro & De Jong active set) is iteration-bound: its warm start — the sign of the unconstrained dense solve — gets ~70–95 of ~1310 passive-set entries wrong on the Delaunay-1250 fiducial, and every wrong entry costs an active-set iteration. This PR (1) instruments the solver, (2) makes the warm start reuse one Cholesky factorisation instead of a dense solve plus a from-scratch refactorisation, and (3) adds a process-local cross-evaluation memo that seeds the next solve from the previous evaluation's final passive set (Settings.nnls_warm_start_memo, default on, kill-switchAUTOARRAY_NNLS_WARM_START=0).Measured on the euclid + hst Delaunay-1250 + MGE-60 fiducial (
autolens_profilingPR to follow;AUTOARRAY_NUMBA_OPERATED_MEMO=0, 30-step random-walk and i.i.d. instance sequences, memo off vs on):Log-likelihood parity memo on vs off: max relative Δ 3e-14; max |Δreconstruction| 1.6e-8; 0 memo-seed failures in 240 evaluations. Solve seconds are ≤ off in every cell (the memo also skips the dense unconstrained solve), so uncorrelated jumps cost iterations but not time. Step-0 re-profile after #453/#497: the solve is now 40% (euclid) / 17% (hst) of an evaluation, not the ~72% the issue was filed on.
Robustness across lens models (32-cell matrix,
results/notes/nnls_warm_start_memo_matrix.md: PowerLaw, NFW subhalo, no / Sersic lens light, rectangular mesh, adaptive regularization, no edge-zeroing, complex source, 600 / 2000-pixel meshes; euclid + hst): parity ≤ 4e-14 in every cell, zero seeded-solve failures, median solve time never worse than memo-off by >2%; every random-walk cell gains (3.2×–118× — rectangular meshes most, their dense-sign start is 44–70% wrong); only i.i.d. sequences lose iterations (worst 0.52×), never time.Wiggle room — a relative fallback guard (
Settings.nnls_warm_start_error_tolerance, default 1.5): each memo entry remembers the error fraction of the last dense-sign-started solve for its key; a seed that comes out worse thantolerance ×that reference is dropped and the next solve restarts dense, refreshing the reference. Relative because the absolute seed error fraction does not separate helpful from harmful seeds (0.048–0.138 overlap) while the seed/dense ratio does (helpful ≤ 0.89, worst 1.42). Fired 3 times in 240 re-run evaluations, all in the one cell at ratio 1.42. Non-finite / non-positive tolerance disables it.A correctness fix rides along: a warm passive set containing entries with a non-positive unconstrained solution used to be clipped only, and an all-True
Pskipped the outer loop entirely and returned that clipped vector. Unreachable from the dense-sign start, reachable from a memo seed — repaired before the outer loop (explicit drop-and-refactor;fix_constraint_choleskyhas no prior feasible iterate there and itsalphais 0/0).API Changes
Additive.
fnnls_choleskygains an optionalstatsdict and accepts a bool mask or index array asP_initial;reconstruction_positive_only_fromgainsfingerprint=;Settingsgainsnnls_warm_start_memoandnnls_warm_start_error_tolerance(configgeneral.inversion.nnls_warm_start_memo, packaged defaulttrue; a shadowing workspacegeneral.yamlwithout the key also resolves totrue); new moduleautoarray.inversion.inversion.nnls_memo. Behaviour change: on the NumPy/numba path the positive-only solve now warm-starts from the previous evaluation's passive set by default — the solution is the unique NNLS optimum and is unchanged to round-off. JAX path untouched.See full details below.
Test Plan
pytest test_autoarray— 1279 passed (new:test_cholesky_inplace.py+7,test_jax_nnls.py+4,test_nnls_memo.pynew,test_settings_dict.py+1)delaunay_numba.pypins PASSED; memo on/off parity 3e-14)Full API Changes (for automation & release notes)
Added
autoarray.util.fnnls.fnnls_cholesky(ZTZ, ZTx, P_initial=, stats: Optional[dict] = None)—statsis filled withouter_iterations,inner_iterations,passive_set,n_passive,warm_start_errors;P_initialmay be a bool mask or an int index arrayautoarray.inversion.inversion.inversion_util.reconstruction_positive_only_from(..., fingerprint=None)— identifies the passive set's index space for the memo;Nonedisables the memo for that callautoarray.Settings(nnls_warm_start_memo: Optional[bool] = None)+ property — NumPy/numba path onlyautoarray/config/general.yaml→inversion.nnls_warm_start_memo: trueautoarray.Settings(nnls_warm_start_error_tolerance: Optional[float] = None)+ property — configinversion.nnls_warm_start_error_tolerance: 1.5; non-finite / ≤ 0 disables the guardautoarray.inversion.inversion.nnls_memo—MemoEntry(passive_set, dense_error_fraction),memo_enabled(),memo_key(n, fingerprint),passive_set_get(key, n) -> Optional[MemoEntry],passive_set_put(key, passive_set, dense_error_fraction),memo_drop(key); FIFO, 8 entries; envAUTOARRAY_NNLS_WARM_START=0disablesreconstruction_positive_only_fromwritesstats["seed_source"]("memo"/"dense") andstats["warm_start_fallback"]after the solveAbstractInversion._nnls_warm_start_fingerprint(ids_to_keep=None)(private)Changed Behaviour
fnnls_cholesky: warm start factorises the initial passive set once and extends it in place (was: densescipy.linalg.solve+ fullcholeskyon the first outer iteration); a warm passive set with non-positive entries is repaired before the outer loop (was: clipped, and returned unsolved whenPwas all-True); near-singular warm-start systems now raiseLinAlgErrorfromcholeskyinstead ofLinAlgWarning+ garbage fromsolvereconstruction_positive_only_from(NumPy path): warm-starts from the memo whensettings.nnls_warm_start_memoand a fingerprint are given; a memo-seeded solve that raises retries once from the dense-sign start before theInversionExceptionpathMigration
aa.Settings(nnls_warm_start_memo=False)orAUTOARRAY_NNLS_WARM_START=0.Generated by the PyAutoLabs agent workflow.