Skip to content

perf: warm-start the positive-only NNLS solve from the previous evaluation's passive set (#498) - #501

Merged
Jammy2211 merged 2 commits into
mainfrom
feature/numba-cpu-nnls-iteration-reduction
Aug 28, 2026
Merged

Jammy2211 merged 2 commits into
mainfrom
feature/numba-cpu-nnls-iteration-reduction

Conversation

@Jammy2211

Copy link
Copy Markdown
Collaborator

Summary

Phase 3a of the numba CPU sparse-operator likelihood speed restoration (epic numba-cpu-likelihood, closes #498). The positive-only reconstruction solve (fnnls_cholesky, Bro & De Jong active set) is iteration-bound: its warm start — the sign of the unconstrained dense solve — gets ~70–95 of ~1310 passive-set entries wrong on the Delaunay-1250 fiducial, and every wrong entry costs an active-set iteration. This PR (1) instruments the solver, (2) makes the warm start reuse one Cholesky factorisation instead of a dense solve plus a from-scratch refactorisation, and (3) adds a process-local cross-evaluation memo that seeds the next solve from the previous evaluation's final passive set (Settings.nnls_warm_start_memo, default on, kill-switch AUTOARRAY_NNLS_WARM_START=0).

Measured on the euclid + hst Delaunay-1250 + MGE-60 fiducial (autolens_profiling PR to follow; AUTOARRAY_NUMBA_OPERATED_MEMO=0, 30-step random-walk and i.i.d. instance sequences, memo off vs on):

instrument sequence median iterations off→on solve s off→on eval s off→on
euclid random walk 69.5 → 7.0 (9.9×) 0.544 → 0.051 1.510 → 0.564
euclid i.i.d. 86 → 94 0.347 → 0.309 0.906 → 0.838
hst random walk 32 → 8 (4.0×) 0.245 → 0.067 1.733 → 1.594
hst i.i.d. 42 → 69.5 0.268 → 0.260 1.687 → 1.867

Log-likelihood parity memo on vs off: max relative Δ 3e-14; max |Δreconstruction| 1.6e-8; 0 memo-seed failures in 240 evaluations. Solve seconds are ≤ off in every cell (the memo also skips the dense unconstrained solve), so uncorrelated jumps cost iterations but not time. Step-0 re-profile after #453/#497: the solve is now 40% (euclid) / 17% (hst) of an evaluation, not the ~72% the issue was filed on.

Robustness across lens models (32-cell matrix, results/notes/nnls_warm_start_memo_matrix.md: PowerLaw, NFW subhalo, no / Sersic lens light, rectangular mesh, adaptive regularization, no edge-zeroing, complex source, 600 / 2000-pixel meshes; euclid + hst): parity ≤ 4e-14 in every cell, zero seeded-solve failures, median solve time never worse than memo-off by >2%; every random-walk cell gains (3.2×–118× — rectangular meshes most, their dense-sign start is 44–70% wrong); only i.i.d. sequences lose iterations (worst 0.52×), never time.

Wiggle room — a relative fallback guard (Settings.nnls_warm_start_error_tolerance, default 1.5): each memo entry remembers the error fraction of the last dense-sign-started solve for its key; a seed that comes out worse than tolerance × that reference is dropped and the next solve restarts dense, refreshing the reference. Relative because the absolute seed error fraction does not separate helpful from harmful seeds (0.048–0.138 overlap) while the seed/dense ratio does (helpful ≤ 0.89, worst 1.42). Fired 3 times in 240 re-run evaluations, all in the one cell at ratio 1.42. Non-finite / non-positive tolerance disables it.

A correctness fix rides along: a warm passive set containing entries with a non-positive unconstrained solution used to be clipped only, and an all-True P skipped the outer loop entirely and returned that clipped vector. Unreachable from the dense-sign start, reachable from a memo seed — repaired before the outer loop (explicit drop-and-refactor; fix_constraint_cholesky has no prior feasible iterate there and its alpha is 0/0).

API Changes

Additive. fnnls_cholesky gains an optional stats dict and accepts a bool mask or index array as P_initial; reconstruction_positive_only_from gains fingerprint=; Settings gains nnls_warm_start_memo and nnls_warm_start_error_tolerance (config general.inversion.nnls_warm_start_memo, packaged default true; a shadowing workspace general.yaml without the key also resolves to true); new module autoarray.inversion.inversion.nnls_memo. Behaviour change: on the NumPy/numba path the positive-only solve now warm-starts from the previous evaluation's passive set by default — the solution is the unique NNLS optimum and is unchanged to round-off. JAX path untouched.
See full details below.

Test Plan

  • pytest test_autoarray — 1279 passed (new: test_cholesky_inplace.py +7, test_jax_nnls.py +4, test_nnls_memo.py new, test_settings_dict.py +1)
  • euclid + hst Delaunay-1250 pinned log-likelihoods unchanged (delaunay_numba.py pins PASSED; memo on/off parity 3e-14)
  • CI green on all legs
Full API Changes (for automation & release notes)

Added

  • autoarray.util.fnnls.fnnls_cholesky(ZTZ, ZTx, P_initial=, stats: Optional[dict] = None) — stats is filled with outer_iterations, inner_iterations, passive_set, n_passive, warm_start_errors; P_initial may be a bool mask or an int index array
  • autoarray.inversion.inversion.inversion_util.reconstruction_positive_only_from(..., fingerprint=None) — identifies the passive set's index space for the memo; None disables the memo for that call
  • autoarray.Settings(nnls_warm_start_memo: Optional[bool] = None) + property — NumPy/numba path only
  • autoarray/config/general.yaml → inversion.nnls_warm_start_memo: true
  • autoarray.Settings(nnls_warm_start_error_tolerance: Optional[float] = None) + property — config inversion.nnls_warm_start_error_tolerance: 1.5; non-finite / ≤ 0 disables the guard
  • autoarray.inversion.inversion.nnls_memo — MemoEntry(passive_set, dense_error_fraction), memo_enabled(), memo_key(n, fingerprint), passive_set_get(key, n) -> Optional[MemoEntry], passive_set_put(key, passive_set, dense_error_fraction), memo_drop(key); FIFO, 8 entries; env AUTOARRAY_NNLS_WARM_START=0 disables
  • reconstruction_positive_only_from writes stats["seed_source"] ("memo"/"dense") and stats["warm_start_fallback"] after the solve
  • AbstractInversion._nnls_warm_start_fingerprint(ids_to_keep=None) (private)

Changed Behaviour

  • fnnls_cholesky: warm start factorises the initial passive set once and extends it in place (was: dense scipy.linalg.solve + full cholesky on the first outer iteration); a warm passive set with non-positive entries is repaired before the outer loop (was: clipped, and returned unsolved when P was all-True); near-singular warm-start systems now raise LinAlgError from cholesky instead of LinAlgWarning + garbage from solve
  • reconstruction_positive_only_from (NumPy path): warm-starts from the memo when settings.nnls_warm_start_memo and a fingerprint are given; a memo-seeded solve that raises retries once from the dense-sign start before the InversionException path

Migration

  • None required. To restore the previous start: aa.Settings(nnls_warm_start_memo=False) or AUTOARRAY_NNLS_WARM_START=0.

Generated by the PyAutoLabs agent workflow.

Jammy2211 and others added 2 commits August 28, 2026 10:06
…ation's passive set (#498)

Phase 3a of the numba CPU likelihood speed restoration. fnnls_cholesky gains
solver diagnostics (stats dict), accepts a mask or index warm start, reuses one
Cholesky factorisation for the warm start instead of a dense solve plus a
from-scratch refactorisation, and repairs an infeasible warm passive set before
the outer loop (an all-True seed previously bypassed the loop and returned the
clipped vector). A process-local FIFO memo (nnls_memo.py) seeds each solve from
the previous evaluation's final passive set, keyed on the index space
(mesh/data shapes + edge-zeroed subset); Settings.nnls_warm_start_memo, default
on, AUTOARRAY_NNLS_WARM_START=0 disables; a failed seeded solve retries once
from the dense-sign start.

Delaunay-1250 fiducial, random-walk sequence: median active-set iterations
69.5 -> 7.0 (euclid), 32 -> 8 (hst); solve 0.54 -> 0.05 s / 0.25 -> 0.07 s;
log-likelihood parity 3e-14. Solve time is never worse than the dense-sign
start, including on uncorrelated i.i.d. jumps.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XySQZP6Npj4adZMkYZ96Bo
Each memo entry now also carries the error fraction of the most recent
dense-sign-started solve for its key; a seeded solve whose own error fraction
exceeds Settings.nnls_warm_start_error_tolerance (config default 1.5) times
that reference is dropped, so the next solve restarts from the dense-sign
start and refreshes the reference. Relative, not absolute: the 32-cell
robustness matrix showed the absolute seed error fraction does not separate
seeds that save iterations from seeds that cost them (0.048-0.138 overlap),
while the seed/dense ratio does (helpful <= 0.89, worst 1.42). Non-finite or
non-positive tolerance disables the guard. reconstruction_positive_only_from
records seed_source and warm_start_fallback in the stats dict.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XySQZP6Npj4adZMkYZ96Bo
@Jammy2211

Copy link
Copy Markdown
Collaborator Author

Workspace leg: PyAutoLabs/autolens_profiling#184 (merges after this PR — library-first gate).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

perf: cut positive-only solve active-set iterations — warm-start memo + diagnostics (numba CPU likelihood, phase 3a)

1 participant