Skip to content

perf: reuse Convolver state, cache operated-matrix dict, hoist linear-func pair loop (#496 phase 1) - #497

Merged
Jammy2211 merged 1 commit into
mainfrom
feature/numba-cpu-mge-batch-convolve-cache
Aug 27, 2026
Merged

Jammy2211 merged 1 commit into
mainfrom
feature/numba-cpu-mge-batch-convolve-cache

Conversation

@Jammy2211

@Jammy2211 Jammy2211 commented Aug 27, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Companion PR: PyAutoLabs/PyAutoGalaxy#588

Phase 1 of the numba CPU sparse-operator likelihood speed restoration (#496, epic numba-cpu-likelihood). Three exact, likelihood-preserving changes on the CPU (use_jax=False) inversion path, plus one companion fix:

  • Convolver.state_from reuses the precomputed state. The old shape test compared the mask to the kernel, so Imaging's psf_setup_state was never returned for convolve_over_sample_size == 1 and every convolution rebuilt ConvolverState (mask resize + rfft2). Now returns it iff ConvolverState.is_for_mask(mask).
  • linear_func_operated_mapping_matrix_dict is a cached_property on AbstractInversionImaging (one build per inversion; the numba subclass already did this via its memo). Companion fix: the numba override reached the parent through .fget, which autonerves.CachedProperty does not have — now .func.
  • Linear-func × linear-func curvature blocks (imaging/sparse.py, imaging_numba/sparse.py): noise-weight each operated matrix once instead of inside both loops, compute only the upper triangle of blocks and mirror the transpose. Exact for any per-pixel (non-constant, non-symmetric) noise map since block(i,j) = M_iᵀ diag(1/σ²) M_j.

Pairs with PyAutoGalaxy PR (batched MGE convolution) — both are independently mergeable; the measured gain needs both.

Measured (memo disabled, fresh FitImaging per call, OMP=1): MGE-60 operated matrix hst 1.11 s → 0.30 s, euclid 0.68 s → 0.125 s; matrices bitwise identical. Breakdown steps "Mapped reconstruction" and "Model data + chi²" ~3.9× faster, "Blurred image" ~1.9× (state reuse). hst pinned log-likelihood 27661.910133664103 unchanged before/after; euclid log-likelihood bit-identical before/after.

API Changes

None — internal changes only. ConvolverState gains source_mask and is_for_mask(); state_from returns the stored state in more cases (identical output).
See full details below.

Test Plan

  • pytest test_autoarray — 1237 passed
  • New: test__state_from__precomputed_state_reused_for_matching_mask, test__precomputed_state__convolution_bit_identical_to_no_state, test__convolved_mapping_matrix_via_real_space_np__matches_image_convolution_per_column (test_convolver.py); test_curvature_matrix_func_list_blocks.py (both sparse inversion classes, random non-symmetric noise map, 1e-12 vs brute force)
  • autolens_profiling likelihood_breakdown/pixelization_numba.py hst pin PASSED (rtol 1e-6), euclid before/after bit-identical
  • CI green on both PRs
Full API Changes (for automation & release notes)

Added

  • autoarray.operators.convolver.ConvolverState.source_mask — the pre-resize mask the state was built from
  • autoarray.operators.convolver.ConvolverState.is_for_mask(mask) — whether the state can be reused for mask

Changed Behaviour

  • Convolver.state_from(mask) — returns the precomputed _state whenever is_for_mask(mask); previously only when the mask shape equalled the kernel shape. Output identical.
  • AbstractInversionImaging.linear_func_operated_mapping_matrix_dict — cached_property (was property); value identical, built once per inversion.

Notes

  • Two formatting-only hunks in convolver.py come from black on pre-existing drift.

Generated by the PyAutoLabs agent workflow.

🤖 Generated with Claude Code

https://claude.ai/code/session_01N6xyNMYmffpHodBrkc2d91

…hoist linear-func pair loop

Phase 1 of the numba CPU likelihood speed restoration (#496).

- Convolver.state_from compared the mask shape to the *kernel* shape, so the
  precomputed ConvolverState that Imaging builds (psf_setup_state) was never
  returned for convolve_over_sample_size == 1 and every convolution rebuilt it
  (mask resize + rfft2). It now returns the stored state iff
  ConvolverState.is_for_mask(mask) (pixel scales, shape and mask array match).
- AbstractInversionImaging.linear_func_operated_mapping_matrix_dict is a
  cached_property (one build per inversion, matching the numba subclass); the
  numba override reaches the parent via .func (autonerves CachedProperty has no
  .fget).
- Linear-func x linear-func curvature blocks (imaging/sparse.py and
  imaging_numba/sparse.py): weight each operated matrix by the noise map once,
  compute only the upper triangle of blocks and mirror the transpose. Exact for
  any per-pixel noise map (block = M_i^T diag(1/sigma^2) M_j).

Tests: state reuse + bit-identical output, mapping-matrix convolution with a
blurring matrix vs per-column image convolution, pair-loop vs brute-force
double loop with a random non-symmetric noise map (both inversion classes).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6xyNMYmffpHodBrkc2d91
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant