Skip to content

perf(decorators): to_grid propagates only an explicit over_sampled; Grid2D.over_sampled short-circuits at sub-size 1 (#514) - #516

Merged
Jammy2211 merged 1 commit into
mainfrom
feature/numpy-deflections-p1
Sep 2, 2026
Merged

Jammy2211 merged 1 commit into
mainfrom
feature/numpy-deflections-p1

Conversation

@Jammy2211

Copy link
Copy Markdown
Collaborator

Summary

Phase 1 of the numpy-deflections-cpu epic (#514). Two zero-numerics fixes on the grid-decorator path that every mass-profile deflection call goes through:

  1. GridMaker.via_grid_2d no longer materialises the over-sampled grid of an already-wrapped Grid2D. It propagated getattr(result, "over_sampled") / getattr(result, "over_sampler"); when the decorated body returns a Grid2D (the chained @to_grid methods of every spherical mass profile in PyAutoGalaxy) those getattrs fire the lazy properties, which run a per-pixel Python loop (over_sample_util.grid_2d_slim_over_sampled_via_mask_from: two np.linspace + one np.meshgrid per pixel) on every call. The value built that way was also mask-derived, so it ignored the profile's centre and rotation; nothing ever read it. Both sites now read the private _over_sampled / _over_sampler, so only a value the caller explicitly set propagates (the two load-bearing callers, Galaxy.traced_grid_2d_from and Tracer.traced_grid_2d_list_from, set it explicitly and round-trip unchanged).
  2. Grid2D.over_sampled short-circuits at uniform sub-size 1 to a copy of the slim grid as a Grid2DIrregular (the loop's output at sub-size 1 is the slim grid in the same order; with a non-zero mask origin the two differ by at most 1 ULP, 2.8e-16). This skips the same loop for every sub-size-1 grid, e.g. the pixelization grid of every CPU likelihood evaluation.

Measured with the new autolens_profiling/scripts/lens/deflections/ cells (hst, 15,361 pixels, OMP_NUM_THREADS=1, direct Grid2D call, before → after): IsothermalSph 699 ms → 0.92 ms, PowerLawSph 701 ms → 1.37 ms, NFWSph 591 ms → 5.8 ms, gNFWSph 1.20 s → 349 ms (the remainder is the MGE, phase 2). Elliptical profiles were never affected (their bodies return a bare array). Every deflection pin passes at rtol 1e-6. Note the direct-Grid2D penalty did not reach the likelihood's ray-trace, which wraps each plane in Grid2DIrregular first; the likelihood-side gain of this phase is the companion tracer change (PyAutoLens / PyAutoGalaxy PRs).

API Changes

None — internal changes only. Grid2D.over_sampled at uniform sub-size 1 now returns a copy of the slim grid (a Grid2DIrregular, as before) instead of the loop's identical output.
See full details below.

Test Plan

  • test_autoarray: 1362 passed (4 new). Negative pin test__to_grid__does_not_materialise_over_sampled_of_wrapped_grid control-tested: fails on main, passes here.
  • Positive pin: an explicit over_sampled= sentinel survives via_grid_2d.
  • Grid2D.over_sampled at sub-size 1 equals the slim grid; at sub-size 2 it has 4× the points and differs.
  • autolens_profiling/scripts/lens/deflections/{total,dark}.py --instrument hst: pins PASSED (rtol 1e-6), timings above.
  • CI green; downstream PyAutoGalaxy / PyAutoLens suites (run in the task worktree against this branch: 1158 / 576 passed).
Full API Changes (for automation & release notes)

Changed Behaviour

  • autoarray.structures.decorators.to_grid.GridMaker.via_grid_2d — propagates only an explicitly-set _over_sampled / _over_sampler from a wrapped Grid2D result; never triggers the lazy properties.
  • autoarray.structures.grids.uniform_2d.Grid2D.over_sampled — at uniform over_sample_size == 1 returns a Grid2DIrregular copy of the slim grid without running the per-pixel over-sampling loop (≤ 1 ULP from the loop's output for a non-zero mask origin; identical otherwise).

Generated by the PyAutoLabs agent workflow. Epic ledger: PyAutoMind/draft/feature/autogalaxy/numpy_deflections_cpu_speedup.md. Companion PRs: PyAutoGalaxy + PyAutoLens (tracer double trace at sub-size 1), autolens_profiling (scripts/lens/ measurement package + before/after).

)

`GridMaker.via_grid_2d` read `result.over_sampled` / `result.over_sampler`
through the public properties. When the wrapped result is already a `Grid2D`
— which it is for every spherical mass profile, whose `@to_grid` decorated
`transformed_to_reference_frame_grid_from` delegates to another `@to_grid`
method — that read *materialises* the over sampled grid, running a per-pixel
Python loop over the whole mask. On a 15k-pixel HST grid that is ~0.5-1.2 s
per deflection-angle call against ~1 ms of actual profile maths, and the
value built is mask-derived, so for a translated/rotated grid it is wrong as
well as expensive. Nothing reads it.

Both call sites now read the private `_over_sampled` / `_over_sampler`, so
only a value a caller explicitly passed in propagates (the load-bearing case:
`Galaxy.traced_grid_2d_from` and `Tracer.traced_grid_2d_list_from`, which
both construct `Grid2D(..., over_sampled=...)`). Anything else is left as
`None` for the new grid to compute lazily, if it is ever asked for.

`Grid2D.over_sampled` additionally short circuits at a uniform sub size of 1,
where every sub pixel is the pixel itself and the over sampled grid is just
the slim grid in the same order — equal to the loop's output bit for bit at
the default origin, and within 1 ULP of it otherwise. The values are copied
so an in-place edit of `over_sampled` cannot write through to the grid.

Measured on hst (15361 image pixels, OMP_NUM_THREADS=1), `Grid2D` s/call:

  IsothermalSph   699.2 ms -> 0.92 ms
  PowerLawSph     701.0 ms -> 1.37 ms
  NFWSph          590.9 ms -> 5.79 ms
  gNFWSph          1.20 s  -> 348.6 ms   (the remainder is real quadrature)

All pinned deflection values still PASS at rtol 1e-6.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HWjPT94MPbEHT45kJmDpDh
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant