Skip to content

perf: numpy deflections phase 1 — scripts/lens baseline, *Sph over-sampled re-materialisation, tracer double trace #514

Description

@Jammy2211

Overview

Phase 1 of the numpy-deflections-cpu epic (PyAutoMind ledger draft/feature/autogalaxy/numpy_deflections_cpu_speedup.md). The numba CPU likelihood route (use_jax=False, apply_sparse_operator_cpu()) now runs an HST rectangular evaluation in 0.33 s, but its mass-profile deflections sit on the plain-numpy xp=np branch that was never tuned for CPU. A probe on the HST grid (15,361 points, OMP_NUM_THREADS=1) shows every *Sph profile costing 0.48–0.83 s per call against ~1 ms of physics, and the Tracer evaluating every Grid2D twice. This phase lands the measurement package the whole epic reports through (autolens_profiling/scripts/lens/deflections/), then takes the two levers that change no numerics at all: the *Sph over-sampled re-materialisation in the to_grid decorator (~500× on four profiles) and the redundant second trace at sub-size 1 (~2× at tracer level on every profile).

Plan

  • Build autolens_profiling/scripts/lens/deflections/ (a new dataset-free, library-component profiling axis) timing all nine numpy mass-profile deflections on the hst + euclid grids with per-profile pins, and commit the baseline artifacts before any library edit.
  • Fix GridMaker.via_grid_2d so wrapping an already-wrapped Grid2D no longer fires the over_sampled property (and its per-pixel Python loop) on every *Sph deflection call; short-circuit Grid2D.over_sampled to the slim grid when every sub-size is 1 (bit-identical).
  • Stop Tracer.traced_grid_2d_list_from and Galaxy.traced_grid_2d_from tracing the over-sampled grid a second time when the over-sampler is uniform size 1.
  • Pin all of it: decorator regression tests, tracer semantics test, per-profile deflection pins rtol 1e-6, existing numba likelihood pins rtol 1e-6.
  • Ship library-first (PyAutoArray → PyAutoGalaxy → PyAutoLens), then the profiling package with the after-numbers and a note.
Detailed implementation plan

Affected Repositories

  • PyAutoArray (primary)
  • PyAutoGalaxy
  • PyAutoLens
  • autolens_profiling

Branch Survey

Repository Current Branch Dirty?
./PyAutoArray main (302d5df) clean
./PyAutoGalaxy main (2ee44d50) clean
./PyAutoLens main (af514d179) clean
./autolens_profiling feature/refs-v1-redo-ruling (021b41b, another task's branch) clean — task worktree branches from origin/main

PyAutoArray is also claimed by numba-vs-jax-sparse as a read-only research verdict on main (no branch, no worktree); file sets are disjoint, so this task runs in its own worktree with a parallel-claim: recorded in active.md.

Suggested branch: feature/numpy-deflections-p1
Worktree: ~/Code/PyAutoLabs-wt/numpy-deflections-p1/

Implementation Steps

Step 0 — measurement package (autolens_profiling), baseline committed before any library edit

scripts/lens/README.md                      axis narrative, links down
scripts/lens/deflections/README.md          <!-- BEGIN auto-table:deflections -->
scripts/lens/deflections/_driver.py         grid construction (as pixelization_numba.py: Mask2D.circular r=3.5",
                                            radial over-sampling [4,2,1]@[0.3,0.6], pixelization sub-size 1),
                                            timing loop, separate cProfile pass, pin check, JSON+PNG write
scripts/lens/deflections/_profiles.py       registry: name -> (ctor, fiducial params, family)
scripts/lens/deflections/total.py           Isothermal, IsothermalSph, PowerLaw, PowerLawSph
scripts/lens/deflections/dark.py            NFW, NFWSph, gNFW, gNFWSph
scripts/lens/deflections/stellar.py         Gaussian (ell + ell_comps=(0,0))
results/lens/deflections/<cell>_summary_<instrument>_v<version>.{json,png}
results/notes/numpy_deflections_cpu.md      before table + epic ledger
  1. Each cell: the standard prologue (ruff.toml root-finder, sys.path, AUTOLENS_PROFILING_SMOKE early exit as pixelization_numba.py:82-93), parse_profile_cli + --instrument {hst,euclid}, resolve_output_paths, device_info_dict. Per profile: median-of-N s/call on Grid2D, on Grid2DIrregular (grids.lp.over_sampled), and through a two-plane Tracer; n_points; cProfile top-12 (cprofile_top, per-call, attribution only, separate pass); pin block {abs_sum, abs_max, sample} with 16 fixed (y, x) arcsec coordinates (8 @ r=1.0", 4 @ 0.15", 4 @ 2.8") on a dedicated Grid2DIrregular, checked rtol 1e-6 via check_pinned + a small check_pinned_vector helper beside _profile_cli.py:313; drift recorded, never adjudicated. --repin requires --repin-reason, refuses shifts > --repin-max-shift (1e-3) without --repin-force, embeds pin_provenance in the JSON.
  2. scripts/misc/tooling/build_readme.py: no scan-root or regex change (artifacts use the summary purpose token; the rglob already walks results/lens/); add _render_deflections_table beside :444-479, register "deflections" in _build_renderers (:762-771), add the new README to TARGET_READMES (:779-788, via REPO_ROOT), fix the comment at :192 and docstring :20-42.
  3. .github/workflows/lint.yml:131-146 smoke gains python scripts/lens/deflections/total.py; workflows README row. README.md:142-144 amended: scripts/lens/ is a second, dataset-free axis.
  4. Run hst + euclid for all three cells; commit artifacts + note (before table). Ruff clean; build_readme.py --check clean.

Step 1 — PyAutoArray: the *Sph bug + the sub-size-1 short-circuit

  1. autoarray/structures/decorators/to_grid.py:17-18 (and :26-27): read _over_sampled / _over_sampler instead of the properties. Identity for the two load-bearing callers (autogalaxy/galaxy/galaxy.py:381, autolens/lens/tracer.py:513 set them explicitly); the lazy recompute elsewhere gives the same answer (same mask / sub-size inputs). Verified in the probe: 0.644 s → 0.0016 s per IsothermalSph call, max abs diff 0.0.
  2. autoarray/structures/grids/uniform_2d.py:211-224 Grid2D.over_sampled: return Grid2DIrregular(self.array) when all sub-sizes are 1 (verified bit-identical to the loop output, same order). Kills the ~1.5 s Python loop (over_sample_util.py:417-429) for every sub-size-1 grid.
  3. Tests in test_autoarray/structures/decorators/test_to_grid.py: (i) negative pin — monkeypatch Grid2D.over_sampled to fail, an IsothermalSph.deflections_yx_2d_from call must return; (ii) positive pin — an explicit over_sampled= sentinel survives via_grid_2d; (iii) sub-size-1 over_sampled equals the slim grid, sub-size-4 still differs.

Step 2 — the double trace

  1. PyAutoLens/autolens/lens/tracer.py:498-501: when np.all(grid.over_sample_size.array == 1) (host-side static bool — the mask is numpy, so this is a static branch under jit, not a traced cond), wrap the already-traced grid_2d_list as the over-sampled list instead of tracing again. Same in PyAutoGalaxy/autogalaxy/galaxy/galaxy.py:376-383.
  2. Test in test_autolens/lens/test_tracer.py: traced .over_sampled equals traced(g.over_sampled) and differs from g.over_sampled (pins that the tracer's explicit value still survives via_grid_2d), for sub-size 1 and 4.

Ship (library-first)

  1. PyAutoArray PR → PyAutoGalaxy PR → PyAutoLens PR → autolens_profiling PR with the after-numbers: re-run total.py/dark.py/stellar.py hst + euclid; re-run pixelization_numba.py + delaunay_numba.py hst + euclid (likelihood pins rtol 1e-6; the "Inversion build (trace+mesh+mapper)" row should drop); re-run the jax_compile warm-compile pins (one fewer duplicate subgraph). test_autoarray, test_autogalaxy, test_autolens green; smoke autolens_workspace/scripts/imaging/features/pixelization/cpu_fast_modeling.py.

Expected: *Sph ≥ 300×, tracer-level cost ≈ 2× cheaper for every profile on sub-size-1 grids; deflections and likelihoods bit-for-bit unchanged.

Key Files

  • PyAutoArray/autoarray/structures/decorators/to_grid.py — GridMaker.via_grid_2d, the getattr(result, "over_sampled") that fires the property (lines 17-18, 26-27)
  • PyAutoArray/autoarray/structures/grids/uniform_2d.py — Grid2D.over_sampled lazy property (211-224), _over_sampled storage (180)
  • PyAutoArray/autoarray/operators/over_sampling/over_sample_util.py — the per-pixel loop (377-431) the fixes avoid
  • PyAutoGalaxy/autogalaxy/profiles/geometry_profiles.py — the double-@to_grid *Sph delegation (158-171, 402-417); unchanged, the fix is one layer down
  • PyAutoGalaxy/autogalaxy/galaxy/galaxy.py — traced_grid_2d_from double trace (376-383)
  • PyAutoLens/autolens/lens/tracer.py — traced_grid_2d_list_from double trace (490-516)
  • autolens_profiling/scripts/imaging/likelihood_breakdown/pixelization_numba.py — the cell whose grid construction and prologue the new cells mirror
  • autolens_profiling/scripts/misc/tooling/build_readme.py — renderer + README registry (444-479, 762-771, 779-788)
  • autolens_profiling/_profile_cli.py — check_pinned (313-345), record_pinned_check (347-360)

Original Prompt

Click to expand starting prompt

Numpy deflections phase 1: measure (scripts/lens), fix the *Sph over-sampled re-materialisation, drop the tracer double trace

Type: feature
Epic: numpy-deflections-cpu
Phase: 1
Target: autoarray
Repos:

  • @PyAutoArray
  • @PyAutoGalaxy
  • @PyAutoLens
  • @autolens_profiling
    Themes:
  • numba-cpu
  • mass-profiles
  • profiling
    Difficulty: large
    Autonomy: supervised
    Priority: high
    Status: draft
    Filed: 2026-09-02
    Issued: 2026-09-02

Phase 1 of the numpy-deflections-cpu epic — ledger
draft/feature/autogalaxy/numpy_deflections_cpu_speedup.md, which holds the measured before-table,
the goal (every one of the nine numpy deflections_yx_2d_from at least 2× faster on the HST grid,
the four *Sph ≥ 100×, gNFW/gNFWSph ≥ 5×), the request's constraints (no new public functions,
keep the xp API single-bodied) and the approved decisions 1–5. This phase lands the measurement
package the whole epic reports through, then takes the two levers that change no numerics at all:
the *Sph over-sampled re-materialisation (~500× on four profiles) and the tracer's double trace
(~2× at tracer level on every profile). Issue on PyAutoArray (the load-bearing code change),
cross-referenced from the other PRs.

Goal

scripts/lens/deflections/ measuring all nine profiles on hst + euclid with committed baseline
artifacts, then two zero-numerics fixes: *Sph ≥ 300×, tracer-level cost ≈ 2× cheaper for every
profile on sub-size-1 grids. Deflections and likelihoods bit-for-bit unchanged (rtol 1e-6 pins).

Branch feature/numpy-deflections-p1, worktree ~/Code/PyAutoLabs-wt/numpy-deflections-p1/.

Step 0 — autolens_profiling/scripts/lens/deflections/ (baseline committed BEFORE any library edit)

scripts/lens/README.md                      axis narrative, links down
scripts/lens/deflections/README.md          <!-- BEGIN auto-table:deflections -->
scripts/lens/deflections/_driver.py         grid construction (as pixelization_numba.py: Mask2D.circular r=3.5",
                                            radial over-sampling [4,2,1]@[0.3,0.6], pixelization sub-size 1),
                                            timing loop, separate cProfile pass, pin check, JSON+PNG write
scripts/lens/deflections/_profiles.py       registry: name -> (ctor, fiducial params, family)
scripts/lens/deflections/total.py           Isothermal, IsothermalSph, PowerLaw, PowerLawSph
scripts/lens/deflections/dark.py            NFW, NFWSph, gNFW, gNFWSph
scripts/lens/deflections/stellar.py         Gaussian (ell + ell_comps=(0,0))
results/lens/deflections/<cell>_summary_<instrument>_v<version>.{json,png}
results/notes/numpy_deflections_cpu.md      before table + epic ledger
  • Each cell: the standard prologue (ruff.toml root-finder, sys.path, AUTOLENS_PROFILING_SMOKE
    early exit as pixelization_numba.py:82-93), parse_profile_cli + --instrument {hst,euclid},
    resolve_output_paths, device_info_dict. Per profile: median-of-N s/call on Grid2D, on
    Grid2DIrregular (grids.lp.over_sampled), and through a two-plane Tracer; n_points; cProfile
    top-12 (cprofile_top, per-call, attribution only, separate pass); pin block
    {abs_sum, abs_max, sample} with 16 fixed (y, x) arcsec coordinates (8 @ r=1.0", 4 @ 0.15",
    4 @ 2.8") on a dedicated Grid2DIrregular, checked rtol 1e-6 via check_pinned + a small
    check_pinned_vector helper added beside _profile_cli.py:313; drift recorded, never adjudicated.
    --repin requires --repin-reason, refuses shifts > --repin-max-shift (1e-3) without
    --repin-force, embeds pin_provenance in the JSON.
  • build_readme.py: no scan-root or regex change (artifacts use the summary purpose token; the
    rglob already walks results/lens/); add _render_deflections_table beside :444-479,
    register "deflections" in _build_renderers (:762-771), add the new README to
    TARGET_READMES (:779-788, via REPO_ROOT), fix the comment at :192 and docstring :20-42.
  • lint.yml:131-146 smoke gains python scripts/lens/deflections/total.py; workflows README row.
  • README.md:142-144 amended: scripts/lens/ is a second top-level axis (library-component
    profiling, dataset-free) beside the dataset-first families (epic decision 4).
  • Ruff clean; build_readme.py --check clean.

Step 1 — PyAutoArray: the *Sph bug + the sub-size-1 short-circuit

The *Sph ~500 ms is a decorator bug, not physics. EllProfile.transformed_to_reference_frame_grid_from
(geometry_profiles.py:402-417, @to_grid) delegates to the also-@to_grid SphProfile method
(:158); the outer GridMaker.via_grid_2d (autoarray/structures/decorators/to_grid.py:17-18)
then does getattr(result, "over_sampled", None) on the wrapped Grid2D, firing the property and
its per-pixel Python loop (over_sample_util.py:417-429) on every call. The materialised value is
also semantically wrong (mask-derived, so not translated by the profile centre) — nothing reads it.

  • autoarray/structures/decorators/to_grid.py:17-18 (and :26-27): read _over_sampled /
    _over_sampler instead of the properties. Identity for the two load-bearing callers
    (galaxy.py:381, tracer.py:513 set them explicitly); lazy recompute elsewhere gives the same
    answer (same mask/sub-size inputs). Verified: 0.644 s → 0.0016 s, max abs diff 0.0.
  • autoarray/structures/grids/uniform_2d.py:211-224 Grid2D.over_sampled: return
    Grid2DIrregular(self.array) when all sub-sizes are 1 (verified bit-identical — same points, same
    order). Kills the ~1.5 s Python loop for every sub-size-1 grid.
  • Tests in test_autoarray/structures/decorators/test_to_grid.py: (i) negative pin —
    monkeypatch Grid2D.over_sampled to fail, IsothermalSph.deflections_yx_2d_from must return;
    (ii) positive pin — an explicit over_sampled= sentinel survives via_grid_2d; (iii) sub-size-1
    over_sampled equals the slim grid; sub-size-4 still differs.

Step 2 — the double trace

Tracer.traced_grid_2d_list_from traces every Grid2D twice (autolens/lens/tracer.py:490-505:
grid, then grid.over_sampled); Galaxy.traced_grid_2d_from (galaxy.py:376-383) does the same.
With a uniform size-1 over-sampler (the pixelization grid in every numba cell) the second trace is
100 % redundant: Isothermal 1.4 → 3.5 ms, PowerLaw 7.4 → 15.5 ms through the tracer.

  • PyAutoLens/autolens/lens/tracer.py:498-501: when np.all(grid.over_sample_size.array == 1)
    (host-side static bool, safe under jit), wrap the already-traced grid_2d_list as the
    over-sampled list instead of tracing again. Same in PyAutoGalaxy/autogalaxy/galaxy/galaxy.py:376-383.
  • Test in test_autolens/lens/test_tracer.py: traced .over_sampled equals traced(g.over_sampled)
    and differs from g.over_sampled (pins that the tracer's explicit value still survives), for
    sub-size 1 and 4.

Ship (library-first)

PyAutoArray PR → PyAutoGalaxy PR → PyAutoLens PR → autolens_profiling PR with the after-numbers:
re-run total.py/dark.py/stellar.py hst + euclid; re-run pixelization_numba.py +
delaunay_numba.py hst + euclid (likelihood pins rtol 1e-6 — hst rectangular 27661.910133664103;
the "Inversion build (trace+mesh+mapper)" row should drop); re-run the jax_compile warm-compile
pins (one fewer duplicate subgraph).

Verification

  • test_autoarray, test_autogalaxy, test_autolens green; smoke
    autolens_workspace/scripts/imaging/features/pixelization/cpu_fast_modeling.py.
  • Baseline artifacts committed before any library edit; after-artifacts in the profiling PR;
    README auto-table regenerated (build_readme.py --check is a lint gate).
  • ruff check + ruff format --check; lint smoke green.
  • Expected: *Sph ≥ 300×, tracer-level cost ≈ 2× for every profile on sub-size-1 grids.

Out of scope

Numerics of any profile (phases 2 and 3); the JAX path's speed; CSE; convergence_2d_from /
potential_2d_from; new mass profiles or public methods.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions