Skip to content

fix: NaN JAX gradient on MGE positive-only solves after #572 - #574

Merged
Jammy2211 merged 4 commits into
mainfrom
feature/mge-nnls-grad-nan
Sep 25, 2026
Merged

Jammy2211 merged 4 commits into
mainfrom
feature/mge-nnls-grad-nan

Conversation

@Jammy2211

Copy link
Copy Markdown
Collaborator

Summary

Fixes NaN JAX gradients on mapper-less (MGE / linear light profile) positive-only inversions, introduced by #572 (the #571 raw-forward PDIP fix). Closes #573.

#572's "raw" mode runs the forward PDIP on the raw system with a loose data-scaled tolerance, which leaves the complementarity s·z about 1e-10 to 2.5e-9. The custom-vjp backward pass then called solve_relaxed_nnls on the Jacobi system at nnls_target_kappa=1e-11 from that iterate. That is a push toward the boundary with z/s around 1e13-1e14; under jit it overshoots (s, z < 0), hits the 50-iteration cap and returns NaN, and diff_nnls passes the NaN to every gradient entry. On main, 4 of 16 perturbation points of autolens_workspace_test/scripts/imaging/jax_grad/mge.py are NaN eagerly, and the jitted value+grad is NaN at the script's own point, which is how it failed Heart Release Integrate (run 36108062907).

Fix: in the backward pass only, polish the mapped iterate with at most RAW_BACKWARD_POLISH_MAX_ITER=10 warm-started PDIP iterations on (Q_pc, q_pc) at jaxnnls's tight tolerance before the relaxed solve. The polished point is kept only if it converged with finite y and s, z > 0; otherwise the unpolished iterate is used, as before. The primal is unchanged, so the forward value and the #571 forward-convergence fix are untouched: the jitted forward HLO is byte-identical to main.

Why not raise the kappa instead: the alternative kappa_eff = max(target_kappa, max(s·z)) was measured and rejected. It left 6/48 SLaM systems NaN and was worse under jit (11/16), and on one system the relaxed solve reported "converged" with s < 0.

check main kappa_eff polish (this PR) pre-#572 jacobi
jax_grad/mge.py, 16 keys finite (eager / jit) 12 / 15 13 / 11 16 / 16 16 / 16
FD agreement (worst max-rel) 1.59e-6 1.63e-6 1.59e-6 1.59e-6
48 SLaM #571 systems, finite gradient 48 42 48 47
backward iterations relaxed 4-6, or 50 → NaN up to 50 polish 4-6 (48/48 converged), relaxed 1 n/a

Corrective PR under Heart RED (human-authorized)

  • Heart RED reason (verbatim): release validation FAILED (stage integrate). Re-read unchanged at ship time.
  • Authorization: the human, live, 2026-09-25: "I authorize on heart red". It is recorded with the causal mapping on fix: NaN JAX gradient on MGE positive-only solves after #572 #573 (comment). Scope: this one PR. Merging is a separate human /prm; no release or rehearsal.
  • What this clears: the autolens_workspace_test imaging/jax_grad/mge.py failure in that integrate run. The run's other two failures (autolens_workspace weak/real_data/a2744.py, cluster/lenstool/modeling.py) were fixed by autolens_workspace#577 (merged).
  • Heart stays RED, and release stays blocked, until a fresh Release Integrate run passes on wheels that include this fix.

API Changes

Additive only; no signature break or default change.

  • solve_nnls gains an optional init=(x, s, z) warm start (default None = unchanged).
  • New raw_forward_backward_status(...) reports the backward pass's convergence (relaxed + polish flags and iterations), and a new RAW_BACKWARD_POLISH_MAX_ITER constant.
  • Changed behaviour: raw-mode (nnls_preconditioning_no_mapper: raw) gradients are now finite on the previously NaN points. Forward values are unchanged.

See full details below.

Test Plan

  • New red-first regression fixture files/mge_grad_nan_systems.npz (6.8 KB, 4 systems captured at the NaN points). The gradient test fails 6/8 on 3de624b5 and passes on the branch. Also added: backward-status tests over 8 SLaM + 4 captured systems.
  • Targeted files: 77 passed. Full PyAutoArray suite: 1702 passed. The existing fix(inversion): JAX PDIP positive-only solve fails to converge on SLaM MGE systems #571 tests are unmodified and green.
  • End-to-end autolens_workspace_test/scripts/imaging/jax_grad/mge.py under the release profile (py3.12, jax 0.10.2, numpy 2.5.3): own point PASS; 16-key sweep 16/16 finite (eager and jit), FD 16/16; logL bit-identical to main.
  • Runtime vs autolens_profiling (local CPU fp64 HST, medians, main vs branch alternated):
    • runtime/mge.py single-JIT forward: 31.22 vs 32.12 ms (noise; HLO identical)
    • vmap per call: 26.43 vs 25.80 ms
    • value+grad: 104.14 vs 101.31 ms (−2.7%)
    • hazards/mge_nnls_capture.py: 48/48 converged on both
  • Independent review: CLEAN (two low docstring notes, fixed in 44762fa).
  • Downstream smoke (pyauto-heart smoke --root <task worktree>, branch PyAutoArray on PYTHONPATH): 159 passed across autofit / autogalaxy / autolens / autolens_workspace_test / euclid / howtolens, scripts and notebooks. 1 failure, pre-existing on main: autolens_workspace_test/scripts/interferometer/jax_likelihood/mge.py vmap likelihood -45560751.36 vs pin -3152.65, identical on main 3de624b5 in the same env (not caused by this PR; flagged separately). autocti / autocti_test did not run (local GSL headers missing, environment only).
  • After merge: fresh Release Integrate run (nightly or re-dispatch).
Full API Changes (for automation & release notes)

Added

  • autoarray.util.jax_nnls.raw_forward_backward_status(Q_pc, q_pc, Q, q, D, target_kappa=1e-3, solver_tol=None, max_iter=50): returns (relaxed_converged, relaxed_iter, polish_converged, polish_iter) of the raw-mode backward pass
  • autoarray.util.jax_nnls.RAW_BACKWARD_POLISH_MAX_ITER = 10

Changed Signature

  • autoarray.util.jax_nnls.solve_nnls(Q, q, solver_tol=None, max_iter=50, init=None): optional (x, s, z) warm start

Changed Behaviour

Migration

  • None required.

Generated by the PyAutoLabs agent workflow.

🤖 Generated with Claude Code

Jammy2211 and others added 4 commits September 25, 2026 17:14
Four 20x20 systems captured from autolens_workspace_test jax_grad/mge.py at
PRNGKey perturbations 2, 10, 12, 14: the raw-mode gradient is NaN on each on
3de624b (eager 4/4, jit 2/4). Adds backward-pass convergence tests over
these and the 8 SLaM #571 systems via raw_forward_backward_status.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…olve (#573)

The raw forward solve (#572) stops at the data-scaled tolerance, leaving
s*z ~1e-10..2.5e-9 >> nnls_target_kappa=1e-11; the relaxed-KKT solve on the
Jacobi system then diverged to NaN at its 50-iteration cap on 4/16
jax_grad/mge.py points. The backward pass now runs <= 10 tight PDIP
iterations on (Q_pc, q_pc) warm-started from the mapped iterate (kept only
if converged), after which the relaxed solve converges in ~1 iteration.
The primal / forward value is unchanged.

solve_nnls gains an init=(x, s, z) warm start; the relaxed solve's
converged flag is kept and exposed with the polish status through the new
raw_forward_backward_status diagnostic. general.yaml comments updated
(values unchanged).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Review finding: on 3de624b the jitted gradient is NaN only on prng10/prng14
(prng2/prng12 pass jitted), and the backward-status test fails on import on
main rather than on convergence, so it is not the red witness.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019jDFQSNoi3ihaeM7ZJhfYL
@Jammy2211 Jammy2211 added the pending-release PR queued for the next release build label Sep 25, 2026
@Jammy2211
Jammy2211 merged commit 5f8a8de into main Sep 25, 2026
3 checks passed
@Jammy2211
Jammy2211 deleted the feature/mge-nnls-grad-nan branch September 25, 2026 18:51
@Jammy2211

Copy link
Copy Markdown
Collaborator Author

Merged by human /prm on 2026-09-25 under Heart RED release validation FAILED (stage integrate); the human's merge command, verbatim: "and I authorize on red heart but I guess keep going to get it out of red." All 3 CI legs green at 44762fa. Heart stays RED until a fresh Release Integrate passes.

@Jammy2211 Jammy2211 removed the pending-release PR queued for the next release build label Sep 26, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix: NaN JAX gradient on MGE positive-only solves after #572

1 participant