Overview
RAL 342411 (ic50 EP ladder, N=50, sweep 4, global factor) died on assert np.isfinite(suff_stats).all() in AbstractMessage.project (autofit/messages/abstract.py:316). factor_step (graphical/expectation_propagation/optimiser.py:150-155) catches ValueError, ArithmeticError, RuntimeError, InitializerException and degrades one factor's bad sweep to its previous message; AssertionError is not caught, so one bad projection killed a 51-factor run. The trigger is stochastic (a local rerun at the same commit converged). Two findings from the plan pass: (i) a ValueError subclass is already recovered by factor_step; (ii) the prior id is not reachable inside AbstractMessage.project (TransformedMessage.project at composed_transform.py:213-236 swallows id_), so context is added at Prior.project and Result.projected_model.
Plan
- Replace the bare assert with a named
ProjectionException(ValueError) that says which input was non-finite.
- Add prior-id and parameter-path context where they are known.
- List the exception explicitly in
factor_step's recovery tuple (already covered by ValueError; listed for intent).
- Do not drop non-finite samples; raise. The designed recovery (
factor_step → StatusFlag.EXCEPTION → max_consecutive_failures=3 → STALE FACTORS warning) handles it loudly.
- Two test files: direct
project cases; the Witness through factor_step and a short EPOptimiser.run.
Detailed implementation plan
Affected Repositories
- PyAutoFit (primary) — the only claimed repo.
- autofit_workspace_test — not claimed, no edits; the graphical smoke scripts exercise the unchanged happy path and are rerun via
/smoke_test at ship.
Branch Survey
| Repository |
Current Branch |
Dirty? |
| ./fit/PyAutoFit |
main |
clean (in sync with origin/main) |
Suggested branch: feature/ep-projection-exception
Worktree root: ~/Code/PyAutoLabs-wt/ep-projection-exception/
Parallel claim
PyAutoFit is claimed in parallel by ep-projection-exception and ep-moment-projection (both registered 2026-09-30). File sets are disjoint — A: messages/abstract.py, mapper/prior/abstract.py, non_linear/result.py, graphical/expectation_propagation/optimiser.py:150-162, exc.py, test_autofit/messages/test_project_nonfinite.py, test_autofit/graphical/functionality/test_factor_failure_recovery.py; B: graphical/laplace/*, graphical/mean_field.py, graphical/declarative/factor/hierarchical.py, graphical/expectation_propagation/diagnostics.py:47, graphical/README.md, test_autofit/graphical/functionality/test_moment_projection.py. A ships first; B rebases. Parallel claim human-approved 2026-09-30 with the plan.
Implementation Steps
autofit/exc.py (~61, next to SearchException): class ProjectionException(ValueError) — not a MessageException subclass (mean_field.py:590, hierarchical.py:150, laplace/newton.py:352 turn that into silent revert / −inf), not SearchException (uncaught, means misconfigured search).
autofit/messages/abstract.py:316: if not np.isfinite(suff_stats).all(): raise exc.ProjectionException(cls._nonfinite_projection_reason(samples, log_weight_list, suff_stats)); import from .. import exc as normal.py:17 does. New classmethod _nonfinite_projection_reason, failure path only, checks in order: (a) non-finite samples (count nan/inf, first index); (b) nan / +inf log weights; (c) all log weights −inf; (d) overflow of T(x)·w (report max |x|). Message prefix f"{cls.__name__}.project: non-finite sufficient statistics {suff_stats}; ". Moment maths untouched.
autofit/mapper/prior/abstract.py:236-241 (Prior.project): wrap self.message.project(...) in try/except exc.ProjectionException as e: raise exc.ProjectionException(f"Projecting prior id={self.id} ({type(self).__name__}) failed: {e}") from e.
autofit/non_linear/result.py:498-504 (projected_model): turn the dict comprehension into a loop, catch-and-reraise adding path=".".join(path). Existing zero-weight ValueError guard at :492-495 stays.
autofit/graphical/expectation_propagation/optimiser.py:150-162: add exc.ProjectionException to the tuple; extend the comment ("a non-finite projection of one factor's samples is a failed sweep update, not a failed graph fit"). Do NOT add AssertionError. Existing path already sets StatusFlag.EXCEPTION (graphical/utils.py:267), updated=False, keeps factor_approx.model_dist, logger.exception, and EPOptimiser.run (:589) counts it toward max_consecutive_failures (:534).
- Tests:
- New
test_autofit/messages/test_project_nonfinite.py: parametrised NormalMessage.project cases — nan sample, inf sample, nan log weight, +inf log weight, all −inf, overflow ([1e200,2,3]) — each pytest.raises(exc.ProjectionException, match=<condition word>), issubclass(..., ValueError), not AssertionError; prior-level cases af.GaussianPrior(0,1).project([1,2,nan], zeros) and af.UniformPrior(0,1).project([.2,.5,nan], zeros) assert str(prior.id) in the message (the second proves the id survives the TransformedMessage path); a finite control (mean ≈ 2). np.errstate(all="ignore") around the calls.
- Append to
test_autofit/graphical/functionality/test_factor_failure_recovery.py (reuse make_shared_variable_approx, ExactFactorFit, the InitializerFailingOptimiser pattern): NonFiniteProjectionOptimiser(n_failures) whose optimise calls the nan projection for the first n calls then exact_fit; test_nonfinite_projection_degrades_factor_step_to_previous_message (flag is EXCEPTION, updated is False, new_dist is factor_approx.model_dist, "samples"/"non-finite" in status.messages[0]); test_nonfinite_projection_does_not_abort_the_ep_run (EPOptimiser(..., factor_optimisers={prior: NonFiniteProjectionOptimiser(1), likelihood: ExactFactorFit()}, paths=False).run(model_approx, max_steps=4) returns, x mean finite, EXCEPTION row for prior.name in optimiser.diagnostics.factor_rows).
- Optional:
test_autofit/graphical/test_unification.py next to test_projected_model (:37) — one ("centre",) value nan → raises with "centre" and the prior id.
- Behaviour-change note for the PR: outside EP,
Result.projected_model callers get a clear ProjectionException where they got AssertionError; no autofit/ code catches AssertionError. Under python -O the old assert vanished silently — strictly worse.
- Ship:
/ship_library (Heart RED override already granted for PR-open; record it in the four sinks), then /smoke_test on the graphical smoke scripts; workspace impact (iii) none.
Key Files
autofit/exc.py — new ProjectionException(ValueError)
autofit/messages/abstract.py — raise instead of assert; _nonfinite_projection_reason
autofit/mapper/prior/abstract.py — Prior.project adds prior id context
autofit/non_linear/result.py — projected_model adds parameter-path context
autofit/graphical/expectation_propagation/optimiser.py — factor_step recovery tuple
test_autofit/messages/test_project_nonfinite.py (new)
test_autofit/graphical/functionality/test_factor_failure_recovery.py (extended)
Verification
pytest test_autofit/messages/test_project_nonfinite.py test_autofit/graphical/functionality/test_factor_failure_recovery.py -q green; new tests red when the raise is reverted to the assert; full pytest test_autofit -x; graphical smoke scripts pass.
Heart RED override (development only)
Authorised by the live human in the Claude Code session 2026-09-30 ~10:55 BST, answer "Override for all three (Recommended)" to the question offering the RED override for PyAutoCortex#50 and "for opening the two PyAutoFit tasks (issue + worktree + plan; no merge, no release)".
Exact Heart RED reasons at the 10:51 BST tick:
- "PyAutoArray: 2 commit(s) behind origin"
- "PyAutoLens: 2 commit(s) behind origin"
- "release validation FAILED (stage integrate)"
Scope: issue + worktree + plan; PR-open permitted; no merge, no release. Plans approved in-session 2026-09-30 ~11:20 BST via Plan Mode.
Original Prompt
Click to expand starting prompt
EP message projection asserts on non-finite sufficient statistics and kills the whole EP run (ic50_workspace N=50 rung, RAL 342411)
Type: bug
Target: PyAutoFit
Repos:
- PyAutoFit
- autofit_workspace_test
Themes:
- graphical-ep
Difficulty: small
Autonomy: supervised
Priority: medium
Status: draft
Consequence: glance
Witness: a unit test in test_autofit/graphical/ in which a factor's search returns samples containing one non-finite value (e.g. NormalMessage.project(np.array([1.0, 2.0, np.nan]), np.zeros(3)), which today raises a bare AssertionError from autofit/messages/abstract.py:316) and, driven through factor_step, the EP run records that factor step as StatusFlag.EXCEPTION, keeps the previous message and carries on, instead of dying. The error message names the prior id and says which input was non-finite (the samples, the log weights, or an overflow). Plus a direct project test: a nan sample is either dropped with a warning or rejected with a ValueError, and never trips an assert.
Review-minutes: 5
Unattended: ready
Epic: graphical-ep
Filed: 2026-09-30
Why
RAL job 342411 (the ic50_workspace EP scale ladder, sim rungs 5/10/25/50, DynestyStatic nlive 150,
max_steps 12, PyAutoFit mirror at 66f9f8d) passed rungs N=5/10/25. It then died at the N=50 rung
right after the 4th EP sweep's global factor search finished (2026-09-09 23:19:27 BST).
The traceback, from hpc/batch_cpu/error/error.342411.err:
ep_sim.py:123 run_ep_fit → util.py:712 factor_graph.optimise
autofit/graphical/declarative/abstract.py:208 optimise → opt.run
autofit/graphical/expectation_propagation/optimiser.py:586 run → self.factor_step
optimiser.py:356 factor_step → factor_step(...)
optimiser.py:135 factor_step → optimiser.optimise(factor_approx)
autofit/non_linear/search/abstract_search.py:401 optimise → result.projected_model.priors
autofit/non_linear/result.py:473 projected_model → prior.project(...)
autofit/mapper/prior/abstract.py:237 project → self.message.project(...)
autofit/messages/abstract.py:316 project → assert np.isfinite(suff_stats).all()
AssertionError
This bug has two parts:
- The projection uses
assert for a data condition. AbstractMessage.project
(autofit/messages/abstract.py:275-319, unchanged on main 5cf687d since the mirror commit)
calculates w = exp(logw − max logw), w /= w.mean(), suff_stats = mean(T(x)·w), then asserts
the result is finite. Probing it on main shows the assertion fires for all of these inputs:
a nan sample, an ±inf sample, a nan log weight, a +inf log weight, all log weights −inf (this
case is guarded upstream in Result.projected_model, result.py:492), or x² overflowing.
A single sample silently returns sigma=0 instead. A negative-precision cavity also reaches it:
GaussianPrior.with_message(NormalMessage(0,1)/NormalMessage(0,0.5)) has sigma=nan, and every
value_for on it returns nan. A bare assert gives no diagnostic (which prior, which input) and
disappears under python -O.
factor_step cannot recover from it. The failure-recovery path
(autofit/graphical/expectation_propagation/optimiser.py:150-155 on main) catches
ValueError, ArithmeticError, RuntimeError, InitializerException. That path exists so that one
factor's bad sweep degrades to its previous message and the run continues, with
max_consecutive_failures as the backstop. AssertionError is not in that list, so one
non-finite projection of one factor ends a 20-minute, 51-factor EP run.
The failure is not deterministic. A local rerun at 66f9f8d on the same dataset (unseeded Dynesty)
converged (Terminating optimisation, EPHistory(kl_tol=1.0)) midway through sweep 3 and never
reached a 4th global projection. So the witness is the unit test above, not a rerun of the rung.
It is still unknown which of the non-finite conditions fired on RAL, because output to disk was
disabled and no samples were kept. The fix should report the condition so the next occurrence
identifies itself.
Scope
- Replace the
assert with a ValueError, or a SearchException subclass that factor_step
catches, naming the prior id and the offending condition. Or drop non-finite samples/weights with
a warning when enough finite, positively weighted samples remain.
- Make sure
factor_step's except tuple covers it.
- Leave the moment-matching maths as it is.
🤖 Generated with Claude Code
Overview
RAL 342411 (ic50 EP ladder, N=50, sweep 4, global factor) died on
assert np.isfinite(suff_stats).all()inAbstractMessage.project(autofit/messages/abstract.py:316).factor_step(graphical/expectation_propagation/optimiser.py:150-155) catchesValueError, ArithmeticError, RuntimeError, InitializerExceptionand degrades one factor's bad sweep to its previous message;AssertionErroris not caught, so one bad projection killed a 51-factor run. The trigger is stochastic (a local rerun at the same commit converged). Two findings from the plan pass: (i) aValueErrorsubclass is already recovered byfactor_step; (ii) the prior id is not reachable insideAbstractMessage.project(TransformedMessage.projectatcomposed_transform.py:213-236swallowsid_), so context is added atPrior.projectandResult.projected_model.Plan
ProjectionException(ValueError)that says which input was non-finite.factor_step's recovery tuple (already covered byValueError; listed for intent).factor_step→StatusFlag.EXCEPTION→max_consecutive_failures=3 → STALE FACTORS warning) handles it loudly.projectcases; the Witness throughfactor_stepand a shortEPOptimiser.run.Detailed implementation plan
Affected Repositories
/smoke_testat ship.Branch Survey
Suggested branch:
feature/ep-projection-exceptionWorktree root:
~/Code/PyAutoLabs-wt/ep-projection-exception/Parallel claim
PyAutoFit is claimed in parallel by
ep-projection-exceptionandep-moment-projection(both registered 2026-09-30). File sets are disjoint — A:messages/abstract.py,mapper/prior/abstract.py,non_linear/result.py,graphical/expectation_propagation/optimiser.py:150-162,exc.py,test_autofit/messages/test_project_nonfinite.py,test_autofit/graphical/functionality/test_factor_failure_recovery.py; B:graphical/laplace/*,graphical/mean_field.py,graphical/declarative/factor/hierarchical.py,graphical/expectation_propagation/diagnostics.py:47,graphical/README.md,test_autofit/graphical/functionality/test_moment_projection.py. A ships first; B rebases. Parallel claim human-approved 2026-09-30 with the plan.Implementation Steps
autofit/exc.py(~61, next toSearchException):class ProjectionException(ValueError)— not aMessageExceptionsubclass (mean_field.py:590,hierarchical.py:150,laplace/newton.py:352turn that into silent revert / −inf), notSearchException(uncaught, means misconfigured search).autofit/messages/abstract.py:316:if not np.isfinite(suff_stats).all(): raise exc.ProjectionException(cls._nonfinite_projection_reason(samples, log_weight_list, suff_stats)); importfrom .. import excasnormal.py:17does. New classmethod_nonfinite_projection_reason, failure path only, checks in order: (a) non-finite samples (count nan/inf, first index); (b) nan / +inf log weights; (c) all log weights −inf; (d) overflow of T(x)·w (report max |x|). Message prefixf"{cls.__name__}.project: non-finite sufficient statistics {suff_stats}; ". Moment maths untouched.autofit/mapper/prior/abstract.py:236-241(Prior.project): wrapself.message.project(...)intry/except exc.ProjectionException as e: raise exc.ProjectionException(f"Projecting prior id={self.id} ({type(self).__name__}) failed: {e}") from e.autofit/non_linear/result.py:498-504(projected_model): turn the dict comprehension into a loop, catch-and-reraise addingpath=".".join(path). Existing zero-weightValueErrorguard at :492-495 stays.autofit/graphical/expectation_propagation/optimiser.py:150-162: addexc.ProjectionExceptionto the tuple; extend the comment ("a non-finite projection of one factor's samples is a failed sweep update, not a failed graph fit"). Do NOT addAssertionError. Existing path already setsStatusFlag.EXCEPTION(graphical/utils.py:267),updated=False, keepsfactor_approx.model_dist,logger.exception, andEPOptimiser.run(:589) counts it towardmax_consecutive_failures(:534).test_autofit/messages/test_project_nonfinite.py: parametrisedNormalMessage.projectcases — nan sample, inf sample, nan log weight, +inf log weight, all −inf, overflow ([1e200,2,3]) — eachpytest.raises(exc.ProjectionException, match=<condition word>),issubclass(..., ValueError), notAssertionError; prior-level casesaf.GaussianPrior(0,1).project([1,2,nan], zeros)andaf.UniformPrior(0,1).project([.2,.5,nan], zeros)assertstr(prior.id)in the message (the second proves the id survives the TransformedMessage path); a finite control (mean ≈ 2).np.errstate(all="ignore")around the calls.test_autofit/graphical/functionality/test_factor_failure_recovery.py(reusemake_shared_variable_approx,ExactFactorFit, theInitializerFailingOptimiserpattern):NonFiniteProjectionOptimiser(n_failures)whoseoptimisecalls the nan projection for the first n calls thenexact_fit;test_nonfinite_projection_degrades_factor_step_to_previous_message(flag is EXCEPTION,updated is False,new_dist is factor_approx.model_dist, "samples"/"non-finite" instatus.messages[0]);test_nonfinite_projection_does_not_abort_the_ep_run(EPOptimiser(..., factor_optimisers={prior: NonFiniteProjectionOptimiser(1), likelihood: ExactFactorFit()}, paths=False).run(model_approx, max_steps=4)returns, x mean finite, EXCEPTION row forprior.nameinoptimiser.diagnostics.factor_rows).test_autofit/graphical/test_unification.pynext totest_projected_model(:37) — one("centre",)value nan → raises with "centre" and the prior id.Result.projected_modelcallers get a clearProjectionExceptionwhere they gotAssertionError; noautofit/code catchesAssertionError. Underpython -Othe old assert vanished silently — strictly worse./ship_library(Heart RED override already granted for PR-open; record it in the four sinks), then/smoke_teston the graphical smoke scripts; workspace impact (iii) none.Key Files
autofit/exc.py— newProjectionException(ValueError)autofit/messages/abstract.py— raise instead of assert;_nonfinite_projection_reasonautofit/mapper/prior/abstract.py—Prior.projectadds prior id contextautofit/non_linear/result.py—projected_modeladds parameter-path contextautofit/graphical/expectation_propagation/optimiser.py—factor_steprecovery tupletest_autofit/messages/test_project_nonfinite.py(new)test_autofit/graphical/functionality/test_factor_failure_recovery.py(extended)Verification
pytest test_autofit/messages/test_project_nonfinite.py test_autofit/graphical/functionality/test_factor_failure_recovery.py -qgreen; new tests red when the raise is reverted to the assert; fullpytest test_autofit -x; graphical smoke scripts pass.Heart RED override (development only)
Authorised by the live human in the Claude Code session 2026-09-30 ~10:55 BST, answer "Override for all three (Recommended)" to the question offering the RED override for PyAutoCortex#50 and "for opening the two PyAutoFit tasks (issue + worktree + plan; no merge, no release)".
Exact Heart RED reasons at the 10:51 BST tick:
Scope: issue + worktree + plan; PR-open permitted; no merge, no release. Plans approved in-session 2026-09-30 ~11:20 BST via Plan Mode.
Original Prompt
Click to expand starting prompt
EP message projection asserts on non-finite sufficient statistics and kills the whole EP run (ic50_workspace N=50 rung, RAL 342411)
Type: bug
Target: PyAutoFit
Repos:
Themes:
Difficulty: small
Autonomy: supervised
Priority: medium
Status: draft
Consequence: glance
Witness: a unit test in
test_autofit/graphical/in which a factor's search returns samples containing one non-finite value (e.g.NormalMessage.project(np.array([1.0, 2.0, np.nan]), np.zeros(3)), which today raises a bareAssertionErrorfromautofit/messages/abstract.py:316) and, driven throughfactor_step, the EP run records that factor step asStatusFlag.EXCEPTION, keeps the previous message and carries on, instead of dying. The error message names the prior id and says which input was non-finite (the samples, the log weights, or an overflow). Plus a directprojecttest: a nan sample is either dropped with a warning or rejected with aValueError, and never trips anassert.Review-minutes: 5
Unattended: ready
Epic: graphical-ep
Filed: 2026-09-30
Why
RAL job 342411 (the ic50_workspace EP scale ladder, sim rungs 5/10/25/50, DynestyStatic nlive 150,
max_steps 12, PyAutoFit mirror at 66f9f8d) passed rungs N=5/10/25. It then died at the N=50 rung
right after the 4th EP sweep's
globalfactor search finished (2026-09-09 23:19:27 BST).The traceback, from
hpc/batch_cpu/error/error.342411.err:This bug has two parts:
assertfor a data condition.AbstractMessage.project(
autofit/messages/abstract.py:275-319, unchanged on main 5cf687d since the mirror commit)calculates
w = exp(logw − max logw),w /= w.mean(),suff_stats = mean(T(x)·w), then assertsthe result is finite. Probing it on main shows the assertion fires for all of these inputs:
a nan sample, an ±inf sample, a nan log weight, a +inf log weight, all log weights −inf (this
case is guarded upstream in
Result.projected_model,result.py:492), orx²overflowing.A single sample silently returns
sigma=0instead. A negative-precision cavity also reaches it:GaussianPrior.with_message(NormalMessage(0,1)/NormalMessage(0,0.5))hassigma=nan, and everyvalue_foron it returns nan. A bareassertgives no diagnostic (which prior, which input) anddisappears under
python -O.factor_stepcannot recover from it. The failure-recovery path(
autofit/graphical/expectation_propagation/optimiser.py:150-155on main) catchesValueError, ArithmeticError, RuntimeError, InitializerException. That path exists so that onefactor's bad sweep degrades to its previous message and the run continues, with
max_consecutive_failuresas the backstop.AssertionErroris not in that list, so onenon-finite projection of one factor ends a 20-minute, 51-factor EP run.
The failure is not deterministic. A local rerun at 66f9f8d on the same dataset (unseeded Dynesty)
converged (
Terminating optimisation,EPHistory(kl_tol=1.0)) midway through sweep 3 and neverreached a 4th global projection. So the witness is the unit test above, not a rerun of the rung.
It is still unknown which of the non-finite conditions fired on RAL, because output to disk was
disabled and no samples were kept. The fix should report the condition so the next occurrence
identifies itself.
Scope
assertwith aValueError, or aSearchExceptionsubclass thatfactor_stepcatches, naming the prior id and the offending condition. Or drop non-finite samples/weights with
a warning when enough finite, positively weighted samples remain.
factor_step'sexcepttuple covers it.🤖 Generated with Claude Code