Skip to content

perf(mapper): frozen-cached prior tuple and instance plan on the instance_from_vector hot path (#1642, PR 3 of 3) - #1645

Merged
Jammy2211 merged 1 commit into
mainfrom
feature/ep-factor-search-overhead-p3
Sep 24, 2026
Merged

Jammy2211 merged 1 commit into
mainfrom
feature/ep-factor-search-overhead-p3

Conversation

@Jammy2211

Copy link
Copy Markdown
Collaborator

Summary

PR 3 of 3 under #1642, stacked on #1644 (base = feature/ep-factor-search-overhead-p2). Retarget to main after #1644 merges (and #1644 after #1643).

Every search maps each sampled vector to a model instance through instance_from_vector. When the likelihood is cheap, that mapper overhead is a real share of the run. The model is frozen for the whole search, so the parts of the mapping that do not depend on the vector can be cached:

  • AbstractPriorModel._vector_priors(): a @frozen_cache tuple of the id-ordered priors. instance_from_vector zips the vector with this tuple instead of rebuilding the prior ordering on every call.
  • Model._instance_plan() + Model._instance_for_arguments_frozen: a @frozen_cache of the vector-independent structure of _instance_for_arguments (which constructor arguments are priors, constants, child models or tuple priors), replayed on each call. It is used only when the model is frozen. The unfrozen path is unchanged.
  • instance_for_arguments checks the model's own assertions before it reads the exception_override config. It uses getattr(self, "_assertions", None) so pickles made before this change still load.

The existing frozen_cache is cleared on unfreeze() and __setstate__, and it is not saved in pickles.

Micro-benchmark: 20k frozen instance_from_vector calls, ABAB ×3, medians

model PR 2 (µs/call) PR 3 (µs/call) speed-up
af.Model(af.ex.Gaussian) 21.27 3.68 5.8×
3-Gaussian af.Collection 58.83 12.81 4.6×

EP harness: n5_profile, --max-steps 3 --nlive 100, interleaved

PR 2 (1cd4a5b) PR 3 (b3deab4)
run_nested calls (factor searches) 15 15
per-search wall, paired runs 0.871 / 1.091 s 0.909 / 1.019 s
centre moments max abs diff vs PR 2 – 0.0 (all runs)
ep_history flags SUCCESS 26 / NO_CHANGE 15 / BAD_PROJECTION 10 identical

The EP wall-clock gain from this PR is within noise on the toy harness. Under cProfile, instance_from_vector runs about 62k times per EP run, and its cumulative time falls from 5.84 s to 1.31 s. Without the profiler that is about 1 s per run, roughly 7 % of search time, which is smaller than the run-to-run noise. The benefit is the lower per-call mapper cost. That matters for every search with a cheap likelihood, not only EP.

Full evidence: #1642 (comment).

API Changes

None public. The change preserves behaviour: the new caches exist only for frozen models, and unfrozen models take the unchanged code path. Evidence: the 10 new tests show frozen and unfrozen models give identical instances across several model shapes (a deliberate mutation to the code was caught); EP moments are bit-identical to PR 2; a frozen model still pickles and unpickles after use; jax.jit and jax.grad through instance_from_vector still work, and the gradient is non-zero.

See full details below.

Test Plan

  • New test_autofit/mapper/test_instance_from_vector_frozen.py (10 tests): frozen and unfrozen models give identical instances across several model shapes. The tests caught a deliberate mutation.
  • Full suite, run serially (pytest test_autofit -p no:xdist): 2895 passed, 2 skipped.
  • Pickle round-trip of a used frozen model; jax.jit / jax.grad through instance_from_vector (non-zero gradient).
  • EP harness: moments bit-identical to PR 2, flags identical, run_nested 15.
  • Workspace smoke on the stacked PR 2 + PR 3 head (pyauto-heart smoke autofit --root <bundle with PyAutoFit = this branch>): 8/8 scripts and 2/2 notebooks PASS. autofit was imported from this branch's worktree. This also covers PR 2 (perf(ep): EP factor searches skip per-search visuals by default (#1642, PR 2 of 3) #1644), whose ship did not run the smoke.
  • Downstream impact (iii) none: PyAutoArray, PyAutoGalaxy, PyAutoLens, autofit_workspace* and HowToFit do not override or call _instance_for_arguments, instance_for_arguments, _vector_priors or _instance_plan, and none subclass Model/Collection. So no subclass skips the frozen path.
Full API Changes (for automation & release notes)

Added

  • AbstractPriorModel._vector_priors() (private, @frozen_cache).
  • Model._instance_plan() and Model._instance_for_arguments_frozen(...) (private; used only when the model is frozen).

Changed Behaviour

  • None observable. instance_from_vector and Model._instance_for_arguments on frozen models use the cached plan and give identical results. instance_for_arguments checks the model's own assertions before it reads the exception_override config.

Migration

  • None required.

Heart RED override (development only)

This PR is opened under the PyAutoBrain/AUTONOMY.md "Human override for Heart RED (development only)". It is a development-shipping override only. It does not claim to fix Heart, and Heart remains RED for release purposes. Merge still needs a separate explicit human /prm with every required check green.

  • Authorisation (live human, 2026-09-24, Claude Code CLI session): asked "Does your override extend to opening PR 2 now as a stacked PR (base = PR 1 branch, retargeted to main after PR 1 merges), and PR 3 the same way once it is green?", the human selected "Yes, same override for PR 2 and PR 3".
  • Exact current RED reasons (pyauto-heart readiness --json, verdict red, score 45, ts 2026-09-24T15:11:12.973142+00:00):
    1. release validation FAILED (stage integrate)
    2. workspace validation not passing (4 failed, cloud#35579888156: autolens notebooks/cluster/modeling.ipynb, autolens notebooks/weak/a2744.ipynb, autolens scripts/cluster/modeling.py, +1 more)
    3. manifest drift: hub organism blurb (organs present) — 7 mismatch(es) vs PyAutoMind/repos.yaml
  • Branch gates passed: full serial test_autofit 2895 passed / 2 skipped on head b3deab4; the 10 new frozen-path tests pass, and they caught a deliberate mutation; EP harness moments bit-identical to PR 2 (evidence: refactor(ep): cut per-factor-search wrapper overhead (dynesty single pass, EP visuals off, mapper fast path) #1642 (comment)); autofit workspace smoke on the stacked PR 2 + PR 3 head passed (8/8 scripts, 2/2 notebooks); downstream workspace impact (iii) none.
  • Permitted scope: commit, push, and opening this stacked pending-release PR. No merge, release or release rehearsal, and no bypassing of tests, checks or branch protection.

Generated by the PyAutoLabs agent workflow.

🤖 Generated with Claude Code

https://claude.ai/code/session_01EQdWrHM9BUx1q2zZRy6Y7N

…ance_from_vector hot path (#1642)

A search freezes its model for the whole fit, so the structure that
instance_from_vector rediscovers on every likelihood call can be cached
with the existing frozen_cache (reset on unfreeze() and __setstate__,
excluded from pickles). Unfrozen models take the unchanged path.

(a) AbstractPriorModel._vector_priors(): @frozen_cache tuple of the
    id-ordered priors; instance_from_vector zips it with the vector instead
    of rebuilding prior_tuples_ordered_by_id name/value wrappers per call.
(b) Model._instance_plan(): @frozen_cache of the value-independent parts of
    _instance_for_arguments (constructor attributes, tuple priors, child
    models, direct priors, deferred flag, Prior-class flag, found class,
    post-construction candidates). _instance_for_arguments_frozen replays it;
    hasattr(result, key) and Constant unwrapping stay per call.
(c) instance_for_arguments tests the model's own assertions list before the
    exception_override config read, so assertion-free models skip it
    (check_assertions only ever inspects self._assertions).

Micro-benchmark, 20k frozen instance_from_vector calls, ABAB x3 medians
(PR 2 head 1cd4a5b -> this commit):
  af.Model(af.ex.Gaussian)             21.27 -> 3.68 us/call (5.8x)
  3-Gaussian af.Collection             58.83 -> 12.81 us/call (4.6x)
EP harness n5_profile (max-steps 3, nlive 100): moments bit-identical
(max |diff| 0.0), run_nested 15, flags unchanged; per-search wall within
run-to-run noise (paired ABAB 0.87/1.09 s PR 2 vs 0.91/1.02 s PR 3). The
~62k likelihood calls per run save ~1 s total (~7 % of search time).

New test_autofit/mapper/test_instance_from_vector_frozen.py: frozen vs
unfrozen instances identical across Gaussian, nested collections, shared
priors, tuple priors, deferred arguments, constants, post-construction
attributes and assertion failures; cache populated and reset on unfreeze.
Full suite: 2895 passed, 2 skipped. Frozen model pickles/unpickles after
use; jax.jit and jax.grad through instance_from_vector work.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EQdWrHM9BUx1q2zZRy6Y7N
@Jammy2211 Jammy2211 added the pending-release PR queued for the next release build label Sep 24, 2026
Base automatically changed from feature/ep-factor-search-overhead-p2 to main September 24, 2026 17:47
@Jammy2211
Jammy2211 merged commit dd9fbe0 into main Sep 24, 2026
4 checks passed
@Jammy2211
Jammy2211 deleted the feature/ep-factor-search-overhead-p3 branch September 24, 2026 17:47
@Jammy2211 Jammy2211 removed the pending-release PR queued for the next release build label Sep 26, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant