Skip to content

feat(aggregation): Add GradNorm - #785

Open
giovannicozzolongo wants to merge 2 commits into
SimplexLab:mainfrom
giovannicozzolongo:issue-665-gradnorm
Open

giovannicozzolongo wants to merge 2 commits into
SimplexLab:mainfrom
giovannicozzolongo:issue-665-gradnorm

Conversation

@giovannicozzolongo

Copy link
Copy Markdown

Refs #665 (GradNorm).

Adds GradNormWeighting and its GradNorm aggregator wrapper. Task weights are nn.Parameters initialised to one. set_losses records the current losses, the forward caches detached gradient norms and returns detached weights, and balancing_loss provides the auxiliary objective for an external optimizer. renormalize restores the weights' sum after each optimizer step.

This follows Algorithm 1's direct weight updates. The balancing target is held fixed during differentiation, and the auxiliary backward cannot change model gradients. The initial loss baseline is saved in the state dict. The number of tasks stays fixed so optimizer references remain valid across calls and resets.

The example computes norms on the last shared layer with autogram.Engine and applies the learned weights to the whole model. No new dependencies are required.

The optimizer placement and API remain open for review. This draft uses an external optimizer and direct weights rather than LibMTL's softmax parameterisation.

Testing

  • Full CPU unit suite with warnings treated as errors: 3,381 passed, 66 skipped, 33 expected failures.
  • 100 new tests passed on CPU and an NVIDIA A100 MIG GPU, in both float32 and float64.
  • New module statement coverage: 100% (106 statements).
  • ty check, ruff check and ruff format --check: passed.
  • Pre-commit hooks on all nine changed files: passed.
  • Sphinx HTML build: passed. Both new GradNorm documentation examples passed.
  • Full documentation doctest: 145 tests, two failures also reproduced on unmodified main (143 tests). The existing engine example references Tensor without importing it on Python 3.12; the existing Lightning example emits a model summary where no output is expected.
  • git diff --check: passed.

Tested with Python 3.12.13 and PyTorch 2.14.0+cu130. The minimum supported Python and PyTorch versions were not tested.

Signed-off-by: Giovanni Cozzolongo <79092266+giovannicozzolongo@users.noreply.github.com>
Signed-off-by: Giovanni Cozzolongo <79092266+giovannicozzolongo@users.noreply.github.com>
@ValerianRey ValerianRey added cc: feat Conventional commit type for new features. package: aggregation labels Sep 29, 2026
@github-actions github-actions Bot changed the title feat(aggregation): add GradNorm feat(aggregation): Add GradNorm Sep 29, 2026
@ValerianRey

Copy link
Copy Markdown
Member

@giovannicozzolongo thanks a lot for the PR! I start reviewing it now. It's a quite hard topic so I may take a few days. Thanks for your patience!

@ValerianRey ValerianRey mentioned this pull request Oct 3, 2026
@ValerianRey

Copy link
Copy Markdown
Member

/opencode:Plan please review this PR. In particular, discuss its API, and compare it with the other possible API where the weight optimizer would be internal, and the usage would look something like:

losses = ...
gramian = ...
gradnorm_weighting.set_losses(losses)
weights = gradnorm_weighting(gramian)
# computes aux loss, updates weights using internal optimizer, and returns previous weights

losses.backward(weights)

Make a table with the pros and cons of both options. If you think about other options, mention them.

@opencode-agent

opencode-agent Bot commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor
Toggle

I've now studied the PR thoroughly: the new _gradnorm.py, its tests, the docs example, the Weighting/Aggregator base contracts, the stateful precedents in the codebase (GradVac, MoDo, SDMGrad), and LibMTL's GradNorm for comparison. (Note: I have no shell access in this environment, so this is a static review — I could not re-run the test suite, but the PR's reported testing is extensive and the tests read as correct.)

Here is my review, focused on the API.


1. What the PR's API actually is

The PR proposes an external-optimizer, split-phase protocol. Per iteration, the user writes:

weight_optimizer.zero_grad()
weighting.set_losses(losses)            # side channel: current unweighted losses
weights = weighting(gramian)            # caches norms; returns current weights (detached clone)
weighting.balancing_loss().backward()   # populates weights.grad only
losses.backward(weights)                # model backward with the *same* (pre-update) weights
model_optimizer.step()
weight_optimizer.step()
weighting.renormalize()                 # restore sum(w) = n_tasks

The alternative under discussion (internal optimizer) collapses the weight update into the forward call:

weighting.set_losses(losses)
weights = weighting(gramian)            # aux loss, backward, internal step, renormalize; returns pre-update weights
losses.backward(weights)

2. A constraint that shapes both designs

Weighting.__call__(stat) and Aggregator.__call__(matrix) take exactly one tensor. There is no way to pass losses through the forward without breaking the interface uniformity between the autojac path (GradNorm) and the autogram path (GradNormWeighting). So both options need the set_losses side channel. The internal-optimizer design only eliminates balancing_loss(), renormalize(), and the optimizer bookkeeping (zero_grad()/step()) — about 4 lines per loop iteration. This bounds how much ergonomics Option B can actually buy.

A second observation, also inherent to both designs: GradNorm's weights are learned state, not a per-step function of the Gramian. The forward is therefore degenerate ("return current weights") and the Gramian only influences future calls. The real computation lives in balancing_loss. This makes the Weighting abstraction a slightly awkward fit for GradNorm no matter what — worth being explicit about in docs, since users coming from UPGrad/MGDA expect the Gramian to determine the weights immediately.

3. Comparison table

Aspect A: External optimizer (this PR) B: Internal optimizer
Per-step user code ~6 lines; ordering constraints (set_losses → forward → balancing_loss → step → renormalize) validated at runtime with clear errors 2 lines; invariants enforced by construction — impossible to forget renormalize or misorder
Silent-failure modes Forgetting renormalize() is silent (weights drift off the sum-to-N constraint); forgetting zero_grad() silently accumulates None — the update is atomic inside __call__
Optimizer choice Any torch.optim.Optimizer; faithful to the paper/reference impls (Adam), but free Requires new API surface: a factory param like optimizer: Callable[[Iterable[Tensor]], Optimizer]
LR scheduling, grad clipping, AMP/GradScaler, freezing weights, grad accumulation over k batches All standard PyTorch patterns work unchanged Each needs bespoke plumbing (scaler-aware step, exposed scheduler hooks, update flag); some become effectively unsupported
Checkpoint / resume Standard state_dict() + optimizer.state_dict() pair everyone knows; tested in the PR nn.Module.state_dict() does not cover optimizer state — the library must invent custom (de)serialization
reset() semantics Parameter identity preserved (good), but the user must reset the external optimizer separately to drop momentum — documented wart reset() can recreate the optimizer state atomically — cleaner
Side effects in __call__ Forward is pure (detached clone out); the aux backward is explicit and inspectable __call__ would perform a backward + optimizer step — surprising for forward hooks/profiling, and requires torch.enable_grad() inside the _NonDifferentiable no-grad wrapper, inverting the mixin's purpose
Transparency / debuggability weights, weights.grad, balancing_loss() value all inspectable; easy to log or skip updates (e.g. eval loops: just don't call balancing_loss) Mechanics hidden; needs extra accessors or train/eval-mode gating (magical semantics for .eval())
Setup-ordering footgun Documented: .to(device/dtype) before creating the optimizer Same PyTorch constraint, but invisible — the optimizer is created in __init__ (or must be created lazily at first forward, more complexity)
Codebase precedent No Weighting owns an optimizer; consistent with "TorchJD provides primitives, users own the loop" Closer to MoDo/SDMGrad (internal update in forward), but those implement a fixed, algorithm-defined scheme (projected SGD with lr as a hyperparameter) — GradNorm's optimizer is a user choice, not part of Algorithm 1
Numerics Exact Algorithm 1: model step and weight step both use w(t); renormalize to sum N Identical — both can be exactly faithful
Evolution path Convenience wrapper (Option C below) can be added later non-breakingly Shipping B first and exposing primitives later is a breaking change

4. The decisive factors

  • Where the complexity lives. Option B's strongest argument is that the orchestration cost is paid once in the library instead of by every user in every loop. That's real. But TorchJD's audience writes explicit training loops already (every example in the docs does), the protocol is runtime-validated, and the two genuinely dangerous mistakes are: forgetting renormalize (silent — this is A's weakest point) and misordering (loud — validated).
  • Ecosystem fit. Checkpointing, AMP, schedulers, clipping, freezing, logging — Option A composes with all of it for free; Option B must rebuild or forgo each one. An nn.Module owning an optimizer is also unprecedented in both this codebase and mainstream PyTorch.
  • set_losses is required either way, so B's win is smaller than it first appears.

5. Other options worth considering

  • C — Hybrid: external optimizer + convenience step. Keep A's primitives and add one helper performing zero_grad → balancing_loss().backward() → optimizer.step() → renormalize() around a user-supplied optimizer (e.g. weighting.step_weights(weight_optimizer)). Captures most of B's ergonomics and fixes the silent-forgotten-renormalize hazard, without owning the optimizer. Non-breaking to add later, so it doesn't need to gate this PR.
  • D — Softmax/logit parameterization (LibMTL-style). Learn unconstrained logits with w = n_tasks · softmax(logits): positivity and the sum constraint hold by construction, renormalize disappears, and the negative-weights dead end (where renormalize raises) can't occur. I verified the PR's claim: LibMTL does use N·softmax(loss_scale) with a first-epoch baseline. The cost is a different update geometry than Algorithm 1 (chain rule through softmax rescales per-coordinate steps) — since this PR explicitly targets Algorithm-1 fidelity, direct weights are the defensible choice, but D could be a future variant if the lr/negativity coupling proves painful in practice.
  • E — Stateless functional core. A pure balancing_loss(weights, norms, losses, initial_losses, alpha) as a building block for power users. Nice-to-have, but as the only API it pushes baseline tracking, device/dtype, and checkpointing onto users. Not recommended as primary.
  • F — Fold renormalize into the next forward. Removes one call while keeping the optimizer external; but it hides in-place parameter mutation inside forward and makes state_dict() contents order-dependent. Marginal.
  • G — Trigger the weight update from the model backward (hook on the losses so losses.backward(weights) also kicks the weight update). Most magical, and fragile with retain_graph, gradient accumulation, and manual .grad handling. Not recommended.

6. Recommendation

Keep Option A (the PR's design). Its costs are ergonomic and recoverable (Option C can be added non-breakingly later if users ask); Option B's costs are structural (checkpointing, AMP/scheduler friction, hidden side effects in __call__, new configuration surface, precedent break) and would be hard to reverse. The implementation itself is high quality: the detached/clone semantics in forward (a clone is required since the optimizer mutates the parameter in place — correct), the all-zero-losses case returning (weights * 0).sum() to yield zero grads instead of None, the proof that the aux backward can't leak into model gradients, the persistent vs non-persistent buffer split, and the exact Algorithm-1 numerics test are all exactly right.

7. Minor, non-blocking observations

  • renormalize() mutates in place; PyTorch convention would suggest renormalize_() (though reset() sets a codebase precedent for mutating methods without the underscore — your call).
  • The README row links to docs/source/docs/aggregation/gradnorm.rst (repo-relative) while all other rows use https://torchjd.org/stable/... URLs — the relative link works on GitHub but will be broken on PyPI.
  • GradNorm (aggregator) computes norms over all Jacobian columns, deviating from the paper's last-shared-layer choice; this is honestly documented, and the GradNormWeighting + Engine path covers the faithful variant — good trade-off.
  • Worth one extra doc sentence: during evaluation, the intended usage is set_losses + forward only (no balancing_loss/step), since that differs from all other weightings where forward is side-effect-free.

Great work overall — the API protocol is unusual for this library, but it's unusual because GradNorm itself doesn't fit the "weights = f(Gramian)" mold, and the chosen design makes the right trade-offs explicit.

New%20session%20-%202026-10-03T10%3A18%3A30.984Z
opencode session  |  github run

@ValerianRey

Copy link
Copy Markdown
Member

Please disregard opencode's comment unless I confirm its concerns in the review.

Also note that opencode is wrong about there being no precedent of having an optimizer inside a nn.Module in TorchJD. We have exactly that in the FAMO scalarizer.

IMO option B is much closer to TorchJD's usual interface, but comes with 4 minor drawbacks:

  1. need extra optimizer_class and optimizer_kwargs params. Not a problem
  2. a bit less expressive: can't do gradient clipping on aux loss' gradient for example. Not a problem IMO, if people really want to drift from the paper's algorithm they can always make a new aggregator
  3. the NonDifferentiableMixin might cause issues
  4. the StateDict will not include the optimizer's state

If 3) and 4) can be solved I would go for option B.

/opencode:Plan please investigate about 3) and 4).

@opencode-agent

opencode-agent Bot commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor
Toggle

Both concerns are investigable from the codebase itself — FAMO (src/torchjd/scalarization/_famo.py) is indeed the perfect precedent, since it already lives with exactly these two constraints. Here is my analysis.

TL;DR

  • Concern 3 (_NonDifferentiable): solvable with an established in-codebase pattern — FAMO's torch.enable_grad() + torch.autograd.grad island (_famo.py:166-176). The only genuinely new design requirement is an eval-mode gate (recommend gating the update on self.training), because an internal update inside forward breaks under torch.inference_mode().
  • Concern 4 (state_dict): solvable with a small state_dict/load_state_dict override (or state-dict hooks) that embeds optimizer.state_dict() under an extra key. It's ~15 lines, keeps strict=True working both ways, and survives torch.save/deepcopy/weights_only=True loading. Alternatively, the FAMO precedent shows the library already tolerates "optimizer state is not checkpointed" as an accepted limitation.

Concern 3 — _NonDifferentiable and the internal update

The mechanism works under the mixin

The mixin (src/torchjd/aggregation/_mixins.py:17-19) wraps __call__ in torch.no_grad(). The internal weight update needs one local gradient computation. FAMO already demonstrates the escape hatch in update():

with torch.enable_grad():
    weights = softmax(self._w.flatten(), dim=0)
    grad = torch.autograd.grad(weights, self._w, grad_outputs=delta.flatten())[0]
self._w.grad = grad
self._optimizer.step()
self._w.grad = None   # don't leak stale grads into a user optimizer

torch.enable_grad() locally overrides an enclosing no_grad — this is documented PyTorch behavior (the standard gradient-penalty idiom). For Option-B GradNorm, forward would be:

def forward(self, gramian):
    # ... validation ...
    w_t = self.weights.detach().clone()          # 1. snapshot BEFORE the step
    self._norms = gramian.detach().diagonal().clamp_min(0).sqrt()
    if self.training:                             # 2. eval gate (see below)
        with torch.enable_grad():                # 3. FAMO-style grad island
            grad = torch.autograd.grad(self._balancing_loss(), self.weights)[0]
        self.weights.grad = grad
        self._optimizer.step()
        self.weights.grad = None
        self.renormalize()                        # 4. invariant enforced atomically
    return w_t

Why each ingredient is safe here:

  • No graph leak: the balancing loss is built from self.weights plus _losses/_norms, which the PR already stores as detached clones (_gradnorm.py:128,142). The local autograd graph touches nothing of the user's model graph — this is exactly what the PR's leak test proves for the external variant, and it carries over unchanged.
  • autograd.grad instead of .backward(): doesn't accumulate into .grad and needs no retain_graph. The assign-step-clear sequence is copied from FAMO (_famo.py:172-176) and additionally eliminates Option A's "forgot zero_grad()" failure mode for free.
  • Clone before step: optimizer.step() mutates the parameter in place, so the pre-update weights w_t must be snapshotted first — the PR already clones for the same reason (_gradnorm.py:143). Returning w_t while stepping to w_{t+1} internally reproduces Algorithm 1's exact numerics (both model update and weight update use w_t), identical to the current implementation.
  • renormalize inside forward: its failure mode (negative weights → raise) stays loud, and forgetting it becomes impossible — removes Option A's one silent footgun.
  • Precedent for in-forward mutation: MoDo/SDMGrad already mutate state inside a _NonDifferentiable forward (analytically, without autograd). GradNorm-B adds only the enable_grad island.

The one real new requirement: an eval gate

torch.inference_mode() is not escapable like no_grad: tensors created inside it are inference tensors and cannot participate in autograd recording (raises RuntimeError). Since Option B updates inside forward, calling the weighting in an inference_mode eval loop would crash — whereas in Option A, plain forward is update-free and safe.

The idiomatic fix is gating the update on self.training (as sketched above): same contract as dropout/BatchNorm, works by default (training=True), and aggregator.eval() propagates to the submodule automatically since GradNorm registers GradNormWeighting as a child module. This is a semantic addition Option A doesn't have, but it's the standard PyTorch mechanism for "stateful module that learns during forward". (Secondary detail: also skip the update if weights.requires_grad is False, so freezing doesn't make autograd.grad raise.)

Concern 4 — state_dict and optimizer state

Three tiers, in increasing effort:

Tier 0 — accept the FAMO status quo

FAMO's internal Adam is a plain lazily-created attribute (_famo.py:135,170-171); its state is in no state dict, and reset() just drops it (self._optimizer = None). TorchJD already ships this. For GradNorm, the important learned state (weights, _initial_losses baseline) is already in the module state dict; only Adam's moments would restart on resume — a brief transient in weight adaptation. So Option B would be no worse than the existing precedent even if we do nothing.

Tier 1 — embed the optimizer state (the real fix, ~15 lines)

_OPT_KEY = "_optimizer_state"

def state_dict(self, *args, **kwargs):
    sd = super().state_dict(*args, **kwargs)
    if self._optimizer is not None:
        sd[_OPT_KEY] = self._optimizer.state_dict()
    return sd

def load_state_dict(self, state_dict, strict=True, assign=False):
    opt_sd = state_dict.pop(_OPT_KEY, None)
    result = super().load_state_dict(state_dict, strict=strict, assign=assign)
    if opt_sd is not None:
        self._ensure_optimizer()
        self._optimizer.load_state_dict(opt_sd)
    return result

This works because of several favorable facts, all verified against this codebase/PyTorch:

  • Strictness both ways: popping the key before delegating avoids the "unexpected key" error, and since the key isn't a registered param/buffer, old checkpoints without it load without "missing key" errors. Backward compatible with the PR's existing checkpoints.
  • Serializability: Optimizer.state_dict() is a plain dict of tensors/ints/floats — survives torch.save/torch.load (including weights_only=True), and the deepcopy used in the PR's checkpoint test.
  • Device/dtype: Optimizer.load_state_dict re-casts state tensors to the parameter's current device/dtype — the standard resume-to-GPU path.
  • Trivially correct param mapping: Optimizer.load_state_dict maps params positionally. Here there is exactly one param group with one parameter whose identity never changes (reset() fills in place) — this is where the PR's fixed-n_tasks design decision pays off.
  • Caveat to document: state_dict() values are conventionally tensors; a nested dict breaks tensor-assuming consumers (safetensors, some checkpointing libs). Fine for TorchJD's torch.save-based audience, but worth one docstring sentence. (A _register_state_dict_hook + _register_load_state_dict_pre_hook pair achieves the same without overriding public methods — equivalent, slightly less readable.)

Tier 2 — expose, don't embed

A weight_optimizer property (force-creating the optimizer if lazy) lets power users checkpoint it themselves, exactly as in Option A. Composes with Tier 0 and can be added non-breakingly at any time.

Eager vs lazy creation (sub-question that falls out of Tier 1)

  • Eager in __init__ is safe despite "move to device before creating the optimizer": Module.to()/.double()/.cuda() mutate param.data in place (Parameter identity is preserved by default), so an optimizer built in __init__ keeps valid references; and Adam/SGD/AdamW initialize their state lazily at first step(), i.e. on the post-move device/dtype. The known breakage cases (old Adagrad allocating state in the constructor; moving after steps have initialized state) apply equally to Option A.
  • Lazy at first forward (FAMO-style) sidesteps even that, at the cost of a _ensure_optimizer() call during loading. Given weights exists from __init__, eager creation gives the simpler Tier-1 story.

Bonus observed during the investigation

Option B makes reset() strictly better than A: recreating the internal optimizer atomically discards momentum, whereas A's reset() cannot touch the external optimizer — a wart the PR had to document and even encode in a test (test_reset_keeps_parameter_and_replaces_baseline, which asserts the external optimizer keeps stale state after reset).

Verdict

Both drawbacks are solvable with patterns already present in (3) or compatible with (4) this codebase:

Drawback Solution Cost
3) _NonDifferentiable FAMO's enable_grad + autograd.grad island Plus one genuinely new but standard piece: a self.training gate for eval/inference-mode
4) state_dict Tier 1 override (~15 lines), or Tier 0 = FAMO status quo One docstring caveat about non-tensor state dict values

Neither is structural. If the maintainer preference is Option B provided 3) and 4) are solvable, I'd say the condition is met — with the caveat that the training-mode gate should be called out explicitly in the docs, since no other TorchJD aggregator currently changes behavior under .eval().

New%20session%20-%202026-10-03T10%3A46%3A16.022Z
opencode session  |  github run

@ValerianRey ValerianRey left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Based on my discussion with OpenCode, I'd like to change the API to something where the weight optimizer is internal only.

For the NonDifferentiableMixin issue, a torch.enable_grad() island, like in FAMO, seems correct. I think having a .eval() and .train() method, with the same interface as those present in some pytorch modules, would also be helpful.

For the StateDict issue, just having an optimizer property to let the user get the optimizer themselves, and save its state dict manually, seems like the best option.

Could you try making the changes to reach this new interface?

Comment thread README.md
| [ExcessMTL](https://torchjd.org/stable/docs/aggregation/excess_mtl#torchjd.aggregation.ExcessMTL) | [ExcessMTLWeighting](https://torchjd.org/stable/docs/aggregation/excess_mtl#torchjd.aggregation.ExcessMTLWeighting) | [Robust Multi-Task Learning with Excess Risks](https://proceedings.mlr.press/v235/he24n.html) |
| [FairGrad](https://torchjd.org/stable/docs/aggregation/fairgrad#torchjd.aggregation.FairGrad) | [FairGradWeighting](https://torchjd.org/stable/docs/aggregation/fairgrad#torchjd.aggregation.FairGradWeighting) | [Fair Resource Allocation in Multi-Task Learning](https://arxiv.org/pdf/2402.15638) |
| [GradDrop](https://torchjd.org/stable/docs/aggregation/graddrop#torchjd.aggregation.GradDrop) | - | [Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign Dropout](https://arxiv.org/pdf/2010.06808) |
| [GradNorm](docs/source/docs/aggregation/gradnorm.rst) | [GradNormWeighting](docs/source/docs/aggregation/gradnorm.rst) | [GradNorm: Gradient Normalization for Adaptive Loss Balancing in Deep Multitask Networks](https://proceedings.mlr.press/v80/chen18a.html) |

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can remove that so that it doesn't appear in the readme until we actually make a release (the release skill will add this row back to the readme).

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We usually have method-specific examples directly in the docstring of the class of the method. Could you move it there?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Very cool example!

model_optimizer.zero_grad()
weight_optimizer.zero_grad()
losses = model(features).square().mean(dim=0)
jacs = jac(losses, list(model.parameters()), retain_graph=True)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

To match the paper a bit better, we could compute only the jacobians w.r.t. the last shared layer here.

Comment on lines +78 to +79
See :doc:`the GradNorm example <../../examples/gradnorm>` for computing the norms
using only the last shared layer while weighting all model parameters.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we do the changes mentioned before, this can be simply removed. We could just have 2 examples in GradNormWeighting: one using autojac to compute jacobians w.r.t. the last shared layer's params, and one using autogram to compute directly the gramian w.r.t. the last shared layer's params.

In GradNorm (the aggregator), we could have just one example, using torchjd.autojac.backward to accumulate jacobians in .jac, then jac_to_grad with the GradNorm aggregator, which will itself do everything.

This will not be equivalent to the examples in GradNormWeighting, because in GradNormWeighting we only consider the last shared layer's params to update the weights, instead of all params. This should be mentioned, maybe with a link to GradNormWeighting to show how to do the alternative.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

cc: feat Conventional commit type for new features. package: aggregation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants