Life sciences · Preprint
arXiv · September 8, 2026
Raises a question worth testing. It does not answer one.
This preprint proposes that equivariant neural networks suffer a learning-rate mismatch when trained with Adam because gradient rank grows with irrep multiplicity, causing different spectral step sizes across blocks. The authors show in controlled toy experiments that blockwise normalization and momentum tuning can make Adam competitive with Muon on rMD17 and MD22 benchmarks, but the work remains unrefereed and lacks quantitative effect sizes and statistical testing.
Preprint; mechanistic analysis with controlled experiments on toy model and empirical evaluation on molecular modelling benchmarks. SO(3)-equivariant neural networks and e3nn interatomic potential models. Intervention: Blockwise normalization of Adam updates per irrep block; independent tuning of momentum coefficients. Compared with: Standard Adam; Muon optimizer.
Gradient of $W_l$ sums $2l+1$ outer product contributions with rank at most $2l+1$, creating block-dependent learning rates under Adam Toy SO(3)-equivariant model shows mismatch grows with width; dense control shows no corresponding growth Block normalization and tuned momentum coefficients combined make Adam competitive with Muon on all datasets
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
This preprint proposes a mechanistic explanation for optimizer performance differences in equivariant networks through controlled experiments, but makes no clinical or hard-outcome claims and remains unrefereed.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Equivariant networks are commonly trained with Adam, yet recent work reports that matrix-structured optimizers such as Muon can perform better on these architectures without explaining why. We identify one source of this difference inside equivariant linear layers. Each irrep block learns a channel-mixing matrix $W_l$ shared across its $2l+1$ components, giving the expanded map $W_l \otimes I_{2l+1}$. For a single application of the layer, the gradient of $W_l$ sums $2l+1$ outer product contributions and has rank at most $2l+1$. Adam rescales stored weights individually without using the irrep boundaries, so one learning rate can produce different spectral step sizes across blocks within a layer. We address this mismatch by normalizing each block update separately, without introducing a new hyperparameter. This changes only the scale of the update, leaving Adam's moment estimates and its direction within each block unchanged. We evaluate the mechanism in a controlled $\mathrm{SO}(3)$-equivariant model with a matched dense control and in an e3nn interatomic potential model trained on rMD17 and MD22. The toy setup isolates a mismatch that grows with width while the dense control shows no corresponding growth. In the interatomic potential model, block normalization and tuning Adam's momentum coefficients independently improve performance, but neither alone matches Muon. Combined, they make Adam competitive with Muon on all datasets, indicating that blockwise step control and momentum accumulation account for much of Muon's advantage.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.