Life sciences · Preprint
arXiv · September 10, 2026
Posted before peer review. The findings may change or fail to hold.
This is an unrefereed preprint describing Musec, a modification to the Muon optimizer for large language model training that replaces spectral flattening with spectral clipping to improve stability. The authors provide theoretical convergence guarantees and empirical evidence that Soft Musec improves stability across learning rates and model sizes, but the work has not been peer reviewed and addresses a computational optimization problem outside the scope of clinical or biomedical evidence.
Preprint.
Musec replaces Muon's spectral flattening with spectral clipping of singular values exceeding a threshold Soft Musec is proposed as an efficient implementation using coupled Newton-Schulz iterations First convergence guarantee reported for Muon-type methods in nonconvex nonsmooth stochastic optimization setting
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
This is an unrefereed preprint presenting a novel optimizer algorithm with theoretical analysis and empirical validation, but it has not undergone peer review and addresses a machine learning optimization problem rather than a clinical or biomedical question.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
What is missing. This record has no reported figures. That is a gap in the analysis, not a judgement about the study.
Muon has emerged as a highly effective optimizer for large language model training, often achieving superior convergence and performance compared with the widely adopted Adam and AdamW optimizers. Nevertheless, Muon is prone to training instability due to its spectral flattening, manifested by loss spikes and unbounded growth of model weights. Existing approaches primarily rely on weight or attention-logit clipping, which require architecture-specific modifications and do not directly address instability across all model components. We propose MomentUm SpEctral Clipping (Musec), which replaces Muon's spectral flattening with spectral clipping: rather than setting all singular values of the momentum matrix to approximately one, Musec clips singular values that exceed a threshold while preserving the underlying spectral structure of the momentum. Our strategy provides an optimizer-level, architecture-agnostic mechanism for stabilizing Muon training. We further develop Soft Musec, an efficient implementation that uses a smooth spectral saturation function approximated by coupled Newton-Schulz iterations. Theoretically, we establish convergence guarantees for Musec in nonconvex nonsmooth stochastic optimization. To the best of our knowledge, this is the first convergence guarantee for Muon-type methods in the nonconvex nonsmooth setting. We provide empirical studies to show that Soft Musec consistently improves training stability over existing Muon variants across a wide range of learning rates and model sizes. Notably, Soft Musec remains stable in settings where existing Muon variants diverge, while matching their performance under well-tuned configurations.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.