Life sciences · Preprint
arXiv · September 8, 2026
Early or partial results. Treat as a signal, not a conclusion.
This preprint proposes unified hyperparameter scaling laws for Mixture-of-Experts models across sparsity levels, derived from 1,800 large-scale pre-training runs. The authors characterize how optimal learning rate and batch size vary with activation ratio and training compute, and validate the scaling form on one held-out ultra-sparse (1/64 activation) model; however, the work remains unpeer-reviewed and validation beyond that single test case is not reported.
Computational empirical study with systematic hyperparameter sweep. Mixture-of-Experts pre-training models; six activated-parameter scale points explored; models up to 6B total non-embedding parameters studied.. Intervention: Systematic variation of activation ratio (sparsity), model size, data volume, and expert granularity across pre-training runs.. Compared with: Alternative hyperparameter scaling functional forms evaluated; prior conflicting findings in literature noted but not directly reproduced or compared in source..
1,800 pre-training runs conducted spanning six activated-parameter scales and models with up to 6B total non-embedding parameters, processing approximately 20 trillion tokens Optimal batch size follows power-law relationship with training tokens D at fixed sparsity Optimal learning rate scales with training compute C and remains robust to allocation between model size and data
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
Large-scale computational study proposing novel hyperparameter scaling laws for sparse models, but lacks peer review, clinical/patient outcomes, and independent validation on held-out test cases beyond one ultra-sparse configuration.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Mixture-of-Experts (MoE) models expand model capacity without a proportional increase in training compute, but increasing sparsity makes reliable hyperparameter transfer challenging. In this work, we show that conventional hyperparameter scaling laws are insufficient for ultra-sparse MoEs: the optimal learning rate and batch size vary with activation ratio, and these shifts cannot be explained by either total or activated parameter count alone. To characterize this dependence, we conduct 1,800 pre-training runs spanning six activated-parameter scales and models with up to 6B total non-embedding parameters, processing approximately 20 trillion tokens at a cost of 200,000 equivalent H800 GPU-hours. Our results reconcile conflicting findings in prior work by revealing two scaling regimes. At fixed sparsity, the optimal batch size follows a power-law relationship with training tokens $D$, whereas the optimal learning rate scales with training compute $C$ and remains robust to the allocation between model size and data. Across sparsity levels, the activation ratio $A$ enters both relationships as an additional multiplicative power-law factor. These observations lead to unified hyperparameter scaling laws that transfer across MoE sparsity levels. Large-scale evaluation shows that the scaling form outperforms alternative functional forms. On a held-out ultra-sparse MoE with 12B total parameters and only 1/64 of its experts activated, the predicted hyperparameters remain close to the observed optima, supporting joint extrapolation across model scale and sparsity. Further experiments demonstrate transfer across expert granularities and isolate the effect of activation ratio from that of total expert count.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.