Life sciences · Preprint
arXiv · September 8, 2026
Early or partial results. Treat as a signal, not a conclusion.
Stochastically perturbed weights (SPW) is a post-hoc method to generate probabilistic ensembles from existing deterministic ML weather models without retraining. At 240 h lead time, SPW ensembles achieve continuous ranked probability skill scores 0.04–0.13 below trained-probabilistic baselines but with zero marginal training cost. However, the method requires model-specific tuning of the noise injection site and exhibits systematic failure modes including coherent whole-field offset.
Methodological ablation study; multi-model comparison. Four deterministic machine-learning weather models; evaluated on global medium-range forecasting benchmarks. Intervention: Stochastically perturbed weights (SPW): injection of noise into network weight tensors at inference time to generate probabilistic ensembles from deterministic checkpoints. Compared with: Trained-probabilistic ensembles (AIFS-ENS, FourCastNet 3, Atlas) and operational ECMWF ensemble (IFS-ENS).
At 240 h (10-day) lead time, SPW ensembles reach CRPSS between 0.04 and 0.13 below the best trained-probabilistic baseline No injection site works uniformly across models; productive tensor group is architecture-specific Main failure mode is coherent whole-field offset that overdisperses domain mean; restricting noise to coarse scales or perturbing initial conditions partially repair this
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
A methodological proof-of-concept study proposing a novel post-hoc uncertainty quantification scheme for existing deterministic models, tested across multiple architectures with modest performance gaps to trained baselines but requiring model-specific tuning and showing known failure modes.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Machine-learning weather models (MLWMs) now match or outperform operational numerical weather prediction (NWP) at global medium-range forecasting, at far lower inference cost. Many deployed MLWMs are deterministic, producing a single forecast with no estimate of its own uncertainty, whereas a growing family of trained-probabilistic models generate calibrated ensembles directly, at the price of a dedicated training run. We ask instead how much uncertainty can be extracted from a deterministic checkpoint that already exists, without retraining it. Where physical ensembles represent model uncertainty by stochastically perturbing parametrisation tendencies, we perturb the network's raw weight tensors at inference time, a scheme we call stochastically perturbed weights (SPW). We also ask whether it works, where and on which scales to inject the noise, and where it fails. A three-phase ablation across four deterministic backbones, Aurora, GraphCast, SFNO, and AIFS, selects one production baseline per model, benchmarked against the trained-probabilistic AIFS-ENS, FourCastNet 3 and Atlas as well as the operational ECMWF ensemble (IFS-ENS) over 112 initialisation times. At a 240 h (10-day) lead time the SPW ensembles reach continuous ranked probability skill scores (CRPSS) between 0.04 and 0.13 below the best trained-probabilistic baseline, at zero marginal training cost. No injection site works across models: the productive tensor group is architecture-specific, so SPW is at present a tuning procedure rather than a plug-and-play recipe. Its main failure mode is a coherent whole-field offset that overdisperses the domain mean, and restricting the noise to coarse scales or perturbing the initial conditions each repair part of it.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.