Life sciences · Preprint
arXiv · September 4, 2026
Early or partial results. Treat as a signal, not a conclusion.
This preprint reports a strong negative correlation between the number of discrete class-separability phase transitions during ResNet training and final test accuracy on clean in-distribution benchmarks (CIFAR-10 r = −0.84, CIFAR-100 r = −0.87), but the relationship weakens substantially under distributional stress. The work is descriptive and exploratory, proposing a training-time proxy for model quality; it does not establish causation, generalize beyond vision tasks, or outperform established competitor metrics on stressed data.
Empirical correlation study across multiple image classification benchmarks. ResNet models trained on CIFAR-10, CIFAR-100, TinyImageNet, and CIFAR-10-C (corruption variant).. Intervention: Measurement of phase transition frequency as a training-time predictor during standard supervised finetuning.. Compared with: Six alternative training-curve signals; no explicit control condition, as the study is observational and correlational.. n = 75.
On CIFAR-10, strong negative correlation between phase transition count and test accuracy: r = −0.84 (p < 10−8, n = 30) On CIFAR-100, strong negative correlation: r = −0.87 (p < 10−5, n = 15) Partial correlation on CIFAR-100 controlling for architecture depth: r_partial = −0.69 (p = 0.007)
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
This is an empirical observational study of a training dynamics proxy in neural networks, with correlational evidence across limited benchmarks and no causal inference or clinical translation; the finding is domain-specific to computer vision and does not address a clinical or health outcome.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
The number of discrete class-separability jumps observed during ResNet finetuning is examined empirically as a predictor of final test accuracy. Across 75 experiments spanning four benchmarks (CIFAR-10, CIFAR-100, TinyImageNet, and CIFAR-10-C) and three architectures (ResNet-18, ResNet-50, and ResNet-101), with five to ten seeds per configuration, a strong within-dataset negative correlation is obtained on standard i.i.d. classification benchmarks: \(r = -0.84\) on CIFAR-10 (\(p < 10^{-8}\), \(n = 30\)) and \(r = -0.87\) on CIFAR-100 (\(p < 10^{-5}\), \(n = 15\)). Under distributional stress, the relationship attenuates: TinyImageNet yields \(r = -0.45\), and the CIFAR-10-C corruption benchmark yields \(r = -0.19\). Two additional analyses discipline the empirical claim. A partial correlation controlling for architecture depth, treated as a linear covariate, shows that on CIFAR-100 the transition count retains statistically significant predictive power (\(r_{\mathrm{partial}} = -0.69\), \(p = 0.007\)); the corresponding result under the stricter categorical conditioning is not established at \(n = 15\). A comparison against six alternative training-curve signals shows that transition count achieved the strongest correlation among the evaluated signals on CIFAR-100 and one of the strongest on CIFAR-10, but is dominated by other signals on the two stressed benchmarks. The comparison is restricted to training-curve-level signals; comparisons against effective rank, Hessian sharpness, Fisher information, margin, and neural-collapse measures, which are the strongest competitors in the current literature, are not part of the present study and remain open. The observation is presented as an in-distribution training-quality probe among a family of candidate probes, and an inexpensive detection procedure suitable for logging alongside a standard training loop is provided.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.