Life sciences · Preprint
arXiv · September 10, 2026
Early or partial results. Treat as a signal, not a conclusion.
This is a computational audit of 263 released batch-normalized machine-learning checkpoints, showing that refitting batch-normalization statistics on retained data alters published unlearning verdicts in 12 cases, with 2 cases clearing a recalibration budget on every replicate. The effect is real but narrow and appears attributable to checkpoint state drift rather than survival of removed data.
Computational audit of released checkpoints. Released batch-normalized machine-learning checkpoints from unlearning studies, described as vision models.. Intervention: Refitting batch-normalization statistics on retained data at bit-identical model weights. Compared with: Original published checkpoints and their variance across method seeds. n = 263.
Refitting batch-normalization statistics moves 47 of 221 released checkpoints beyond the variance shown by the method's own seeds 12 verdicts cross when batch-normalization statistics are refitted on retained data 2 verdict crossings clear a measured recalibration budget on every replicate
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
An uncontrolled audit of machine learning checkpoints documenting a technical artifact in batch-normalization statistics that affects a small fraction of published unlearning verdicts, without experimental intervention or comparison group.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
An unlearning audit reads its verdict off numbers that an unlearned model and its retrained reference each publish, and both also ship batch-normalization statistics that no gradient step wrote and no release records. Refitting them on kept data at bit-identical weights moves 47 of 221 released checkpoints past the spread their own release's seeds show, several inside a method whose average does not move: what moves is the checkpoint's property, not its method's. What does the moving is not the removed data surviving in the state: exchanging kept records for removed ones inside a fixed fitting pool moves a published cell by almost nothing, while how far a checkpoint's shipped state has drifted from any refit does track it. The consequence for a published decision is real but narrow: twelve verdicts cross, four clear a measured recalibration budget, two clear it on every replicate, and a population we trained and sited near its own criterion yields none. A release should therefore name the fitting convention beside the number, on the batch-normalized vision models where this channel exists.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.