Life sciences · Preprint
arXiv · September 9, 2026
Early or partial results. Treat as a signal, not a conclusion.
This computational study demonstrates that small Transformers trained with annealable soft positional priors can retain circuit functionality after the prior is removed, but only if the prior is gradually faded during training rather than abruptly switched or removed post-hoc. The finding is reproducible across two synthetic retrieval tasks and shows that training trajectory, not final architecture alone, determines whether learned circuits remain functional.
Controlled computational experiment with systematic manipulation of training trajectory. Small Transformer models on synthetic retrieval and in-context learning tasks. Intervention: Smooth fade-to-zero training of soft positional prior. Compared with: Forced-zero training, hard switching, and post-hoc continuation; baseline of unforced model with prior active.
Unforced models with active prior achieve 0.772 ± 0.020 accuracy on associative recall but collapse to 0.095 ± 0.009 at zero gate Smooth fade-to-zero training preserves high zero-gate accuracy at 0.734 ± 0.028, whereas forced-zero training, hard switching, and post-hoc continuation fail Pattern replicates on Markov induction task
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
Mechanistic study of neural network circuit learning using controlled architectural interventions and ablations; reports reproducible effects but limited to synthetic tasks with small models and no comparison to established baselines or clinical relevance.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Soft positional priors can help small Transformers learn retrieval circuits, but it is unclear whether the resulting circuits remain functional once the prior is removed. We test this with an annealable soft-prior Transformer whose attention biases can be learned, faded, or zeroed during training and evaluation. On associative recall, unforced models perform well with the prior active ($0.772 \pm 0.020$) but collapse at zero gate ($0.095 \pm 0.009$). Smooth fade-to-zero training preserves high zero-gate accuracy ($0.734 \pm 0.028$), whereas forced-zero training, hard switching, and post hoc continuation fail to recover the same effect. The pattern also appears on Markov induction. Linear regression ICL provides a boundary case because zero-gate training can learn that task directly. Mechanistic traces show that circuit consolidation occurs after the gate reaches zero, even though the responsible heads vary across seeds. These results suggest that circuit removability in small discrete retrieval tasks depends on the training trajectory, not just the final architecture.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.