Life sciences · Preprint
arXiv · September 9, 2026
Early or partial results. Treat as a signal, not a conclusion.
This unreviewed computational study tests sparse autoencoders on a controlled synthetic problem and reports an unexpected result: universal convergence to a diffuse phase rather than the predicted dictionary recovery or feature merging. The learned features remain distant from ground truth (median cosine similarity 0.5–0.7 vs. 0.95 criterion) and codes are an order of magnitude denser than true sparse codes, suggesting that gradient-trained SAEs may not reach global optima of their stated objective.
Controlled computational experiment on synthetic data. Synthetic sparse coding instances with known nested-feature dictionary; no human or animal subjects.. Intervention: Sparse autoencoder training with variable nesting fraction γ, sparsity penalty λ, and dictionary size M. Compared with: Ground truth synthetic dictionary and sparse codes; theoretical predictions of phase diagram (recovery vs. merging).
Zero full-dictionary recoveries across 200 independent fits over ten grid cells Zero merges observed; every run converges reproducibly to diffuse phase Median best cosine similarity 0.5–0.7 against 0.95 recovery criterion, indicating learned atoms remain far from true features
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
This is an empirical study of sparse autoencoders addressing a defined open problem, but it reports unexpected null results (no recoveries, no merges) and characterizes a previously undescribed diffuse phase rather than validating the proposed phase diagram; the work is unreviewed and requires confirmation.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Sparse autoencoders (SAEs) are increasingly used to recover interpretable features from neural-network activations, yet systematic feature co-occurrence can cause distinct features to be absorbed or merged. The MAIS-O43 open problem proposes a controlled experiment to characterize when recovery of a true synthetic dictionary gives way to feature merging as the nesting fraction $γ$, sparsity penalty $λ$, and dictionary size $M$ vary. We implement the specified protocol and evaluate 200 independently initialized fits across ten of the 165 grid cells. We observe zero full-dictionary recoveries and zero merges. Instead, every run converges to a reproducible diffuse phase: reconstruction is nearly perfect, but learned atoms typically remain far from the true features (median best cosine 0.5-0.7 against a 0.95 recovery criterion) and learned codes are an order of magnitude denser than the ground truth. This behavior persists under robustness checks and across the full 165-cell grid using standard minibatch Adam (3,300 additional fits). Since the global optimum of the exact sparse-coding objective is known to merge nested features in the two-feature case, these results suggest that trained SAEs need not reach the corresponding minima, and that the phase diagram of trained models may differ fundamentally from that of objective minimizers.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.