Life sciences · Preprint
arXiv · September 8, 2026
Raises a question worth testing. It does not answer one.
This preprint introduces SAEScientist-Bench, a benchmark for evaluating whether AI agents can autonomously discover interpretable features in neural networks using Sparse Autoencoders. Frontier agents show partial capability in concept separation but substantial gaps remain in causal steering and measurement interpretation, establishing a measurement framework rather than demonstrating reliable autonomous discovery.
Benchmark evaluation study. AI agent systems tasked with autonomous feature discovery in mechanistic interpretability; reference baseline from expert annotators on Neuronpedia. Intervention: AI agents conducting autonomous mechanistic discovery using Sparse Autoencoders and contrastive probe design. Compared with: Expert baseline reference features curated on Neuronpedia.
Benchmark evaluates 10 agent configurations against 20 tasks using 131K+ features in Gemma-2-9B-IT dictionary Frontier agents approach expert levels on separating target concepts from contrastive controls Agents lag substantially in causal generation steering compared to expert baseline
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
This is a benchmarking study of AI agent capabilities in mechanistic interpretability research, establishing measurement methodology but not generating evidence about model safety, alignment, or real-world discovery outcomes.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
While research on recursive self-improvement (RSI) has predominantly automated model training pipelines, reliable autonomous development demands a missing pillar: post-hoc monitoring and auditing to understand what models learn and ensure safe alignment. Mechanistic interpretability tools are essential to bridge this gap, among which Sparse Autoencoders (SAEs) serve as a cornerstone by isolating interpretable features for model inspection and steering. In this paper, we introduce SAEScientist-Bench to evaluate whether AI agents can act as scientists utilizing SAE tools for autonomous mechanistic discovery. Given a target concept, an agent designs contrastive probes and navigates a Gemma Scope dictionary of 131K+ features in Gemma-2-9B-IT to discover the optimal feature, evaluated against curated expert reference features anchored on Neuronpedia across activation rank, concept selectivity on contrastive texts, and causal steering. Across 10 agent configurations and 20 tasks, frontier agents demonstrate genuine discovery capabilities and lead different evaluation dimensions, but remain well behind the expert baseline, approaching expert levels on separating target concepts from contrastive controls while lagging substantially in causal generation steering. Further analysis reveals that although agents can design contrasts to rule out spurious candidates, they frequently misinterpret experimental measurements. These results establish experimental model understanding as a measurable capability for closed-loop autonomous AI R&D. Our code is available at https://github.com/Trae1ounG/SAEScientist.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.