Life sciences · Preprint
arXiv · September 8, 2026
Early or partial results. Treat as a signal, not a conclusion.
This preprint identifies two specific failure modes in ConSim, an automated simulatability evaluation protocol for explanations: LLM simulators can bypass explanations by solving tasks directly from class names, and can exploit class anonymization to infer hidden labels. The findings suggest that simulator predictions rely primarily on task priors rather than explanation content, undermining the validity of automated simulatability as a proxy for human-judged explanation utility.
Qualitative methodological critique and replication study. Explanation evaluation protocols and LLM simulators applied to classification tasks with meaningful class names. Intervention: Classes-as-concepts baseline and class anonymization procedure. Compared with: ConSim protocol and original explanation methods.
Simulators obtain high simulatability by solving classification tasks directly without relying on explanations when class names are meaningful Class anonymization can reward explanations for leaking hidden label mappings Simulator predictions mainly rely on task priors, while explanations produce small changes
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
A methodological critique identifying systematic vulnerabilities in an automated evaluation protocol, based on qualitative replication and new baselines, but lacking empirical quantification of the magnitude or prevalence of the shortcuts identified.
As stated by the source record.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
What is missing. This record has no reported figures. That is a gap in the analysis, not a judgement about the study.
Simulatability is an evaluation protocol for explanations that quantifies their usefulness by how well they help a user predict a task model's outputs. Since human evaluation is costly, automated simulatability replaces human explainees with LLM simulators, as proposed in ConSim (Poché et al., 2025) for large-scale experiments. We qualitatively replicate and extend ConSim's ranking of explanation methods across the tested datasets, explanation families, and simulator LLMs, and identify two limitations. First, when class names are meaningful, simulators can obtain high simulatability by solving the classification task directly, without relying on the explanations. Second, class anonymization can reward explanations for leaking the hidden label mapping, a limitation we expose with a new classes-as-concepts baseline. These results are consistent with a shortcut hypothesis: in the tested settings, simulator predictions mainly rely on task priors, while explanations produce small changes. We derive recommendations for more robust automated simulatability evaluations.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.