Life sciences · Preprint
arXiv · September 10, 2026
Raises a question worth testing. It does not answer one.
CausalArena is a unified benchmark framework for evaluating causal discovery methods across diverse structural causal model families and evaluation protocols. Experiments reveal substantial performance variability across benchmark regimes, suggesting that method rankings are not stable across different SCM structures and mechanisms, and that foundation model pretraining–evaluation overlap complicates interpretation of fixed-benchmark results.
Benchmarking and comparative evaluation study. Causal discovery methods (classical, neural, and foundation model-based) evaluated on diverse structural causal models. Intervention: CausalArena benchmark framework with multiple SCM families and unified evaluation protocol. Compared with: Classical causal discovery methods versus neural and pretrained foundation model approaches across diverse benchmark regimes.
Strong performance in one benchmark regime does not reliably transfer to others Ranking shifts occur across SCM families and protocols for classical, neural, and pretrained methods Benchmark diversity and pretraining–evaluation overlap identified as central evaluation challenges
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
This is a benchmarking and methodological paper that raises questions about evaluation practices in causal discovery rather than testing a clinical or intervention hypothesis with empirical outcome data.
As stated by the source record.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
What is missing. This record has no reported figures. That is a gap in the analysis, not a judgement about the study.
Causal discovery aims to uncover causal structures from data and is fundamental to scientific reasoning and intervention-based decision making. Its evaluation relies heavily on structural causal models (SCMs), which specify a causal graph together with the mechanisms that generate data, yet existing studies differ substantially in graph families, mechanisms, and evaluation protocols. The emergence of causal discovery foundation models (CDFMs) further complicates evaluation: performance may reflect not only causal discovery ability, but also overlap between pretraining environments and test SCMs, making results on fixed synthetic benchmarks difficult to interpret. We introduce CausalArena, a unified and evolvable benchmark for causal discovery under a common protocol. Synthetic SCMs supply controlled breadth over structures and mechanisms; semantic operational SCMs provide human-auditable, semantically grounded environments beyond standard synthetic generators; and formula-grounded SCMs test discovery under explicit scientific mechanisms. Public real-world datasets provide an additional external-validity check. Experiments across classical, neural, and pretrained methods reveal substantial ranking shifts across SCM families and protocols, showing that strong performance in one benchmark regime does not reliably transfer to others. These results highlight benchmark diversity and pretraining--evaluation overlap as central challenges for evaluating causal discovery in the foundation model era.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.