Life sciences · Preprint
arXiv · September 9, 2026
Early or partial results. Treat as a signal, not a conclusion.
This unreviewed preprint reports a controlled empirical study of adversarial transferability across 60 deepfake detectors, systematically isolating the effects of backbone architecture, pretraining regime, and training data on black-box attack success. The authors report that source–target compatibility substantially shapes transfer rates, with attack-method-dependent patterns: exact backbone compatibility dominates under AutoAttack (7.21% mean ASR), while shared pretraining and training data dominate under Carlini–Wagner EOT (19.52% mean ASR). A multi-source oracle combining both attacks achieves 64.48% mean ASR, indicating that single-source evaluations substantially underestimate vulnerability.
Controlled computational study with systematic factorial manipulation. 60 deepfake detectors, not human subjects.. Intervention: Adversarial examples generated via AutoAttack (AA) and Carlini–Wagner attack with Expectation over Transformation (CW–EOT).. Compared with: Non-target source models versus multi-source oracle; single-source averaging versus source-averaged attack success.. n = 60.
Mean attack success rate (ASR) 7.21% under AutoAttack (AA) for non-target source averaging. Mean ASR 19.52% under Carlini–Wagner with Expectation over Transformation (CW–EOT) for non-target source averaging. Multi-source oracle combining both attacks attains 64.48% mean ASR after excluding exact backbone and training-data matches.
Findings are in silico and do not address real-world deepfake detection, deployment, or user harm.
The source did not state who this applies to in practice.
This is an unreviewed computational study characterizing adversarial transferability across deepfake detectors using controlled experiments, contributing methodological insight but lacking clinical or real-world validation data.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Deepfake detectors remain vulnerable to transfer-based black-box attacks, in which adversarial examples are generated on a source surrogate model and transferred to a target model, unknown to the attacker. Yet how source--target compatibility shapes attack success remains poorly understood. Prior studies evaluate limited detector pools and rarely disentangle architectural from training factors. We conduct a controlled evaluation of adversarial transferability across 60 detectors spanning six backbones, two pretraining regimes, and five training-data configurations, using two attack procedures: AutoAttack (AA) and the Carlini--Wagner attack with Expectation over Transformation (CW--EOT). Matched comparisons reveal significantly higher transfer when source and target share an exact backbone, architecture family, pretraining regime, or training data. This compatibility structure is attack-dependent: exact backbone compatibility has the largest effect under AA, whereas shared pretraining and training data have the largest effects under CW--EOT. When transfer is averaged across non-target sources, mean attack success rate (ASR) is $7.21\%$ under AA and $19.52\%$ under CW--EOT. By contrast, a multi-source oracle combining both attacks attains a \(64.48\%\) mean ASR after excluding exact backbone and training-data matches, showing that source averaging can substantially understate target vulnerability. We release 240,000 adversarially perturbed images, complete pairwise transfer results, detector configurations, and evaluation code. These findings establish source--target compatibility and source-model selection as central dimensions of credible transfer-based black-box robustness evaluation.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.