Life sciences · Preprint
arXiv · August 18, 2026
Early or partial results. Treat as a signal, not a conclusion.
The authors propose an optimal-transport rectified flow for generating label-free explanations of clinical AI predictions and test whether the resulting heatmaps localize disease. On tabular biomarkers, the method achieves AUROC 0.91 and modest agreement with supervised attribution (r ~0.5), but on chest X-rays it reveals a critical synthetic-to-real gap: heatmaps localize artificial lesions (pointing game 0.52) but collapse to chance on real radiologist annotations, whereas supervised Grad-CAM remains above chance. This finding questions whether generative explanation heatmaps offer genuine localization without supervision.
Methodological study with controlled synthetic and real-world benchmarking. Tabular tumour biomarker data from Breast Cancer Wisconsin; chest X-ray imagery with radiologist annotations from RSNA.. Intervention: Optimal-transport rectified flow for generative explanations of clinical distributions. Compared with: Logistic regression (tabular), supervised Grad-CAM (imaging).
Unsupervised malignancy score on tabular tumour biomarkers: AUROC 0.91 (0.93 ± 0.01 across five seeds) Label-free attribution agrees with supervised classifier at r ≈ 0.5 on tabular data On synthetic chest X-rays, transport heatmap localizes lesions at pointing game 0.52
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
Clinicians and AI developers should be cautious about trusting label-free explanation heatmaps from generative models, particularly in imaging: compelling visualizations on synthetic data do not guarantee real disease localization. The work highlights the importance of external validation and stress-testing before adopting such explanations for clinical decision support.
An early-stage methodological study developing and stress-testing a novel optimal-transport explanation framework for clinical AI, with mixed results and a synthetic-to-real gap that limits immediate clinical utility.
As stated by the source record.
Quoted from the source exactly as published.
Clinicians and AI developers should be cautious about trusting label-free explanation heatmaps from generative models, particularly in imaging: compelling visualizations on synthetic data do not guarantee real disease localization. The work highlights the importance of external validation and stress-testing before adopting such explanations for clinical decision support.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Generative models promise a route to explainable clinical AI: rather than probe a classifier, model the distributions of healthy and diseased patients and read explanations off the geometry between them. We build such a system - an optimal-transport rectified flow trained between two clinical distributions - and use it to ask a pointed question the field too rarely tests: do the resulting explanation heatmaps actually localize disease? On tabular tumour biomarkers (Breast Cancer Wisconsin) a single flow yields per-patient counterfactuals, an unsupervised malignancy score (AUROC 0.91; 0.93 +/- 0.01 across five seeds), and a label-free attribution that agrees with a supervised classifier (r ~ 0.5) - a compact, honest interpretability engine, though it never out-predicts logistic regression. Moving to chest X-rays, we show the transport heatmap is a population-level signal, not a localiser; a reconstruction-based, identity-preserving variant does localize synthetic lesions (pointing game 0.52), yet on real RSNA radiologist boxes it collapses to chance while only supervised Grad-CAM stays above it. The central result is a synthetic-to-real gap: label-free heatmaps that look compelling on planted lesions are not evidence of real localisation. We contribute a reusable optimal-transport recipe for generative explanations and a controlled benchmark for stress-testing whether they localize.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.