Life sciences · Preprint
arXiv · August 18, 2026
Raises a question worth testing. It does not answer one.
This is a preprint presenting MemCatalyst, a data poisoning technique designed to amplify membership inference attacks on vision-language models by forcing over-learning of image-text inconsistencies. The work demonstrates computational proof-of-concept across multiple VLM architectures with improved membership inference AUC scores and cross-model transferability, but lacks peer review and addresses a machine-learning security problem rather than a clinical or health outcome.
Computational methods development with empirical evaluation across multiple architectures. Vision-Language Models (VLMs); two prominent architectures evaluated. Intervention: MemCatalyst data poisoning with Poisoning Text (PT) and Poisoning Image (PI) strategies. Compared with: Baseline membership inference auditing without poisoning.
MemCatalyst markedly enhances MI AUC scores with a minimal budget of poisoned samples Poisoned samples transfer effectively across different VLM architectures in black-box settings Model performance is maintained with negligible impact despite data poisoning
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
This is a methods paper proposing a novel data poisoning technique for amplifying membership inference attacks on vision-language models; it presents a computational approach without clinical, patient, or real-world health outcomes.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Vision-Language models (VLMs) achieve outstanding performance largely due to the amount of training data available on the internet. At the same time, data holders (e.g., artists) urgently need to determine whether their data has been used for model training without authorization, which concerns both intellectual property rights and personal privacy. Data auditing, particularly through membership inference (MI), has attracted attention as a direct tool. This work proposes MemCatalyst, a set of data poisoning tools, aiming to amplify the data auditing performance on VLMs. MemCatalyst employs two strategies: Poisoning Text (PT) and Poisoning Image (PI). MemCatalyst forces VLMs to over-learn specific inconsistencies between image features and textual semantics during training, thereby increasing their susceptibility to membership information auditing. Crucially, the transferability of poisoned samples across different VLM architectures is demonstrated to be effective in the black-box setting. Extensive evaluations using five state-of-the-art data audits on two prominent VLMs demonstrate that MemCatalyst markedly enhances MI AUC scores with a minimal budget of poisoned samples, while maintaining a negligible impact on model performance.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.