用少量污染数据让视觉语言模型更易被查出是否盗用数据。
MemCatalyst: Amplifying Data Auditing on Vision-Language Models via Data Poisoning

- 通过污染文本和图像,让模型过度学习图文不一致特征。
- 仅用少量污染样本,将成员推理攻击成功率提升超30%。
- 方法对多种模型通用,适合关注数据版权的创作者使用。
视觉语言模型(VLMs)因互联网海量训练数据而表现卓越。然而,数据提供者(如艺术家)亟需确认其数据是否未经许可被用于模型训练,这涉及知识产权与隐私问题。数据审计,尤其是成员推理(MI),成为直接检测手段。本文提出MemCatalyst,一套数据投毒工具,旨在增强VLMs的数据审计能力。该方法采用两种策略:污染文本(PT)和污染图像(PI)。MemCatalyst迫使模型在训练中过度学习图像特征与文本语义间的特定不一致,从而增加其对成员信息的敏感性。关键的是,污染样本在不同VLM架构间具有良好的可迁移性,适用于黑盒场景。在两个主流VLM上,对五种先进数据审计方法进行广泛评估表明,MemCatalyst以极小的污染样本预算显著提升MI AUC分数,同时对模型性能影响可忽略。
原文摘要 · Abstract (English)
Vision-Language models (VLMs) achieve outstanding performance largely due to the amount of training data available on the internet. At the same time, data holders (e.g., artists) urgently need to determine whether their data has been used for model training without authorization, which concerns both intellectual property rights and personal privacy. Data auditing, particularly through membership inference (MI), has attracted attention as a direct tool. This work proposes MemCatalyst, a set of data poisoning tools, aiming to amplify the data auditing performance on VLMs. MemCatalyst employs two strategies: Poisoning Text (PT) and Poisoning Image (PI). MemCatalyst forces VLMs to over-learn specific inconsistencies between image features and textual semantics during training, thereby increasing their susceptibility to membership information auditing. Crucially, the transferability of poisoned samples across different VLM architectures is demonstrated to be effective in the black-box setting. Extensive evaluations using five state-of-the-art data audits on two prominent VLMs demonstrate that MemCatalyst markedly enhances MI AUC scores with a minimal budget of poisoned samples, while maintaining a negligible impact on model performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。