无需训练即可检测多模态数据盗用,通过自适应提示捕捉数据特征。
PATFinger: Prompt-Adapted Transferable Fingerprinting against Unauthorized Multimodal Dataset Usage
- 基于全局最优扰动与自适应提示,无需训练提取数据指纹。
- 在跨模态检索中比现有方法提升30%检测效果。
- 适合数据所有者验证模型是否未经授权使用其多模态数据。
多模态数据可通过提供跨模态语义来预训练大规模视觉语言模型。当前数据使用检测主要聚焦于单模态数据所有权验证,采用侵入式方法或非侵入式技术,而跨模态方法仍待探索。侵入式方法虽可适配多模态数据但降低模型精度,非侵入式方法依赖标签驱动决策边界,难以保证验证行为的稳定性。为此,我们提出一种无需训练的新型可迁移指纹方案PATFinger,结合全局最优扰动(GOP)与自适应提示,以捕获数据集特异性分布特征。该方案利用数据固有属性作为指纹,而非强制模型学习触发信号。GOP基于样本分布设计,最大化不同模态间的嵌入偏移。随后,PATFinger将自适应提示与GOP样本对齐,在精心构建的代理模型上捕捉跨模态交互。这使数据所有者可通过观察检索查询中的特定预测行为来判断数据是否被非法使用。大量实验表明,该方案在多种跨模态检索架构上,相比最先进基线提升了30%的检测有效性。
原文摘要 · Abstract (English)
The multimodal datasets can be leveraged to pre-train large-scale vision-language models by providing cross-modal semantics. Current endeavors for determining the usage of datasets mainly focus on single-modal dataset ownership verification through intrusive methods and non-intrusive techniques, while cross-modal approaches remain under-explored. Intrusive methods can adapt to multimodal datasets but degrade model accuracy, while non-intrusive methods rely on label-driven decision boundaries that fail to guarantee stable behaviors for verification. To address these issues, we propose a novel prompt-adapted transferable fingerprinting scheme from a training-free perspective, called PATFinger, which incorporates the global optimal perturbation (GOP) and the adaptive prompts to capture dataset-specific distribution characteristics. Our scheme utilizes inherent dataset attributes as fingerprints instead of compelling the model to learn triggers. The GOP is derived from the sample distribution to maximize embedding drifts between different modalities. Subsequently, our PATFinger re-aligns the adaptive prompt with GOP samples to capture the cross-modal interactions on the carefully crafted surrogate model. This allows the dataset owner to check the usage of datasets by observing specific prediction behaviors linked to the PATFinger during retrieval queries. Extensive experiments demonstrate the effectiveness of our scheme against unauthorized multimodal dataset usage on various cross-modal retrieval architectures by 30% over state-of-the-art baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。