arXiv:2504.21044cs.CRcs.AI2025-04被引 2

用隐蔽触发器保护多模态模型版权,防篡改且难检测。

AGATE: Stealthy Black-box Watermarking for Multimodal Model Copyright Protection

  • 从普通数据生成隐蔽对抗性触发器,保持视觉真实并引发语义变化。
  • 通过后处理模块缩小图像与文本嵌入距离,避免输出异常被识破。
  • 两阶段验证机制可有效识别盗用模型,适用于版权保护场景。

大规模多模态AI模型日益成为基础性技术,也面临模型盗用风险。现有方法使用分布外(OoD)数据作为后门水印并重新训练原模型进行版权保护,但易被恶意检测和伪造,导致水印逃逸。本文提出模型无关的黑盒后门水印框架AGATE,解决隐蔽性与鲁棒性问题。具体地,我们设计一种对抗性触发器生成方法,从常规数据中生成隐蔽触发器,在保持视觉保真度的同时诱导语义偏移;为缓解模型输出异常带来的检测风险,引入后处理模块,通过缩小对抗触发图像嵌入与文本嵌入之间的距离来修正输出;随后提出两阶段水印验证机制,通过对比启用与禁用该模块时的结果判断模型是否侵权。实验表明,AGATE在五个数据集上的多模态图文检索与图像分类任务中持续优于当前最优方法。此外,在两种对抗攻击场景下验证了其鲁棒性。

原文摘要 · Abstract (English)

Recent advancement in large-scale Artificial Intelligence (AI) models offering multimodal services have become foundational in AI systems, making them prime targets for model theft. Existing methods select Out-of-Distribution (OoD) data as backdoor watermarks and retrain the original model for copyright protection. However, existing methods are susceptible to malicious detection and forgery by adversaries, resulting in watermark evasion. In this work, we propose Model-\underline{ag}nostic Black-box Backdoor W\underline{ate}rmarking Framework (AGATE) to address stealthiness and robustness challenges in multimodal model copyright protection. Specifically, we propose an adversarial trigger generation method to generate stealthy adversarial triggers from ordinary dataset, providing visual fidelity while inducing semantic shifts. To alleviate the issue of anomaly detection among model outputs, we propose a post-transform module to correct the model output by narrowing the distance between adversarial trigger image embedding and text embedding. Subsequently, a two-phase watermark verification is proposed to judge whether the current model infringes by comparing the two results with and without the transform module. Consequently, we consistently outperform state-of-the-art methods across five datasets in the downstream tasks of multimodal image-text retrieval and image classification. Additionally, we validated the robustness of AGATE under two adversarial attack scenarios.

模型版权水印技术多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。