arXiv:2502.16593cs.AIcs.LG2025-02ICLR被引 7

通过对抗样本让大模型自曝版权,无需修改原模型。

Tracking the Copyright of Large Vision-Language Models through Parameter Learning Adversarial Images

  • 用对抗图像诱导原模型学习特定触发信号。
  • 在多种微调后模型上仍能有效识别原始版权归属。
  • 可部署于已发布模型,适合版权保护场景。

大型视觉语言模型(LVLMs)在图像理解与对话任务中表现出色,但其广泛可用性引发未经授权使用和版权侵权问题,用户可通过微调公开模型构建自己的版本。本文提出参数学习攻击(PLA)方法,在不修改原模型的前提下追踪其版权。通过针对原模型的定向攻击生成对抗图像,使其输出特定内容;同时允许原模型在对抗攻击过程中反向更新参数,以学习这些触发图像。该方法可在模型发布后应用,不影响原性能。为模拟真实场景,我们采用多种策略和数据集对原模型进行微调,生成多样化的变体用于版权验证。大量实验表明,相比基线方法,本方法能更有效地识别微调模型的原始版权来源。该工作为追踪版权和检测非授权使用提供了有力工具。

原文摘要 · Abstract (English)

Large vision-language models (LVLMs) have demonstrated remarkable image understanding and dialogue capabilities, allowing them to handle a variety of visual question answering tasks. However, their widespread availability raises concerns about unauthorized usage and copyright infringement, where users or individuals can develop their own LVLMs by fine-tuning published models. In this paper, we propose a novel method called Parameter Learning Attack (PLA) for tracking the copyright of LVLMs without modifying the original model. Specifically, we construct adversarial images through targeted attacks against the original model, enabling it to generate specific outputs. To ensure these attacks remain effective on potential fine-tuned models to trigger copyright tracking, we allow the original model to learn the trigger images by updating parameters in the opposite direction during the adversarial attack process. Notably, the proposed method can be applied after the release of the original model, thus not affecting the model's performance and behavior. To simulate real-world applications, we fine-tune the original model using various strategies across diverse datasets, creating a range of models for copyright verification. Extensive experiments demonstrate that our method can more effectively identify the original copyright of fine-tuned models compared to baseline methods. Therefore, this work provides a powerful tool for tracking copyrights and detecting unlicensed usage of LVLMs.

版权追踪对抗攻击视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。