测试大模型能否识别生成图像的版权侵权,发现容易误判。
Can Large Vision-Language Models Detect Images Copyright Infringement from GenAI?
- 构建包含侵权与模糊非侵权样本的基准数据集
- 主流大模型存在过拟合,误将非侵权图判为侵权
- 适合关注AI版权检测与模型可靠性的研究者
生成式AI模型虽能合成高质量内容,但引发对非法生成受版权保护素材的担忧。现有研究虽提出多种解决方案,但大型视觉语言模型(LVLMs)在版权侵权检测方面的能力仍缺乏系统评估。本文针对最新LVLMs,使用多样图像样本开展评测。鉴于缺乏涵盖侵权与模糊非侵权样本的全面数据集,我们通过高级提示工程构建了一个基准数据集:包含知名知识产权人物的侵权正样本,以及外观相似但不构成侵权的负样本。随后对领先LVLMs进行评估。实验结果表明,LVLMs易出现过拟合,导致部分负样本被错误分类为侵权。最后分析失败案例,并提出缓解过拟合问题的潜在方案。
原文摘要 · Abstract (English)
Generative AI models, renowned for their ability to synthesize high-quality content, have sparked growing concerns over the improper generation of copyright-protected material. While recent studies have proposed various approaches to address copyright issues, the capability of large vision-language models (LVLMs) to detect copyright infringements remains largely unexplored. In this work, we focus on evaluating the copyright detection abilities of state-of-the-art LVLMs using a various set of image samples. Recognizing the absence of a comprehensive dataset that includes both IP-infringement samples and ambiguous non-infringement negative samples, we construct a benchmark dataset comprising positive samples that violate the copyright protection of well-known IP figures, as well as negative samples that resemble these figures but do not raise copyright concerns. This dataset is created using advanced prompt engineering techniques. We then evaluate leading LVLMs using our benchmark dataset. Our experimental results reveal that LVLMs are prone to overfitting, leading to the misclassification of some negative samples as IP-infringement cases. In the final section, we analyze these failure cases and propose potential solutions to mitigate the overfitting problem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。