测试大模型对版权内容的识别与尊重能力,发现主流模型仍严重不足。
Bridging the Copyright Gap: Do Large Vision-Language Models Recognize and Respect Copyrighted Content?
- 构建5万组多模态数据集,评估模型在有无版权标识下的表现
- 即使有明确版权标识,顶级闭源模型仍频繁违规生成内容
- 提出新防御框架,在各种场景下显著降低侵权风险
大型视觉语言模型(LVLMs)在多模态推理任务中取得显著进展,但其广泛可用性引发版权侵权的严重担忧。当模型处理用户输入或检索到的内容时,是否能准确识别并遵守版权规定?若未能遵守,可能带来法律与伦理后果,尤其是在基于版权材料(如书籍摘录、新闻报道)生成回应时。本文系统评估多种LVLM对版权内容的处理能力,涵盖书籍摘录、新闻文章、音乐歌词和代码文档等视觉输入。为此,我们构建了一个包含50,000个多模态查询-内容对的大规模基准数据集,用于衡量模型在潜在侵权情境下的合规性。数据集覆盖两种场景:有版权标识和无版权标识;对于有标识情况,还涵盖四种常见版权标注类型。评估结果显示,即使是当前最先进的闭源模型,在存在版权标识的情况下仍表现出显著缺陷。为解决此问题,我们提出一种新型工具增强型防御框架,有效降低所有场景下的侵权风险。研究强调了开发具备版权意识的LVLM的重要性,以确保版权内容的负责任与合法使用。
原文摘要 · Abstract (English)
Large vision-language models (LVLMs) have achieved remarkable advancements in multimodal reasoning tasks. However, their widespread accessibility raises critical concerns about potential copyright infringement. Will LVLMs accurately recognize and comply with copyright regulations when encountering copyrighted content (i.e., user input, retrieved documents) in the context? Failure to comply with copyright regulations may lead to serious legal and ethical consequences, particularly when LVLMs generate responses based on copyrighted materials (e.g., retrieved book experts, news reports). In this paper, we present a comprehensive evaluation of various LVLMs, examining how they handle copyrighted content -- such as book excerpts, news articles, music lyrics, and code documentation when they are presented as visual inputs. To systematically measure copyright compliance, we introduce a large-scale benchmark dataset comprising 50,000 multimodal query-content pairs designed to evaluate how effectively LVLMs handle queries that could lead to copyright infringement. Given that real-world copyrighted content may or may not include a copyright notice, the dataset includes query-content pairs in two distinct scenarios: with and without a copyright notice. For the former, we extensively cover four types of copyright notices to account for different cases. Our evaluation reveals that even state-of-the-art closed-source LVLMs exhibit significant deficiencies in recognizing and respecting the copyrighted content, even when presented with the copyright notice. To solve this limitation, we introduce a novel tool-augmented defense framework for copyright compliance, which reduces infringement risks in all scenarios. Our findings underscore the importance of developing copyright-aware LVLMs to ensure the responsible and lawful use of copyrighted content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。