测试大模型能否识别并遵守用户输入中的版权信息。
Do LLMs Know to Respect Copyright Notice?
- 设计实验验证大模型处理带版权内容时的合规性。
- 发现多数模型未有效尊重版权信息,易生成侵权内容。
- 适合关注AI版权合规与安全对齐的研究者阅读。
已有研究表明大模型可能生成侵犯版权的内容。本文探讨一个更关键却未被充分研究的问题:当用户输入包含受版权保护材料时,大模型是否会识别并遵循版权信息?该问题至关重要,因为若答案为否定,意味着大模型可能成为版权侵权行为的主要推动者。我们通过一系列实验,使用多种语言模型、用户提示和受版权保护材料(包括书籍、新闻文章、API文档及电影剧本),评估了大模型在处理此类输入时可能引发的版权侵权程度。研究结果提供了一个保守但重要的评估,强调需要进一步探索并确保大模型在处理用户输入时尊重版权法规,以防止受保护内容的未经授权使用或复制。我们还发布了一个基准数据集,用于测试大模型的侵权行为,并呼吁未来加强模型对齐。
原文摘要 · Abstract (English)
Prior study shows that LLMs sometimes generate content that violates copyright. In this paper, we study another important yet underexplored problem, i.e., will LLMs respect copyright information in user input, and behave accordingly? The research problem is critical, as a negative answer would imply that LLMs will become the primary facilitator and accelerator of copyright infringement behavior. We conducted a series of experiments using a diverse set of language models, user prompts, and copyrighted materials, including books, news articles, API documentation, and movie scripts. Our study offers a conservative evaluation of the extent to which language models may infringe upon copyrights when processing user input containing protected material. This research emphasizes the need for further investigation and the importance of ensuring LLMs respect copyright regulations when handling user input to prevent unauthorized use or reproduction of protected content. We also release a benchmark dataset serving as a test bed for evaluating infringement behaviors by LLMs and stress the need for future alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。