arXiv:2503.05713cs.CYcs.CL2025-03被引 2

发现大模型对多语言版权内容保护存在偏见

Beyond English: Unveiling Multilingual Bias in LLM Copyright Compliance

  • 构建四语歌曲歌词数据集,测试七款大模型跨语言版权响应
  • 不同语言的版权内容被泄露风险差异显著,提示语言影响内容生成
  • 揭示现有模型在多语言版权保护上的不均衡性,适合关注AI伦理的研究者

大型语言模型在版权保护方面引发广泛关注。以往研究多聚焦英语内容,忽视多语言维度。本文通过构建英文、法文、中文和韩文流行歌曲歌词数据集,系统测试七款大模型在四种语言提示下的版权内容生成行为。结果表明,大模型在处理不同语言的版权内容时存在显著不平衡,既体现在被保护内容的语言类型上,也体现在提示语言的影响上。这说明当前模型在跨语言版权合规方面存在系统性偏差,亟需发展更鲁棒、语言无关的版权保护机制,以实现全球范围内公平一致的版权保护。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have raised significant concerns regarding the fair use of copyright-protected content. While prior studies have examined the extent to which LLMs reproduce copyrighted materials, they have predominantly focused on English, neglecting multilingual dimensions of copyright protection. In this work, we investigate multilingual biases in LLM copyright protection by addressing two key questions: (1) Do LLMs exhibit bias in protecting copyrighted works across languages? (2) Is it easier to elicit copyrighted content using prompts in specific languages? To explore these questions, we construct a dataset of popular song lyrics in English, French, Chinese, and Korean and systematically probe seven LLMs using prompts in these languages. Our findings reveal significant imbalances in LLMs' handling of copyrighted content, both in terms of the language of the copyrighted material and the language of the prompt. These results highlight the need for further research and development of more robust, language-agnostic copyright protection mechanisms to ensure fair and consistent protection across languages.

大模型版权保护多语言偏见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。