arXiv:2412.09460cs.CL2024-12中稿 · NoDaLiDa/Baltic-HL…被引 8

研究挪威语大模型训练中版权材料的影响,发现非虚构类内容提升性能。

The Impact of Copyrighted Material on Large Language Models: A Norwegian Perspective

  • 构建评估框架,测试出版商控制的版权语料对挪威语大模型的影响。
  • 加入书籍和报纸数据可提升模型表现,但小说类反而降低效果。
  • 结果可为作者贡献补偿机制提供实证依据,适合关注AI伦理者阅读。

在训练语言模型时使用受版权保护的材料引发了重要的法律与伦理问题。本文提出一个评估框架,并报告了针对挪威语生成式大语言模型(LLMs)的实证研究结果,考察出版社控制的版权语料对其性能的影响。在多样化的任务上进行评估后发现,将书籍和报纸加入模型训练数据混合中通常能提升性能,而添加小说类作品则可能产生负面影响。该研究结果可为作者在人工智能发展中的贡献制定合理补偿机制提供依据。

原文摘要 · Abstract (English)

The use of copyrighted materials in training language models raises critical legal and ethical questions. This paper presents a framework for and the results of empirically assessing the impact of publisher-controlled copyrighted corpora on the performance of generative large language models (LLMs) for Norwegian. When evaluated on a diverse set of tasks, we found that adding both books and newspapers to the data mixture of LLMs tend to improve their performance, while the addition of fiction works seems to be detrimental. Our experiments could inform the creation of a compensation scheme for authors whose works contribute to AI development.

大模型版权问题挪威语AI伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。