推动LLM训练数据开源,解决法律与技术难题
Towards Best Practices for Open Datasets for LLM Training
- 提出跨领域协作框架,系统解决开放数据集构建的法律与技术障碍
- 指出当前无大规模公开许可语料库可用,因元数据不全、数字化成本高
- 适合关注AI伦理、政策制定及数据治理的研究者与从业者
许多AI公司未经版权方许可训练大语言模型(LLMs),其合法性因国家而异:欧盟和日本在特定条件下允许,而美国则存在法律模糊性。尽管如此,创作者担忧引发多起重大版权诉讼,导致企业和公共利益相关方普遍减少对训练数据集信息的披露。这种信息限制损害了透明度、问责制与创新,使研究者、审计人员及受影响个体无法获取理解模型所需信息。虽可转向使用开放获取或公有领域数据训练模型,但截至本文撰写时,尚无在有意义规模下训练的此类模型,主要因汇编语料库面临显著的技术与社会挑战,包括元数据不完整不可靠、实体资料数字化成本高、以及需跨领域法律与技术能力保障数据的相关性与责任性。实现未来基于负责任地策展与治理的开放数据训练AI系统,需法律、技术与政策领域的协作,以及对元数据标准、数字化投入与开放文化培育的投资。
原文摘要 · Abstract (English)
Many AI companies are training their large language models (LLMs) on data without the permission of the copyright owners. The permissibility of doing so varies by jurisdiction: in countries like the EU and Japan, this is allowed under certain restrictions, while in the United States, the legal landscape is more ambiguous. Regardless of the legal status, concerns from creative producers have led to several high-profile copyright lawsuits, and the threat of litigation is commonly cited as a reason for the recent trend towards minimizing the information shared about training datasets by both corporate and public interest actors. This trend in limiting data information causes harm by hindering transparency, accountability, and innovation in the broader ecosystem by denying researchers, auditors, and impacted individuals access to the information needed to understand AI models. While this could be mitigated by training language models on open access and public domain data, at the time of writing, there are no such models (trained at a meaningful scale) due to the substantial technical and sociological challenges in assembling the necessary corpus. These challenges include incomplete and unreliable metadata, the cost and complexity of digitizing physical records, and the diverse set of legal and technical skills required to ensure relevance and responsibility in a quickly changing landscape. Building towards a future where AI systems can be trained on openly licensed data that is responsibly curated and governed requires collaboration across legal, technical, and policy domains, along with investments in metadata standards, digitization, and fostering a culture of openness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。