arXiv:2502.19413cs.LGcs.AI2025-02被引 2

用大模型将论文转为无风格的知识单元,突破版权限制

Project Alexandria: Towards Freeing Scientific Knowledge from Copyright Burdens via LLMs

  • 用大模型提取实体、属性和关系,生成不依赖文风的结构化知识单元
  • 在四个领域中保留约95%的事实信息,通过多选题测试验证
  • 结合德国版权法与美国合理使用原则,提供法律可行性依据

付费墙、许可证和版权规则常阻碍科学知识的广泛传播与重用。我们认为,从学术文本中提取科学知识在法律和技术上均是可行的。现有方法如文本嵌入难以可靠保留事实内容,而简单改写可能面临法律风险。我们提出一种新思路:利用大模型将学术文档转化为名为知识单元(Knowledge Units)的、保持事实但无关文风的表示形式。这些单元以结构化数据捕获实体、属性与关系,不包含风格内容。我们提供证据表明,知识单元(1)基于德国版权法与美国合理使用原则,具备法律可辩护性;(2)在四个研究领域中,通过原稿事实的多选题测试,保留了约95%的事实知识。释放受版权保护的科学知识,将为科研与教育带来变革性影响,使语言模型得以重用重要事实。为此,我们开源了将研究文档转换为知识单元的工具。总体而言,本工作论证了在尊重版权的前提下,实现科学知识民主化的可行性。

原文摘要 · Abstract (English)

Paywalls, licenses and copyright rules often restrict the broad dissemination and reuse of scientific knowledge. We take the position that it is both legally and technically feasible to extract the scientific knowledge in scholarly texts. Current methods, like text embeddings, fail to reliably preserve factual content, and simple paraphrasing may not be legally sound. We propose a new idea for the community to adopt: convert scholarly documents into knowledge preserving, but style agnostic representations we term Knowledge Units using LLMs. These units use structured data capturing entities, attributes and relationships without stylistic content. We provide evidence that Knowledge Units (1) form a legally defensible framework for sharing knowledge from copyrighted research texts, based on legal analyses of German copyright law and U.S. Fair Use doctrine, and (2) preserve most (~95\%) factual knowledge from original text, measured by MCQ performance on facts from the original copyrighted text across four research domains. Freeing scientific knowledge from copyright promises transformative benefits for scientific research and education by allowing language models to reuse important facts from copyrighted text. To support this, we share open-source tools for converting research documents into Knowledge Units. Overall, our work posits the feasibility of democratizing access to scientific knowledge while respecting copyright.

知识提取版权合规LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。