用法律结构增强检索生成,帮创作者判断版权合理使用
Incorporating Legal Structure in Retrieval-Augmented Generation: A Case Study on Copyright Fair Use
- 按美国版权法四要素构建法律知识图谱与引用加权网络
- 检索结果更贴合法律要件,提升判例相关性
- 适合法律AI、内容平台合规工具开发者参考
本文提出一种面向美国版权合理使用原则的领域特定检索增强生成(RAG)系统。针对版权侵权下架频发且创作者缺乏法律支持的问题,设计结合语义搜索、法律知识图谱与法院判例引用网络的方法,提升检索质量与推理可靠性。模型在法定因素层面(如使用目的、作品性质、使用量、市场影响)对判例进行建模,并采用引用权重图结构优先筛选具有权威性的法律来源。通过链式思维推理与交错检索步骤,更贴近真实法律推理过程。初步测试表明,该方法显著提升检索结果的法律要件相关性,为未来大语言模型驱动的法律辅助工具评估与部署奠定基础。
原文摘要 · Abstract (English)
This paper presents a domain-specific implementation of Retrieval-Augmented Generation (RAG) tailored to the Fair Use Doctrine in U.S. copyright law. Motivated by the increasing prevalence of DMCA takedowns and the lack of accessible legal support for content creators, we propose a structured approach that combines semantic search with legal knowledge graphs and court citation networks to improve retrieval quality and reasoning reliability. Our prototype models legal precedents at the statutory factor level (e.g., purpose, nature, amount, market effect) and incorporates citation-weighted graph representations to prioritize doctrinally authoritative sources. We use Chain-of-Thought reasoning and interleaved retrieval steps to better emulate legal reasoning. Preliminary testing suggests this method improves doctrinal relevance in the retrieval process, laying groundwork for future evaluation and deployment of LLM-based legal assistance tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。