从训练结构切入,用因果分析评估大模型版权风险
Interrogating LLM design under a fair learning doctrine
- 通过因果与相关性分析,检验训练决策对模型记忆的影响
- 发现特定训练数据选择会显著提升模型记忆程度
- 为法官提供可操作的版权判罚标准,适合法律与AI交叉研究者
当前关于大语言模型(LLMs)与版权的讨论多聚焦于模型输出是否与训练数据高度相似,但这种“行为”视角难以算法化定义且不足以覆盖全部版权风险。本文提出互补的“结构”视角,关注模型训练过程本身。我们通过测量训练决策是否显著影响模型记忆,将“公平学习”概念具体化。以开源模型Pythia为例,采用因果与相关性分析,对训练过程进行拆解,得出可验证的事实结论。通过将记忆分析与法律标准关联,揭示法官如何通过判决推动版权法目标实现。最后探讨公平学习标准如何通过更规则化和引入外部技术指南来增强清晰度。
原文摘要 · Abstract (English)
The current discourse on large language models (LLMs) and copyright largely takes a "behavioral" perspective, focusing on model outputs and evaluating whether they are substantially similar to training data. However, substantial similarity is difficult to define algorithmically and a narrow focus on model outputs is insufficient to address all copyright risks. In this interdisciplinary work, we take a complementary "structural" perspective and shift our focus to how LLMs are trained. We operationalize a notion of "fair learning" by measuring whether any training decision substantially affected the model's memorization. As a case study, we deconstruct Pythia, an open-source LLM, and demonstrate the use of causal and correlational analyses to make factual determinations about Pythia's training decisions. By proposing a legal standard for fair learning and connecting memorization analyses to this standard, we identify how judges may advance the goals of copyright law through adjudication. Finally, we discuss how a fair learning standard might evolve to enhance its clarity by becoming more rule-like and incorporating external technical guidelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。