arXiv:2604.18649cs.CRcs.AI2026-04

生成式AI训练中非法数据使用无法通过事后清理免责

Position: No Retroactive Cure for Infringement during Training

  • 非法数据摄入即构成侵权,模型权重等同于固定拷贝
  • 合同与反不正当竞争规则可独立限制使用,绕过版权抗辩
  • 侵权收益可能被追回,甚至需处置模型本身

随着生成式AI面临日益严峻的法律挑战,机器学习界越来越依赖事后的缓解措施——尤其是机器遗忘和推理时的防护机制——来论证合规性。本文认为,此类事后缓解方法无法追溯性地消除因非法获取和训练数据而产生的法律责任,因为合规的核心在于数据来源链条,而非输出结果。首先,未经授权的复制或数据摄入可构成完整的法律行为,模型权重可能作为保留训练衍生表达价值的固定副本,使后期过滤失去意义。其次,合同条款、服务协议以及反搭便车原则等民法规则可独立限制访问与使用,常能绕过著作权抗辩(如合理使用或数据挖掘例外)。第三,由于受保护输入的价值可能持续存在于模型权重中,不当得利和没收收益等救济手段可能要求收回全部收益,甚至触及模型本身。因此,我们主张从‘事后净化’转向可验证的‘事前流程合规’。

原文摘要 · Abstract (English)

As generative AI faces intensifying legal challenges, the machine learning community has increasingly relied on post-hoc mitigation -- especially machine unlearning and inference-time guardrails -- to argue for compliance. This paper argues that such post-hoc mitigation methods cannot retroactively cure liability from unlawful acquisition and training, because compliance hinges on data lineage, not the outputs. Our argument has three parts. First, unauthorized copying/ingestion can be a legally complete completed act, and model weights may operate as fixed copies that retain training-derived expressive value, making later filtering beside the point for infringement. Second, contract and tort/unfair-competition rules -- via licenses, terms of service, and anti-free-riding principles -- can independently restrict access and use, often bypassing copyright defenses (e.g., fair use or TDM exceptions). Third, since value from protected inputs can persist in weights, remedies such as unjust enrichment and disgorgement may require stripping gains and, in some cases, reaching the model itself. We therefore argue for a shift from Post-Hoc Sanitization to verifiable Ex-Ante Process Compliance.

AI合规数据侵权模型责任

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。