arXiv:2602.08816cs.LGcs.AI2026-02KDD被引 4

96%的开源AI数据集和模型缺失许可证文本,存在法律风险。

Permissive-Washing in the Open AI Supply Chain: A Large-Scale Audit of License Integrity

  • 大规模审计12万条AI供应链,验证许可证合规性
  • 仅2.3%的数据集和3.2%的模型满足完整许可要求
  • 上游版权信息极少传递至下游,法律效力难以保障

MIT、Apache-2.0等宽松许可证在开源AI中占主导,表明模型、数据集和代码可自由使用、修改和分发。但这些许可证有强制要求:必须包含完整许可证文本、提供版权声明并保留上游署名,而这些在实践中未被大规模验证。未能满足条件会使再利用脱离许可证保护范围,导致默认版权约束,使下游使用者面临诉讼风险。我们称此现象为「宽松许可洗白」:将AI产物标记为免费使用,却省略使该标签有效的法律文件。为评估其在AI供应链中的普遍性,我们对跨Hugging Face与GitHub的124,278条数据集→模型→应用链路进行实证审计,涵盖3,338个数据集、6,664个模型和28,516个应用。结果显示:96.5%的数据集和95.8%的模型缺少必要的许可证文本;仅2.3%的数据集和3.2%的模型同时满足许可证文本与版权声明要求;即使上游提供完整授权证据,署名也极少向下传递:仅27.59%的模型保留合规的数据集声明,仅5.75%的应用保留合规的模型声明(仅有6.38%保留任何上游关联声明)。实践者不可假设宽松标签即代表合法权利——许可证文件与声明,而非元数据,才是法律效力的依据。为支持后续研究,我们公开完整审计数据集与可复现的分析流程。

原文摘要 · Abstract (English)

Permissive licenses like MIT, Apache-2.0, and BSD-3-Clause dominate open-source AI, signaling that artifacts like models, datasets, and code can be freely used, modified, and redistributed. However, these licenses carry mandatory requirements: include the full license text, provide a copyright notice, and preserve upstream attribution, that remain unverified at scale. Failure to meet these conditions can place reuse outside the scope of the license, effectively leaving AI artifacts under default copyright for those uses and exposing downstream users to litigation. We call this phenomenon ``permissive washing'': labeling AI artifacts as free to use, while omitting the legal documentation required to make that label actionable. To assess how widespread permissive washing is in the AI supply chain, we empirically audit 124,278 dataset $\rightarrow$ model $\rightarrow$ application supply chains, spanning 3,338 datasets, 6,664 models, and 28,516 applications across Hugging Face and GitHub. We find that an astonishing 96.5\% of datasets and 95.8\% of models lack the required license text, only 2.3\% of datasets and 3.2\% of models satisfy both license text and copyright requirements, and even when upstream artifacts provide complete licensing evidence, attribution rarely propagates downstream: only 27.59\% of models preserve compliant dataset notices and only 5.75\% of applications preserve compliant model notices (with just 6.38\% preserving any linked upstream notice). Practitioners cannot assume permissive labels confer the rights they claim: license files and notices, not metadata, are the source of legal truth. To support future research, we release our full audit dataset and reproducible pipeline.

开源合规许可证审计AI供应链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。