追踪23万条AI供应链,发现62%的链条存在无许可证文件,许可义务几乎全丢失。
Don't Trust the Label: License Laundering in AI Supply Chains

- 追踪数据集→模型→应用的23万条供应链路径,检测许可证传递情况。
- 62.3%的链条中出现无许可证文件的环节,义务型许可证存活率不足7%。
- 宽松型许可证存活率达95.1%,提示平台需加强许可证溯源与合规机制。
AI资源在多个平台间流动,涵盖Hugging Face上的数据集与模型、GitHub上的应用。尽管每个资源都应携带可追溯的许可证,但尚无研究验证这些许可义务是否在传播过程中得以保留。本文追踪了232,270条从数据集到模型再到应用的供应链路径,量化了两种许可证清洗现象:一是在下游为无明确许可证的资源添加标签;二是原有许可证类别被替换。研究发现,62.3%的路径经过至少一个未声明许可证的环节(集中在少数基础数据集);所有带有义务约束的许可证类别在端到端链路中的存活率均低于7%,而宽松类许可证存活率达95.1%。基于此,本文向从业者、模型发布者、权利方及平台方提出可操作建议。
原文摘要 · Abstract (English)
AI artifacts move through a multi-platform supply chain, spanning datasets and models on Hugging Face and applications on GitHub. While each artifact carries a license whose obligations should propagate through redistribution, no study has yet measured whether those obligations survive the chain or are stripped and replaced as artifacts move downstream. We trace 232,270 dataset$\rightarrow$model$\rightarrow$application chains and quantify two forms of license laundering: when artifacts with no declared license acquire definitive labels downstream, and when one declared license category replaces another during redistribution. We find that 62.3% of chains pass through at least one artifact with no declared license (concentrated in a small set of foundational datasets), and that every obligation-bearing license category falls below 7% end-to-end survival while the Permissive category reaches 95.1%. Based on these findings, we provide actionable recommendations for practitioners, model publishers, rights holders, and platform owners.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。