arXiv:2509.09873cs.SEcs.AI2025-09被引 3

首次系统审计开源AI生态许可证风险,发现超三分之一模型被改用宽松许可。

From Hugging Face to GitHub: Tracing License Drift in the Open-Source AI Ecosystem

  • 构建端到端审计框架,追踪数据集、模型与代码仓库的许可证流转。
  • 35.5%的模型在集成到项目时移除限制性条款,加剧合规风险。
  • 开发可扩展规则引擎,解决86.4%的软件许可证冲突,适合合规研究者使用。

开源AI生态中的隐性许可证冲突带来严重的法律与伦理风险,使组织面临诉讼威胁,用户承担未知风险。然而,当前领域缺乏对冲突发生频率、源头及影响范围的数据驱动理解。本文首次对Hugging Face上的数据集与模型及其下游在GitHub上的开源应用进行全链路审计,涵盖36.4万份数据集、160万模型和14万份项目。实证分析揭示系统性不合规现象:35.5%的模型-应用转换中,限制性条款被移除,转为宽松许可。同时,我们构建了一个可扩展的规则引擎,编码近200条SPDX及模型特定条款,可解决86.4%的软件许可证冲突。为支持未来研究,我们公开数据集与原型工具。本研究凸显开源AI治理中许可证合规的关键挑战,并提供实现规模化、智能化合规所需的数据与工具。

原文摘要 · Abstract (English)

Hidden license conflicts in the open-source AI ecosystem pose serious legal and ethical risks, exposing organizations to potential litigation and users to undisclosed risk. However, the field lacks a data-driven understanding of how frequently these conflicts occur, where they originate, and which communities are most affected. We present the first end-to-end audit of licenses for datasets and models on Hugging Face, as well as their downstream integration into open-source software applications, covering 364 thousand datasets, 1.6 million models, and 140 thousand GitHub projects. Our empirical analysis reveals systemic non-compliance in which 35.5% of model-to-application transitions eliminate restrictive license clauses by relicensing under permissive terms. In addition, we prototype an extensible rule engine that encodes almost 200 SPDX and model-specific clauses for detecting license conflicts, which can solve 86.4% of license conflicts in software applications. To support future research, we release our dataset and the prototype engine. Our study highlights license compliance as a critical governance challenge in open-source AI and provides both the data and tools necessary to enable automated, AI-aware compliance at scale.

许可证合规开源生态AI治理数据审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。