arXiv:2607.17242cs.SEcs.AI2026-07

检测了9.75万模型的透明度,发现文档大多不完整。

A Large-Scale Measurement of AI Bill of Materials Completeness in Hugging Face Models

论文配图:A Large-Scale Measurement of AI Bill of Materials Completeness in Hugging Face Models
图 1 · 摘自论文原文
  • 基于Hugging Face模型库,量化评估AI物料清单完整性
  • 结构字段全有,但模型卡、限制、风险等关键信息缺失
  • 适合关注AI治理与可追溯性的开发者和研究者

预训练机器学习模型使开发者无需从头训练即可构建复杂系统,但模型仓库常缺少关于模型来源、许可、数据集、局限性及外部引用的机器可读文档,造成人工智能供应链中的透明度与治理缺口。人工智能物料清单(AIBOM)通过记录模型、元数据、许可证、数据集、模型卡信息及外部引用等,填补这一空白。本文以公开的Hugging Face模型库为案例,实证研究了AIBOM的完整性,即仓库提供可机器解析的AI供应链文档的程度。我们分析了约97,500个AIBOM相关实体,评估生成的AIBOM在:(i) 是否包含必需的结构与元数据字段;(ii) 是否准确表示模型身份、许可与外部引用;(iii) 是否涵盖模型卡内容如数据集、局限性、安全风险评估与环境信息;(iv) 在任务类型、许可可用性、数据集声明、模型族与论文引用等方面的文档覆盖差异。结果表明,生成的AIBOM在必要结构上覆盖完整,但在模型卡、元数据、负责任使用、环境、局限性和有意义描述等关键字段上仍表现薄弱或缺失。研究呼吁改进模型卡实践、加强仓库级可追溯性,并推动自动化AIBOM验证,以促进更完整的AIBOM生成与采纳。

原文摘要 · Abstract (English)

Pretrained machine learning (ML) models help developers build ML-intensive software systems without training models from scratch. However, model repositories often provide incomplete machine-readable documentation about model provenance, licenses, datasets, limitations, and external references, creating transparency and governance gaps across the AI supply chain. Artificial Intelligence Bills of Materials (AIBOMs) address these gaps by documenting AI artifacts, including models, metadata, licenses, datasets, model-card information, and external references. Taking public Hugging Face (HF) model repositories as a case study, this paper empirically investigates AIBOM completeness, defined as the extent to which repositories provide AIBOM-relevant information for machine-readable AI supply-chain documentation. We examine approximately 97.5K AIBOM artifacts to assess the extent to which generated AIBOMs: (i) contain required structural and metadata fields, (ii) represent model identity, license, and external-reference information, (iii) capture model-card documentation such as datasets, limitations, safety-risk assessment, and environmental information, and (iv) vary in documentation coverage across repository and artifact characteristics such as task, license availability, dataset declaration, model family, and paper reference. Results indicate that generated AIBOMs provide complete coverage of required AIBOM structure but limited AI-specific documentation completeness. Required fields are fully represented, but model-card, metadata, responsible-use, environmental, limitation, and meaningful-description fields remain weakly represented or missing across generated artifacts. Our findings motivate improved model-card practices, repository-level traceability, and automated AIBOM validation to advance the generation and adoption of more complete AIBOMs.

AI治理模型透明度Hugging Face物料清单

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。