arXiv:2512.23385cs.SEcs.AI2025-12中稿 · the 48th IEEE/ACM …被引 2

分析31万条开发者讨论,揭示AI供应链真实安全问题与应对方案。

Securing the AI Supply Chain: What Can We Learn From Developer-Reported Security Issues and Solutions of AI Projects?

  • 通过关键词+微调distilBERT识别开源平台中的安全讨论
  • 发现32类安全问题,模型与数据类问题缺乏有效解决方案
  • 为开发者提供基于真实实践的AI安全防御参考

人工智能模型与应用的快速发展带来了日益复杂的安全挑战。开发者不仅面临传统软件供应链问题,还需应对AI特有的安全威胁。然而,当前对实际中常见安全问题及其解决方式仍知之甚少,制约了各环节安全措施的有效制定。本文基于Hugging Face和GitHub上的开发者讨论,开展实证研究,构建结合关键词匹配与最优微调distilBERT分类器的识别管道,在多种深度学习与大语言模型中表现最佳,最终生成包含312,868条安全讨论的数据集,揭示了AI项目的安全报告实践。从数据集中随机抽取753篇帖子进行主题分析,提炼出涵盖系统与软件、外部工具与生态、模型、数据四大主题的32类安全问题及24种解决方案。研究发现,多数问题源于AI组件的复杂依赖关系与黑箱特性,尤其模型与数据相关问题普遍缺乏具体应对措施。研究结果可为开发者与研究人员提供基于现实场景的安全防护依据。

原文摘要 · Abstract (English)

The rapid growth of Artificial Intelligence (AI) models and applications has led to an increasingly complex security landscape. Developers of AI projects must contend not only with traditional software supply chain issues but also with novel, AI-specific security threats. However, little is known about what security issues are commonly encountered and how they are resolved in practice. This gap hinders the development of effective security measures for each component of the AI supply chain. We bridge this gap by conducting an empirical investigation of developer-reported issues and solutions, based on discussions from Hugging Face and GitHub. To identify security-related discussions, we develop a pipeline that combines keyword matching with an optimal fine-tuned distilBERT classifier, which achieved the best performance in our extensive comparison of various deep learning and large language models. This pipeline produces a dataset of 312,868 security discussions, providing insights into the security reporting practices of AI applications and projects. We conduct a thematic analysis of 753 posts sampled from our dataset and uncover a fine-grained taxonomy of 32 security issues and 24 solutions across four themes: (1) System and Software, (2) External Tools and Ecosystem, (3) Model, and (4) Data. We reveal that many security issues arise from the complex dependencies and black-box nature of AI components. Notably, challenges related to Models and Data often lack concrete solutions. Our insights can offer evidence-based guidance for developers and researchers to address real-world security threats across the AI supply chain.

AI安全供应链开发者实践

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。