构建可验证的AI安全审计框架,让模型风险可控可查。
AI Bill of Materials and Beyond: Systematizing Security Assurance through the AI Risk Scanning (AIRS) Framework
- 基于威胁建模自动生成机器可读的安全证据
- 实测验证量化模型加载安全与后门检测能力
- 扩展传统软件清单,专为大模型安全设计
当前人工智能系统安全保障分散在软件供应链、对抗性机器学习和治理文档中。现有透明度机制如模型卡、数据表和软件物料清单(SBOM)虽能追踪来源,但缺乏可验证的机器可读安全证据。本文提出AI风险扫描(AIRS)框架,基于威胁建模生成证据,通过三轮试点研究(Smurf、OPAL、Pilot C)将AI文档从描述性披露转向可测量、有证据支撑的验证。该框架对齐MITRE ATLAS对抗性机器学习分类体系,自动生成结构化产物,涵盖模型完整性、打包与序列化安全、结构适配器及运行时行为。目前聚焦大语言模型(LLM)级保证,未来可扩展至多模态与系统层威胁(如应用层滥用、工具调用)。在量化版GPT-OSS-20B模型上的概念验证显示,可执行安全加载策略、分片哈希校验,并在受控运行环境中完成污染与后门探测。与SPDX 3.0和CycloneDX 1.6标准对比发现,两者在身份与评估元数据上一致,但在表示AI特有保障字段方面存在关键缺口。AIRS框架因此将SBOM实践延伸至AI领域,结合威胁建模与自动化可审计证据生成,为标准化、可信且机器可验证的AI风险文档提供原则性基础。
原文摘要 · Abstract (English)
Assurance for artificial intelligence (AI) systems remains fragmented across software supply-chain security, adversarial machine learning, and governance documentation. Existing transparency mechanisms - including Model Cards, Datasheets, and Software Bills of Materials (SBOMs) - advance provenance reporting but rarely provide verifiable, machine-readable evidence of model security. This paper introduces the AI Risk Scanning (AIRS) Framework, a threat-model-based, evidence-generating framework designed to operationalize AI assurance. The AIRS Framework evolved through three progressive pilot studies - Smurf (AIBOM schema design), OPAL (operational validation), and Pilot C (AIRS) - that reframed AI documentation from descriptive disclosure toward measurable, evidence-bound verification. The framework aligns its assurance fields to the MITRE ATLAS adversarial ML taxonomy and automatically produces structured artifacts capturing model integrity, packaging and serialization safety, structural adapters, and runtime behaviors. Currently, the AIRS Framework is scoped to provide model-level assurances for LLMs, but it could be expanded to include other modalities and cover system-level threats (e.g. application-layer abuses, tool-calling). A proof-of-concept on a quantized GPT-OSS-20B model demonstrates enforcement of safe loader policies, per-shard hash verification, and contamination and backdoor probes executed under controlled runtime conditions. Comparative analysis with SBOM standards of SPDX 3.0 and CycloneDX 1.6 reveals alignment on identity and evaluation metadata, but identifies critical gaps in representing AI-specific assurance fields. The AIRS Framework thus extends SBOM practice to the AI domain by coupling threat modeling with automated, auditable evidence generation, providing a principled foundation for standardized, trustworthy, and machine-verifiable AI risk documentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。