arXiv:2607.16414cs.CRcs.AI2026-07

提出信号分级风险框架,帮机构看清模型被攻击的暴露程度。

Signal-based Model Access Risk Analysis for AI System Operations Security

论文配图:Signal-based Model Access Risk Analysis for AI System Operations Security
图 1 · 摘自论文原文
  • 按系统返回信号丰富度分类攻击者权限,突破传统黑白盒划分。
  • 不同信号暴露层级导致攻击能力差异显著,影响安全策略设计。
  • 适合关注模型部署安全的工程师、采购决策者阅读。

人工智能系统已广泛应用于安全、金融、医疗、消费科技和大规模云服务等领域,每日处理海量数据并做出重要决策。这种普及带来了广泛的攻击面,攻击者可通过操纵、逃避、信息提取等方式破坏已部署模型。根据系统设计和暴露程度,攻击者可能仅获最终决策结果,或可获取置信度分数、中间表示,甚至完整模型参数。以往研究多基于攻击者对模型内部知识(架构、参数、梯度)的了解程度,将逃避攻击分为白盒、灰盒和黑盒,但该分类常混淆不同部署场景下输出信号的本质差异,导致同为“黑盒”却具备根本不同的攻击路径。理解逃避攻击如何随部署系统返回的信息信号变化而调整,对组织的采购与部署决策至关重要。为此,本文提出信号基础模型访问风险分类框架(SMART),一种以部署为导向的分类体系,依据已部署AI系统输出信息信号的性质与丰富度对攻击者权限进行划分。利用该框架,我们系统梳理了从低到高信息暴露水平下的逃避攻击策略,揭示部署接口如何影响攻击能力,为更安全的AI部署与采购决策提供依据。

原文摘要 · Abstract (English)

Artificial intelligence (AI) systems are now ubiquitous across domains such as security, finance, healthcare, consumer technology, and large-scale cloud services, where they process massive volumes of data and make consequential decisions daily. This widespread adoption has created a broad attack surface through which adversaries can manipulate, evade, extract information from, or otherwise subvert deployed models. Depending on system design and exposure, attackers may have very different forms of access: some observe only final decisions, while others receive confidence scores, intermediate representations, or even full model parameters. While previous surveys typically organize evasion attacks into white-box, gray-box, and black-box categories based on the attacker's knowledge of model internals (architecture, parameters, gradients), this taxonomy often conflates different deployment scenarios that provide vastly different output signals, all labeled as ``black-box'' despite enabling fundamentally different attack strategies. Understanding how evasion attack strategies adapt to the specific information signals returned by deployed systems is critical for organizations making procurement and deployment decisions. To address this gap, we introduce the Signal-based Model Access Risk Taxonomy (SMART), a deployment-oriented framework that classifies attacker access according to the nature and richness of the information signals available from deployed AI systems. Using this taxonomy, we provide a structured overview of evasion attacks across progressively richer levels of information exposure, highlighting how deployment interfaces influence attack capabilities and informing more secure AI deployment and procurement decisions.

模型安全攻击分析部署风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。