从攻击行为反推攻击者特征,助力系统防御与优化。
Identifying Adversary Characteristics from an Observed Attack
- 提出无领域依赖框架,从攻击现象推断最可能的攻击者
- 证明无额外信息时攻击者无法唯一识别,需引入概率推理
- 适用于多种学习器,提升防御策略的针对性与效果
在自动化决策系统中,机器学习模型易受数据操纵攻击。现有防御机制或作用于模型本身(如对抗正则化),或作用于系统整体(如异常检测)。本文另辟蹊径,不关注攻击行为本身,而是聚焦于攻击者,提出一个从可观测攻击中识别攻击者特征的框架。我们证明,在缺乏额外知识的前提下,攻击者不可识别(多个潜在攻击者可能产生相同攻击行为)。为应对这一挑战,提出一种领域无关的框架,用于识别最可能的攻击者。该框架可帮助防御方实现双重目标:其一,通过外部手段(如调整决策系统或限制攻击者能力)进行外源性缓解;其二,在实施直接影响学习过程的防御方法(如对抗正则化)时,利用具体攻击者信息提升防御性能。文中详述框架设计,并通过多种学习器实例验证其适用性。
原文摘要 · Abstract (English)
When used in automated decision-making systems, machine learning (ML) models are vulnerable to data-manipulation attacks. Some defense mechanisms (e.g., adversarial regularization) directly affect the ML models while others (e.g., anomaly detection) act within the broader system. In this paper we consider a different task for defending the adversary, focusing on the attacker, rather than the attack. We present and demonstrate a framework for identifying characteristics about the attacker from an observed attack. We prove that, without additional knowledge, the attacker is non-identifiable (multiple potential attackers would perform the same observed attack). To address this challenge, we propose a domain-agnostic framework to identify the most probable attacker. This framework aids the defender in two ways. First, knowledge about the attacker can be leveraged for exogenous mitigation (i.e., addressing the vulnerability by altering the decision-making system outside the learning algorithm and/or limiting the attacker's capability). Second, when implementing defense methods that directly affect the learning process (e.g., adversarial regularization), knowledge of the specific attacker improves performance. We present the details of our framework and illustrate its applicability through specific instantiations on a variety of learners.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。