arXiv:2503.04480stat.MLcs.LG2025-03被引 3

攻击者通过删改数据,操控贝叶斯模型的后验分布,实现精准误导。

Poisoning Bayesian Inference via Data Deletion and Replication

  • 利用数据删减与复制,操纵贝叶斯后验分布
  • 仅需少量操作即可显著改变模型信念
  • 可精准污染特定推断,不影响其他结果

对抗机器学习研究已揭示统计模型对恶意数据篡改的脆弱性。然而,尽管贝叶斯机器学习取得进展,多数对抗研究仍集中于传统方法。本文将白盒模型投毒范式扩展至通用贝叶斯推断,揭示其在对抗环境中的脆弱性。提出一套攻击方法,通过战略性删除和复制真实观测数据,即使仅有后验采样访问权,也能引导贝叶斯后验趋向目标分布。理论上证明了算法性质,并在合成与真实场景中验证其性能。攻击者以较低成本显著改变模型信念,且通过承担更高风险,可完全按意愿塑造信念。通过精心构建对抗后验,实现精准投毒,仅影响目标推断,其余推断几乎不受干扰。

原文摘要 · Abstract (English)

Research in adversarial machine learning (AML) has shown that statistical models are vulnerable to maliciously altered data. However, despite advances in Bayesian machine learning models, most AML research remains concentrated on classical techniques. Therefore, we focus on extending the white-box model poisoning paradigm to attack generic Bayesian inference, highlighting its vulnerability in adversarial contexts. A suite of attacks are developed that allow an attacker to steer the Bayesian posterior toward a target distribution through the strategic deletion and replication of true observations, even when only sampling access to the posterior is available. Analytic properties of these algorithms are proven and their performance is empirically examined in both synthetic and real-world scenarios. With relatively little effort, the attacker is able to substantively alter the Bayesian's beliefs and, by accepting more risk, they can mold these beliefs to their will. By carefully constructing the adversarial posterior, surgical poisoning is achieved such that only targeted inferences are corrupted and others are minimally disturbed.

贝叶斯攻击数据投毒后验操控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。