研究物理打印对抗样本如何篡改选举结果,揭示数字与现实攻击效果差异。
Analyzing Physical Adversarial Example Threats to Machine Learning in Election Systems
- 构建概率模型计算需打印多少对抗选票可翻转选举
- 实验验证物理场景下l1和l2攻击比数字场景更有效
- 适用于关注选举安全的机器学习研究者与政策制定者
机器学习在投票系统中的应用展现出良好性能(选票分类准确率超99%),但也面临对抗样本攻击风险。本文分析攻击者通过打印物理对抗选票干扰美国选举的可能性。首先,建立概率框架,基于候选人胜率与总票数,推导出翻转选举所需打印的对抗选票数量。其次,评估六种对抗攻击类型在物理打印场景下的有效性:l_infinity-APGD、l2-APGD、l1-APGD、l0 PGD、l0 + l_infinity PGD、l0 + sigma-map PGD。通过打印并扫描14.4万张对抗选票,测试四种不同机器学习模型。实验证明,数字域最有效的攻击(l2、l_infinity)在物理域表现不佳,而物理域中以l1和l2攻击为主,具体效果因模型而异。该研究结合概率选举模型与数字/物理攻击评估,突破以往仅关注接近选举的局限,首次明确量化对抗选票操纵改变选举结果的条件与可能性。
原文摘要 · Abstract (English)
Developments in the machine learning voting domain have shown both promising results and risks. Trained models perform well on ballot classification tasks (> 99% accuracy) but are at risk from adversarial example attacks that cause misclassifications. In this paper, we analyze an attacker who seeks to deploy adversarial examples against machine learning ballot classifiers to compromise a U.S. election. We first derive a probabilistic framework for determining the number of adversarial example ballots that must be printed to flip an election, in terms of the probability of each candidate winning and the total number of ballots cast. Second, it is an open question as to which type of adversarial example is most effective when physically printed in the voting domain. We analyze six different types of adversarial example attacks: l_infinity-APGD, l2-APGD, l1-APGD, l0 PGD, l0 + l_infinity PGD, and l0 + sigma-map PGD. Our experiments include physical realizations of 144,000 adversarial examples through printing and scanning with four different machine learning models. We empirically demonstrate an analysis gap between the physical and digital domains, wherein attacks most effective in the digital domain (l2 and l_infinity) differ from those most effective in the physical domain (l1 and l2, depending on the model). By unifying a probabilistic election framework with digital and physical adversarial example evaluations, we move beyond prior close race analyses to explicitly quantify when and how adversarial ballot manipulation could alter outcomes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。