黑盒攻击神经上下文老虎机,通过模拟对手策略实现高效干扰。
Learning to Attack: A Bandit Approach to Adversarial Context Poisoning
- 将攻击建模为连续动作的贝叶斯优化问题,自适应学习并扰动目标策略。
- 在三个真实数据集上,使受害者累积损失超过现有方法30%以上。
- 无需模型内部信息,适合研究对抗安全或系统漏洞的人员参考。
神经上下文老虎机易受对抗攻击,细微的奖励、动作或上下文扰动可导致次优决策。我们提出AdvBandit,一种无需访问受害者内部参数、奖励函数或梯度信息的黑盒自适应攻击方法。该方法将上下文投毒建模为连续动作的贝叶斯强化学习问题,利用最大熵逆强化学习模块从观察到的上下文-动作对构建代理模型,并通过投影梯度下降优化扰动。采用上置信界感知的高斯过程指导动作选择,并引入攻击预算控制机制以降低被检测风险和计算开销。理论分析表明,攻击者后悔值为亚线性,受害者后悔值在攻击次数上呈线性下界。在Yelp、MovieLens和Disin三个真实数据集上对多种受害者上下文老虎机的实验显示,本方法获得的受害者累积后悔值高于当前最优基线。
原文摘要 · Abstract (English)
Neural contextual bandits are vulnerable to adversarial attacks, where subtle perturbations to rewards, actions, or contexts induce suboptimal decisions. We introduce AdvBandit, a black-box adaptive attack that formulates context poisoning as a continuous-armed bandit problem, enabling the attacker to jointly learn and exploit the victim's evolving policy. The attacker requires no access to the victim's internal parameters, reward function, or gradient information; instead, it constructs a surrogate model using a maximum-entropy inverse reinforcement learning module from observed context-action pairs and optimizes perturbations against this surrogate using projected gradient descent. An upper confidence bound-aware Gaussian process guides arm selection. An attack-budget control mechanism is also introduced to limit detection risk and overhead. We provide theoretical guarantees, including sublinear attacker regret and lower bounds on victim regret linear in the number of attacks. Experiments on three real-world datasets (Yelp, MovieLens, and Disin) against various victim contextual bandits demonstrate that our attack model achieves higher cumulative victim regret than state-of-the-art baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。