用大模型生成恶意奖励,诱导强化学习模型做出错误决策。
Policy Disruption in Reinforcement Learning:Adversarial Attack with Large Language Models and Critical State Identification
- 用大模型生成针对性奖励,干扰目标策略决策。
- 识别关键脆弱状态,使攻击在少数位置即造成显著性能下降。
- 无需修改环境,适合真实场景中的安全测试。
强化学习在机器人和自动驾驶等领域取得显著进展,但针对RL系统的对抗攻击仍具挑战性。现有方法多依赖修改环境或策略,实用性受限。本文提出一种新攻击方法:利用环境中已有智能体引导目标策略输出次优动作,不改变环境本身。通过构建奖励迭代优化框架,利用大语言模型(LLMs)生成针对目标智能体弱点的特定对抗奖励,有效提升诱导其做出次优决策的能力。同时设计关键状态识别算法,定位目标智能体最脆弱的状态,这些状态下次优行为会导致整体性能大幅下降。在多种环境中的实验表明,该方法优于现有攻击方案。
原文摘要 · Abstract (English)
Reinforcement learning (RL) has achieved remarkable success in fields like robotics and autonomous driving, but adversarial attacks designed to mislead RL systems remain challenging. Existing approaches often rely on modifying the environment or policy, limiting their practicality. This paper proposes an adversarial attack method in which existing agents in the environment guide the target policy to output suboptimal actions without altering the environment. We propose a reward iteration optimization framework that leverages large language models (LLMs) to generate adversarial rewards explicitly tailored to the vulnerabilities of the target agent, thereby enhancing the effectiveness of inducing the target agent toward suboptimal decision-making. Additionally, a critical state identification algorithm is designed to pinpoint the target agent's most vulnerable states, where suboptimal behavior from the victim leads to significant degradation in overall performance. Experimental results in diverse environments demonstrate the superiority of our method over existing approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。