PIMbot通过调控奖励与策略,实现对多机器人协作的自适应攻击。
PIMbot: A Self-Adaptive Attack Framework for Adversarial Manipulation of Multi-Robot Reinforcement Learning

- 利用奖励与自身策略双重杠杆在线调节协作行为。
- 在仿真与真实嵌入式设备上均验证了高效操控能力。
- 适合研究多机器人系统安全漏洞的研究者参考。
近期研究表明强化学习在多机器人协作中具有潜力,尤其在个体利益与集体收益冲突的社会困境场景中。然而,通信错误和对抗性机器人等环境因素会影响协作效果,因此探索如何操纵多机器人通信以改变结果至关重要。本文提出PIMbot框架,通过两种互补机制实现操控:(i) 奖励通道的激励操纵,(ii) 代理自身行为的策略操纵。一个自适应多目标控制器在线平衡这两类杠杆。本工作引入一种新颖的操纵方法,应用于基于独特激励函数的多智能体强化学习社会困境场景。通过PIMbot机制,单个机器人可有效操控社会困境环境。大量实验结果表明该方法在Gazebo模拟的多机器人环境中有效。此外,在NVIDIA Jetson Orin Nano上的真实嵌入式设备案例研究量化了系统开销,并验证了PIMbot在真实自主嵌入式系统场景中的有效性。这些结果使PIMbot成为检验多机器人协作任务关键脆弱性的严格压力测试工具。
原文摘要 · Abstract (English)
Recent research has demonstrated the potential of reinforcement learning in effective multi-robot collaboration, particularly in social dilemmas where robots face a trade-off between self-interest and collective benefits. However, environmental factors such as miscommunication and adversarial robots can impact cooperation, making it crucial to explore how multi-robot communication can be manipulated to achieve different outcomes. This paper presents PIMbot, a framework that manipulates outcomes via two complementary levers: (i) incentive manipulation of the reward channel and (ii) policy manipulation of an agent's own actions. An adaptive multi-objective controller balances these levers in an online manner. Our work introduces a novel approach to manipulation in recent multi-agent RL social dilemmas that utilize a unique reward function for incentivization. By utilizing our proposed PIMbot mechanisms, a robot is able to manipulate the social dilemma environment effectively. Comprehensive experimental results demonstrate the effectiveness of our proposed methods in the Gazebo-simulated multi-robot environment. Moreover, a real embedded device case study on NVIDIA Jetson Orin Nano quantifies system cost and validates PIMbot's effectiveness on realistic autonomous embedded systems scenarios beyond simulation. Together, these results position PIMbot as a rigorous stress-test tool exposing critical vulnerabilities in multi-robot cooperative tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。