让智能体在决策时主动误导对手,隐藏真实目标且保持高效。
Deceptive Sequential Decision-Making via Regularized Policy Optimization
- 通过三种正则化策略实现欺骗性决策:误导、定向误导、模糊误导
- 欺骗时仍能保持98%以上的最优累积奖励
- 适用于对抗环境下需隐藏真实目标的自主系统
自主系统在面对对手时,可能因行为被观察而泄露敏感信息。本文提出一种欺骗性序贯决策框架,不仅隐藏敏感信息,还主动误导对手对系统奖励函数的认知。将自主系统建模为马尔可夫决策过程,对手使用逆强化学习推断奖励函数。为此,提出三种策略:‘转移欺骗’使对手得出任意错误结论;‘定向欺骗’引导对手形成特定错误认知;‘模糊欺骗’使对手难以区分真实与虚假奖励。分析表明,每种欺骗方式均可融入策略优化,并可严格界定其带来的累积奖励损失。在多智能体环境中验证,三类欺骗均成功引导对手产生错误信念,同时维持至少98%的最优非欺骗累积奖励。
原文摘要 · Abstract (English)
Autonomous systems are increasingly expected to operate in the presence of adversaries, though adversaries may infer sensitive information simply by observing a system. Therefore, present a deceptive sequential decision-making framework that not only conceals sensitive information, but actively misleads adversaries about it. We model autonomous systems as Markov decision processes, with adversaries using inverse reinforcement learning to recover reward functions. To counter them, we present three regularization strategies for policy synthesis problems that actively deceive an adversary about a system's reward. ``Diversionary deception'' leads an adversary to draw any false conclusion about the system's reward function. ``Targeted deception'' leads an adversary to draw a specific false conclusion about the system's reward function. ``Equivocal deception'' leads an adversary to infer that the real reward and a false reward both explain the system's behavior. We show how each form of deception can be implemented in policy optimization problems and analytically bound the loss in total accumulated reward induced by deception. Next, we evaluate these developments in a multi-agent setting. We show that diversionary, targeted, and equivocal deception all steer the adversary to false beliefs while still attaining a total accumulated reward that is at least 98% of its optimal, non-deceptive value.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。