arXiv:2503.24284cs.LGcs.AI2025-03被引 2

提出新方法让智能体在对手干预下隐藏真实目标,靠降低对手信息价值实现欺骗。

Value of Information-based Deceptive Path Planning Under Adversarial Interventions

  • 基于马尔可夫决策过程建模,用信息价值引导路径设计
  • 实验显示在网格世界中有效欺骗对手干预,优于现有方法
  • 适合对抗性环境中的隐蔽路径规划,如军事或安全场景

现有欺骗性路径规划(DPP)方法仅针对被动观察者设计,无法应对具备主动干扰能力的对手。本文提出一种基于马尔可夫决策过程(MDP)的新模型,并设计新的信息价值(VoI)目标函数,使路径规划代理通过选择对对手信息价值低的轨迹,诱使其做出次优干扰决策。利用与线性规划理论的关联,推导出高效的策略合成算法。实验表明,该方法在典型网格世界任务中有效实现对抗性干扰下的欺骗性路径规划,性能优于现有DPP方法及保守路径规划策略。

原文摘要 · Abstract (English)

Existing methods for deceptive path planning (DPP) address the problem of designing paths that conceal their true goal from a passive, external observer. Such methods do not apply to problems where the observer has the ability to perform adversarial interventions to impede the path planning agent. In this paper, we propose a novel Markov decision process (MDP)-based model for the DPP problem under adversarial interventions and develop new value of information (VoI) objectives to guide the design of DPP policies. Using the VoI objectives we propose, path planning agents deceive the adversarial observer into choosing suboptimal interventions by selecting trajectories that are of low informational value to the observer. Leveraging connections to the linear programming theory for MDPs, we derive computationally efficient solution methods for synthesizing policies for performing DPP under adversarial interventions. In our experiments, we illustrate the effectiveness of the proposed solution method in achieving deceptiveness under adversarial interventions and demonstrate the superior performance of our approach to both existing DPP methods and conservative path planning approaches on illustrative gridworld problems.

路径规划对抗性策略信息价值

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。