针对可学习的追踪者,设计能持续欺骗的路径规划方法。
Repeated Deceptive Path Planning against Learnable Observer

- 分两层优化:短期调整策略应对当前追踪者,长期学习追踪者更新规律。
- 相比现有方法,欺骗成功率显著提升,路径成本保持相近。
- 适合军事、物流等需长期隐蔽行动的场景。
我们研究欺骗性路径规划(DPP)问题,即智能体试图隐藏其真实目的地免受外部观察者窥探。现有工作假设观察者为静态非学习型,但在现实场景如关键物资运输或军事行动中,对手可通过学习历史轨迹动态适应。为此,我们提出重复欺骗路径规划(RDPP),显式建模可学习的观察者。实验表明,现有DPP方法在该设定下失效,因无法适应不断演化的敌方预测。虽将观察者历史预测纳入更新可实现部分自适应,但增量更新导致累积延迟,削弱欺骗效果。为此,我们提出欺骗元规划(DeMP),一种双层优化框架:通过剧集级适应实现短期策略调整以对抗更新后的观察者,同时通过元级更新利用跨剧集反馈,捕捉观察者模型更新规律,加速未来剧集的适应速度。该机制有效缓解适应滞后积累,实现对学习型观察者的持续欺骗。多环境实验表明,DeMP在RDPP任务中显著优于现有方法,同时保持具有竞争力的路径成本。结果凸显了建模与可学习对手反复交互的重要性,为多智能体系统中的欺骗与隐私保护提供了新洞见。
原文摘要 · Abstract (English)
We study the problem of deceptive path planning (DPP), where an agent aims to conceal its true destination from external observers. While existing work assumes static, non-learning observers, real-world adversaries-such as in critical goods transportation or military operations-can adapt by learning from historical trajectories. To address this gap, we introduce Repeated Deceptive Path Planning (RDPP), a new formulation that explicitly models learnable observers. We show that existing DPP methods fail under this setting, as they cannot adapt to evolving adversarial predictions. While incorporating observer previous predictions into updates enables some adaptation, such incremental updates cause accumulative lag that degrades deception. To this end, we propose Deceptive Meta Planning (DeMP), a two-level optimization framework that combines episode-level adaptation, which enables short-term policy adjustment to counter updated observer, and meta-level updates, which leverage cross-episode feedback to capture how observers update their models and accelerate adaptation in future episodes. In this way, DeMP mitigates the accumulation of adaptation lag, enabling sustained deception against a learning observer. Experiments across environments demonstrate that DeMP significantly outperforms existing approaches in RDPP while maintaining competitive path cost. Our results highlight the importance of modeling repeated interactions with learnable adversaries, providing new insights into deception and privacy in multi-agent systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。