提出在线方法,让智能体在未知对手下主动获取信息并优化行动轨迹。
Online Competitive Information Gathering for Partially Observable Trajectory Games
- 基于粒子估计联合状态空间,设计可在线计算的轨迹规划方法。
- 在追逃和仓库取货场景中实现主动探查,性能优于被动对手。
- 适用于多智能体、含障碍物等复杂环境,支持实时部署。
博弈论智能体需制定最优策略以获取对手信息。此类问题通常建模为部分可观测随机博弈(POSGs),但在完全连续的POSG中进行规划需大量离线计算或对玩家信念顺序做强假设,难以实现。本文提出一种有限历史/时域的POSG精化模型,支持在轨迹空间中产生竞争性信息获取行为。通过一系列近似,我们构建了一种在线方法,利用粒子对联合状态空间进行估计,并执行随机梯度博弈,以计算理性轨迹计划。同时提供个体智能体部署所需的调整。方法在连续追逃与仓库取货场景中验证,扩展至N>2玩家及包含视觉与物理障碍的复杂环境,表现出主动信息获取行为,且性能优于被动对手。
原文摘要 · Abstract (English)
Game-theoretic agents must make plans that optimally gather information about their opponents. These problems are modeled by partially observable stochastic games (POSGs), but planning in fully continuous POSGs is intractable without heavy offline computation or assumptions on the order of belief maintained by each player. We formulate a finite history/horizon refinement of POSGs which admits competitive information gathering behavior in trajectory space, and through a series of approximations, we present an online method for computing rational trajectory plans in these games which leverages particle-based estimations of the joint state space and performs stochastic gradient play. We also provide the necessary adjustments required to deploy this method on individual agents. The method is tested in continuous pursuit-evasion and warehouse-pickup scenarios (alongside extensions to $N > 2$ players and to more complex environments with visual and physical obstacles), demonstrating evidence of active information gathering and outperforming passive competitors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。