让机器人引导人类到关键决策点,更快猜出对方目标。
Robots Influencing Humans to Reveal their Goals during Collaboration and Competition

- 通过设计关键决策点引导人类行为
- 在模拟和真实机器人上提前准确推断目标
- 适用于协作与对抗场景的快速目标推理
我们提出一种统一策略,实现人机交互中的快速目标推断。核心思想是将人类引导至关键决策点(CDPs)——即不同人类策略在此处会做出不同动作的状态,从而最大化暴露其目标。通过目标条件下的策略分歧度定义CDPs,并将其融入滚动时域规划器,该规划器在优化任务进展与信息增益的联合成本函数时,探索未来动作序列。我们在仿真和真实机器人上评估了该方法,在协作的烹饪任务和对抗的藏匿-搜寻游戏中均表现更优:相比基线策略,本方法能更早、更准确地推断出人类目标。
原文摘要 · Abstract (English)
We propose a unified strategy for fast goal inference in human-robot interaction. The core idea is to drive the human toward Critical Decision Points (CDPs)-states where competing human strategies prescribe different next actions and thus maximally reveal the goal. We formalise CDPs using a goal-conditioned policy divergence measure and incorporate them into a Receding-Horizon Planner that explores future action sequences while optimizing a cost function balancing task progress and information gain. We evaluate this approach in both a collaborative, fully observable cooking task and a competitive, partially observable hide-and-seek game, each in simulation and on real robots. In both scenarios, our method infers human goals more accurately and earlier than baseline strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。