让机器人一次互动就猜对人意图,提升协作效率
Corrigible Assistance in One Round: Pragmatic-Pedagogic Best Response

- 用实用-教学式推理,在一步内解决目标困惑
- 相比传统方法,能突破认知上限,更好对齐人类意图
- 适合需要快速理解用户目标的智能助手场景
协助游戏形式化了信息不对称下的人机协作:人类知晓目标,而机器人需通过观察与交互推断目标以有效协助。通常在线计算最优策略不可行,因精确解需在部分可观马尔可夫决策过程(POMDP)中规划。本文识别出一类协助游戏,其中实用-教学式推理可在单步内解决目标不确定性,使整个时域游戏可通过可处理的最佳响应程序精确求解。在此类游戏中,主流逆向最优控制存在推理瓶颈,阻碍对齐;而实用-教学式推理通过行动设计,使任务执行看似等效的行为实现即时目标消歧。最后,我们在一个简单的协作积木搭建例子上验证了理论结果与所提方法的有效性。
原文摘要 · Abstract (English)
Assistance games formalize human-robot collaboration under asymmetric information: the human knows the goal, while the robot must infer it from observation and interaction in order to assist effectively. In general, computing optimal assistance game strategies online is intractable, since exact solutions require planning in a POMDP. We identify a class of assistance games in which pragmatic-pedagogic reasoning resolves goal uncertainty in a single time step, rendering the full-horizon game exactly solvable by a tractable best-response procedure. Within this class, we show that mainstream inverse optimal control exhibits an inference ceiling that hinders alignment, while pragmatic-pedagogic reasoning overcomes this barrier by immediately disambiguating goals through actions that look equivalent under task execution alone. Finally, we validate our theoretical results and proposed method on a simple collaborative block-building example.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。