arXiv:2512.04453cs.ROcs.AI2025-12中稿 · ACM/IEEE Internati…被引 2

让机器人通过语言和动作推断人类模糊目标,减少误判。

Open-Ended Goal Inference through Actions and Language for Human-Robot Collaboration

  • 融合语言与动作信息,在规划树中双向推理目标
  • 比基线方法错误率显著降低,预测更稳定
  • 只在信息增益大于中断成本时才提问,适合真实协作

为实现人机协作,机器人需推断那些常模糊、难以表达或不在固定集合中的目标。以往方法受限于预定义目标集,仅依赖观察动作或完全依赖明确指令,导致在真实交互中表现脆弱。本文提出BALI(双向动作-语言推理)方法,将自然语言偏好与人类行为结合,在滚动视野规划树中进行目标预测。BALI融合语言与动作线索,仅当预期信息增益超过中断成本时才发起澄清提问,并选择与推断目标一致的辅助动作。我们在协作烹饪任务中评估该方法,目标对机器人而言可能是全新的且无边界。相比基线,BALI实现了更稳定的预测和显著更少的错误。

原文摘要 · Abstract (English)

To collaborate with humans, robots must infer goals that are often ambiguous, difficult to articulate, or not drawn from a fixed set. Prior approaches restrict inference to a predefined goal set, rely only on observed actions, or depend exclusively on explicit instructions, making them brittle in real-world interactions. We present BALI (Bidirectional Action-Language Inference) for goal prediction, a method that integrates natural language preferences with observed human actions in a receding-horizon planning tree. BALI combines language and action cues from the human, asks clarifying questions only when the expected information gain from the answer outweighs the cost of interruption, and selects supportive actions that align with inferred goals. We evaluate the approach in collaborative cooking tasks, where goals may be novel to the robot and unbounded. Compared to baselines, BALI yields more stable goal predictions and significantly fewer mistakes.

人机协作目标推断多模态推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。