机器人通过对话和非语言信号主动化解协作中的不确定性。
Knowing When to Ask: Resolving Uncertainty in Human-Robot Joint Planning via Explicit Dialogue and Implicit Intent Cues
- 双通道通信:关键时主动提问,无需时读取人体姿态等隐含意图。
- 对话成本降低51.9%,任务成功率保持100%;非语言协作提速25.4%。
- 适用于需高效人机协同的无人机、服务机器人等开放场景。
在开放世界环境中实现有效的人机协作,需要在任务、环境及人类队友状态不确定的情况下进行联合规划。沟通是解决此类不确定性最直接的方式,但现有系统大多仅支持单向通信:机器人倾听并执行,将人类视为被动监督者,而非可双向对话的协作伙伴。本文提出一个统一的人机联合规划系统,机器人通过两种互补的通信渠道主动化解不确定性。当不确定性影响决策时,系统启动不确定性缓解联合规划模块,借助大模型辅助的主动问询机制澄清模糊指令,通过假设增强型A*搜索枚举可通行性假设,并用动态规划计算代价最优的提问策略,确保仅询问对计划真正重要的问题。当显式对话不必要或不可行时,实时意图感知协作模块则从空间与方向信号中推断人类潜在任务意图,建立概率信念,实现无通信开销的协调式任务选择。我们在Gazebo仿真与真实无人机部署中验证了该系统,集成语音对话接口与基于视觉-语言模型(VLM)的3D语义感知流水线。实验结果表明,代价最优的澄清对话使交互成本降低51.9%,同时保持100%任务成功率;隐式意图识别相较基线将协作任务执行时间减少25.4%。
原文摘要 · Abstract (English)
Effective human-robot collaboration in open-world environments requires joint planning under uncertainty about the task, the environment, and the human teammate. Communication is the most direct means of resolving such uncertainty, yet most existing systems support only one-way communication: robots listen and act, treating humans as passive supervisors rather than conversational teammates capable of two-way dialogue. We propose a unified human-robot joint planning system in which the robot actively resolves uncertainty through two complementary communication channels. When uncertainty is decision-critical, an uncertainty-mitigation joint planning module engages the human in clarification dialogue: it grounds ambiguous instructions via an LLM-assisted active elicitation mechanism, enumerates traversability hypotheses through a hypothesis-augmented A* search, and computes a cost-optimal querying policy via dynamic programming, so that the robot asks only the questions whose answers actually matter for the plan. When explicit dialogue is unnecessary or impractical, a real-time intent-aware collaboration module instead reads implicit, nonverbal cues, maintaining a probabilistic belief over the human's latent task intent from spatial and directional signals to enable coordination-aware task selection without any communication overhead. We validate the proposed system in both Gazebo simulations and real-world UAV deployments, integrated with a voice dialogue interface and a Vision-Language Model (VLM)-based 3D semantic perception pipeline. Experimental results show that cost-optimal clarification dialogue cuts the interaction cost by 51.9% while maintaining a 100% task success rate, and implicit intent reading reduces the cooperative task execution time by 25.4% compared to the baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。