让机器人在模糊指令下精准找物,兼顾语言手势与环境不确定性。
LEGS-POMDP: Language and Gesture-Guided Object Search in Partially Observable Environments
- 构建模块化POMDP框架,融合语言、手势与视觉信息
- 模拟中多模态融合成功率达89%,显著优于单一模态
- 适用于真实机器人,适合开放世界复杂任务场景
为在开放世界环境中协助人类,机器人需理解模糊指令以定位目标物体。基于基础模型的方法在多模态对齐上表现优异,但缺乏对长时任务不确定性的系统建模能力。相反,部分可观测马尔可夫决策过程(POMDP)提供了不确定性下的规划框架,但常受限于模态支持和严苛环境假设。本文提出语言与手势引导的开放世界物体搜索系统LEGS-POMDP,其显式建模两类部分可观测性:目标物体身份的不确定性及其空间位置的不确定性。在仿真中,多模态融合显著优于单模态基线,跨挑战性环境与物品类别平均成功率达89%。最后,我们在四足移动操作机上验证了完整系统,真实实验定性证明了鲁棒的多模态感知与模糊指令下的不确定性降低能力。
原文摘要 · Abstract (English)
To assist humans in open-world environments, robots must interpret ambiguous instructions to locate desired objects. Foundation model-based approaches excel at multimodal grounding, but they lack a principled mechanism for modeling uncertainty in long-horizon tasks. In contrast, Partially Observable Markov Decision Processes (POMDPs) provide a systematic framework for planning under uncertainty but are often limited in supported modalities and rely on restrictive environment assumptions. We introduce LanguagE and Gesture-Guided Object Search in Partially Observable Environments (LEGS-POMDP), a modular POMDP system that integrates language, gesture, and visual observations for open-world object search. Unlike prior work, LEGS-POMDP explicitly models two sources of partial observability: uncertainty over the target object's identity and its spatial location. In simulation, multimodal fusion significantly outperforms unimodal baselines, achieving an average success rate of 89\% across challenging environments and object categories. Finally, we demonstrate the full system on a quadruped mobile manipulator, where real-world experiments qualitatively validate robust multimodal perception and uncertainty reduction under ambiguous instructions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。