让机器人在看不见用户时也能听懂‘把那个拿给我’这类指令。
Take That for Me: Multimodal Exophora Resolution with Interactive Questioning for Ambiguous Out-of-View Instructions
- 结合声音定位与视觉语言模型,定位不可见用户的指向目标。
- 用户不可见时效果提升2倍,可见时提升1.3倍。
- 通过GPT-4o主动提问澄清歧义,适合真实家庭场景应用。
日常助手机器人需理解包含指示代词(如“把那个杯子拿给我”)的模糊口语指令,即使物体或用户不在机器人视野内。现有方法主要依赖视觉数据,在对象或用户不可见时失效。本文提出多模态交互式外指消解框架MIEL,融合声源定位(SSL)、语义地图、视觉语言模型(VLMs)与GPT-4o驱动的交互式提问。该方法首先构建环境语义地图,并结合用户骨骼数据从语言查询中生成候选目标。利用SSL使机器人朝向初始不在视野内的用户,从而准确识别手势与指向方向。当仍存在歧义时,机器人主动发起交互,由GPT-4o生成澄清问题。真实环境实验表明,用户可见时性能较无SSL与交互提问的方法提升约1.3倍,不可见时提升达2.0倍。项目主页:https://emergentsystemlabstudent.github.io/MIEL/
原文摘要 · Abstract (English)
Daily life support robots must interpret ambiguous verbal instructions involving demonstratives such as ``Bring me that cup,'' even when objects or users are out of the robot's view. Existing approaches to exophora resolution primarily rely on visual data and thus fail in real-world scenarios where the object or user is not visible. We propose Multimodal Interactive Exophora resolution with user Localization (MIEL), which is a multimodal exophora resolution framework leveraging sound source localization (SSL), semantic mapping, visual-language models (VLMs), and interactive questioning with GPT-4o. Our approach first constructs a semantic map of the environment and estimates candidate objects from a linguistic query with the user's skeletal data. SSL is utilized to orient the robot toward users who are initially outside its visual field, enabling accurate identification of user gestures and pointing directions. When ambiguities remain, the robot proactively interacts with the user, employing GPT-4o to formulate clarifying questions. Experiments in a real-world environment showed results that were approximately 1.3 times better when the user was visible to the robot and 2.0 times better when the user was not visible to the robot, compared to the methods without SSL and interactive questioning. The project website is https://emergentsystemlabstudent.github.io/MIEL/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。