用3D语义场景图实现机器人实时问答,提升探索效率与准确率。
GraphEQA: Using 3D Semantic Scene Graphs for Real-time Embodied Question Answering
- 基于实时3D度量-语义图与图像构建多模态记忆
- 在两个数据集上成功率更高,规划步数更少
- 适用于家庭和办公室等真实场景的智能机器人
在具身问答(EQA)任务中,智能体需探索未知环境并建立语义理解以准确回答情境问题。该任务在机器人领域仍具挑战性,主要源于难以获取有效语义表征、在线更新表征困难,以及缺乏先验世界知识支持高效规划与探索。为此,我们提出GraphEQA,利用实时3D度量-语义场景图(3DSGs)和任务相关图像作为多模态记忆,为视觉-语言模型(VLMs)提供接地支持,实现在未见环境中的EQA任务。采用分层规划策略,利用3DSGs的层次结构实现结构化规划与语义引导探索。我们在仿真环境中对HM-EQA和OpenEQA两个基准数据集进行评估,结果表明,GraphEQA在完成任务的成功率上优于多个基线方法,且所需规划步数更少。进一步实验验证了其在多个真实家庭与办公环境中的有效性。
原文摘要 · Abstract (English)
In Embodied Question Answering (EQA), agents must explore and develop a semantic understanding of an unseen environment to answer a situated question with confidence. This problem remains challenging in robotics, due to the difficulties in obtaining useful semantic representations, updating these representations online, and leveraging prior world knowledge for efficient planning and exploration. To address these limitations, we propose GraphEQA, a novel approach that utilizes real-time 3D metric-semantic scene graphs (3DSGs) and task relevant images as multi-modal memory for grounding Vision-Language Models (VLMs) to perform EQA tasks in unseen environments. We employ a hierarchical planning approach that exploits the hierarchical nature of 3DSGs for structured planning and semantics-guided exploration. We evaluate GraphEQA in simulation on two benchmark datasets, HM-EQA and OpenEQA, and demonstrate that it outperforms key baselines by completing EQA tasks with higher success rates and fewer planning steps. We further demonstrate GraphEQA in multiple real-world home and office environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。