为3D点云序列定位引入时序推理模块,提升多步指令理解能力
GroundFlow: A Plug-in Module for Temporal Reasoning on 3D Point Cloud Sequential Grounding
- 设计可插拔的GroundFlow模块,捕捉历史步骤信息
- 在5个数据集上提升基线模型准确率7.5%~10.2%
- 适合需要理解多步语言指令的3D场景理解任务
3D点云序列定位(SG3D)指根据包含详细步骤的日常活动文本指令,定位一系列物体。现有3D视觉定位方法将多步指令整体处理,未提取每一步的时序信息。然而,SG3D指令常使用代词如'it'、'here'和'同样的',要求模型理解上下文并回溯前步信息以准确定位。由于缺乏有效的历史信息聚合模块,当前最优3DVG方法在SG3D任务中面临显著挑战。为此,我们提出GroundFlow——一个用于3D点云序列定位的时序推理插件模块。实验表明,集成GroundFlow使3DVG基线方法在SG3D基准上准确率大幅提升(+7.5%和+10.2%),甚至超越在多数据集预训练的3D大语言模型。该模块基于相关性选择性提取短时与长时步骤信息,实现对历史信息的全面感知,并在步骤增多时保持时序理解优势。本工作为现有3DVG模型引入时序推理能力,在五个数据集上达到当前最优性能。
原文摘要 · Abstract (English)
Sequential grounding in 3D point clouds (SG3D) refers to locating sequences of objects by following text instructions for a daily activity with detailed steps. Current 3D visual grounding (3DVG) methods treat text instructions with multiple steps as a whole, without extracting useful temporal information from each step. However, the instructions in SG3D often contain pronouns such as "it", "here" and "the same" to make language expressions concise. This requires grounding methods to understand the context and retrieve relevant information from previous steps to correctly locate object sequences. Due to the lack of an effective module for collecting related historical information, state-of-the-art 3DVG methods face significant challenges in adapting to the SG3D task. To fill this gap, we propose GroundFlow -- a plug-in module for temporal reasoning on 3D point cloud sequential grounding. Firstly, we demonstrate that integrating GroundFlow improves the task accuracy of 3DVG baseline methods by a large margin (+7.5\% and +10.2\%) in the SG3D benchmark, even outperforming a 3D large language model pre-trained on various datasets. Furthermore, we selectively extract both short-term and long-term step information based on its relevance to the current instruction, enabling GroundFlow to take a comprehensive view of historical information and maintain its temporal understanding advantage as step counts increase. Overall, our work introduces temporal reasoning capabilities to existing 3DVG models and achieves state-of-the-art performance in the SG3D benchmark across five datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。