首个3D视觉查询定位基准,支持多模态数据与精准标注。
Towards Visual Query Localization in the 3D World

- 构建3D多模态视觉查询定位新基准3DVQL,含2002段序列和17万帧数据。
- 提出LaF融合算法,显著优于现有基线模型,在复杂场景中定位更准确。
- 适合关注3D视觉理解、多模态融合与时空定位的研究者使用。
视觉查询定位(VQL)旨在根据查询预测序列中最近发生的时空响应。当前研究主要集中于2D视频中的VQL,而3D空间中的对应任务尚未受到足够关注。本文首次尝试解决3D世界中的视觉查询定位问题,提出了一个新基准3DVQL。该基准包含2,002个序列,约17万帧图像和6,400个响应轨迹片段,覆盖38个物体类别。每个序列提供点云、RGB图像和深度图等多种模态数据,支持灵活研究。为确保标注质量,所有序列均经过多轮人工标注与验证。据我们所知,3DVQL是首个面向3D多模态视觉查询定位的基准。为促进后续研究,我们基于点云和RGB图像实现了多个代表性3D多模态VQL基线模型。实验表明,现有方法在不同融合模块上表现差异显著。为此,我们提出一种名为LaF的提升-注意力融合算法,显著优于现有基线模型。相关基准与模型将公开发布于https://github.com/wuhengliangliang/3DVQL。
原文摘要 · Abstract (English)
Visual query localization (VQL) aims to predict the spatio-temporal response of the most recent occurrence in a sequence given a query. Currently, most research focuses on visual query localization in 2D videos, while its counterpart in 3D space has received little attention. In this paper, we make the first attempt to address visual query localization in the 3D world by introducing a novel benchmark, dubbed 3DVQL. Specifically, 3DVQL contains 2,002 sequences with around 170,000 frames and 6.4K response track segments from 38 object categories. Each sequence in 3DVQL is provided with multiple modalities, including point clouds, RGB images, and depth images, to support flexible research. To ensure high-quality annotations, each sequence is manually annotated with multiple rounds of verification and refinement. To the best of our knowledge, 3DVQL is the first benchmark for 3D multimodal visual query localization. To facilitate comparison in subsequent research, we implement a series of representative 3D multimodal VQL baselines using point clouds and RGB images. The experimental results show that existing methods exhibit significant performance variations across different fusion modules. To encourage future research, we propose a lift-and-attention fusion algorithm named LaF, which significantly outperforms existing baseline models. Our benchmark and model will be publicly released at https://github.com/wuhengliangliang/3DVQL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。