用语音定位3D场景中的物体,效果媲美文本方法。
Audio-3DVG: Unified Audio -- Point Cloud Fusion for 3D Visual Grounding
- 拆解语音为物体提及检测与注意力融合两部分,提升推理结构化。
- 在ScanRefer等数据集上达到新最好结果,性能接近文本基方法。
- 适合研究语音与3D视觉融合、多模态交互的学者参考。
3D视觉定位(3DVG)旨在根据自然语言描述在3D点云中定位目标物体。尽管已有工作利用文本描述取得进展,但基于语音的3D视觉定位——即利用口语进行定位——仍鲜有探索且极具挑战。受自动语音识别(ASR)和语音表征学习进展的启发,本文提出Audio-3DVG,一个简单而高效的框架,用于融合音频与空间信息以增强定位能力。不同于将语音视为整体输入,我们将其分解为两个互补组件:首先引入(i)物体提及检测,一种多标签分类任务,显式识别音频中提到的物体,从而实现更结构化的音-景推理;其次提出(ii)音频引导注意力模块,建模目标候选与提及物体之间的交互,在复杂3D环境中增强区分能力。为支持基准测试,我们(iii)为标准3DVG数据集(包括ScanRefer、Sr3D和Nr3D)合成语音描述。实验表明,Audio-3DVG不仅在语音基定位任务中达到新的最先进水平,还与文本基方法性能相当,凸显了将口语融入3D视觉任务的巨大潜力。
原文摘要 · Abstract (English)
3D Visual Grounding (3DVG) involves localizing target objects in 3D point clouds based on natural language. While prior work has made strides using textual descriptions, leveraging spoken language-known as Audio-based 3D Visual Grounding-remains underexplored and challenging. Motivated by advances in automatic speech recognition (ASR) and speech representation learning, we propose Audio-3DVG, a simple yet effective framework that integrates audio and spatial information for enhanced grounding. Rather than treating speech as a monolithic input, we decompose the task into two complementary components. First, we introduce (i) Object Mention Detection, a multi-label classification task that explicitly identifies which objects are referred to in the audio, enabling more structured audio-scene reasoning. Second, we propose an (ii) Audio-Guided Attention module that models the interactions between target candidates and mentioned objects, enhancing discrimination in cluttered 3D environments. To support benchmarking, we (iii) synthesize audio descriptions for standard 3DVG datasets, including ScanRefer, Sr3D, and Nr3D. Experimental results demonstrate that Audio-3DVG not only achieves new state-of-the-art performance in audio-based grounding, but also competes with text-based methods, highlight the promise of integrating spoken language into 3D vision tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。