arXiv:2506.14495cs.CV2025-06

让语音识别出错时也能准确定位3D场景中的物体

I Speak and You Find: Robust 3D Visual Grounding with Noisy and Ambiguous Speech Inputs

  • 用语音信号自身特征生成互补定位分数,减少对错误文字的依赖
  • 通过对比学习对齐语音与错误文本特征,提升抗噪能力
  • 适用于语音嘈杂或口音重的真实场景,适合做智能交互系统

现有3D视觉定位方法依赖精确文本提示来定位3D场景中的物体。语音作为一种自然直观的模态,具有替代潜力。然而,真实语音常因口音、背景噪声和语速差异导致转录错误,限制了现有3DVG方法的应用。为此,我们提出SpeechRefer——一种新型3DVG框架,旨在应对噪声和模糊的语音转录输入。该框架可无缝集成至现有3DVG模型,并引入两项关键创新:第一,语音互补模块利用音素相关词间的声学相似性,捕捉细微差异,从语音信号生成互补候选框得分,降低对潜在错误转录的依赖;第二,对比互补模块采用对比学习,将错误文本特征与对应语音特征对齐,确保在转录错误主导情况下仍具鲁棒性。在SpeechRefer和SpeechNr3D数据集上的大量实验表明,SpeechRefer显著提升现有3DVG方法的性能,凸显其弥合噪声语音与可靠3DVG之间差距的潜力,推动更直观、实用的多模态系统发展。

原文摘要 · Abstract (English)

Existing 3D visual grounding methods rely on precise text prompts to locate objects within 3D scenes. Speech, as a natural and intuitive modality, offers a promising alternative. Real-world speech inputs, however, often suffer from transcription errors due to accents, background noise, and varying speech rates, limiting the applicability of existing 3DVG methods. To address these challenges, we propose \textbf{SpeechRefer}, a novel 3DVG framework designed to enhance performance in the presence of noisy and ambiguous speech-to-text transcriptions. SpeechRefer integrates seamlessly with xisting 3DVG models and introduces two key innovations. First, the Speech Complementary Module captures acoustic similarities between phonetically related words and highlights subtle distinctions, generating complementary proposal scores from the speech signal. This reduces dependence on potentially erroneous transcriptions. Second, the Contrastive Complementary Module employs contrastive learning to align erroneous text features with corresponding speech features, ensuring robust performance even when transcription errors dominate. Extensive experiments on the SpeechRefer and peechNr3D datasets demonstrate that SpeechRefer improves the performance of existing 3DVG methods by a large margin, which highlights SpeechRefer's potential to bridge the gap between noisy speech inputs and reliable 3DVG, enabling more intuitive and practical multimodal systems.

3D视觉定位语音识别多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。