直接用语音定位图像中的物体,比先转成文字再处理更鲁棒。
Layover or Direct Flight: Rethinking Audio-Guided Image Segmentation
- 跳过文字转录,直接对齐音频与图像实现物体定位。
- 在多种口音下,直接语音方法优于传统文字转录方法。
- 适用于需要抗语言差异的机器人交互场景。
理解人类指令是实现人机顺畅交互的关键。本文聚焦于物体定位任务,即根据口头指令在视觉场景(如图像)中定位目标物体。尽管近期已有进展,主流方法仍依赖文本作为中间表示:先将语音转为文字,提取关键物体词,再使用在大规模图文数据集上预训练的模型进行定位。然而,我们质疑这种转录式流程的效率与鲁棒性。核心问题在于:能否不依赖文本,直接实现音频-视觉对齐?为此,我们简化任务,专注于单字语音指令的定位。构建了一个覆盖广泛物体类别和多样人类口音的新型音频定位数据集,并适配与基准测试多个来自紧密相关领域的模型。结果表明,直接从音频进行定位不仅可行,且在应对语言变体时表现更优,部分情况下甚至超越转录基线。研究鼓励重新关注直接语音定位,为构建更鲁棒、高效的多模态理解系统铺平道路。
原文摘要 · Abstract (English)
Understanding human instructions is essential for enabling smooth human-robot interaction. In this work, we focus on object grounding, i.e., localizing an object of interest in a visual scene (e.g., an image) based on verbal human instructions. Despite recent progress, a dominant research trend relies on using text as an intermediate representation. These approaches typically transcribe speech to text, extract relevant object keywords, and perform grounding using models pretrained on large text-vision datasets. However, we question both the efficiency and robustness of such transcription-based pipelines. Specifically, we ask: Can we achieve direct audio-visual alignment without relying on text? To explore this possibility, we simplify the task by focusing on grounding from single-word spoken instructions. We introduce a new audio-based grounding dataset that covers a wide variety of objects and diverse human accents. We then adapt and benchmark several models from the closely audio-visual field. Our results demonstrate that direct grounding from audio is not only feasible but, in some cases, even outperforms transcription-based methods, especially in terms of robustness to linguistic variability. Our findings encourage a renewed interest in direct audio grounding and pave the way for more robust and efficient multimodal understanding systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。