用语音指令精准定位图像中的物体,提升机器人视觉理解能力。
You Only Speak Once to See

- 结合预训练音频与视觉模型,通过对比学习对齐多模态信息。
- 实验证明语音引导可有效提升物体定位精度与鲁棒性。
- 适合希望增强机器人交互能力的研究者和开发者。
利用视觉线索进行图像中物体定位是计算机视觉中的成熟方法,但音频作为物体识别与定位的模态仍鲜有探索。我们提出 YOSS(You Only Speak Once to See),通过将预训练音频模型与视觉模型结合,采用对比学习与多模态对齐技术,直接将语音指令或描述映射到图像中的对应物体。实验表明,音频引导可有效应用于物体定位任务,说明引入音频有助于提升现有物体定位方法的精度与鲁棒性,并改善机器人系统及计算机视觉应用的表现。这一发现为更先进的物体识别、场景理解以及更直观、高效的机器人系统开发开辟了新路径。
原文摘要 · Abstract (English)
Grounding objects in images using visual cues is a well-established approach in computer vision, yet the potential of audio as a modality for object recognition and grounding remains underexplored. We introduce YOSS, "You Only Speak Once to See," to leverage audio for grounding objects in visual scenes, termed Audio Grounding. By integrating pre-trained audio models with visual models using contrastive learning and multi-modal alignment, our approach captures speech commands or descriptions and maps them directly to corresponding objects within images. Experimental results indicate that audio guidance can be effectively applied to object grounding, suggesting that incorporating audio guidance may enhance the precision and robustness of current object grounding methods and improve the performance of robotic systems and computer vision applications. This finding opens new possibilities for advanced object recognition, scene understanding, and the development of more intuitive and capable robotic systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。