从多人对话中精准提取目标说话人指令,实现更准确的语音图像检索。
Listening for "You": Enhancing Speech Image Retrieval via Target Speaker Extraction
- 通过目标说话人感知对比学习,融合音频与视觉模型。
- 在2人和3人场景下召回率分别达36.3%和29.9%,显著优于现有方法。
- 适用于助手机器人、多模态交互等真实场景中的语音定位任务。
基于语音线索的图像检索已成为多模态感知的重要方向,但在多人对话场景中利用语音仍具挑战。本文提出新的目标说话人语音-图像检索任务及框架,通过目标说话人提取与检索模块,在存在多个说话人的情况下建模图像与语音信号间的关系。该方法结合预训练自监督音频编码器与视觉模型,采用目标说话人感知的对比学习策略,使系统能从目标说话人处提取语音指令并与其对应图像对齐。在SpokenCOCO2Mix和SpokenCOCO3Mix数据集上的实验表明,所提方法在2人和3人场景下的Recall@1分别达到36.3%和29.9%,大幅超越单说话人基线与现有最先进模型。该方法在辅助机器人与多模态交互系统中具有实际应用潜力。
原文摘要 · Abstract (English)
Image retrieval using spoken language cues has emerged as a promising direction in multimodal perception, yet leveraging speech in multi-speaker scenarios remains challenging. We propose a novel Target Speaker Speech-Image Retrieval task and a framework that learns the relationship between images and multi-speaker speech signals in the presence of a target speaker. Our method integrates pre-trained self-supervised audio encoders with vision models via target speaker-aware contrastive learning, conditioned on a Target Speaker Extraction and Retrieval module. This enables the system to extract spoken commands from the target speaker and align them with corresponding images. Experiments on SpokenCOCO2Mix and SpokenCOCO3Mix show that TSRE significantly outperforms existing methods, achieving 36.3% and 29.9% Recall@1 in 2 and 3 speaker scenarios, respectively - substantial improvements over single speaker baselines and state-of-the-art models. Our approach demonstrates potential for real-world deployment in assistive robotics and multimodal interaction systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。