首个支持自然指令的语音检索基准,让语音搜索更懂用户意图。
INSPIRE: A Benchmark for Instruction-Aware Speech Retrieval
- 用自然语言指令动态定义检索标准,涵盖语义、说话人、音色等多维度
- 现有模型在不同指令下表现参差,无一能稳定应对所有类型检索
- 适合语音交互、智能助手研发者,推动可理解语音搜索发展
现有语音检索系统依赖固定相似性匹配,无法适应多样化的用户意图。我们提出INSPIRE,首个面向指令感知语音检索的基准数据集,其中自然语言指令动态指定相关性标准,包括语义内容、说话人身份、说话风格、环境声音及其组合。我们评估了四种检索范式:大音频-语言模型、级联流水线、自监督语音模型和对比音频-语言模型。结果表明,当前任何方法均无法稳健处理所有检索意图。基于文本的方法在语义检索中表现相对较好,但在语音属性(如音色)上表现不佳;基于语音的模型在捕捉声学特性方面中等表现,但难以准确遵循指令。这些发现凸显了统一架构在实现指令感知语音检索中的迫切需求。
原文摘要 · Abstract (English)
Existing speech retrieval systems rely on fixed similarity matching and cannot adapt to diverse user intents. We introduce INSPIRE, the first benchmark for instruction-aware speech retrieval, in which natural-language instructions dynamically specify relevance criteria, including semantic content, speaker identity, speaking style, environmental sounds, and their combinations. We evaluate four retrieval paradigms: large audio-language models, cascaded pipelines, self-supervised speech models, and contrastive audio-language models. Our results reveal that no current method robustly handles all retrieval intents. Text-based approaches perform relatively better at semantic retrieval but struggle with paralinguistic attributes, while speech-based models are moderately better at capturing acoustic properties but falter at following instructions. These findings highlight the need for unified architectures capable of instruction-aware speech retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。