arXiv:2511.09282cs.SDcs.CL2025-11AAAI被引 3

提出端到端语音检索模型CLSR,提升长音频问答的精准定位能力。

End-to-end Contrastive Language-Speech Pretraining Model For Long-form Spoken Question Answering

  • 将声学特征转为类文本表示,再进行跨模态对齐,更好融合语音与文本
  • 在4个数据集上优于现有端到端和流水线方法,显著提升长音频片段检索效果
  • 适合需要处理长时语音的问答系统开发者使用

近年来,语音问答(SQA)取得显著进展。然而,许多现有方法,包括大型音频语言模型,在处理长音频时仍存在困难。受检索增强生成启发,语音相关检索器在预处理长音频方面展现出潜力,但当前性能仍不理想。为此,我们提出端到端对比语言-语音检索模型CLSR,能高效从长音频中提取与问题相关的片段,用于下游SQA任务。不同于传统语音-文本对比模型,CLSR在对齐前增加一步:将声学特征转换为类文本表示,从而更有效地弥合模态差距。在四个跨模态检索数据集上的实验表明,CLSR在性能上超越了现有的端到端语音检索器及结合语音识别与文本检索的流水线方法,为实际长音频问答应用提供了坚实基础。

原文摘要 · Abstract (English)

Significant progress has been made in spoken question answering (SQA) in recent years. However, many existing methods, including large audio language models, struggle with processing long audio. Follow the success of retrieval augmented generation, a speech-related retriever shows promising in help preprocessing long-form speech. But the performance of existing speech-related retrievers is lacking. To address this challenge, we propose CLSR, an end-to-end contrastive language-speech retriever that efficiently extracts question-relevant segments from long audio recordings for downstream SQA task. Unlike conventional speech-text contrastive models, CLSR incorporates an intermediate step that converts acoustic features into text-like representations prior to alignment, thereby more effectively bridging the gap between modalities. Experimental results across four cross-modal retrieval datasets demonstrate that CLSR surpasses both end-to-end speech related retrievers and pipeline approaches combining speech recognition with text retrieval, providing a robust foundation for advancing practical long-form SQA applications.

语音问答跨模态检索长音频处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。