arXiv:2608.26194cs.CLcs.AI2026-08中稿 · Interspeech 2026

提出STeReO模型,统一调度语音与文本检索结果,提升问答准确率。

A Reranker for Orchestrating Heterogeneous Speech and Text Retrievers

论文配图:A Reranker for Orchestrating Heterogeneous Speech and Text Retrievers
图 1 · 摘自论文原文
  • 构建跨模态检索融合框架,用统一评分器排序异构检索结果。
  • 在混合模态场景下,问答准确率显著优于单一模态检索。
  • 专为多模态数据库设计,适合需要语音+文本联合检索的系统。

检索增强生成(RAG)系统因能缓解大语言模型的幻觉问题而受到广泛关注。尽管RAG的知识库正日益多样化,涵盖语音与文本等多种模态,但针对多模态数据库的研究仍较有限。本文提出STeReO(Speech and Text Reranking Orchestrator),一种基于语音与文本检索器的重排序器,用于聚合不同模态的知识库。为解决专用训练数据匮乏问题,我们首先构建了一个包含查询、混合模态证据及其相关性排名的数据集。随后训练该重排序器,并在单模态与混合模态场景下评估其有效性。结果表明,所提算法能有效选出最相关证据,显著提升下游问答性能。

原文摘要 · Abstract (English)

Retrieval-Augmented Generation (RAG) systems have attracted significant interest for their ability to mitigate hallucinations in Large Language Models (LLMs). Although knowledge databases for RAG are increasingly diversifying to include various modalities such as speech and text, research on handling such multi-modal database scenarios remains limited. In this paper, we propose STeReO (Speech and Text Reranking Orchestrator), a reranker based on speech and text retrievers that aggregates disparate modality databases. To address the lack of specialized training data, we first curate a dataset comprising queries, mixed-modality evidence, and their corresponding relevance ranks. We then train the reranker and evaluate its effectiveness in both single-modality and mixed-modality scenarios. Our results demonstrate that the proposed algorithm excels at selecting the most relevant evidence, thereby significantly improving downstream question-answering performance.

多模态检索RAG语音处理重排序

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。