arXiv:2602.12783cs.IRcs.AI2026-02中稿 · SIGIR 2026被引 1

构建噪声下语音检索的基准测试,评估系统在真实环境中的鲁棒性。

SQuTR: A Robustness Benchmark for Spoken Query to Text Retrieval under Acoustic Noise

  • 构建含3.7万条语音查询的大规模噪声数据集,覆盖中英文多领域。
  • 在不同信噪比下测试发现,模型性能随噪声增强显著下降。
  • 提供可复现评估协议,适合研究语音检索鲁棒性的团队使用。

语音查询检索是现代信息检索的重要交互方式,但现有评测数据集多限于简单查询和受限噪声条件,难以评估系统在复杂声学扰动下的鲁棒性。为此,我们提出SQuTR,一个面向语音查询到文本检索的鲁棒性基准,包含大规模数据集与统一评估协议。SQuTR整合了六个常用英汉文本检索数据集中的37,317条唯一查询,覆盖多个领域和多样化查询类型。采用200名真实说话人语音样本,混合17类真实环境噪声,在可控信噪比(SNR)下合成语音,实现从安静到高噪声条件的可复现评估。基于统一协议,对代表性级联式与端到端检索系统进行大规模评测。结果表明,随着噪声增加,检索性能普遍下降,不同系统间降幅差异显著;即使大型检索模型在极端噪声下仍表现不佳,说明鲁棒性仍是关键瓶颈。SQuTR为鲁棒性评测与诊断分析提供了可复现的测试平台,推动未来研究发展。

原文摘要 · Abstract (English)

Spoken query retrieval is an important interaction mode in modern information retrieval. However, existing evaluation datasets are often limited to simple queries under constrained noise conditions, making them inadequate for assessing the robustness of spoken query retrieval systems under complex acoustic perturbations. To address this limitation, we present SQuTR, a robustness benchmark for spoken query retrieval that includes a large-scale dataset and a unified evaluation protocol. SQuTR aggregates 37,317 unique queries from six commonly used English and Chinese text retrieval datasets, spanning multiple domains and diverse query types. We synthesize speech using voice profiles from 200 real speakers and mix 17 categories of real-world environmental noise under controlled SNR levels, enabling reproducible robustness evaluation from quiet to highly noisy conditions. Under the unified protocol, we conduct large-scale evaluations on representative cascaded and end-to-end retrieval systems. Experimental results show that retrieval performance decreases as noise increases, with substantially different drops across systems. Even large-scale retrieval models struggle under extreme noise, indicating that robustness remains a critical bottleneck. Overall, SQuTR provides a reproducible testbed for benchmarking and diagnostic analysis, and facilitates future research on robustness in spoken query to text retrieval.

语音检索鲁棒性噪声数据集信息检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。