arXiv:2512.16395eess.AS2025-12中稿 · ICASSP 2026

提升语音检索的准确与效率,增强抗噪能力。

BEST-STD2.0: Balanced and Efficient Speech Tokenizer for Spoken Term Detection

  • 通过噪声和混响增强训练提升分词器鲁棒性。
  • 引入最优传输正则化实现分词使用均衡,提高效率。
  • 结合TF-IDF加速检索,适合实时语音搜索场景。

快速准确的语音内容检索对语音搜索等应用至关重要。查询式语音片段检测(Query-by-Example Spoken Term Detection, STD)旨在给定一段语音查询时,从音频数据库中检索出匹配片段。基于分词的STD系统利用离散语音表示实现高效检索,但在噪声和混响环境下表现不佳,且分词使用效率低。本文提出一种噪声与混响增强训练策略,提升分词器鲁棒性;引入基于最优传输的正则化方法,确保分词使用均衡,提升分词效率;并采用TF-IDF-based检索机制进一步加速搜索。实验证明,所提方法在多种失真条件下均优于现有基线,同时保持高检索效率。

原文摘要 · Abstract (English)

Fast and accurate spoken content retrieval is vital for applications such as voice search. Query-by-Example Spoken Term Detection (STD) involves retrieving matching segments from an audio database given a spoken query. Token-based STD systems, which use discrete speech representations, enable efficient search but struggle with robustness to noise and reverberation, and with inefficient token utilization. We address these challenges by proposing a noise and reverberation-augmented training strategy to improve tokenizer robustness. In addition, we introduce optimal transport-based regularization to ensure balanced token usage and enhance token efficiency. To further speed up retrieval, we adopt a TF-IDF-based search mechanism. Empirical evaluations demonstrate that the proposed method outperforms STD baselines across various distortion levels while maintaining high search efficiency.

语音检索分词器噪声鲁棒

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。