融合测试时自适应与语言模型重打分,提升语音识别鲁棒性
SUTA-LM: Bridging Test-Time Adaptation and Language Model Rescoring for Robust ASR
- 基于熵最小化实现可控自适应,自动选择最优调整步长
- 在18个数据集上平均字错误率降低3.2%,跨领域性能更稳定
- 适合实际部署中应对语音分布差异的场景,如远场、口音等
尽管端到端语音识别取得进展,真实场景中的域差异仍导致性能下降。测试时自适应(TTA)通过推理阶段调整模型来缓解此问题。近期研究尝试将TTA与外部语言模型结合,采用束搜索重打分或生成式纠错等方法。本文发现,TTA可能干扰语言模型重打分,揭示两者结合存在非平凡挑战。为此,提出SUTA-LM——一种对基于熵最小化的TTA方法SUTA的简单而有效的扩展,引入语言模型重打分。SUTA-LM首先利用声学与语言信息联合引导的自动步长选择机制进行受控适应,再通过语言模型重打分优化输出。在18个多样化语音识别数据集上的实验表明,SUTA-LM在广泛域上均表现稳健。
原文摘要 · Abstract (English)
Despite progress in end-to-end ASR, real-world domain mismatches still cause performance drops, which Test-Time Adaptation (TTA) aims to mitigate by adjusting models during inference. Recent work explores combining TTA with external language models, using techniques like beam search rescoring or generative error correction. In this work, we identify a previously overlooked challenge: TTA can interfere with language model rescoring, revealing the nontrivial nature of effectively combining the two methods. Based on this insight, we propose SUTA-LM, a simple yet effective extension of SUTA, an entropy-minimization-based TTA approach, with language model rescoring. SUTA-LM first applies a controlled adaptation process guided by an auto-step selection mechanism leveraging both acoustic and linguistic information, followed by language model rescoring to refine the outputs. Experiments on 18 diverse ASR datasets show that SUTA-LM achieves robust results across a wide range of domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。