arXiv:2602.04217cs.SDcs.CL2026-02中稿 · ICASSP 2026

提出前端语音分词增强方法,提升噪声环境下语音识别效果。

Frontend Token Enhancement for Token-Based Speech Recognition

  • 用波形直接预测干净语义分词,跳过中间特征表示
  • 在CHiME-4数据集上,波形到分词模型效果最佳
  • 适合噪声环境下的语音识别系统优化

语音信号的离散化表示(如基于自监督学习模型聚类得到的语义或音素分词)是自动语音识别(ASR)和语音语言模型等任务的有效替代方案。然而,这类分词易受环境噪声影响,导致后端任务性能下降。本文提出一种前端系统,从含噪语音中估计出干净语音分词,并在使用语义分词的ASR后端上进行评估。考虑四种输入输出域不同的增强模型:波形到波形、分词到分词、连续SSL特征到分词、波形到分词。这些模型与ASR后端独立训练。在CHiME-4数据集上的实验表明,波形到分词增强模型在各类前端中表现最优,且多数情况下优于基于连续SSL特征的ASR系统。

原文摘要 · Abstract (English)

Discretized representations of speech signals are efficient alternatives to continuous features for various speech applications, including automatic speech recognition (ASR) and speech language models. However, these representations, such as semantic or phonetic tokens derived from clustering outputs of self-supervised learning (SSL) speech models, are susceptible to environmental noise, which can degrade backend task performance. In this work, we introduce a frontend system that estimates clean speech tokens from noisy speech and evaluate it on an ASR backend using semantic tokens. We consider four types of enhancement models based on their input/output domains: wave-to-wave, token-to-token, continuous SSL features-to-token, and wave-to-token. These models are trained independently of ASR backends. Experiments on the CHiME-4 dataset demonstrate that wave-to-token enhancement achieves the best performance among the frontends. Moreover, it mostly outperforms the ASR system based on continuous SSL features.

语音识别前端增强分词噪声鲁棒

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。