arXiv:2601.22480cs.SDeess.AS2026-01中稿 · ICASSP 2026

通过语音语义信息优化降噪模型特征聚合,提升识别准确率。

Rethinking Speech Representation Aggregation in Speech Enhancement: A Phonetic Mutual Information Perspective

  • 用音素互信息衡量噪声下语音表征的语义保持能力
  • 预训练特征聚合层最大化音素信息,再冻结用于降噪
  • 解耦训练使语义信息更鲁棒,显著降低词错误率

近期语音增强(SE)模型越来越多地利用自监督学习(SSL)表征以获取丰富的语义信息。通常通过轻量级适配模块将中间特征聚合为单一表示。然而,大多数SSL模型并未针对噪声鲁棒性训练,可能导致语义表征受损。此外,适配模块与SE模型联合训练,可能优先关注声学细节而非语义信息,违背初衷。为此,我们从信息论角度分析了SSL模型在含噪语音上的表现,具体测量受损的SSL表征与对应音素标签之间的互信息(MI),聚焦语言内容的保留程度。基于此分析,我们提出语言聚合层,该层预先训练以最大化与音素标签的互信息(支持动态聚合),随后在SE训练中冻结。实验表明,这种解耦方法在词错误率(WER)上优于联合优化基线,证明显式对齐适配模块与语言内容的优势。

原文摘要 · Abstract (English)

Recent speech enhancement (SE) models increasingly leverage self-supervised learning (SSL) representations for their rich semantic information. Typically, intermediate features are aggregated into a single representation via a lightweight adaptation module. However, most SSL models are not trained for noise robustness, which can lead to corrupted semantic representations. Moreover, the adaptation module is trained jointly with the SE model, potentially prioritizing acoustic details over semantic information, contradicting the original purpose. To address this issue, we first analyze the behavior of SSL models on noisy speech from an information-theoretic perspective. Specifically, we measure the mutual information (MI) between the corrupted SSL representations and the corresponding phoneme labels, focusing on preservation of linguistic contents. Building upon this analysis, we introduce the linguistic aggregation layer, which is pre-trained to maximize MI with phoneme labels (with optional dynamic aggregation) and then frozen during SE training. Experiments show that this decoupled approach improves Word Error Rate (WER) over jointly optimized baselines, demonstrating the benefit of explicitly aligning the adaptation module with linguistic contents.

语音增强自监督学习音素信息特征聚合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。