arXiv:2601.21084cs.CLcs.SD2026-01中稿 · ICASSP 2026

解决语音增强模型微调时位置敏感问题,提升降噪效果

Position-invariant Fine-tuning of Speech Enhancement Models with Self-supervised Speech Representations

  • 用零填充和软DTW损失实现位置无关微调
  • 软DTW方法收敛更快,下游任务性能更好
  • 适合做语音增强与自监督语音模型结合的研究者

将前端语音增强(SE)模型与基于自监督学习(SSL)的语音模型结合,在噪声环境下对下游任务有效。通常使用均方误差(MSE)损失在增强语音与干净语音之间进行微调,但该方法容易利用SSL模型中的位置嵌入,使目标通过位置相关性而非内容信息最小化。本文将此问题视为自监督表示微调的一般局限,并通过表示引导的语音增强进行研究。提出两种策略:(1)零填充,此前在SSL预训练中探索过,现首次应用于微调场景;(2)速度扰动结合软动态时间规整(soft-DTW)损失。实验表明,软DTW方法具有更快收敛速度和更优下游性能,凸显了在基于SSL的语音建模中采用位置无关微调的重要性。

原文摘要 · Abstract (English)

Integrating front-end speech enhancement (SE) models with self-supervised learning (SSL)-based speech models is effective for downstream tasks in noisy conditions. SE models are commonly fine-tuned using SSL representations with mean squared error (MSE) loss between enhanced and clean speech. However, MSE is prone to exploiting positional embeddings in SSL models, allowing the objective to be minimised through positional correlations instead of content-related information. This work frames the problem as a general limitation of self-supervised representation fine-tuning and investigates it through representation-guided SE. Two strategies are considered: (1) zero-padding, previously explored in SSL pre-training but here examined in the fine-tuning setting, and (2) speed perturbations with a soft-DTW loss. Experiments show that the soft-DTW-based approach achieves faster convergence and improved downstream performance, underscoring the importance of position-invariant fine-tuning in SSL-based speech modelling.

语音增强自监督学习位置不变性软DTW

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。