arXiv:2512.00482eess.AS2025-12

探究语音增强模型在噪声和混响下的内部适应位置

Where Does Speech Enhancement Adapt? Probing Study Under Controlled Degradation

  • 通过可控降质测试模型各层激活,分析内部表征变化
  • 编码器保持噪声不变表征,解码器随降质程度线性适应
  • 结果与性能指标(如PESQ)相关,适用于模型优化研究

语音增强(SE)模型发展迅速,但输入信号退化如何影响其内部表征仍不明确。本文提出一种探测方法,通过控制信噪比(SNR)和混响时间C50,在MUSE SE模型中提取各层激活,并使用中心核对齐(CKA)度量层间表征相似性,回归降质程度,获得稳健的自适应表征轮廓。编码器层保持噪声不变表征,解码器层适应性强,敏感度在模块内单调上升,跳跃连接处出现最显著转变。该结构在混响下同样存在,且独立于架构差异(在MP-SENet和Demucs中重现),表明适应机制由增强目标驱动而非特定设计。结果揭示了模型适应退化的内在位置,并关联内部表征与输出性能(如PESQ)。

原文摘要 · Abstract (English)

Speech enhancement (SE) models advance rapidly, yet it remains underexplored how degradation of input signals affects their internal representations. We introduce a probing process, aimed at modeling the behavior of internal representations in SE models under controlled degradations to input signals. We apply it to the MUSE SE model by extracting its layer activations under controlled Signal-to-Noise Ratio (SNR) and reverberation C50. We measure layer-wise representational similarity to clean input references using Centered Kernel Alignment (CKA) and regress it against the degradation level, yielding compact, robustness-adaptive profiles. Encoder layers maintain noise-invariant representations while decoder layers adapt strongly, with sensitivity increasing monotonically within blocks and skip-connection boundaries marking the sharpest transitions. The same structure emerges under reverberation and is reproduced independently by MP-SENet and Demucs, two structurally distinct architectures, suggesting that the tradeoff is induced by the enhancement objective rather than a particular model design. Together, these results characterize where SE models adapt to degradation. We then offer insight into how internal representations correlate with output-level performance metrics, e.g., PESQ.

语音增强表征分析降质探测PESQ

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。