arXiv:2606.09557eess.AS2026-06中稿 · Interspeech 2026

U-Net模型暗中学习房间声学特征,可直接用于声学还原。

Your U-Net Dereverberation Model is Secretly an RIR Encoder

  • 利用自监督对比学习预训练房间冲激响应嵌入,显式引导网络
  • 条件化后推理步数减少50%以上,去混响性能显著提升
  • 适用于需要快速高效语音还原的实时场景

本文研究基于NCSN++ U-Net的音频去混响模型在中间表示中捕捉全局房间特性能力。通过实证分析先进扩散模型与判别式模型,发现深层特征编码了结构化的、依赖于房间冲激响应(RIR)的嵌入。该隐式房间表征的判别能力与去混响性能在客观指标上高度相关。受此启发,我们提出一种训练策略:显式地将预训练的RIR嵌入(通过自监督对比学习获得)引入网络。实验表明,该方法提升了表示质量,加速收敛,并显著增强去混响效果,同时使扩散模型在推理阶段所需反向扩散步数大幅减少。

原文摘要 · Abstract (English)

In this work, we analyze the ability of NCSN++ U-Net based audio dereverberation models to capture global room characteristics in their intermediate representations. Through an empirical study of both a state-of-the-art diffusion-based model and a discriminative counterpart, we show that deeper layers encode structured room impulse response (RIR)-dependent embeddings. Moreover, the discriminative ability of this implicit room representation correlates with dereverberation performance across objective metrics. Motivated by this observation, we propose a training strategy that explicitly conditions the network on pre-trained RIR embeddings, obtained via self-supervised contrastive learning. Incorporating RIR conditioning improves representation quality, accelerates convergence, and enhances dereverberation performance, while significantly reducing the number of reverse diffusion steps required by the diffusion-based model during inference.

去混响自监督扩散模型声学建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。