用扩散模型无监督还原混响语音,提升多麦克风语音清晰度。
Unsupervised Multi-channel Speech Dereverberation via Diffusion
- 基于扩散模型后验采样,以纯净语音为先验指导去混响。
- 每步采样估计各麦克风的冲激响应,约束多通道混合一致性。
- 结合优化与解析方法高效估计多通道冲激响应,适合语音增强研究者。
针对多通道单说话人盲去混响问题,提出无监督语音去混响方法USD-DPS。该方法利用无条件纯净语音扩散模型作为强先验,通过后验采样求解。在每一步扩散采样中,估计所有麦克风通道的房间脉冲响应(RIR),并用于施加多通道混合一致性约束以指导扩散过程。对于多通道RIR估计,参考通道的RIR通过优化子带RIR信号模型的参数得到,采用Adam优化器;非参考通道的RIR则通过前向卷积预测(FCP)解析估计。该组合在采样效率与RIR先验建模间取得良好平衡,在无监督去混响方法中表现更优。音频演示页见 https://usddps.github.io/USDDPS_demo/。
原文摘要 · Abstract (English)
We consider the problem of multi-channel single-speaker blind dereverberation, where multi-channel mixtures are used to recover the clean anechoic speech. To solve this problem, we propose USD-DPS, {U}nsupervised {S}peech {D}ereverberation via {D}iffusion {P}osterior {S}ampling. USD-DPS uses an unconditional clean speech diffusion model as a strong prior to solve the problem by posterior sampling. At each diffusion sampling step, we estimate all microphone channels' room impulse responses (RIRs), which are further used to enforce a multi-channel mixture consistency constraint for diffusion guidance. For multi-channel RIR estimation, we estimate reference-channel RIR by optimizing RIR parameters of a sub-band RIR signal model, with the Adam optimizer. We estimate non-reference channels' RIRs analytically using forward convolutive prediction (FCP). We found that this combination provides a good balance between sampling efficiency and RIR prior modeling, which shows superior performance among unsupervised dereverberation approaches. An audio demo page is provided in https://usddps.github.io/USDDPS_demo/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。