arXiv:2602.12701cs.SD2026-02中稿 · 2026 IEEE Internat…

分离语音表征,让模型能通用修复各种噪声下的语音。

DisSR: Disentangling Speech Representation for Degradation-Prior Guided Cross-Domain Speech Restoration

  • 用退化先验提取与说话人无关的干扰特征,指导修复过程。
  • 跨域对齐训练使模型在未见场景下仍保持良好表现。
  • 适合需要通用语音修复能力的研究与应用者。

以往语音修复(SR)多聚焦于单任务修复(SSR),难以应对多样化退化问题。为不同失真训练专用模型耗时且缺乏泛化性。此外,多数研究忽视模型在未见域上的适应能力。为此,我们提出DisSR,一种基于解耦语音表征的通用语音修复模型,具备两大特性:1)退化先验引导,通过提取与说话人无关的退化表征,指导基于扩散模型的语音修复;2)跨域适配,设计跨域对齐训练策略,增强模型在跨域数据上的适应性与泛化能力。实验表明,该方法在多种失真条件下均能生成高质量修复语音。音频样本可访问 https://itspsp.github.io/DisSR。

原文摘要 · Abstract (English)

Previous speech restoration (SR) primarily focuses on single-task speech restoration (SSR), which cannot address general speech restoration problems. Training specific SSR models for different distortions is time-consuming and lacks generality. In addition, most studies ignore the problem of model generalization across unseen domains. To overcome those limitations, we propose DisSR, a Disentangling Speech Representation based general speech restoration model with two properties: 1) Degradation-prior guidance, which extracts speaker-invariant degradation representation to guide the diffusion-based speech restoration model. 2) Domain adaptation, where we design cross-domain alignment training to enhance the model's adaptability and generalization on cross-domain data, respectively. Experimental results demonstrate that our method can produce high-quality restored speech under various distortion conditions. Audio samples can be found at https://itspsp.github.io/DisSR.

语音修复扩散模型跨域适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。