arXiv:2505.05077cs.SDeess.AS2025-05被引 2

让语音修复保留混响特征并可调控,提升真实感。

ReverbMiipher: Generative Speech Restoration meets Reverberation Characteristics Controllability

  • 用专用编码器提取混响特征,控制语音重建
  • 在去噪同时保留原始混响,性能优于传统方法
  • 支持混响风格迁移与新效果生成,适合音频创作

混响包含声源环境的空间信息,但传统语音修复(SR)通常完全消除混响。我们提出 ReverbMiipher,一种扩展参数重合成框架的语音修复模型,可在去噪的同时保留并控制混响特性。ReverbMiipher 引入专用的 ReverbEncoder,从含噪输入中提取混响特征向量,该特征用于条件化声码器重建语音信号,实现去噪并保留原始混响。训练中采用随机零向量替换策略,确保特征仅编码混响,与其它语音属性解耦。该可学习表示支持通过特征插值、替换或潜空间采样等方式实现混响控制。客观与主观评估表明,ReverbMiipher 有效保留混响、去除噪声,优于传统两阶段 SR 和卷积模拟混响响应的方法。进一步验证其可通过特征操作生成新颖混响效果。

原文摘要 · Abstract (English)

Reverberation encodes spatial information regarding the acoustic source environment, yet traditional Speech Restoration (SR) usually completely removes reverberation. We propose ReverbMiipher, an SR model extending parametric resynthesis framework, designed to denoise speech while preserving and enabling control over reverberation. ReverbMiipher incorporates a dedicated ReverbEncoder to extract a reverb feature vector from noisy input. This feature conditions a vocoder to reconstruct the speech signal, removing noise while retaining the original reverberation characteristics. A stochastic zero-vector replacement strategy during training ensures the feature specifically encodes reverberation, disentangling it from other speech attributes. This learned representation facilitates reverberation control via techniques such as interpolation between features, replacement with features from other utterances, or sampling from a latent space. Objective and subjective evaluations confirm ReverbMiipher effectively preserves reverberation, removes other artifacts, and outperforms the conventional two-stage SR and convolving simulated room impulse response approach. We further demonstrate its ability to generate novel reverberation effects through feature manipulation.

语音修复混响控制声码器特征解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。