轻量模型一键修复手机录音的多类失真问题
Smule Renaissance Small: Efficient General-Purpose Vocal Restoration
- 直接在复数短时傅里叶域端到端修复,支持大分析窗提升频谱精度
- iPhone 12 CPU 实现 10.5 倍实时推理,48kHz 下表现超越多个基线
- 新数据集 EDB 模拟真实严苛场景,适合语音修复与音视频应用研究
消费设备录制的歌声常受噪声、混响、带宽限制和削波等多重退化影响。我们提出 Smule Renaissance Small(SRS),一个紧凑的单阶段模型,直接在复数短时傅里叶变换(STFT)域实现端到端歌声修复。通过引入相位感知损失,SRS 在保持高频率分辨率的同时,可在 iPhone 12 CPU 上以 48 kHz 采样率实现 10.5 倍实时推理速度。在 DNS 5 挑战赛盲测集上,尽管未进行语音训练,其性能仍优于强基线 GAN 模型,并接近计算开销高昂的流匹配系统。为评估真实多退化场景下的表现,我们构建了极端退化基准测试集 EDB:包含 87 条在恶劣声学条件下录制的歌唱与语音样本。在 EDB 上,SRS 在歌唱修复中超越所有开源基线,媲美商用系统;在语音修复上虽无专门训练,仍保持竞争力。SRS 与 EDB 已开源,采用 MIT 许可证。
原文摘要 · Abstract (English)
Vocal recordings on consumer devices commonly suffer from multiple concurrent degradations: noise, reverberation, band-limiting, and clipping. We present Smule Renaissance Small (SRS), a compact single-stage model that performs end-to-end vocal restoration directly in the complex STFT domain. By incorporating phase-aware losses, SRS enables large analysis windows for improved frequency resolution while achieving 10.5x real-time inference on iPhone 12 CPU at 48 kHz. On the DNS 5 Challenge blind set, despite no speech training, SRS outperforms a strong GAN baseline and closely matches a computationally expensive flow-matching system. To enable evaluation under realistic multi-degradation scenarios, we introduce the Extreme Degradation Bench (EDB): 87 singing and speech recordings captured under severe acoustic conditions. On EDB, SRS surpasses all open-source baselines on singing and matches commercial systems, while remaining competitive on speech despite no speech-specific training. We release both SRS and EDB under the MIT License.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。